Two-Key Engineering: Maker-Checker for Software That AI Writes

Agents write the code, but nothing ships until two keys turn: an approved requirement, and test evidence produced independently of the code it judges.

Anyone who has worked in a bank knows maker-checker. One person initiates the entry. A different person authorises it. Not because the first person is careless, but because one pair of eyes on a consequential action is a control failure waiting to happen. Defence and nuclear operations run the same idea as the two-person rule: two keys have to turn before anything happens.

Software delivery has quietly let go of that principle in the last two years.

AI agents now write a large share of production code. In April 2026, Google said 75 percent of its new code is AI-generated and approved by engineers. Read the second half of that sentence carefully: every line still needs a person to say it is right. That approval step is where the strain has landed, and it is where most teams have only one key turning.

Two-Key Engineering: agents write the code, but nothing ships unless two keys turn. The first is an approved requirement. The second is independent evidence that what was built matches it.

Why one key is not enough

The evidence on AI coding is not that it makes teams worse. It is that the gains leak away downstream unless verification scales with them.

Output went up. Checking did not.
75%of new code at Google is AI-generated, and approved by engineers. Every line still needs a person to say it is right.Google, April 2026
+98%more pull requests merged on teams with high AI adoption.Faros AI, 10,000+ developers
+91%longer spent in review on those same teams. The queue moved to the checker.Faros AI, same study
19%sloweron real tasks for 16 experienced developers, working in repositories they already knew with early-2025 tools, while they believed they had been about 20% faster.METR randomised trial
None of these are tiramai measurements. The Faros figures are vendor telemetry from Faros AI, not independent research; the others come from Google and published studies. Generation scaled; the checking stage did not, so the queue moved to the reviewer.

The pattern holds across studies. Google’s DORA research finds AI adoption now goes with higher delivery throughput, and also with more delivery instability. In Stack Overflow’s 2025 developer survey, more developers distrusted the accuracy of AI output than trusted it, and their top frustration was solutions that are almost right, but not quite. Output went up. Confidence went up. Checking capacity did not.

The failure that is easy to miss: the agent marking its own homework

There is one risk in agentic delivery that deserves a plain name. When the same agent writes the code and the tests, the tests tend to describe what the code does, not what it was supposed to do. If the agent misread the requirement, the code is wrong and the tests confirm the wrong behaviour. Everything goes green and nobody learns anything.

The agent marking its own homework
One key: the maker checks itself
Approved requirement REQ-207

A refund above ₹50,000 needs a checker's approval before it posts.

Code the agent wroteif refund > 5_00_000:
require_checker()
Tests the same agent wrote, from the codeneeds_checker(6_00_000) pass
needs_checker(1_00_000) is False pass
All green. A ₹1 lakh refund posts with no checker, and nothing says so.
Two keys: tests from the requirement
Approved requirement REQ-207

A refund above ₹50,000 needs a checker's approval before it posts.

Code the agent wroteif refund > 5_00_000:
require_checker()
Tests written from REQ-207needs_checker(50_001) fail
needs_checker(1_00_000) fail
Red before release. The misread limit is caught where it is cheap.
An illustrative example. The code misreads ₹50,000 as ₹5 lakh. Tests written from that code agree with it; tests written from the requirement do not.

This is not hypothetical. A December 2024 University of Waterloo study of LLM-based test generators found that tools which build tests by observing existing code can end up validating bugs, and discarding the very tests that would have exposed them, because a failing test is assumed to be the test’s fault. In maker-checker language: the maker signed the checker’s box.

Safety-critical engineering settled this long ago through independence. DO-178C in aviation requires traceability from each requirement to the code and the tests that verify it, with certain verification carried out independently of development. ISO 26262 in automotive and IEC 62304 in medical device software require traceable, documented verification, and the rigour they expect scales with the safety or risk class, with ISO 26262 setting explicit levels of independence for the most critical functions. The principle is simple: the thing being built and the thing doing the checking cannot share a brain.

What if the requirement came from the code?

On legacy systems, the requirement often does not exist. Nobody wrote it down, so it is recovered from the code: someone, or something, reads the logic and describes what it does.

That is useful work, but be clear about what it produces. A recovered requirement is a description of behaviour, not a statement of intent. Tests built from it can only prove that the behaviour has not changed. If the code was wrong on the day the requirement was recovered, the tests will protect the mistake.

The second key only turns once a subject matter expert has reviewed that recovered requirement and approved it, because that is the moment human intent enters the chain. Recovery gives you the draft. Approval gives you the key.

Where the requirement came from decides what the tests prove
StateWhere it came fromWhat tests from it prove
DerivedRecovered from code, unreviewedBehaviour has not changed
ApprovedRecovered, then reviewed and signed by a subject matter expertThe code matches intent
AuthoredWritten from business or regulatory intentThe code matches intent, independently

This belongs in the release report too. A report that separates regression evidence, the behaviour has not changed, from conformance evidence, the code matches intent, is more trustworthy than one that implies everything was verified against intent.

Key one: the approved requirement

You cannot verify against intent you never wrote down. The first key is a requirement that is clear, approved and connected to the code it governs. Approval means a named person signs, typically a product owner, a subject matter expert or a risk owner depending on the change, and that person should not be the one who ran the agent. Without it, an agent builds from whatever it happens to find, which is how context poisoning and drift get in.

This is the job of ProductSphere, agentic product intelligence for enterprise systems. It connects requirements, code, defects, documentation and releases into one governed knowledge layer, where every answer cites the exact line or artefact it came from. For teams sitting on decades of legacy that matters more than any coding assistant: a copilot sees the file in front of it, while ProductSphere understands the product around it.

  • Requirements intelligence. Turning unclear requirements into buildable statements, and taking a BRD through to use cases and development-ready inputs.
  • Requirements-to-code traceability, in both directions, with gap detection and ISO, IEC and DO traceability evidence on demand.
  • Change and release impact. Seeing what a change breaks before anyone makes it, across the product and downstream systems.
  • Governed autonomy. Configurable action autonomy, recommend, draft or execute, with human approval on every critical action, granular access control and full audit trails.

That last point is deliberate. The platform is not human-free. It is maker-checker built into the tooling, rather than bolted on as a policy document. It is also what makes context coding possible: agents building from approved requirements rather than retyped prompts.

Key two: independent evidence

The second key is proof that what was built matches what was approved, produced by something other than the thing that built it.

This is AionQ, tiramai’s AI quality and test automation platform, and one design choice matters above all: AionQ generates tests from requirements, not from the code. Its agents read your requirements and return tests that are written, automated and ready to run across the UI, API and data layers. The checker is anchored to intent, not to the implementation it is meant to judge.

  • Validation before testing. Every requirement is checked for atomicity, ambiguity and completeness first. Vague in, flagged; clear in, tested. A requirement too loose to test is too loose to build from, and catching that at the start is cheaper than at the end.
  • Bidirectional traceability. Every requirement is mapped to its tests, scripts and results, and every test links back to why it exists. Audit answers take seconds, not sprints, and release readiness becomes a number you can stand behind instead of a judgement call.
  • Governed self-healing. When the UI shifts and a locator breaks, AionQ proposes a repair with a confidence score, and every fix is proposed, reviewed and logged. A checker that silently repairs itself is not a checker. One that shows its work is.

How the two keys turn together

How the two keys turn together
  1. 1. Requirement approvedWritten, clarified and signed off, then held in the product knowledge layer and linked to the code it governs. Key 1ProductSphere
  2. 2. Agent builds from itThe coding agent works from the approved requirement, not from a prompt someone retyped from memory.
  3. 3. Tests derived independentlyGenerated and run from the same requirement, never from the code the agent produced. Key 2AionQ
  4. 4. Evidence links backRequirement to code to test to result, traceable in both directions.
  5. 5. Release on evidenceIf either key is missing, nothing ships.
Two keys, one trace. The requirement that authorised the change and the test that proved it point at each other.

Both products run on the tiramai AI platform, with the same governance, audit trails and deployment control, including fully on-premise with your own LLM for teams whose code and requirements cannot leave the building.

Does this slow everything down?

Two keys does not mean two teams, or twice the work. It means the evidence comes from somewhere other than the code. Scale the ceremony to the risk. A copy change on a help page does not need the same treatment as a refund limit, a tariff calculation or a dosing rule. Start with the rules where being wrong is expensive, and let low-risk changes keep moving the way they do today.

Why regulators will land here first

Indian regulators have already been clear that responsibility does not transfer to the machine. The Reserve Bank of India’s FREE-AI Committee report states that accountability rests with the entities deploying AI and cannot be delegated to the model and the underlying algorithm. SEBI’s June 2025 consultation on AI and ML in securities markets proposes governance, testing and auditability expectations, with oversight from senior management.

Those documents address AI systems rather than coding agents, and the RBI report is advisory while the SEBI paper was a consultation. But the principle travels without modification. If an agent wrote it, your organisation still owns the outcome, and you will be asked to show who approved what, and what proved it.

Maker-checker was never about distrust. It was about evidence. That is what Two-Key Engineering restores.

A seven-point check of your own setup

Answer honestly about how code reaches production in your team today.

Are you running on one key?
  1. 1Can an agent start work without an approved requirement attached?
  2. 2Are your requirements testable, not just readable?
  3. 3Are your tests written by something other than the agent that wrote the code, and without seeing that code first?
  4. 4Can you trace a shipped change back to the requirement that authorised it, and forward to the test that proved it?
  5. 5When a test heals itself, is there a record of what changed and who approved it?
  6. 6Is your release decision a number backed by evidence, rather than a meeting?
  7. 7Would all of this survive an auditor asking for it on a Friday afternoon?
0 of 7 answered
More than two uncomfortable answers and you are running on one key. Your answers stay in your browser.

If more than two answers were uncomfortable, here is where to start.

  1. List the ten rules in your system where being wrong costs the most.
  2. For each one, check whether a current, approved requirement exists. Most teams find fewer than half.
  3. Write tests for those ten from the requirement, not from the code, and see what goes red.

Something always goes red, and that is the finding.

Closing

Agents will keep writing more of the code. That is not the risk. The risk is letting the one that wrote it be the one that signs it off.

Two keys. One trace. Nothing ships on a single signature.

See how the first key works in ProductSphere, how the second turns in AionQ, or talk to our team about putting both on your delivery pipeline.

Sources

  • #maker-checker
  • #AI coding agents
  • #test automation
  • #traceability
  • #BFSI
  • #governance
Share
Written bytiramai team
Follow tiramai on LinkedIn
Talk to us

Working through this for your organisation? Our team can walk you through it on your own use case.

See the platform

How tiramai runs governed AI in the cloud, your private cloud or fully on-premises.

Explore ProductSphere
FAQ

Questions About This Topic

Two-Key Engineering applies the banking maker-checker control to software that AI agents write. Agents write the code, but nothing ships unless two keys turn: an approved requirement, and independent test evidence that what was built matches it, joined by one trace.