Context Coding: What Comes After Vibe Coding for Banks and Regulated Teams

Context coding is AI coding from approved requirements passed machine to machine, not retyped prompts, so shipped code traces back to what was signed off.

If you run engineering, risk or change management inside a bank, an insurer or any data-regulated team, you already know the tension. Your business wants the speed of AI coding. Your auditors want to know exactly which approved requirement produced which line of code, and who signed off. Right now those two goals pull against each other. This article is about closing that gap.

What vibe coding is, and where it breaks

In February 2025, Andrej Karpathy, a co-founder of OpenAI and former AI director at Tesla, described a new way of working with AI. He called it vibe coding: you talk to the model, you accept what it gives you, and you "forget that the code even exists." It was a good description of something real. The term spread so fast that Collins named it Word of the Year on 6 November 2025.

Karpathy was careful, though. He said it was "not too bad for throwaway weekend projects." He did not say it was fine for a core banking system.

That caveat matters. Vibe coding is one human describing intent to a machine, one prompt at a time. It is intuitive and fast for a prototype. It is also inefficient and hard to trust at scale. As a conversation grows, the model carries more and more history in its context window, which burns tokens and can crowd out the detail that actually mattered. The output varies from run to run.

And the security record is not reassuring. Veracode's 2025 GenAI Code Security Report, which tested more than 100 large language models across 80 tasks, found that when a model had a choice between a secure and an insecure way to write code, it chose the insecure option 45 percent of the time. Wiz Research, studying applications built on vibe-coding platforms, found a pattern of common, high-impact misconfigurations and reported that 1 in 5 organizations building on these platforms were inadvertently exposing themselves to risk.

The deeper problem for regulated teams: drift

Your requirements were reviewed, approved and baselined in a document. Then someone rewrites them, informally, as a prompt. The version that gets built is now a paraphrase of a paraphrase. When an auditor asks you to trace a control back to its approved requirement, the thread has been cut.

A paraphrase of a paraphrase
Vibe coding: retyped by hand
Approved requirement REQ-114

Hold any transfer above ₹10,00,000 until a second approver signs off.

Prompt, day 1

“add an approval step for big transfers”

Prompt, day 9

“large transfers, so 1 lakh and up, need approval”

Shipped code if amount > 100_000:
require_approval()
Trace broken. Limit is 10× too low, one approver missing.
Context coding: handed machine to machine
Approved requirement REQ-114

Hold any transfer above ₹10,00,000 until a second approver signs off.

Machine-readable artefact REQ-114{ "trace": "REQ-114",
"limit_inr": 1000000,
"approvers": 2 }
Shipped code REQ-114# REQ-114
if amount > 1_000_000:
require_second_approver()
Trace intact. Same limit, same approvers.
An illustrative example. Each retelling sounds reasonable on its own, but together they move the limit from 10 lakh to 1 lakh and drop the second approver, and nothing links the shipped code back to REQ-114.

What context coding is

Context coding is the opposite handoff. Instead of a human retyping approved intent into a chat box, the approved context is passed machine to machine. Coding agents work directly from structured, approved requirements rather than from a fresh prompt written from memory.

Put simply: vibe coding is human to machine, one prompt at a time. Context coding is machine to machine, from an approved source of truth. The human still reviews and approves the requirement. What changes is that the agent no longer needs anyone to re-describe it.

Where it fits with context engineering and spec-driven development

This is not a rejection of the good ideas already in the field. It builds on them, and two adjacent ideas deserve credit.

The first is context engineering. Shopify CEO Tobi Lütke posted on 18 June 2025 that he preferred the term over prompt engineering, calling it "the art of providing all the context for the task to be plausibly solvable by the LLM." Karpathy amplified it a week later, and Anthropic formalized it on 29 September 2025 as "the set of strategies for curating and maintaining the optimal set of tokens during LLM inference." In short: give a model the right information, in the right form, at the right time, rather than just wording a single prompt well.

The second is spec-driven development, seen in tools like GitHub's open-source Spec Kit and Amazon's Kiro, where an approved specification, not a chat prompt, is the thing the agent builds from.

Built on what came before
  1. 1Prompt engineeringWord one prompt well.The starting point
  2. 2Context engineeringCurate everything the model sees, not just the prompt.Tobi Lütke, June 2025. Anthropic, September 2025
  3. 3Spec-driven developmentMake an approved spec the source of truth.GitHub Spec Kit, Amazon Kiro
  4. 4Context codingThe spec is reviewed, approved, baselined and traceable, and reaches the agent automatically.For regulated teams
Context engineering says: curate what the model sees. Spec-driven development says: make the spec the source of truth. Context coding adds that, in a regulated setting, the spec must be approved and traceable, and the handoff automatic.

Why the machine-to-machine handoff is more efficient

Three reasons.

Tokens

Agentic coding sessions are token-hungry by nature. Independent analyses note that agents can use several times more tokens than a simple chat, largely because context accumulates turn after turn. When an agent starts from a clean, structured artefact instead of reconstructing intent through a long conversation, there is simply less to re-read and re-explain. Research on shared, structured context has shown token reductions of over 50 percent in some setups. We will not put a number on ProductSphere yet, but the direction is well established.

What the model has to carry
Vibe coding session
  1. 1Describe the feature
  2. 2Fix the field names
  3. 3“No, the other approval”
  4. 4Paste the error
  5. 5Re-explain the rule
  6. 6Fix the regression
Context re-read each turn
Context coding handoff
  1. 1Approved artefact REQ-114, read once
  2. 2“Implement REQ-114”
Context re-read each turn
Illustrative, not measured. In a chat, each turn is added to the context and re-read on the next one. A structured artefact is read once.

Consistency

A structured artefact removes the guesswork that produces different code on different runs. Peer-referenced research reports error reductions of up to 50 percent when agents work from human-refined specs, and a product-context benchmark found that agents given structured, approved context followed team-specific decisions 95 percent of the time, up from 46 percent for a codebase-only baseline.

Traceability

This is the one regulated teams care about most. Requirements traceability, the ability to follow a requirement forward into code and tests and backward to its origin, has been a formal expectation in safety-critical and regulated work for decades. If the same approved artefact drives both the human review and the machine build, the trace is intact by design rather than reconstructed after the fact.

What the research says
Decision compliance of coding agents
Codebase only46%
With structured, approved context95%
Product-context benchmark, arXiv:2605.08112
45%of the time, models picked the insecure way to write code when a secure one was available.Veracode 2025 GenAI Code Security Report, 100+ models, 80 tasks
1 in 5organizations building on vibe-coding platforms had high-impact misconfigurations.Wiz Research
70 to 85%of rework cost traces back to requirements errors.Leffingwell 1997, cited by Karl Wiegers. An older, directional estimate.
Third-party figures that support the general principle. None of them is a ProductSphere measurement, and results from specific experimental setups do not transfer automatically to any given deployment.

What context coding takes in a regulated setting

Speed is not the hard part anymore. Control is. For BFSI and data-regulated teams, context coding only works if a few things are true.

  • An approved source of truth. Requirements go through review, approval and baselining before any agent touches them.
  • Trace IDs. One trace anchor links the human-readable artefact and the machine-readable version, so the reviewed intent and the built code are provably the same thing.
  • On-premise and zero data retention. Under RBI, SEBI, IRDAI and the DPDP Act, where and how data lives is not a preference. Sensitive requirements and code should not leave your environment, and nothing should be retained where it should not be.
  • Model choice. The ability to bring your own LLM keeps you from being locked to one vendor's risk posture.

How tiramai does context coding

This is the gap ProductSphere is built to close. ProductSphere is product intelligence on the tiramai AI platform. It generates requirement artefacts, BRDs, use cases and user stories, in a standard human-readable template that your reviewers can actually read and approve. Internally, it converts that same artefact into a machine-readable format that coding agents can use directly. One trace ID links both versions. The artefact moves through review, approve and baseline before anything is built.

One trace ID, draft to test
  1. Draft
  2. Review
  3. Approve
  4. Baseline
  5. Agent builds
  6. Tests trace back
What reviewers approve

REQ-114. Hold any transfer above ₹10,00,000 until a second approver signs off. Owner: Payments. Status: Approved, baselined v3

same trace ID
What the agent reads{
"trace": "REQ-114",
"baseline": "v3",
"rule": "hold_transfer",
"limit_inr": 1000000,
"approvers": 2
}
The trace ID links the version people approve to the version the agent builds from, so the review and the build are provably about the same requirement.

Because tiramai is security-first, model-agnostic and fully on-premise, with zero data retention and bring-your-own-LLM, context coding can happen inside your walls, under your controls, mapped to your regulator's expectations. See how the platform handles this on the Security & Governance page.

The approved requirement, not a retyped prompt, is what the machine builds from. The vibe stays in the prototype. The context runs the production line.

Sources

  • Andrej Karpathy, original vibe coding post on X, 2 February 2025
  • Collins Dictionary, Word of the Year 2025 announcement, 6 November 2025
  • Veracode, 2025 GenAI Code Security Report, announced 30 July 2025
  • Wiz Research, "Wiz Research Finds Risks in 20% of Vibe-Coded Apps"
  • Tobi Lütke on X, 18 June 2025. Anthropic, "Effective context engineering for AI agents", 29 September 2025
  • GitHub Spec Kit; Amazon Kiro; Martin Fowler, "Understanding Spec-Driven Development"
  • arXiv:2605.08112, "Context-Augmented Code Generation" (decision compliance 46 to 95 percent)
  • arXiv:2602.00180, "Spec-Driven Development: From Code to Contract" (up to 50 percent error reduction)
  • Karl Wiegers, Leffingwell 1997 and Charette 2005 on requirements rework; Donald Firesmith, Carnegie Mellon SEI, on the cost of requirements errors
  • Grant Thornton India, RingSafe and DPO India on RBI, SEBI, IRDAI and the DPDP Act
  • #context coding
  • #vibe coding
  • #spec-driven development
  • #traceability
  • #BFSI
Share
Written bytiramai team
Follow tiramai on LinkedIn
Talk to us

Working through this for your organisation? Our team can walk you through it on your own use case.

See the platform

How tiramai runs governed AI in the cloud, your private cloud or fully on-premises.

Explore ProductSphere
FAQ

Questions About This Topic

Context coding is a way of working where AI coding agents build directly from structured, approved requirements passed machine to machine, instead of from prompts a person retypes from memory. A human still reviews and approves the requirement; the agent no longer needs anyone to re-describe it.