Your Knowledge Base Needs a Gate: Context Poisoning and What Your AI Is Allowed to Believe

Context poisoning rarely needs an attacker. One unapproved document is enough, because retrieval ranks by similarity, not by what your business signed off.

It is a Tuesday afternoon. A delivery lead uploads a vendor’s integration spec to the shared project folder so the team can "have it handy". The spec says the payment retry limit is five attempts. The approved business requirements document, signed off three weeks ago by risk and operations, says three. Nobody notices. Nobody has done anything wrong.

Two sprints later, a coding agent asked to build the retry logic retrieves both documents. The vendor spec is newer, longer and more detailed, so it ranks higher. The agent writes code for five retries. The reviewer sees clean code that matches "the spec". The approved requirement has been quietly overruled by a document nobody approved.

That is context poisoning. No attacker, no breach, just the wrong words reaching the model at the wrong moment.

Two documents, one wrong answer
Approved BRD, baselined v3

Payment retry limit: 3 attempts

Signed off by risk and operations, three weeks ago
Vendor integration spec, never approved

Payment retry limit: 5 attempts

Uploaded to the shared folder on Tuesday
The agent asks: what is the payment retry limit?
1Vendor specnewer, longer, more detailed
2Approved BRDapproved, but ranked second
Ranked by similarity to the question. Approval is not part of the score.
Shipped codeMAX_PAYMENT_RETRIES = 5The reviewer sees clean code that matches "the spec".
An approved requirement was overruled by a document nobody approved, and nothing in the chain objected.
With a gate, the vendor spec is never authoritative context: it was never reviewed or baselined.
Retrieval ranks by similarity to the question, not by whether anyone approved the document. Newer, longer and more detailed usually wins.

What context poisoning actually means

The term was popularised in June 2025 by Drew Breunig in his essay "How Long Contexts Fail". He defined context poisoning as "when a hallucination or other error makes it into the context, where it is repeatedly referenced", alongside three related failures: context distraction, context confusion and context clash, where parts of the context disagree.

He took the name from Google DeepMind’s Gemini 2.5 technical report. A Gemini agent playing Pokémon would sometimes hallucinate about the game state, and DeepMind wrote that parts of its context, such as its goals and summary, could be "poisoned" with misinformation that "can often take a very long time to undo", leaving the model "fixated on achieving impossible or irrelevant goals".

Since then the meaning has widened. Elastic describes it as occurring "when compromised, outdated, or irrelevant information enters an LLM’s context window". Redis puts it more bluntly: once bad information gets in, "the agent treats it as ground truth". The OWASP GenAI Security Project now lists Memory and Context Poisoning as ASI06 in its Top 10 for Agentic Applications, released in December 2025.

Why it compounds in agentic systems

A chatbot that gives one wrong answer is a nuisance. An agent that writes a wrong fact into its notes, memory or plan will reuse it at every following step. The error does not fade; it gets cited. In a software delivery chain, a poisoned requirement becomes a poisoned user story, then poisoned code, then a test that passes because it checks the wrong behaviour.

Three ways poison gets in

  • Stale versions. The superseded policy, the draft BRD that was never retired, last year’s rate card. Retrieval systems rank by similarity, not by validity, so an outdated document that closely matches the question can outrank the current one.
  • Contradictions between sources. Xu and colleagues, in their EMNLP 2024 survey of knowledge conflicts, group them into three kinds: retrieved text disagreeing with what the model learned in training, retrieved documents disagreeing with each other, and the model’s own knowledge being inconsistent. Models are poor referees. In IBM Research’s WikiContradict benchmark (NeurIPS 2024), built on over 3,500 human judgements across five models, "all models struggle to generate answers that accurately reflect the conflicting nature of the context, especially for implicit conflicts requiring reasoning".
  • Unverified or adversarial content. This is where quality becomes security: prompt injection hidden inside retrieved documents, poisoned entries written into agent memory, and instructions buried in tool descriptions. Invariant Labs named "tool poisoning" in April 2025 after showing that hidden instructions in Model Context Protocol tool metadata could make agents leak data.
How little poison it takes
90%attack success from five malicious texts per question, in a knowledge base of millions.PoisonedRAG, USENIX Security 2025
80%+average attack success against agent memory, at a poison rate below 0.1%.AgentPoison, NeurIPS 2024
250 docswere enough to backdoor models from 600M to 13B parameters. Corpus size was no protection.Anthropic, UK AI Security Institute and the Alan Turing Institute, 2025A narrow pretraining backdoor, not a retrieval system.
3500+human judgements found that models struggle to flag two passages that contradict each other.WikiContradict, IBM Research, NeurIPS 2024
These come from controlled research settings, in some cases from the authors of the attacks. Read them as evidence that the bar is low, not as expected breach rates.

Why repetition looks like consensus

Here is the subtle part. Many retrieval and knowledge graph systems treat frequency as a signal. Microsoft’s GraphRAG paper describes relationships extracted from many chunks being aggregated into graph edges, "where the number of duplicates for a given relationship becomes edge weights". Frequently mentioned entities become more prominent in the graph. That is useful when repetition reflects genuine agreement. It is dangerous when the same claim appears ten times because one wrong document was copied into ten folders.

Language models show a similar pull. Researchers at Johns Hopkins describe a majority bias, trusting the answer supported by more frequent evidence across documents, and found that with repeated context the remaining mistakes disproportionately favour answers that appear more often. Deduplication helps, but it is not enough: in the PoisonedRAG study, duplicate-text filtering failed because each malicious text was worded differently. Ten paraphrases of one bad claim still look like ten sources.

Repetition is not consensus
One wrong claimin a document nobody approved
copied and reworded across folders
“retry limit is five”“up to five attempts”“five retries permitted”“max 5 payment retries”“allow five tries”“retries: five”“five attempts max”“permits 5 retries”“retry cap of five”“five, per the spec”
Duplicate filterten different texts, so nothing is removed
retry limitedge weight 10five
Duplicates become edge weight, as in GraphRAG
The model sees a majority. Ten mentions of one mistake outweigh the single approved document that says three.
One claim, ten wordings. Duplicate filters see ten different texts; the graph sees an edge ten times heavier; the model sees a majority.

What good looks like

The fix is not a smarter model. It is a gate in front of the knowledge base.

A gate in front of the knowledge base
Anything anyone uploads
The gate
Conflict checkCompare claims against what is already approved, and flag contradictions for a person
Valid from, valid untilSupersede, never delete, so you can say what the system believed on a given day
Sign-offA named expert approves before ingest
Provenance and trace IDEvery chunk carries its source, its approver and its version
Pipeline you controlIf you cannot see what feeds the model, you cannot certify what it read
Approved knowledge baseWhat agents are allowed to treat as authoritative
Five controls, all of them before retrieval. Anything that has not passed them is not context an agent should treat as authoritative.
  • Conflict check at ingest. Before a document becomes retrievable, compare its claims against what is already approved and flag contradictions for a human. OWASP’s guidance on vector and embedding weaknesses names "data federation knowledge conflict errors" as a risk and advises reviewing combined datasets thoroughly.
  • Versioning with valid-from and valid-until. Treat time as data. Bi-temporal models track two clocks: when a fact was true in the world, and when the system recorded it. The Zep and Graphiti temporal knowledge graph stores validity intervals on every relationship and, when a new fact contradicts an old one, invalidates the old one rather than deleting it. Supersede, never delete. That is what lets you answer an auditor’s question: what did the system believe on the day it made that decision?
  • Sign-off by a named expert. A person approves before ingest. This is slower than drag and drop, and that is the point.
  • Provenance and trace IDs. Every chunk should carry where it came from, who approved it and which version it belongs to. NIST’s Generative AI Profile asks organisations to "employ methods to trace the origin and modifications of digital content" and notes that robust version control can track changes across the AI lifecycle.
  • Control of the pipeline. If you cannot see or control what feeds your model, you cannot certify what it read.

Why this is a compliance problem, not just a quality problem

In regulated industries, "the AI got it wrong" is not an answer an auditor accepts. The Reserve Bank of India’s FREE-AI Committee report of August 2025 states that accountability cannot be delegated to the model and the underlying algorithm, and recommends an AI audit framework covering data inputs, model and algorithm, and decision outputs. SEBI’s June 2025 consultation on AI and ML in securities markets proposes governance, testing and auditability expectations. IRDAI’s cyber security guidelines require ICT logs to be kept for a rolling 180 days within India. The DPDP Act allows penalties of up to ₹250 crore for failing to maintain reasonable security safeguards.

The RBI report is advisory and the SEBI paper was a consultation, so read them as direction of travel rather than binding rules. The direction is unmistakable: when the context feeding an AI is uncontrolled, every one of these obligations becomes harder to evidence.

How ProductSphere applies this

*This section describes our own approach. It is not third-party research.*

ProductSphere is product intelligence on the tiramai AI platform. It treats requirements as governed context rather than loose files.

  • One artefact, two faces. ProductSphere generates requirement artefacts such as BRDs, use cases and user stories in a standard human-readable template for reviewers, then converts the same artefact into a machine-readable format for coding agents. One trace ID links both, so the version a human approved is the version an agent reads.
  • A lifecycle with gates. Artefacts move through review, approve and baseline states, and baselined context is what goes to coding agents. We call that context coding: a machine-to-machine handoff of approved, structured context, rather than an agent working from whatever it happens to find.
  • Traceability by default. The trace anchor connects an artefact back to the story, the use case and the approved BRD it came from, which is the chain an auditor asks for.
  • Your infrastructure, your model. The platform deploys fully on-premise, with zero data retention and bring-your-own-LLM, so the gate sits inside your perimeter.

In the retry example, the vendor spec was never reviewed or baselined against the approved BRD, so it was never authoritative context. An agent working from baselined artefacts would have built the requirement that risk and operations actually signed.

A five-minute check of your own setup

Seven questions. Answer them honestly about the knowledge base your AI retrieves from today.

Does your knowledge base have a gate?
  1. 1Can anyone add a document to the knowledge base your AI retrieves from, without approval?
  2. 2When two documents disagree, does anything flag it before a person or an agent sees both?
  3. 3Does every document carry a valid-from and valid-until date, with superseded versions kept rather than deleted?
  4. 4Can you trace an AI output back to the exact document versions it used?
  5. 5Does your retrieval or graph layer weight claims by how often they appear? If so, what stops copies passing as consensus?
  6. 6Are tool descriptions and MCP servers reviewed like code?
  7. 7Could you show a regulator what your system believed on a specific date?
0 of 7 answered

Closing

Context poisoning is rarely dramatic. It is usually an honest upload, a forgotten draft, a copy of a copy. The organisations that get agentic AI right will not be the ones with the biggest context windows.

They will be the ones that decide, deliberately, what their AI is allowed to believe.

If you would like to see how ProductSphere gates requirements before they reach your coding agents, talk to our team.

Sources

  • #context poisoning
  • #RAG
  • #knowledge base
  • #governance
  • #agentic AI
Share
Written bytiramai team
Follow tiramai on LinkedIn
Talk to us

Working through this for your organisation? Our team can walk you through it on your own use case.

See the platform

How tiramai runs governed AI in the cloud, your private cloud or fully on-premises.

Explore ProductSphere
FAQ

Questions About This Topic

Context poisoning is when information that should not be trusted gets treated as trusted by an AI system, and every later step builds on it. It covers stale documents, contradictions between sources, and content planted deliberately.