Your Knowledge Base Needs a Gate: Context Poisoning and What Your AI Is Allowed to Believe
Context poisoning rarely needs an attacker. One unapproved document is enough, because retrieval ranks by similarity, not by what your business signed off.
It is a Tuesday afternoon. A delivery lead uploads a vendor’s integration spec to the shared project folder so the team can "have it handy". The spec says the payment retry limit is five attempts. The approved business requirements document, signed off three weeks ago by risk and operations, says three. Nobody notices. Nobody has done anything wrong.
Two sprints later, a coding agent asked to build the retry logic retrieves both documents. The vendor spec is newer, longer and more detailed, so it ranks higher. The agent writes code for five retries. The reviewer sees clean code that matches "the spec". The approved requirement has been quietly overruled by a document nobody approved.
That is context poisoning. No attacker, no breach, just the wrong words reaching the model at the wrong moment.
Payment retry limit: 3 attempts
Payment retry limit: 5 attempts
MAX_PAYMENT_RETRIES = 5What context poisoning actually means
The term was popularised in June 2025 by Drew Breunig in his essay "How Long Contexts Fail". He defined context poisoning as "when a hallucination or other error makes it into the context, where it is repeatedly referenced", alongside three related failures: context distraction, context confusion and context clash, where parts of the context disagree.
He took the name from Google DeepMind’s Gemini 2.5 technical report. A Gemini agent playing Pokémon would sometimes hallucinate about the game state, and DeepMind wrote that parts of its context, such as its goals and summary, could be "poisoned" with misinformation that "can often take a very long time to undo", leaving the model "fixated on achieving impossible or irrelevant goals".
Since then the meaning has widened. Elastic describes it as occurring "when compromised, outdated, or irrelevant information enters an LLM’s context window". Redis puts it more bluntly: once bad information gets in, "the agent treats it as ground truth". The OWASP GenAI Security Project now lists Memory and Context Poisoning as ASI06 in its Top 10 for Agentic Applications, released in December 2025.
Why it compounds in agentic systems
A chatbot that gives one wrong answer is a nuisance. An agent that writes a wrong fact into its notes, memory or plan will reuse it at every following step. The error does not fade; it gets cited. In a software delivery chain, a poisoned requirement becomes a poisoned user story, then poisoned code, then a test that passes because it checks the wrong behaviour.
Three ways poison gets in
- Stale versions. The superseded policy, the draft BRD that was never retired, last year’s rate card. Retrieval systems rank by similarity, not by validity, so an outdated document that closely matches the question can outrank the current one.
- Contradictions between sources. Xu and colleagues, in their EMNLP 2024 survey of knowledge conflicts, group them into three kinds: retrieved text disagreeing with what the model learned in training, retrieved documents disagreeing with each other, and the model’s own knowledge being inconsistent. Models are poor referees. In IBM Research’s WikiContradict benchmark (NeurIPS 2024), built on over 3,500 human judgements across five models, "all models struggle to generate answers that accurately reflect the conflicting nature of the context, especially for implicit conflicts requiring reasoning".
- Unverified or adversarial content. This is where quality becomes security: prompt injection hidden inside retrieved documents, poisoned entries written into agent memory, and instructions buried in tool descriptions. Invariant Labs named "tool poisoning" in April 2025 after showing that hidden instructions in Model Context Protocol tool metadata could make agents leak data.
Why repetition looks like consensus
Here is the subtle part. Many retrieval and knowledge graph systems treat frequency as a signal. Microsoft’s GraphRAG paper describes relationships extracted from many chunks being aggregated into graph edges, "where the number of duplicates for a given relationship becomes edge weights". Frequently mentioned entities become more prominent in the graph. That is useful when repetition reflects genuine agreement. It is dangerous when the same claim appears ten times because one wrong document was copied into ten folders.
Language models show a similar pull. Researchers at Johns Hopkins describe a majority bias, trusting the answer supported by more frequent evidence across documents, and found that with repeated context the remaining mistakes disproportionately favour answers that appear more often. Deduplication helps, but it is not enough: in the PoisonedRAG study, duplicate-text filtering failed because each malicious text was worded differently. Ten paraphrases of one bad claim still look like ten sources.
What good looks like
The fix is not a smarter model. It is a gate in front of the knowledge base.
- Conflict check at ingest. Before a document becomes retrievable, compare its claims against what is already approved and flag contradictions for a human. OWASP’s guidance on vector and embedding weaknesses names "data federation knowledge conflict errors" as a risk and advises reviewing combined datasets thoroughly.
- Versioning with valid-from and valid-until. Treat time as data. Bi-temporal models track two clocks: when a fact was true in the world, and when the system recorded it. The Zep and Graphiti temporal knowledge graph stores validity intervals on every relationship and, when a new fact contradicts an old one, invalidates the old one rather than deleting it. Supersede, never delete. That is what lets you answer an auditor’s question: what did the system believe on the day it made that decision?
- Sign-off by a named expert. A person approves before ingest. This is slower than drag and drop, and that is the point.
- Provenance and trace IDs. Every chunk should carry where it came from, who approved it and which version it belongs to. NIST’s Generative AI Profile asks organisations to "employ methods to trace the origin and modifications of digital content" and notes that robust version control can track changes across the AI lifecycle.
- Control of the pipeline. If you cannot see or control what feeds your model, you cannot certify what it read.
Why this is a compliance problem, not just a quality problem
In regulated industries, "the AI got it wrong" is not an answer an auditor accepts. The Reserve Bank of India’s FREE-AI Committee report of August 2025 states that accountability cannot be delegated to the model and the underlying algorithm, and recommends an AI audit framework covering data inputs, model and algorithm, and decision outputs. SEBI’s June 2025 consultation on AI and ML in securities markets proposes governance, testing and auditability expectations. IRDAI’s cyber security guidelines require ICT logs to be kept for a rolling 180 days within India. The DPDP Act allows penalties of up to ₹250 crore for failing to maintain reasonable security safeguards.
The RBI report is advisory and the SEBI paper was a consultation, so read them as direction of travel rather than binding rules. The direction is unmistakable: when the context feeding an AI is uncontrolled, every one of these obligations becomes harder to evidence.
How ProductSphere applies this
*This section describes our own approach. It is not third-party research.*
ProductSphere is product intelligence on the tiramai AI platform. It treats requirements as governed context rather than loose files.
- One artefact, two faces. ProductSphere generates requirement artefacts such as BRDs, use cases and user stories in a standard human-readable template for reviewers, then converts the same artefact into a machine-readable format for coding agents. One trace ID links both, so the version a human approved is the version an agent reads.
- A lifecycle with gates. Artefacts move through review, approve and baseline states, and baselined context is what goes to coding agents. We call that context coding: a machine-to-machine handoff of approved, structured context, rather than an agent working from whatever it happens to find.
- Traceability by default. The trace anchor connects an artefact back to the story, the use case and the approved BRD it came from, which is the chain an auditor asks for.
- Your infrastructure, your model. The platform deploys fully on-premise, with zero data retention and bring-your-own-LLM, so the gate sits inside your perimeter.
In the retry example, the vendor spec was never reviewed or baselined against the approved BRD, so it was never authoritative context. An agent working from baselined artefacts would have built the requirement that risk and operations actually signed.
A five-minute check of your own setup
Seven questions. Answer them honestly about the knowledge base your AI retrieves from today.
- 1Can anyone add a document to the knowledge base your AI retrieves from, without approval?
- 2When two documents disagree, does anything flag it before a person or an agent sees both?
- 3Does every document carry a valid-from and valid-until date, with superseded versions kept rather than deleted?
- 4Can you trace an AI output back to the exact document versions it used?
- 5Does your retrieval or graph layer weight claims by how often they appear? If so, what stops copies passing as consensus?
- 6Are tool descriptions and MCP servers reviewed like code?
- 7Could you show a regulator what your system believed on a specific date?
Closing
Context poisoning is rarely dramatic. It is usually an honest upload, a forgotten draft, a copy of a copy. The organisations that get agentic AI right will not be the ones with the biggest context windows.
They will be the ones that decide, deliberately, what their AI is allowed to believe.
If you would like to see how ProductSphere gates requirements before they reach your coding agents, talk to our team.
Sources
- Drew Breunig, How Long Contexts Fail, June 2025
- Elastic, Context poisoning in LLMs; Redis, Context Poisoning: How Bad Data Breaks Agent Reasoning
- OWASP GenAI Security Project, Top 10 for Agentic Applications (ASI06) and LLM08: Vector and Embedding Weaknesses
- Zou and colleagues, PoisonedRAG, USENIX Security 2025
- Chen and colleagues, AgentPoison, NeurIPS 2024
- Dong and colleagues, MINJA, and the 2026 follow-up under realistic conditions
- Anthropic, UK AI Security Institute and the Alan Turing Institute, A small number of samples can poison LLMs, October 2025
- Hou and colleagues, WikiContradict, NeurIPS 2024; Xu and colleagues, Knowledge Conflicts for LLMs, EMNLP 2024
- Edge and colleagues, GraphRAG, Microsoft Research; Rasmussen and colleagues, Zep and Graphiti
- Invariant Labs, MCP tool poisoning, April 2025
- NIST, AI 600-1 Generative AI Profile; RBI FREE-AI Committee report, August 2025; SEBI consultation paper, June 2025; IRDAI Information and Cyber Security Guidelines 2023; DPDP Act 2023 and Rules 2025
Working through this for your organisation? Our team can walk you through it on your own use case.
How tiramai runs governed AI in the cloud, your private cloud or fully on-premises.
Explore ProductSphere


