Grad shape
Grad shape
Blog

When Trusted AI Systems Are Quietly Poisoned

Last updated: 9 July 2026

When Trusted AI Systems Are Quietly Poisoned
When Trusted AI Systems Are Quietly Poisoned

Even if an AI system explains its reasoning and cites sources, it may nonetheless reach the incorrect conclusion for the wrong reason.

Consider an internal decision making model. It is based on reports, technical notes, company policies, and other authorised papers. People eventually come to accept it because it makes sense, it cites sources, shows a clear line of reasoning, and seems to comprehend not only what it is saying but also why.

It then offers a suggestion that it shouldn't have. Not a blatant error. Something believable, attractive, and just a little off, just enough to pass without being questioned. When prompted to explain itself, it provides a concise, rational explanation and refers to a document in the knowledge base. People frequently express a desire for AI to exhibit precisely that kind of behaviour. A system that can demonstrate how it operates, not a black box.

The source itself was incorrect, which is the issue. A document with subtle flaws or poisoning had been added to the knowledge base and was regarded as authoritative. The system took it in without making any distinctions, applied it to its logic, and then presented the outcome as if the problem were simple. The failure was not made clear by the explanation. It eased the acceptance of the failure.

 

Why explanation feels like trust

There is a good reason why explainability became important. For many years, AI systems generated results with minimal insight into the process that led to them. A system is easier to evaluate and, thus, more trustworthy if it can provide references and explain its logic. That can be helpful at times. A clear explanation provides context and enables the user to evaluate the response against their own assessment.

Trust and explanation are two different things. A persuasive explanation shows what the system says it did. It doesn't always demonstrate what really affected the choice, whether those impacts were suitable, or whether the underlying inputs were worthy of the weight they received. Explanation and cause are two different things. It is possible for a system to honestly state, "This recommendation came from this document," while concealing the more crucial question of why that document was initially permitted to have such a big influence on the response.

 

The real issue

Certain types of errors are reduced when AI is grounded on approved knowledge bases, internal documents, and carefully selected content. It can reduce the likelihood of fabrication and increase the auditability of the output.

This shifts the trust issue rather than solving it though. The system still needs to make decisions about how to handle the content it receives, what authority to grant it, and how to distinguish between instruction and information, reliable guidance and harmful influence, or trusted content and content that only seems trusted. This gets more difficult as knowledge bases expand. Documents overlap which can result in conflicts in guidance, emphasis varies, information is out-of-date, unclear, and subtly incorrect.

A human reviewer may enquire as to whether a document is indeed appropriate for the purpose, which sources should be prioritised, and what should be disregarded. These questions are usually not answered in the same way by a model. It combines what it discovers to create a reaction. The outcome is frequently not a clear mistake. It's something more perilous: a result that makes sense, is well-supported, and is explained, but was influenced by signals that shouldn't have had such a significant impact. That is what makes it hard to challenge.

 

What this means for trust in AI

It is possible for a system to be simpler to comprehend without being more reliable. Explaining incorrect behaviour can sometimes foster misguided trust by making it appear rational and ordered. The conclusion was still influenced by inputs that the system was ill-prepared to evaluate, even though the reasoning is obvious, the sources are visible, and the structure is sound.

An obvious inaccuracy may not be as convincing as one that is clearly explained. Therefore, explanation is insufficient to build true trust. It must originate with how judgements are made in the first place: how the system manages information that has been retrieved, how it distinguishes between influences that are trusted and those that are not, and how cautiously it restricts what might influence its behaviour.

Explanations are still important. For inspection, they are helpful. However, they cannot take the place of methodical system design. The system in the aforementioned example performed just what it was designed to do: retrieve data, apply it, and clearly explain the result. The issue was that something had already gotten into the chain that it shouldn't have trusted.

What it did might be explained by it. The gap that trustworthy AI systems need to fill is that it was unable to explain why it shouldn't have done it.

 


 

This article is part of our Building Trustworthy AI Systems series exploring how reliability, control, and system design are important for engineering AI systems.

These ideas inform the development of Cyclone Sage, our AI assistant for MCNP being built with a focus on structured, reliable system design.

If you’re working on similar challenges or exploring how AI can be applied reliably in technical environments, feel free to get in touch at support@orthrussoftware.com

Related blogs