Answers

Our AI retrieves the right document and still answers with the outdated line from the top of it. Why does retrieval look fine when the answer is wrong?

Because retrieval and assembly are two different steps, and a green number on the first says nothing about the second. What reaches the model here is not a document for some later layer to excerpt: it is a statement, delivered as one authored line that carries the operative point, with the full body coming back only when someone asks for that memory by name. There is no top of a file to truncate into and no paragraph to pick wrongly. A statement that was replaced is not a lower-ranked neighbour of its replacement — it is retired in the same act that accepts the replacement, with the reason and the author kept on its archived line. And a statement someone saw contradicting its source arrives marked as contradicted rather than arriving clean. Which version is current is settled once, by a person, in the record; it is not re-derived on every call by whatever gets the last word on the prompt.

Last updated September 22, 2026

Recall answers the question one step before the one that hurts

Retrieval metrics score the search: was the right document findable, and how high did it rank. The thing that produces the wrong answer sits further along — the text that actually went into the model after something decided what would fit. Between those two points there is a layer that truncates, summarises, or keeps the first N characters of what it was handed, and it makes that choice with a notion of similarity and no notion of currency. Documents grow by appending, so the correction is usually at the bottom and the original claim is at the top; a layer that keeps the head keeps the oldest version of every file it touches. Every dashboard along that path stays green, because each layer did the job it is measured on. The failure exists only in the gap between them, which is exactly where nobody is looking.

A statement is served, not a document to excerpt

Each memory here is one statement, and it travels as one authored line — never the opening sentence of a longer body, which is the shortcut that brings the same problem back under a new name. The line states the operative point; the full body is recoverable, but only when someone asks for that memory by name. Ranking therefore chooses which statements come back, not which paragraph of a large one gets in. That is a narrower job, and it is one that fails visibly: the worst case is a relevant memory left out, which the read says out loud, instead of a relevant memory included by its stale half, which nothing says at all.

Currency is decided once, in the record, not on every call

When something is replaced, the replacement names what it contradicts, and accepting it retires the earlier statement in the same act, with the reason kept on the archived line. The old version is not a sibling that could out-rank its own successor on a differently worded query — it is out of circulation and comes back only as history. Where a divergence has been noticed but not yet resolved, the record carries that too: whoever opened the source and saw it disagree writes down what the source actually says, and from then on the memory arrives marked instead of arriving as though nothing had happened. All of it is decided by people, once, and read by every session afterwards. The alternative — deciding which of two versions is current inside prompt assembly, on every call, from similarity alone — is the arrangement that produces confident answers from March.

What it looks like in practice

A team keeps its release process in a context file. The section was written in March. In August somebody appends a new heading at the bottom: releases are no longer tagged by hand. An assistant is asked how releases go out. Retrieval works perfectly — the right file, at position one, with a confident score. The assembly step then has a budget, the file is longer than the budget, and it passes along the first part. The model answers with March, in detail, in the tone of something that was looked up. The person asking has no way to see that the answer came from a paragraph that a later paragraph of the same file contradicts. Nothing in that pipeline reports a failure. Retrieval recall is the number the team watches, and it is correct. Whoever wrote the August paragraph did the right thing in the right place. The layer that dropped it was doing what it exists to do. With statements the shape is different, because there is no file to cut. The March claim and the August one are two statements; the second names the first as the one it replaces; accepting it retired the first. The read returns the August statement as a line, with nothing above it to truncate into and no older sibling to lose to.

Questions people ask about this

Isn't this a chunking problem? Better splitting would fix it.
Better splitting improves the odds on any given call and leaves the shape untouched. The assembler still has to choose, every time, and it still chooses on similarity — which cannot tell a current paragraph from the one it replaced, since both are about the same subject, which is precisely why both score well. You can make the wrong pick rarer and you cannot make it visible, because a truncated document does not report what it dropped. What removes the class of failure is not choosing better at read time; it is not having a document to choose from.
How do we know a statement we get back is the current one and not an old one that ranked well?
Because being replaced is an act with a record, not a matter of ranking. A replacement has to name what it contradicts, and on acceptance the earlier statement is archived with the reason and the author on it, so it is no longer something a query can return. Two statements that are both in force and genuinely disagree are a different case, and that one is handled as a conflict to surface rather than a tie to break quietly. What you never get back is a previous version presented as though it were the present one.
We measure retrieval recall and it is high. What should we be measuring instead?
Measure what reached the model, against what was in force at that moment. Recall answers whether the right thing was findable. The question that predicts wrong answers is whether the right thing — and nothing older than it — was in the text the model actually read. The two come apart as soon as any layer between the store and the model is allowed to shorten what it was handed. If you keep only one number, take it from the served payload rather than from the search result.
Install ArrowaySee how it works →

Where this is verifiable

Product documentation on this site (How it works, Install), the answers on why RAG does not replace sanctioned memory, on out-of-date memory being read as current and on handoffs that bloat every agent reading them next, and the authored-essence, supersede, verification and conflict rules in the sanctioned product spec. Everything described here is behaviour the tools apply today, not roadmap.

https://www.arroway.app/en/answers/retrieval-succeeds-but-assembly-uses-the-stale-part

Continue exploring

See all answers