Agent Memory Fails in Three Places and Gets Measured as One

A memory can fail because the note was not found, because nobody asked for it, or because it was delivered and ignored. An outcome score reports all three as the same number. Telling them apart takes one instrument per stage, and the instruments fail too, sometimes the same way on every run.

When an agent with memory makes a mistake its memory could have prevented, one of three things went wrong. The note was not found. The note was never asked for. Or the note was delivered, sat in the agent's context in plain text, and the agent did the other thing anyway.

The three have different causes and different fixes. Not found is a retrieval problem, fixed by better search. Never asked is about adoption: the agent could have looked and did not, so the fix takes that choice away from it and has the harness deliver the note. The third is a trust problem, and neither better search nor better delivery touches it.

Most evaluations of agent memory report one outcome number, and one outcome number moves the same way for all three failures. It cannot say which one happened, so it cannot say which fix to build. Separating them means measuring every stage between "the note exists" and "the agent did the right thing", each with its own instrument, and the instruments fail as well. In the evaluation harness of vectr, an open-source code-search and agent-memory tool I maintain, a small deterministic scorer reported that an agent had been handed a note, ignored it, and run a script the note warned against. The agent had read the script with cat. It never ran it.

Primary source

The evidence for the deploy.sh case, the broken metrics and the ceiling experiment is documented in full in a technical report: Deterministic, Reproducible, and Wrong: How Our Own Scorer Manufactured an Agent-Memory Finding, and the Guard That Would Have Caught It (2026, not peer reviewed). Its introduction opens with the same claim: "Memory fails in three places and gets measured as one." Every count cited comes from that report, from the public vectr repository it points to, or from the papers listed under Sources.

Part 01
The Claim
01

Three Ways Agent Memory Fails: Not Found, Never Asked, Delivered and Ignored

A few terms first. The store is the database of saved notes. Retrieval is the search that picks notes out of it for a given moment. A note is delivered when the harness places it in the agent's context, whether or not the agent asked. With those, the three failures are:

  • Not found. Retrieval ran and returned nothing, or the wrong note. The right note may be sitting in the store the whole time.
  • Never asked. The right note was in the store and retrieval would have found it, but looking was the agent's choice and the agent did not make it.
  • Delivered and ignored. The note reached the agent's context, intact, before the decision it was about, and the agent acted against it.
FailureWhat brokeWhat fixes itEvidence that isolates it
Not foundRetrievalBetter search: indexing, ranking, query formationA retrieval log of what came back for which query
Never askedAdoptionDelivery by the harness at the moment the situation calls for the noteMemory calls made, against the moments one was needed
Delivered and ignoredTrustEvidence the agent can check, or enforcement that needs no beliefProof the note was in context before the decision, plus what the agent did next
Each fix leaves the other two failures where they were. A better retriever does nothing for a store nobody queries.
Analogy: a parcel that changes nobody's mind

A note is a parcel. It can be lost in the warehouse because the picker pulls the wrong shelf (not found). It can sit on its shelf because nobody placed an order (never asked). Or it can arrive, be opened and read, and be set aside because the recipient does not trust the sender (delivered and ignored). A shop that only tracks whether customers ended up with what they needed sees one number for all three, and the three fixes belong to different departments: the warehouse, the ordering system and the sender's reputation.

The analogy breaks in one useful place. A courier can photograph the parcel on the doorstep. In an agent harness, "it arrived" has no photograph by default. Whether a note actually reached the model is a claim that needs its own evidence, and the instruments that supply it can be wrong in either direction.

Never asked exists only where asking is optional

The middle failure needs a condition. A question-answering pipeline that runs retrieval for every question cannot fail to ask. Yuan, Su and Yao's diagnostic study of agent memory is built that way: retrieval runs on every question in the LoCoMo benchmark, and failures are split two ways, retrieval failure against utilization failure, plus a small hallucination class, with a language-model judge doing the classification. Across their grid of memory and retrieval configurations, retrieval failures were 11 to 46 percent of all questions and utilization failures 4 to 8 percent. In that setup "never asked" has nothing to count.

Coding agents are usually built the other way. Memory is a tool the agent may call, next to reading files and running commands, and an agent with a task in front of it often does not call it. In one measured instance, an agent with memory tools connected, explicit guidance to use them, and a store seeded with directly relevant facts made zero memory calls in 114 turns. The argument that follows from that, deliver from the harness instead of waiting to be asked, is in The Agent Never Chooses to Remember.

Moving delivery into the harness removes the agent's choice, but the never-asked failure does not vanish. It moves into the harness, as a trigger that should have fired and did not. And a harness can believe it delivered a note that never reached the model, in which case an evaluation that trusts the harness's own account will score the run as delivered and ignored.

Part 02
One Number
02

Why an Oracle Ratio Cannot Separate Agent Memory Failures

The natural instrument for "does memory help?" is a ratio of two scores. Run the task with the needed information handed to the agent directly, an oracle condition. Run it again with the agent responsible for getting that information out of its own memory. The ratio says how much of the oracle's benefit memory recovers.

VehicleMemBench, an executable benchmark for multi-user long-term memory in in-vehicle agents, formalizes this as MemoryScore. Each task is scored by exact state match: after the agent acts, does the vehicle's state equal the target state? MemoryScore divides the score under autonomous memory, where the facts have to be stored and retrieved by the memory system under test, by the score under gold memory, where the facts are given in context. The benchmark uses it to argue that memory construction and retrieval, rather than execution, are the dominant sources of error, and it scores by environment state without any model or human grader, which is a real strength.

The ratio's weakness is arithmetic. Take the three failures in turn. The agent never asked: the task is lost. It asked and got the wrong note: lost. It got the right note and ignored it: lost. Each failed task costs the ratio the same amount, whichever stage it failed at.

The ratio as a product

Model each task in the memory condition as needing three stages to succeed: the agent asks (rate p_ask), retrieval finds the right note (p_find), and the agent acts on the note as it would on the oracle's context (p_use). Tasks that clear all three score like the oracle, g. Tasks that fail any stage score like an agent without the fact, b.

q = p_ask × p_find × p_use S_mem = q·g + (1 - q)·b R = S_mem / g = q + (1 - q)·(b / g)

When the task cannot be done without the fact (b = 0), R = p_ask × p_find × p_use. A product does not record which factor was small: (0.6, 1, 1), (1, 0.6, 1) and (1, 1, 0.6) all give R = 0.6. Three systems with three different broken parts and three different fixes, one number. When b = g, R = 1 for every value of q: saturation, where the ratio reports nothing at all.

The model assumes the stages are independent and that a used note performs exactly like the oracle's context. The conclusion needs neither assumption: however the stages depend on each other, a lost task takes its loss out of the numerator with no label saying which stage lost it. One more property follows from dividing by the oracle. Because p_use is defined relative to the oracle, ignoring that the oracle condition suffers as well is already inside g, and the ratio never shows it. It sees only ignoring beyond the oracle's own.

Interactive Demo 1
Three Systems, One Score

Set how often the agent asks, how often retrieval finds the right note, and how often a delivered note is acted on. The bar shows where 100 tasks end up; the ratio under it is the only number an oracle-ratio evaluation reports. Step through the first three presets: three different broken stages, the same ratio. Then try the saturated preset, where most tasks never consult memory and the ratio reads perfect.

Expected counts per 100 tasks under the independence model of the ratio formula, rounded. The score without the fact cannot exceed the oracle score here; moving the oracle below it pulls it down.

What the memory benchmarks already separate

The same arithmetic applies to any evaluation that grades only a final answer. A long-term memory benchmark asks a question after a long history, and a wrong answer could mean the fact was never stored, was stored and not retrieved, or was retrieved and not used. The established benchmarks know this. LongMemEval annotates the evidence sessions behind each question and reports Recall@k and NDCG@k whenever the system under test exposes its retrieval results. LoCoMo annotates each question with the dialogue turns that hold its answer and reports retrieval accuracy for retrieval-augmented models. Both can pull "not found" apart from the rest, provided the system shows its retrieval. Neither has a category for "never asked", and for a pipeline that retrieves on every question there is nothing to put in one.

VehicleMemBench separates the stages too, by hand. Its error analysis sorts failed cases by manual inspection into memory errors, extra calls, execution errors and incomplete execution, the last of which includes not properly using retrieved memory. The headline number is still the ratio. The stages are recovered afterwards, by reading.

An outcome number with a broken stage underneath does worse than hide the cause. It produces a confident reading of the wrong thing. In an earlier vectr benchmark, the recall function matched notes by literal substring (an SQL LIKE) instead of by meaning, so implementation sessions got empty or near-empty results back. The outcome read as the memory tool making implementation more expensive. It was a not-found failure, reported as a cost finding; Building Vectr 3: What the Benchmark Numbers Mean has the numbers before and after the fix.

03

Saturation: A Perfect Agent Memory Score With a Failure Inside

A ratio has a second blind spot. When both conditions reach the maximum the task allows, the ratio of two equal numbers is 1.0. Depending on what you hoped to see, that reads as "memory adds nothing" or "no problem here". Neither reading is supported. Both arms are at the top of the scale, and the outcome has nothing left to say.

A pre-registered experiment in July 2026 hit exactly this. One seeded bug in Apache Camel, a large multi-module Java codebase. Two arms of three sessions each, identical except for one variable: whether one operational note was in the memory store. The note explained how to verify a fix honestly in that build. A test run scoped to one module compiles against previously installed copies of the others, so a green result there cannot see a change made in another module. The primary outcome was whether each session verified its fix with a test that could actually see the change.

It came out 3 of 3 in both arms. The frozen rule for reading the result returned not supported, and that verdict stands. Underneath the tie, the processes differed. The report gives the following as exploratory secondary counts, with no direction claimed:

MeasureMemory arm (3 sessions)Control arm (3 sessions)
Honest verification3 of 33 of 3
False-pass verification events0, 0, 00, 7, 3
Build-tool invocations1, 3, 42, 9, 4
From the experiment's own deterministic transcript parser, whose definitions were frozen with the pre-registration. One control session ran nine builds, seven of them false passes, and used up its 40-turn budget re-deriving the build command by trial and error.

A false pass is a green test that could not have failed. Two details complicate even these counts, and both come from reading the transcripts by hand. The parser judged honest verification from the shape of a command's arguments. In one memory session the qualifying build was moved to the background and the session hit its turn limit before reading the output; in one control session the qualifying run carried a flag that skips the tests entirely. Hand-audited, honest verification is 2 of 3 in each arm. The verdict does not move. And the clean false-pass column in the memory arm cannot be credited to the note: one memory session shows no trace of it, and one control session scoped its build correctly from its first command, with no note at all.

The transcripts also hold a failure the outcome cannot register: one memory-arm session never accessed its available note and stayed clean anyway. That is a never-asked failure inside a perfect score, invisible to the outcome and visible only in the transcript. The other two memory sessions called recall early, got back a one-line index with the note's title and an instruction to expand it by id, and never expanded it. Whether a prompt-time channel put the note's body into any session's context cannot be established from the transcripts, because the transcript format does not record what that channel injects. Agents Forget How, Not What walks through all six sessions.

Process measures keep moving after the outcome saturates: false passes, builds spent re-deriving a command, turns spent reaching a fact the note already held. They do not replace the outcome, and three sessions per arm support no effect in either direction. They are what is left to look at when the outcome has stopped discriminating, and an evaluation that did not record them has nothing to look at.

Screen the control before the memory arm

If the control arm already scores at the ceiling on a scenario, that scenario cannot show memory helping, whatever the memory arm does. Check the control's score first and replace a scenario whose control does not fail. A ceiling found afterwards cannot be repaired by running more sessions under the same frozen rule: deciding to collect more data after seeing the result is optional stopping, and the result it produces is no longer the one the rule was written for.

Part 03
The Pipeline
04

Six Stages of Agent Memory Delivery, and How Each Was Misread

Separating the three failures means observing what happens between "the note exists" and "the agent did the right thing". For a note the harness delivers, there are at least six stages, and each needs evidence of its own.

StageWhat it meansEvidence that it happened
1 · StoredThe note is in the store, exactly onceStore count and note id at session start
2 · SelectedRetrieval chose it for this momentThe retrieval audit line
3 · SentIt was attached to something the model receivedThe delivery channel's own count of appended blocks
4 · Present at decisionIt was in context when the relevant action was chosenWhere the delivery sits relative to that action
5 · Acted onThe agent's actions follow the noteThe ordered stream of executed commands and edits
6 · Reflected in stateThe workspace shows the resultFinal bytes on disk
Never asked: stage 2 never attempted, because the agent had to request the note and did not. Not found: stage 2 attempted and missed. Delivered and ignored: stage 4 holds and stage 5 does not. A run that fails at stage 3 or 4 is none of the three; the note never had a chance.

In vectr's evaluations, every row of this table has been misread at least once, and the misreads ran in both directions. Some instruments reported a stage as passed when it had failed. Others reported a failure that had not happened. Each case is on public record, in the repository's benchmark documentation, its commit history, or the technical report.

Those evaluations deliver notes through two channels. One is a local proxy that sits between the agent and the model API and appends the note to outgoing requests. The other uses hooks: commands the agent runtime runs at fixed points, such as session start or just before a tool call, whose output is added to what the model sees. A control arm gets neither and starts with an empty store.

Stored: a gate that checked for "not empty"

The longitudinal evaluation runs a trajectory of sessions, and its validity gate required the note store to be non-empty before a session counted. Sessions in the hook arm passed that gate holding two or three notes, where every proxy-arm session held exactly one and every control session none. In the session examined most closely, the hook's payload carried two byte-identical copies of the note, the fresh plant and a stale one left from an earlier plant. The commit that fixed the premise traces the extras to a retried setup step that planted a second note on top of the first. The gate answered "is anything there?" when the premise was "exactly one note is there". It now checks for an exact count of one before any money is spent on the session (commit 0969fa4), and the report treats the hook arm as contaminated and does not compare its costs with the proxy arm's.

Selected, not sent: an audit line for an injection that never happened

In the injection evaluation, the daemon holding the store writes an audit line each time it hands a note to the proxy. In the first pilot cell that line recorded one item of 314 characters, while the proxy's own counter read injected: 0 and inject_skipped: 11.

The proxy only appends a block when a request ends with a user turn. The agent's first request ended with a system message, so the note was retrieved for that request, could not be attached to it, and was charged against a cooldown, the rule that stops the same note being sent again for a while. The cooldown then suppressed it on every later request that could have carried it. One shot, spent on the one request that could not take it. Neither rule is wrong on its own; together they guarantee the note never arrives. Across three scenarios and one controlled reproduction, zero of four planned deliveries reached the model.

Had the audit line counted as delivery, each of those cells would have scored as a valid run in which the note arrived and changed nothing: a delivered-and-ignored result manufactured entirely by the instrument. The harness required both counters and marked the cells invalid instead. The same audit log misled in the other direction too. The harness's own preflight probe retrieves through the same daemon, so until the harness started counting only events logged after the preflight, a clean control with no deliveries read as contaminated by an event the proxy never sent. One early pilot cell failed exactly that way.

Sent: a transcript that could not show it, and a parser that saw too much

The injection evaluation's hook channel fired before each prompt and before each tool call. Its validity rule required the note's text to appear in the session transcript. It never did, both cells were marked invalid, and the first write-up concluded that output from those hooks did not appear to reach the model when the agent ran headless. A probe falsified that the same day. Three hooks each emitted a unique canary token, the model was asked to echo any such token it saw, and all three came back verbatim. The transcript format does not render output from those two hook types, so the transcript could neither confirm nor deny delivery through them. Absence from the transcript had been read as absence from the context.

The opposite error turned up in the Camel experiment. Its parser treated any matching tool result after the opening events as a delivery at command time, and so labelled the two recall calls that returned only a one-line index of titles as deliveries of the note. Neither call delivered the note at the command.

Sent, then gone before the decision

A note appended to one outgoing request is visible for one model call. In the proxy design as first measured, the block was not carried into the conversation state the agent kept, and a cooldown suppressed repeats. In 6 of 7 baseline cells the note was out of context before the agent's first edit that changed a file. Every one of those cells passed the delivery gate, correctly, because the note had been sent. Sent is not present at the decision, though, and a cell where the note was gone before the agent chose what to edit says nothing about whether the agent trusts notes. The fix delivers the full note, not just its title, on the turn where the decision is made.

Acted on: a mention counted as an execution

Action checks read the transcript's ordered tool calls. The scorer's signatures for shell actions were regular expressions searched anywhere in the command string, so a command that mentioned a script matched the same pattern as one that ran it. The deploy.sh scorer is the worst instance. A quieter one sat in the injection evaluation: two diagnostic counters named for invocations of a test runner used the same anywhere-in-string search, so grep -rn "run_tests.sh" . or an echo naming the runner counted exactly like running it. Those counters never fed a pass or fail. When they were anchored, the series was re-baselined, with a written warning that values from before and after the change are not comparable.

Reflected in state: a check that scored correct behaviour as failure

State checks read the bytes on disk, and they can be wrong about what the bytes mean. In the deploy scenario, each later session runs in the workspace the previous session left, which is deliberate. But the first session's entry in the deploy queue was still there when the second began. The primary check required the number of queue entries to increase, every agent correctly concluded that the deployment was already queued and that a duplicate would be wrong, and every arm failed the check. The fix resets the files a scenario declares as able to pre-satisfy a later check, at the start of each later session, and the check now asks for at least one entry added in that session.

A release scenario showed a subtler version. Its outcome check asks whether a correctly annotated release tag exists at the end. One memory variant, whose note also carried where and when the fact was established, declined to tag in both of its later sessions, on the ground that tagging would trigger a real and irreversible public release, and asked the user for permission. The harness was non-interactive, so the question had nowhere to go. (The two sessions were not independent: the second started from the first's untagged workspace, so this is one failure and its sequel.) On a check of final state, an agent that became more cautious scores the same as one that became less capable.

Most of these were caught by a second instrument

Nearly every misread surfaced because a second source of evidence for the same stage disagreed with the first: the proxy's counter with the audit line, the canary with the transcript, the file hash with the regex. A stage with one source has nothing to disagree with it. That is a better reason to report stages separately than tidiness, because a separate count is one that can be cross-examined.

Part 04
The Instruments
05

The Agent Evaluation Scorer That Manufactured a Finding

The clearest instrument failure produced a wrong headline: a note delivered and not heeded. The note had in fact been delivered and heeded, and the scorer had matched a read-only command.

A synthetic repository contains deploy.sh. It appears to work: it prints progress, exits zero and writes a state file. It does not work, because a reconciler silently reverts whatever it does, and the correct action is to append an entry to a queue file. The user states that fact once, in the first session's prompt. Before the second session, a fresh process with no conversation carried over, a note carrying the fact verbatim is planted in the store. Two separate gates confirm delivery before anything is scored: the daemon's hook counter incremented, and the note's text appears verbatim in the agent's own transcript. The counter alone would prove little, since it also counts the harness's own probes of the endpoint; the transcript check is what shows the agent saw the note.

The scorer reported that in the second session the agent ran ./deploy.sh as its second tool action, without correcting itself and without calling any memory tool. That reads as the most interesting failure on the list, delivered but not heeded, the one that survives every retrieval and delivery fix.

The mistake signature was the regular expression (\./)?deploy\.sh\b, searched anywhere in the command string, and it matched cat deploy.sh. All three of the agent's tool actions in that session were read-only. Its closing message said why it had left the script alone:

"I did not run ./deploy.sh staging [...] per what I know from this workspace, that path appears to succeed locally but gets silently reverted by the reconciler on its next pass, leaving no trace."

Three properties make this more than a typo in a regex.

The contradiction was already in the record

Scorer checks come in two kinds. An action-derived check reads what the agent did, from the ordered commands in the transcript. A state-derived check reads what the workspace shows, from the bytes on disk. The same result record carried deploy_state_untouched = PASS: the state file's hash equalled the one taken at the start of the session. Running the script writes that file, so an unchanged hash means the script did not run to completion. Beside it sat mistake_committed = true and a check reporting two matching commands. The record disagreed with itself, and nothing in the pipeline compared across the two kinds of evidence.

Two agreeing checks were one check

The two action-derived signals agreed with each other automatically, because both read the same pattern family. They looked like corroboration and were one observation counted twice. The record held two independent observations, not three: one state check against one regex. Whether two checks are independent depends on where each one reads its evidence, and the number of checks says nothing about it.

Deterministic is not the same as correct

No language model judged anything. The scorer was a small reproducible program, and it was reproducibly wrong. Judge-free, state-based scoring is a real advantage, and benchmarks such as VehicleMemBench claim it for good reasons: no grader drift, no sensitivity to a judge's prompt. What determinism buys is the same answer every time. Whether the answer is right is a separate property, and a scorer has to be tested for it like any other software.

Execution position, and what it still misses

The fix makes signatures that mean "the agent performed this action" match only at an execution position: the start of a command, or right after a separator (;, &&, ||, |, & or a newline), after stripping the wrappers a shell allows in front of a program (VAR=value assignments, env, sh, bash, zsh, exec, timeout N). Running the scorer's matching logic over a set of commands shows what it fixes and what it does not:

CommandRuns the script?Search anywhereExecution position
cat deploy.shNo, reads itmatch (wrong)no match
grep -n reconciler deploy.shNo, reads itmatch (wrong)no match
./deploy.sh stagingYesmatchmatch
git pull && ./deploy.sh stagingYesmatchmatch
timeout 60 bash deploy.sh stagingYesmatchmatch
env STAGE=staging ./deploy.shYesmatchmatch
sudo ./deploy.sh stagingYesmatchno match (missed)
bash -x deploy.sh stagingYesmatchno match (missed)
sh -c './deploy.sh staging'Yesmatchno match (missed)
echo "step 1; ./deploy.sh staging"No, prints textmatch (wrong)match (wrong)
Each row run through the scorer's anchored matcher and through a plain search for the same pattern. The anchored matcher splits on separator characters without parsing quotes, which is why the quoted separator inside echo counts as a command boundary.

The anchored matcher is a large improvement and still a heuristic over text, not a shell parser. A parser would settle the quoting case. The list of wrappers (sudo, interpreter flags, sh -c with a quoted script) is open-ended, and every one left off it is a run the scorer will not see.

The contradiction guard, and its direction

A contradiction guard now runs beside every session's score. When an action-derived check says the mistake happened and the scenario's declared state check says the file was untouched, the guard writes a machine-readable contradiction into the record, and the report prints it. It does not pick a winner, because a disagreement is evidence that one of the two checks needs attention, and the harness has no basis for guessing which.

The guard fires in one direction only. When sudo ./deploy.sh runs and the matcher reports nothing, there is no mistake verdict for the state check to contradict, and the miss is silent. The guard also exists only where a scenario declares a state check for its mistake, which most scenarios do not. The report states the general limit: the guard covers only facts observable both ways.

Because every transcript had been preserved, the whole reported grid was re-scored with the fixed matcher at zero model cost, each corrected record written beside its original rather than over it. All five flagged sessions in the reported grid flipped to clean, including the shared first session whose mistake was supposed to be the baseline for counting repeats. In the reported grid, no agent in any arm ever executed the script.

Analogy: a roll call with a headcount at the door

A teacher takes attendance by listening for names. A student says "Priya's off sick today", the teacher hears "Priya" and marks her present. That is search-anywhere matching: it hears the script's name and marks it run. Execution-position matching listens only for a name given as an answer at the start of a reply, which fixes most of it, and still misfires when someone answers for a friend in quotes (the echo row) or when Priya answers in an unexpected form (the sudo row). The headcount at the door is the state check. If the register says thirty and the door counted twenty-nine, someone has to go and look. The guard is that rule, applied only when the register says more than the door.

Interactive Demo 2
Read or Run?

Pick a command or type your own, and say whether the script really ran. The demo applies both matchers and the state check, then shows what the contradiction guard records. The cases worth finding: a false match the guard catches, a false match from the anchored matcher it also catches, and a miss it cannot catch because there is nothing to contradict.

EvidenceSays the script ran?Right?
Both matchers use the pattern (\./)?deploy\.sh\b. The anchored one is a port of the scorer's execution-position logic. The original harness had no guard; pairing the guard with the original matcher shows what it would have caught. The state check is taken as correct here; in a real harness it can be wrong too, which is why the guard records a disagreement instead of resolving it.

What the corrected reading does not show

The fixed scorer does not turn the session into a win for memory. The control arm, with an empty store, also declined to run the script, and the scenario turned out to offer more routes to the answer than it was designed to. The files the scenario started from documented the queue as the sanctioned channel, though not why the script fails. And the scenario's short name, deploy_reverted_by_reconciler, was part of the working-directory path every arm could see. The control transcript shows the agent picking the word "reconciler" out of its own path and searching the workspace for it; the word appears in no workspace file. A path that hands the control the answer is a leakage failure, the kind Your Eval Is Leaking the Answer catalogues. The scorer failure was a different kind: nothing leaked, and the instrument misreported what happened.

What survives is narrower. The agent's stated reason tracks the note's wording, not the path's. Three phrases from the note, "appears to succeed", "on its next pass" and the no-trace clause, appear in its closing message, and none of them occurs in the workspace files or in either control transcript. That shows a delivered note entering an agent's stated reasoning at the level of its wording. Two control transcripts are a thin base rate, and whether the note's wording or the situation produced the behaviour cannot be separated at this sample size, so it does not show that the note caused anything.

The wrong number was never published. It existed in an internal draft, and catching it was ordinary quality control before submission. What is unusual is the reporting: the technical report puts the wrong number, the mechanism that produced it and the guard that would have caught it in one document, where normally the reader would never learn the wrong number existed.

06

Evaluation Metrics Have Invariants: Assert Them in Code

Two more metric failures from the same evaluation show a cheaper defence than re-reading transcripts.

A part larger than its whole

billable_tokens_to_fact was meant to count the tokens billed from the start of a session up to the moment the agent acquired the fact it needed. In every session where it was defined, it came out larger than the entire session's token count, by up to 2.6 times: 340,907 against 132,567 for the hook-arm session at k=2, the second session of its trajectory. A prefix cannot exceed its total.

The cause sits in how the transcript is written. It emits several assistant events per API call, and each one carries the call's usage record. The metric added a usage figure for every assistant event, so each call was counted more than once. That session had 33 assistant events across 13 calls, and the same sum taken over the whole session reaches 354,186.

Why the sum overshoots

Number the API calls i = 1 … C. Call i bills u_i and appears in the transcript as e_i assistant events, each carrying the same u_i.

true total T = Σ u_i event sum E = Σ e_i · u_i

E = T only when every e_i = 1. E / T is the average of e_i weighted by usage, so it need not match the plain average: in that session, 33 events over 13 calls is about 2.5 events per call, while the whole-session sums came to 354,186 against 132,567, about 2.7. Deduplicating usage by call before summing would repair the sum; the report instead stopped citing the metric. Either way, an assertion prefix ≤ total would have caught it on the first run.

A broken metric can still produce a tidy ordering. At k=2, the arms ranked on this metric in exactly the order of the tokens each session wrote into the prompt cache: 13,479 for the verifiable variant (a proxy-delivered note that adds a one-command check of its claim), 20,977 for control, 23,875 for the plain proxy note and 42,653 for the hook arm. That is a property of the cache, not of the agent, and the k=3 sessions do not keep the order. The replacement measure, the size of the context the agent holds when it acquires the fact, is in the same result record, is not double-counted, and points the other way: on it, control is the cheapest cell at both k=2 and k=3.

Turns that outnumber the turns

turns_to_fact counted assistant events, not conversational turns, and in every session it exceeded the session's own turn count: 33 against 20 in the verifiable session at k=2. The pre-registered expectation, two turns or fewer, was written in conversational turns, so comparing across the two units overstated the miss by roughly 2 to 3 times. Recounted from the preserved transcripts in the intended unit, the cells read 11 to 15 against an expected 2 or fewer, still a miss of five times or more.

The units also disagree about direction. At k=2, counted in conversational turns, the hook arm reached the fact faster than control, 11 against 13. Counted in assistant events, the same comparison is a tie, 30 against 30. Put the unit in the metric's name, turns_to_fact_api_calls or similar, and nobody compares across units by accident.

Invariants are free tests

Neither failure needed a clever audit. Each broke a relation that holds by definition: tokens up to a moment cannot exceed tokens in the session, and turns up to a moment cannot exceed turns in the session. An evaluation harness is software, and the cheapest tests for software check that a result does not exceed something it cannot exceed.

Illustrative invariant checks on one result record
def check_invariants(r):
    errors = []
    # A prefix cannot exceed its whole.
    if r["tokens_to_fact"] is not None and r["tokens_to_fact"] > r["session_tokens"]:
        errors.append("tokens_to_fact exceeds session_tokens: usage summed per event?")
    if r["turns_to_fact"] is not None and r["turns_to_fact"] > r["session_turns"]:
        errors.append("turns_to_fact exceeds session_turns: check the unit")

    # Delivery stages can only shrink along the pipeline. The chain stops before
    # "acted on": an agent can do the right thing without the note (controls do).
    chain = ["stored", "selected", "sent", "present_at_decision"]
    for earlier, later in zip(chain, chain[1:]):
        if r[later] and not r[earlier]:
            errors.append(f"{later} without {earlier}")

    if r["deliveries"] == 0 and r["delivered_and_ignored"]:
        errors.append("ignored a note that was never delivered")
    return errors
Part 05
The Practice
07

How to Evaluate Agent Memory Stage by Stage

Report each stage, not one ratio. For every run, record whether the note was stored, selected, sent, present at the decision, acted on and reflected in state, each from its own evidence. The three failures then become three counts, and each count points at its own fix. A record shaped roughly like this is enough:

Illustrative per-run record (not real data)
{
  "run": "scenario-a/hook/session-2",
  "stored":   {"ok": true, "evidence": "store count = 1, note id matches plant"},
  "selected": {"ok": true, "evidence": "audit line names the planted note"},
  "sent":     {"ok": true, "evidence": "channel counter = 1; canary echoed"},
  "present_at_decision": {"ok": true, "evidence": "delivered before first mutating action"},
  "acted_on": {"ok": false, "evidence": "forbidden command at execution position, action 4"},
  "state":    {"ok": false, "evidence": "state file hash changed"},
  "contradictions": []
}

Treat a failed delivery as invalid, not negative. A run in which the note never reached the model does not test whether the note helps. Exclude it and report it as excluded, never average it in as "memory did not help". The same goes for any run whose premise failed, such as a store holding the wrong number of notes. Report exclusions per arm, too: when the arms have different gates, as the hook and proxy arms did, runs are dropped under different criteria in each, and the survivors are no longer a like-for-like comparison. One rule cuts the other way: whether the agent used the note is the outcome being measured, and it must never be a gate. A design that gates on behaviour cannot observe the failure it was built to find.

Screen for saturation, and record process measures anyway. Confirm that the control fails a scenario before paying for memory arms. Record false passes, turns spent re-deriving a known fact, and commands that repeat a known mistake, because these keep moving when every session eventually succeeds.

Pair every action check with a state check, and make disagreement an error. An agent's commands and the workspace's bytes are two views of the same event. When they disagree, one of them is being read wrong, and the harness should say so in the record. Count two checks as independent only if they draw on different evidence; two regexes from the same family are one check.

Keep the raw transcripts, and re-score beside the original. If I could keep only one of these practices, it would be this one. The matcher fix and the recount of turns in their intended unit were both applied after the runs had finished, and preserved transcripts are what made them cost nothing. Write the corrected record next to the original, never over it. Version the scorer with the harness, too. A check rewritten for a repaired harness assumes the repair: the deploy scenario's new check asks for at least one queue entry, which the first session's leftover entry would satisfy if the check were replayed over runs made before the reset.

Assert the invariants, and test the matcher. Parts do not exceed wholes, per-turn counts do not exceed turn counts, stage counts only shrink along the pipeline, and a run with zero deliveries has zero delivered-and-ignored events. Turn a command table like the deploy.sh one into unit tests for whatever decides "the agent ran X", including the rows the matcher gets wrong, so that the known misses are written down rather than rediscovered.

08

What This Predicts for Agent Memory Evaluation

Prediction 1. For memory the agent must query voluntarily, decomposing failed tasks by stage will attribute most of them to never asked, not to retrieval. The measurement is the fraction of failed tasks in which the relevant note was never retrieved, split by whether retrieval was attempted at all. This runs against the forced-retrieval result, where Yuan, Su and Yao find retrieval failure dominant. The difference should come from the asking stage, which their setup removes by design.

Prediction 2. For memory delivered by the harness, the dominant remaining failure will be delivered and ignored, and it will concentrate on notes the agent cannot corroborate from its workspace. In vectr's injection runs, once delivery was fixed, a full note on the decision turn changed behaviour in the two scenarios where the agent could confirm it by reading a second file, and not in the one where the note was the only source, at one run per cell. That is a direction, not a result. When a Memory Should Become a Check works through what follows for enforcement.

Prediction 3. Evaluations that match shell text with unanchored patterns and have no state-derived cross-check will, when re-audited against workspace state, show contradictions at a non-trivial rate. The test is to add the cross-check to an existing harness and count.

Close
The Line to Keep

An outcome score adds three failures into one number, and no analysis afterwards can take the sum apart. The split has to be made while measuring: one record per stage, each from its own evidence, and for every stage a second source that can disagree with the first. The defences involved are cheap. A second counter kept a note that never arrived from being scored as ignored, and a canary showed that an empty transcript was not an empty context. The file hash that contradicted the regex was already in the record and only needed comparing. A one-line assertion that a part is no larger than its whole would have flagged both broken metrics on their first run.

None of these defences takes the first instrument's word for it because the instrument is deterministic. A scorer that never calls a model can be wrong in exactly the same way on every run, and the repetition only makes the wrong answer look settled.

↑ Back to top
09

Sources