When a Memory Should Become a Check
Delivered as text, a memory is advice, and advice works only if the agent believes it. Compiled into a check, it needs no belief, but a rule that is always enforced is never tested again. The design question is which memories should become checks, and what a check has to keep so that it can still be retired.
A memory delivered as text is advice. It works only if the agent believes it, and agents are sceptical of text they cannot check. In a set of controlled runs, an agent that received a correct note, in full and on the turn where it mattered, still declined to act on it, and said why: the claim had not come from the user, and nothing in the repository backed it up.
A memory compiled into a check does not need belief. The check runs, and the action passes or it does not. Recent work that turns a user's corrections into runtime checks cut repeat violations well below what the same corrections achieved as stored memory.
That seems to settle the matter: compile every memory that can be compiled. It does not, because enforcement carries a cost that advice does not. A rule that is always enforced can never be tested again. If the agent can never take the forbidden action, nobody observes what the action would do now, and a rule that has outlived its reason stays in force with nothing to retire it.
The question that remains is narrower than advice versus enforcement: which memories should become checks, and what a check has to carry so that it can still be retired.
Advice Has a Belief Problem: Why Agents Decline Correct Memory Notes
The runs that exposed the problem were built to separate belief from the other ways a memory can fail. They come from an evaluation harness in vectr, an open-source code-search and agent-memory tool I maintain. Each scenario is a small synthetic repository with one trap and one planted note that describes the trap:
- A generated config.
config/limits.confis produced bytools/gen_limits.py, so a hand edit is silently reverted the next time the file is regenerated. - A superseded API. The file being edited still calls
format_currency_legacy(); the correct call isformat_currency(). - A flaky test.
./run_tests.shincludes a test that fails at random, and./run_tests.sh --coreis the deterministic subset to run instead.
Scoring is mechanical. The scorer looks at the final bytes of files and at the commands the agent actually executed, never at its prose, so an agent that says it will use the new function and then does not scores the same as one that says nothing. Every scenario also runs with the note withheld, as a control, and in the first full round those controls walked into the trap in all seven trap scenarios tried. In these runs, "believing" a note means only that it changed which files the agent wrote or which commands it ran.
The first rounds failed for mechanical reasons. At first the note was retrieved and never delivered, because the one request it was spent on could not carry it. Then it was delivered as a title only, on a single early turn, and in six of seven baseline runs it had left the agent's context before the first edit. Those were delivery bugs in the harness, and they were fixed: the full note, delivered on the turn where the decision is made.
With delivery fixed, the note changed what the agent did in two of the three scenarios. In the third, the flaky test, it did not, and it did not through a second, independent delivery channel either: a hook that hands the note to the agent as it is about to use a tool, instead of text appended to its request. An earlier round had already produced the sharpest version of the same result. With the note on the decision turn, in full, and stripped of the cautionary framing it was normally wrapped in, the agent declined to act on it and gave its reasons: the claim had not come from the user, and nothing in the repository corroborated it. In that round, changing the wrapper, sending the full body, and presenting the note as the user's own each left the outcome where it was, at two runs per variant.
What separated the two successes from the failure was not the channel or the wording. In both successes, the agent could confirm the note by reading a second file: the generator script in one case, the replacement function's definition in the other. The transcript of one configuration run shows the agent opening tools/gen_limits.py and establishing from it that the config file is generated. In the flaky-test scenario the note was the only source for its claim. Nothing in that repository recorded that a test was flaky, and seeing it takes repeated runs of the suite. A real project might have a quarantine list or a flaky marker on the test; the scenario left both out on purpose, so that the note stood alone. On the second channel the agent applied the note's other piece of advice, a rounding fix in the code, while still declining the test command it had no way to check.
Each of these results is one run per scenario and channel. They show a pattern, not a rate, and the harness's own write-up says so.
The general reading: a delivered note competes with the agent's own judgment, and the judgment wins when the note is uncorroborated and asks for something the agent would not otherwise do. Running the whole test suite is what a careful engineer does. A note from nowhere asking to skip part of it is the kind of text a careful agent ought to distrust.
A contractor starts on Monday and finds two sticky notes on the desk. One says the rate-limit config is generated, so edit the generator and not the file. There is a script called gen_limits.py in the tools folder, and thirty seconds of reading confirms it. The other says one of the tests is flaky and to run only the core suite. Nothing on disk says so, and the note is unsigned. A good contractor follows the first note and runs the full suite anyway, because an anonymous request to skip tests is what a careful person is supposed to refuse.
The analogy breaks in one place. The contractor can walk over and ask someone. An agent running unattended cannot, so a note it cannot check is simply outvoted.
An agent that does whatever unverified text in its context tells it to is the security problem, not the fix: instructions planted in auto-loaded memory files get obeyed often enough to be a named attack surface. Making delivered notes more persuasive across the board would reopen that hole, and Deterministic, Reproducible, and Wrong, my technical report on corroborating memory notes, calls it a security regression outright. Whatever fixes the belief problem has to work without asking the agent to believe more.
Agents Over-Trust Instructions and Under-Trust Corrections works through why planted instructions get obeyed while true corrections get ignored, and why the deciding factor is the shape of the text rather than its source. For deciding where enforcement belongs, one consequence is enough: louder wording does not close the gap, because what the agent lacks is evidence.
A Check Does Not Need to Be Believed: Compiling User Corrections Into Runtime Enforcement
The strongest recent evidence for the alternative is TRACE, for Test-time Rule Acquisition and Compiled Enforcement (Zhou et al., 2026). It watches a user's conversation with a coding agent for corrections, such as "You left debug files again. Clean them before stopping." Each correction is rewritten as an atomic rule with an applicability condition that says when it should and should not fire, and compiled into three parts:
- an applicability check, which decides whether the rule is in play for a given task;
- a behaviour instruction, text placed in the agent's context while the rule is active;
- a verifier, attached to a hook that fires at a prompt, a tool call, a file write or the end of the task. If the verifier fails, the hook interrupts the event and hands the violation back to the agent, which has to fix it before the task is allowed to finish.
So the compiled form still carries the advice. What it adds is a gate that does not care whether the advice was believed. The retry loop is bounded, at three retries per task in the paper's diagnostic, after which the run ends with the violation logged.
The evaluation removes the target preference from every test prompt and freezes each condition's rule or memory store after a shared training stream, so the agent can satisfy a preference only through what it acquired earlier. An agent with no memory therefore violates every one, which is why each result below is a reduction from 100%. On held-out tasks from ClawArena, a coding-agent benchmark, TRACE cut violations to 37.6% in distribution and to 2.0% on task families it had never seen, with task pass in the same band as having no memory at all. The strongest memory baseline, Mem0, left 51.1% violated in distribution. On memory-intensive tasks derived from MemoryArena, violations fell from 100.0% to 60.5%, while task pass matched or exceeded the strongest memory baseline.
The cleanest comparison is a diagnostic earlier in the paper, because it holds the knowledge fixed and changes how it reaches the agent. It uses 19 held-out tasks built from real correction cases, with 29 preference checks among them, and averages over six models.
| Condition | What the agent gets | Compliance |
|---|---|---|
| No rules | The task, with the preference removed | 31.6% |
| Mem0 | Past corrections retrieved from a memory store | 42.5% |
| Relevant rules | The rules that apply to this task, as text | 54.0% |
| All rules | The user's whole rule set, as text | 55.0% |
| Compiled rules | The rules that apply, as checks that gate completion | 70.1% |
The like-for-like pair is relevant rules against compiled rules: the same applicable rules reached 54.0% as text and 70.1% as checks that gate completion. Part of that gap is the retry loop, since only the compiled condition gets its failures fed back and up to three more attempts. That is not a confound to subtract. Feedback on a failed check is most of what a check is. Putting the rules in context bought 22 points over nothing; turning them into checks bought 16 more.
What the numbers do and do not show
Three things bound them. The users in the main benchmarks are simulated: a model built from a real user's correction history decides when to correct the agent, validated against held-out real correction moments at a decision F1 of 0.906. The diagnostic draws on one researcher's two months of transcripts, deliberately, so that preferences do not conflict across users. And enforcement covers only what has been compiled. Compiled rules still left 29.9% of the diagnostic's checks unmet, and on unseen families of memory tasks the violation rate stayed at 97.0%, which the authors attribute to constraints no compiled rule covered.
What changed is who has to do the believing. A check turns a claim the agent must accept on trust into an observation the harness makes. The agent does not need to be persuaded that the rule is right, only told that the check failed and why. The flaky-test scenario failed on that point: the harness can see which command ran, and a check on the command would never have required the agent to take anyone's word for anything.
Two different things could be checked, and they are easy to run together. One is whether an action breaks the rule: did this command, this diff, this final workspace violate it? The other is whether the rule still matches the world: is the reason it was written still true? "The agent's diff touches config/limits.conf" decides the first. "config/limits.conf is still generated by tools/gen_limits.py" decides the second.
A compiled check needs the first predicate to work at all. Nothing in the enforcement loop needs the second, so nothing in it evaluates the second.
If enforcement were free, the conclusion would be to compile everything that has a predicate. It has two costs.
An Enforced Rule Cannot Be Tested Again
A Directive Is Not a Memory makes the underlying argument: obeying a directive destroys the evidence that would test it. A rule like "never run the tests with the system interpreter" is followed, and following it produces a record of compliance, not of correctness. The tests pass in the virtual environment, CI uses the virtual environment, the docs say to use it. Every one of those signals is a consequence of the rule having been obeyed, and none of them says whether the rule is still needed.
Advice leaves a gap in that loop. An agent that ignores advice sometimes generates the counterfactual: it runs the forbidden command, and the result shows whether the rule still holds. The flaky-test agent that ran the full suite anyway was producing that kind of observation, the forbidden action together with its outcome. A blocking check closes the gap. The more reliably a check is enforced, the less anyone learns about whether it should still exist.
A rule library has selective labels
Statistics has a name for data whose outcomes exist only where an earlier decision allowed them. In the selective labels problem, the standard example is bail: whether a defendant fails to appear in court is observed only if the judge released them (Lakkaraju et al., 2017). The outcomes that would show whether a detention was necessary are the ones detention prevents. A library of blocking checks is selectively labelled by construction. Every blocked action is an outcome nobody observes, and those are the only outcomes that could say whether the block is still needed.
Why a moving codebase makes it expensive
Reinforcement learning calls the underlying trade the exploration problem. An agent that always takes the action it currently rates best, the greedy action, stops learning about the others. Sutton and Barto note that when the true values of actions change over time, exploration is needed even if rewards are perfectly predictable, "to make sure one of the nongreedy actions has not changed to become better than the greedy one." A rule library lives in that kind of nonstationary world. The package the system interpreter lacked gets installed globally; the generator script is deleted in a refactor; the deprecated function is finally removed and the lint rule that guarded it now checks for a name nobody can call. A rule's reason is a fact about a codebase that keeps moving.
The remedy A Directive Is Not a Memory proposes is a deliberate, scoped test of the forbidden action: take it once, in a sandbox, and record what happened. A check system with no path for that has made its rules permanent without anyone deciding to.
Retirement by complaint
TRACE does retire rules. When a new correction contradicts an existing rule, its lifecycle resolver supersedes the old one: a new rule is installed, and the old one is archived and kept on disk for audit and manual rollback. That catches the stale rules that hurt someone enough to prompt a correction. It misses the quiet ones. A user corrects what goes wrong in front of them, and a rule that has become merely unnecessary rarely does anything wrong in front of anyone. It forces a slower path, an extra step, a redundant command, and its cost is paid in small amounts that never show up as an event. A check that blocks an action also hides what the agent would have done, so even a user who is watching sees the constraint and never the case against it.
Anyone who has inherited a pre-commit configuration knows the result. Every hook in it was added after something went wrong, few of them say what went wrong, and removing one feels like a bet nobody wants to own. Agents are about to write rules like these faster than people ever did.
In words: each month a new rules arrive, a fraction h of the live ones lose their reason, anchors catch a share c of those, and the rest stay enforced until something re-tests them. L counts enforced rules whose reason still holds. S counts enforced rules whose reason has lapsed without anything noticing. With a re-test path (r > 0) the library settles, and the stale share f* depends only on how fast reasons lapse, how much of that the anchors catch, and how often the rest is re-tested. Give reasons a half-life of one year (h ≈ 0.058 per month), let anchors catch 60% of lapses, and re-test each uncaught stale rule about once a year (r ≈ 0.083 per month): f* ≈ 0.22.
With r = 0 there is no steady state. Live rules level off at a/h while uncaught stale rules keep accumulating, so the stale share climbs toward 1 however good the anchors are, unless they catch every lapse. With the same one-year half-life and no anchors, stale rules outnumber live ones after about 27 months; anchors that catch 60% push that to just under five years. Anchors slow the inflow of stale rules, and something else has to drain what they miss. Retirement by complaint is a small r that only reaches the rules that hurt someone; a deliberate re-test path is the general one. Every input here is an assumption rather than a measurement. Reasons do not lapse at a constant rate, anchors also fire on changes that leave the reason intact (each false flag costs an evaluation, not a rule), and a stale rule is not always harmful. The structure is what carries over.
Set how fast rules are added, how long their reasons last, how many lapses the anchors catch, and how often uncaught rules are re-tested. The chart splits enforced rules over five years into those whose reason still holds and those enforced after their reason is gone. Start with re-tests at zero and watch the stale band grow without limit, then add a small re-test rate and watch it level off. Rules added per month changes the size of the library but not the share.
Compiled Checks Are Wrong in Their Own Ways
The second cost is that a check can be wrong, and it is wrong with more authority than a note.
A mistaken memory becomes a mistaken block
A check compiled from a mistaken memory blocks correct work, and the agent has no standing to argue. A note can be weighed against other evidence. A failed verifier ends the event, and the rule's only voice at that point is its failure message.
Satisfied in letter, violated in spirit
A check that is too literal gets satisfied in letter and violated in spirit. ImpossibleBench (Zhong, Raghunathan and Carlini, 2025) measures what coding agents do when a task's unit tests contradict its specification, so that any pass requires a shortcut. With the tests writable and the instructions saying not to modify them, frontier models passed impossible tasks four ways: by editing the tests anyway, by overloading comparison operators so that equality checks always succeed, by recording state so that the same input returns different answers on different calls, and by special-casing the exact inputs the tests use. Making the tests read-only stopped the edits and left other routes in use, special-casing and operator overloading among them.
The tasks were impossible on purpose, so the rates describe behaviour under pressure rather than on an ordinary day. The structure carries over to compiled rules. A check guards one route to an outcome, and an agent pushed hard toward the outcome can take a route the check does not watch.
A check is a program, and programs have bugs
My own evaluation scorer, a deterministic program with no language model in it, matched shell text anywhere in a command. It counted cat deploy.sh as running the deploy script, and so reported an agent committing the exact mistake a delivered note warned against. The transcript shows the agent reading the script, declining to run it, and citing the note. Deterministic, Reproducible, and Wrong documents the defect and the guard that would have caught it.
The same class of error appears in the worked example in TRACE's appendix. Rule F2 comes from a user's request to clean up log files before a project closes. It compiles to a hook on every shell command, and its applicability check is four regular expressions meant to detect a command that writes to a file with a log-style name. As printed, they also match cat run_log_20260915.log and python analyze.py run_log_20260915.log, which read a log rather than write one, because the patterns accept the file name anywhere after cat or python. The appendix does not report how often patterns like these misfire in practice, and a library of 47 rules is small enough to audit by hand. None of this makes the paper careless. Matching a command's text is the obvious first way to build a command check, and text alone cannot tell naming a file from acting on it.
| Command | Runs deploy.sh? | Pattern (\./)?deploy\.sh\b |
|---|---|---|
| ./deploy.sh staging | yes | fires |
| cd ops && ./deploy.sh staging | yes | fires |
| timeout 60 bash deploy.sh staging | yes | fires |
| cat deploy.sh | no | fires |
| grep -n reconcile deploy.sh | no | fires |
| echo 'do not run deploy.sh directly' | no | fires |
A check on shell commands needs negative cases in its test suite: reads, searches, listings and quoted mentions of the same file, all of which must not fire. Where the outcome is visible in the workspace, compare the two signals as well. If the check says the deploy ran and the deploy state is untouched, one of them is wrong, and the record should fail loudly instead of quietly picking one. That comparison is the contradiction guard the technical report adds, and it would have caught the scorer's false headline without anyone reading a transcript.
Which Agent Memories Should Become Runtime Checks
Four conditions, all of which should hold.
1. The rule has an action predicate
Something mechanical can decide whether an action violates it. "Do not hand-edit config/limits.conf; it is generated" has one: flag any change to the file, or regenerate it and compare with the committed version. "Use format_currency(), not format_currency_legacy()" has one: a lint rule. "Deploys to staging go through the queue file, because the reconciler reverts direct deploys" has a weaker one: detect the deploy script at execution position. "The reconciler reverts direct deploys" has none that can run inside a session. It is a fact about a process the agent cannot observe, and it has to stay advice.
Deterministic, Reproducible, and Wrong calls this class TOLD, writing its class labels in capitals: facts that are "uncorroborable: nothing in the workspace records them, the fact enters once from the user." The report draws the consequence as well: where a claim can be compiled into a check, enforcement answers the compliance problem completely, and TOLD facts are the claims that cannot be compiled, because there is nothing to check them against. The deploy memory shows that one note can hold both kinds. Its prescriptive half, use the queue, has a predicate. Its causal half, the reconciler reverts direct deploys, does not.
2. The violation is silent
If breaking the rule produces an immediate, visible failure, a compile error or a red test, the environment already enforces it, and a check adds cost without adding signal. Checks earn their place where the failure is invisible: a reverted config, a green test that could not see the change, a deploy that looks successful and is undone later.
A loud but misleading failure counts as silent. A flaky test's random red is visible, but it points at the wrong cause and says nothing about a rule being broken. The same rule can also land on either side depending on the project. "Never run the tests with the system interpreter" announces itself if the missing package fails at import. If the suite guards that package with pytest.importorskip, the system interpreter produces a green run that quietly tested less, and the violation is silent.
3. The claim is trustworthy enough to enforce
Enforcement amplifies whatever it enforces, including mistakes. A rule the user stated, or one verified against the workspace, can be compiled. A rule the agent inferred from one observation should not be, because compiling it launders a guess into policy. A Guess Becomes a Fact on the Next Read describes the same laundering for plain notes; a check does it with a gate attached. The trust classes in Agent Memory Needs a Trust Ladder are the input this condition needs, and they have to be recorded when the memory is written, because they cannot be recovered afterwards.
TRACE takes its rules only from user corrections, which puts every rule on the right side of this line at its source. There is still one hop between the user and the check: the rule text and its applicability condition are a model's rewrite of what the user said. In the F2 example the user's words were about what is left behind when a project closes, and the compiled applicability condition is about the moment a log file is written. Whether a rewrite keeps the user's meaning is itself a claim, and keeping the user's words next to the check is what makes it possible to find out.
4. The reason is anchored to something observable
Every rule exists because some condition held when it was written: the file is generated, the interpreter lacks a package, the old API is deprecated. If that condition can be tied to an artifact (the generator script, the dependency manifest, the deprecation marker), the check can be flagged automatically when the artifact changes. The anchor belongs on the condition, not on the rule. A Directive Is Not a Memory points out that anchoring a rule to a file is decorative, because the rule was never a claim about that file. "config/limits.conf is generated by tools/gen_limits.py" is a claim about those two files, so a change to either is evidence. If no artifact carries the condition, the check needs a named human owner and a review date, or it will outlive its reason.
Memories that fail any of these stay advice. That covers most of what an agent learns: explanations, locations, design history, preferences with exceptions, and facts about systems outside the workspace.
The Middle Rung: A Memory Note That Carries a Verify Command
Between advice and enforcement there is a rung that keeps something of both: advice that carries its own verification. The note does not demand belief and blocks nothing. It offers a one-command check the agent can run to confirm the claim, together with where the claim came from and when it was last confirmed.
"Hand-editing config/limits.conf is reverted; to confirm, regenerate it into a scratch file with tools/gen_limits.py and diff the two" is still advice. But it turns the trust question into an observation the agent makes itself, so the judgment that rejected the bare note now has evidence to weigh. It also preserves the test. The agent runs the check, and if the rule no longer holds, the check says so. In the configuration scenario the agent went looking for that evidence on its own. A verify command hands it over instead of hoping the agent goes looking. It has to be safe to run, though: read-only, or writing only to a scratch path, because an agent that runs it halfway through an edit should not lose the edit to a regenerated file.
Agents Over-Trust Instructions and Under-Trust Corrections lays out these corroboration affordances (an origin, dates, a verify hint, a content anchor) and why honest claims can carry them while planted ones cannot. Whether they change behaviour is an open question. My own pilot was designed to test this and could not answer it, and Deterministic, Reproducible, and Wrong reports why. The report also records a bound that no sample size removes: a check on a claim's content can only be built where the workspace records the fact. The flaky test shows the limit in practice. Some claims have no cheap verification, and for those this rung is empty.
The full ladder, from least to most forceful, is my own synthesis rather than a published taxonomy.
| Rung | Needs belief? | Preserves the test? | Fits |
|---|---|---|---|
| Note on request | yes | yes | reference material, history |
| Delivered note | yes | yes | told facts with no predicate |
| Delivered note with a verify command | partly: evidence replaces trust | yes | checkable claims where a block would be premature |
| Warning at the action, non-blocking | yes, at the moment of action | yes | high-cost actions with a detectable signature |
| Blocking check before the action or before completion | no | no, unless a scoped override exists | trusted, silent-failure rules with an observable anchor |
For the flaky test, my own choice would be the warning at the action plus a flaky marker in the test file itself, which gives the rule the anchor it lacked: when someone fixes the test and deletes the marker, the anchor changes and the warning is re-examined with it. The option I would rule out is a block with no anchor and no owner, which is the one a system that compiles every correction would produce.
Pick a memory and see where the four conditions put it on the ladder above, from a note on request to a blocking check, and which condition decided it. Then flip the conditions one at a time to see which one is doing the work. The flaky test and the deploy rule are the instructive cases: trusted, silent, checkable at the action, and still without anything that would say when their reason is gone.
What a Compiled Check Must Carry to Stay Retirable
If a memory does become a check, the check should keep what the memory knew, not just the rule.
- The condition, not only the rule. The rule is an imperative. The condition ("this file is generated by that script") is a proposition that is true or false today. Store it in a form that can be evaluated, and evaluate it.
- The anchor. A content hash of each artifact the condition depends on. When a hash changes, flag the check as possibly stale instead of enforcing it silently. A changed anchor means the reason may have gone, never that it has, which is how Agent Memory Needs a Trust Ladder treats staleness for notes; vectr, for one, flags a note whose anchored file has changed and keeps delivering it with the flag attached rather than dropping it. For a check, the flag should trigger an evaluation of the condition: re-arm with the new hash if it still holds, suspend and escalate if it does not, and fall back to a warning if it cannot be evaluated.
- The provenance. Who stated the rule, in what words, and when, and what turned those words into a predicate. A check that blocks correct work can then be traced to its source and argued with.
- A failure message that teaches. When the check fails, its message is the memory. "Blocked:
config/limits.confis generated bytools/gen_limits.py; edit the generator's input instead" does the job the note was supposed to do, at the moment it matters. - A logged override. A scoped way to take the forbidden action deliberately, which records what happened. Each logged override is a re-test, the r in the stale-share formula. Where the forbidden action is irreversible, the override is sandbox-only or absent, and the check should declare itself untestable by design rather than look like the others.
# illustrative: the hash, dates, words and generator flag are made up
check: limits-conf-is-generated
rule: Do not hand-edit config/limits.conf.
predicate: the agent's diff touches config/limits.conf
condition:
claim: config/limits.conf is generated by tools/gen_limits.py
verify: python tools/gen_limits.py -o /tmp/limits.conf && diff -q /tmp/limits.conf config/limits.conf
anchors:
- path: tools/gen_limits.py
sha256: 9c41e0d27b3a8f15
last_confirmed: 2026-09-12
provenance:
stated_by: user
words: "limits.conf gets regenerated, never edit it by hand"
stated_at: 2026-08-30
compiled_by: model, confirmed by user
on_violation: >
Blocked: config/limits.conf is generated by tools/gen_limits.py,
so a hand edit is reverted on the next regeneration. Edit the
generator's input instead. To confirm, run the verify command.
on_anchor_change: run verify; re-arm if it passes, suspend and flag if not
override:
allowed_in: sandbox
logged: true
last_override: never
Most of the record is ordinary configuration. The fields that keep it retirable, the condition, the anchors and the override, are the ones the enforcement loop never reads, which is why they have to be required rather than left to whoever compiles the rule. To be clear about what exists: vectr stores notes and flags them; it does not compile anything into checks, and the record above is a design, not something I have shipped.
What This Predicts About Enforcement in Agent Memory
Prediction 1. For claims with a cheap verification command, a note carrying that command will be followed at a rate close to a blocking check's, without the check's cost to testability. For claims without one, only a check will change behaviour reliably. The measurement is compliance by rung on matched scenarios, split by whether a one-command verification exists. The runs with planted notes point this way at one run per cell, which is a direction and not a result.
Prediction 2. Checks without anchors will accumulate. The set of enforced rules will only grow, apart from the rules a user happens to contradict, and a rising fraction will block actions whose original reason no longer holds. Checks with anchors will be flagged when their reason changes, and that fraction will be much smaller, though it will still rise unless something re-tests what the anchors miss. The measurement: take a long-lived rule library, evaluate each rule's condition today, and split the stale share by whether the rule was anchored and whether it was ever overridden.
Prediction 3. Compiling low-provenance memories (agent-inferred, unverified) into checks will produce more blocked-correct-action events than compiling user-stated or verified ones. The measurement is the rate at which enforced rules are overridden or removed, split by the provenance of the memory they came from.
Advice fails when the agent does not believe it, and in the runs with planted notes, what moved the agent was evidence it could check, not wording. A check removes the need for belief, and with it the only route by which a rule ever met a counterexample. Every blocking check is a bet that its reason will hold for as long as the check is enforced.
Compile a memory into a check when it has a predicate, fails silently, comes from a trustworthy source, and is anchored to something that will say when its reason is gone, and give the check a logged way to be broken on purpose. Otherwise give the agent the check to run, not the check to obey.
Sources
- Zhou, Y., Guo, K., Zhuang, H., Wang, X., Huang, Y., Liang, Z., Chen, P.-Y., Gao, T., Moniz, N., Chawla, N. V., Zhang, X. Getting Better at Working With You: Compiling User Corrections into Runtime Enforcement for Coding Agents. arXiv:2606.13174, 2026.
- Saha, S. Deterministic, Reproducible, and Wrong: How Our Own Scorer Manufactured an Agent-Memory Finding, and the Guard That Would Have Caught It. Technical report, 2026. doi:10.5281/zenodo.22079952.
- vectr,
benchmarks/injection_utility/README.md. - vectr,
agent/working_context_store/_store.py,check_staleness(). - Lakkaraju, H., Kleinberg, J., Leskovec, J., Ludwig, J., Mullainathan, S. "The Selective Labels Problem: Evaluating Algorithmic Predictions in the Presence of Unobservables." KDD 2017. doi:10.1145/3097983.3098066; open copy at PMC5958915.
- Sutton, R. S., Barto, A. G. Reinforcement Learning: An Introduction, 2nd ed. MIT Press, 2018.
- Zhong, Z., Raghunathan, A., Carlini, N. ImpossibleBench: Measuring LLMs' Propensity of Exploiting Test Cases. arXiv:2510.20270, 2025.