Answers
Shouldn't rules for AI agents be enforced in code instead of written down as memory?
Some of them, yes — and those should never have been instructions in the first place. A rule that holds for as long as the system is built the way it is, like one deployment at a time, belongs in a lock the agent cannot skip. But a guardrail only knows how to enforce a rule while it holds; it has no way to say that the rule stopped holding, who decided that, or what replaced it. Most rules a team gives its agents are of that second kind: decisions someone made and someone will revisit — what may go to a client, which branch ships, what a price is. Arroway holds those. Each one carries the condition that ends it, a replacement points at the rule it retires, and only a person makes a rule binding. Put the invariant in code and the decision in memory; a rule that is both gets the lock and the record of why the lock exists.
Last updated September 24, 2026
A lock is the right answer for a rule that never changes
The objection starts from something true. An instruction is text an agent reads and weighs against everything else in front of it, and under pressure it sometimes loses. When two builds collide because a written rule said not to run them at once, turning that sentence into a mutex is the correct fix: the collision becomes impossible instead of discouraged. Anything that can be made mechanically impossible should be, and a memory that asked you to keep such rules as prose would be asking you to accept a weaker guarantee for nothing in return.
A guardrail has no field for “this stopped being true”
The trouble starts the day the reason behind a guardrail goes away and the guardrail does not. Code enforces with the same force on its last day as on its first, and nothing in it says what would end it. So it goes one of two ways: the lock keeps blocking work that is now correct and people route around it, or someone deletes it — and the history shows a changed line, not the decision, who had the authority to make it, or whether it was ever meant to be permanent. In Arroway the ending is part of the rule: every memory states what would retire it, and when a new decision replaces an old one, the old one leaves the read with the reason and the name of whoever decided.
Sort rules by how they die, not by how strict they are
The useful question is not whether a rule matters enough to enforce, but what would make it stop being true. If only a change in how the system is built would end it, it is an invariant, and it belongs in code. If it ends when someone decides otherwise — a client changes terms, a policy is revised, an experiment finishes — it is a decision, and it needs an owner, an ending and a way to be replaced without anyone hunting for copies. Plenty of rules are both. Then the lock does the enforcing, and the memory carries what the lock cannot: why it is there, who put it there, and when it should come out.
What it looks like in practice
A team runs several agents against one staging environment. Two deployments land at the same time and corrupt each other, so the team replaces the written rule — never deploy while another deployment is running — with a lock. It works, and nobody thinks about it again. Four months later staging is split into one environment per branch. Deployments can no longer collide, but the lock still serialises all of them, and agents queue for forty minutes behind work that has nothing to do with theirs. The engineer who added the lock has moved teams. The commit message says “add deploy lock”. Nobody knows whether removing it is safe, so nobody removes it. Had the rule also lived as a decision — enforced by the lock, recorded with “ends when staging stops being shared” — that condition would have been sitting next to the rule in every read, and the day staging was split, retiring both would have been one line for a person to approve.
Questions people ask about this
- So should we move our rules out of code?
- No. Anything you can make mechanically impossible should stay impossible, and memory does not compete with that. What moves is the part code was never good at holding: the reason, the owner and the ending. A guardrail with a recorded decision behind it is stronger than either alone — the lock cannot be argued with, and the record says when it has outlived its purpose.
- Couldn't we just leave a comment next to the guardrail saying when to remove it?
- You could, and it beats nothing. But a comment is an ending nobody is ever asked about: it is read only by whoever happens to open that file, it cannot tell whether it is still accurate, and changing it takes the same access as changing the code. In Arroway the ending travels with the rule into every read of the project, by every assistant in every connected tool, and replacing the decision is a visible act with a name and a date on it — not a diff someone has to go looking for.
- Agents ignore instructions. Why would they respect a rule just because it is in memory?
- Memory is not enforcement, and it should not pretend to be. What it changes is whether the agent had the rule at all. The opening read is made a required first step by the tools rather than left to the agent's discretion, so the rule arrives before the work starts instead of depending on whether some file happened to be loaded. For the rules that must never be broken, that is not enough on its own — which is exactly why those belong in code as well.
Where this is verifiable
Product documentation on this site (How it works, Install), the answers on memory without a write policy, on out-of-date memory and on rules that slip a few messages later, and the expiry-condition, supersession and sanction rules in the sanctioned product spec. The line between invariants and decisions is this page's argument, not a product feature. Everything described here is behaviour the tools apply today, not roadmap.
https://www.arroway.app/en/answers/guardrails-in-code-dont-expire-sanctioned-rules-do