Answers
If every session has to read from a service before acting, what happens when that service has a bad second?
The read is attempted again instead of failing. A database that scales to zero refuses the first call after an idle period while its compute wakes, and it says so: the refusal carries an explicit marker from the provider meaning the call never reached a running compute and can be repeated. That marker used to be ignored here, and the refusal travelled all the way to the caller as a failed read — which, for a session whose first act is to check what the team has already decided, means carrying on without it. Since September 2026 a refusal carrying that marker is retried, up to three attempts, costing on the order of three quarters of a second in the worst case. Two limits keep this from being a licence: only the provider's own declaration counts, never our reading of an error message, and only statements that are unambiguously reads are repeated. Anything that writes is attempted exactly once, as before.
Last updated September 22, 2026
A check that can fail is a check people stop running
Read before acting only survives if reading is dependable. The first time the opening read errors in the middle of real work, the work carries on without it, because the work is what someone is waiting for. The second time, somebody wraps it and moves on. By the third it is decoration: still in the code, still in the routine's description, and no longer load-bearing. The cost of an unreliable check is never the failed call itself — it is that the rule quietly stops being a rule, and nothing announces the day that happened.
The provider says the call can be repeated; we never guess
The decision to retry is not ours to infer. A database that sleeps when idle answers the first call after a quiet period with a refusal that explicitly flags itself as repeatable, because the call never reached a running compute. That flag is what gets read — never the prose of the message, which is written for humans and gets reworded without warning. A refusal that declares nothing stays final: silence from the provider is not permission. That direction matters more than it looks, because retrying something the provider did not mark means guessing that an operation did not take effect, and that is the kind of guess that eventually applies something twice.
Reads repeat; writes are attempted once
The second limit is the statement itself, and the test fails closed: anything not recognised with certainty as a pure read is not repeated. The asymmetry is the whole reason this is safe to have. Mistaking a read for a write costs exactly the old behaviour — the call is not retried, and nobody is worse off than before. Mistaking a write for a read costs a write applied twice. So the test is deliberately narrow and deliberately suspicious: a statement that opens as a read but carries a write anywhere inside it is treated as a write. And the retry has a ceiling — three attempts, short waits, roughly three quarters of a second in the worst case — because a read that hangs is worse than one that fails, since nothing on the calling side can tell it apart from a read that is working.
What it looks like in practice
A scheduled routine opens at three in the morning. It is the first thing to touch the project in several hours, so the database compute is asleep, and waking it takes a fraction of a second. Its first statement is the opening read: what has this team decided, what is in flight, what did last night leave behind. That is the call that meets the compute on its way up, and it is refused — with the provider's marker saying the call never landed and can be repeated. Under the old behaviour that surfaced as an error. The routine did not stop, because a routine that stops on a cold database is a routine that stops most nights. It carried on and did the work: opened files, made changes, wrote its result. What it did not do was read the rules, and nothing in its output said so. The expensive part is not the failed call — it is the hour of work built on top of it. Now the same refusal costs a wait measured in hundreds of milliseconds, the second attempt meets a compute that is up, and the read arrives. The routine never learns that anything happened, which is the correct outcome: a wake-up is infrastructure, not news.
Questions people ask about this
- Doesn't retrying just paper over a real outage?
- It papers over exactly the refusals that declare themselves temporary, and nothing else. A genuine outage does not produce that declaration — it produces a connection that never answers, or an error carrying no such marker — and there the attempts run out and the read fails visibly, reporting how many tries it took. The line between the two is drawn by the provider in its own response, not by us reading tea leaves in an error string, and that is what keeps a retry from turning into a way of not noticing.
- How long can a caller be left waiting?
- Three attempts, with short waits between them, adding up to roughly three quarters of a second in the worst case before the failure surfaces. The ceiling is small on purpose. Waiting longer would trade a visible failure for an invisible one: a caller cannot distinguish a slow read from a working read, and a session that hangs while checking the rules is worse for everyone than a session that is told the check failed.
- We would rather our agent fail hard than act on something incomplete. Does this change what it gets back?
- No. A retry returns the same read the first attempt would have returned; nothing is dropped, summarised or filled in to make the call succeed. And the failure you are guarding against is the one the old behaviour produced: an agent that asked what the team had decided, was told the database was busy, and proceeded anyway. If the read genuinely cannot be delivered it still fails, and it fails loudly enough to be recorded rather than absorbed.
Where this is verifiable
Product documentation on this site (How it works, Install), the answers on reads that stop costing the same when nothing changed and on scheduled AI routines losing the thread, and the retry, read-only classification and write-safety rules in the sanctioned product spec. The behaviour described here reached production on 12 September 2026. Everything described here is behaviour the tools apply today, not roadmap.
https://www.arroway.app/en/answers/reads-that-recover-from-a-retryable-database-error