An unreadable ledger was read as empty, and every comment was allowed again

What happened

A permission ledger failed open. A power cut left it unreadable, and the code treated “unreadable” as “nothing recorded yet.” With nothing recorded, every comment ever written was authorised again.

No attacker was involved and no exploit was needed. The ledger exists to answer one question: has this already happened? When it could not tell, it answered no, so the system went ahead. The safe default for “I cannot tell” had been chosen to be “go ahead.”

That is the whole incident, and it is worth a post because the mechanism is common and easy to miss. Two different states were collapsed into one.

Two states, one code path

A ledger can be in at least two conditions that look alike from the caller’s side:

  • Empty. The file loaded and holds no entries. Nothing has happened yet, so proceeding is correct.
  • Unreadable. The file did not load. The system has no idea what has happened, so proceeding is a guess.

The first is information. The second is the absence of information. If both return an empty collection, the caller cannot tell them apart, and the guess gets made silently on its behalf.

The usual way this happens is a few lines of error handling. As an illustration, not a quote from any real codebase:

try:
    sent = load_ledger()
except Exception:
    sent = {}

That except branch reads as tidy and defensive. It also converts every failure of the storage layer into a statement about the world: nothing has been sent. The power cut changed the file, and the code changed the file’s silence into permission.

Every fallback encodes a bet on which error is cheaper

Whoever writes a fallback is choosing which mistake to risk. Proceeding when the answer was unknown risks acting twice. Stopping when the answer was unknown risks acting late or not at all. Those costs are not equal, and which one is larger depends on what the system does.

For a system that publishes comments, the costs are lopsided. A stopped job leaves a gap that someone can fill the next morning. A re-authorised history means every comment ever written is eligible to go out again, and a published comment cannot be taken back as if it had never appeared.

Fallbacks get chosen by default because nobody reviews them the way they review features. They sit in the error-handling path, which only runs when something is already going wrong, so they are read less and tested less than any other code. The cheapest line to write for a failed read is the one that returns an empty value, so that is the line that tends to ship.

Agents make the same mistake more expensive

A script that loses its ledger repeats its own work. An agent with tool access that loses its record of what it may do can repeat anything it is able to reach.

The paper “Agentao: A Governed Local-First Runtime for Tool-Using LLM Agents” names the risk class: “over-privileged actions, weak auditability, prompt injection, tool poisoning, and uncontrolled side effects.”

Over-privilege is usually described as a grant problem: someone gave the agent too much access. The ledger failure is a second route to the same outcome. Access can be correct on paper and still widen at runtime, because the check that narrows it depends on a record that can disappear. The permission was only as narrow as the file was readable.

Weak auditability compounds this. From outside, the check looked fine, because an empty ledger is a valid state and the system behaved as an empty ledger demands. A log that records actions but not the state that authorised them would show a burst of activity with no visible cause.

The position: unknown has to resolve to stop

Any autonomous system that takes outward-facing actions should treat “I cannot tell” as a refusal. In practice that means four things:

  1. Give “empty” and “unreadable” different code paths. Empty means proceed. Unreadable means halt and report. If one return value covers both, the design is already wrong.
  2. Make the safe branch the shortest one. If halting takes more code than carrying on, carrying on will win the next time someone tidies the function.
  3. Make the halt loud. A halt nobody sees is a silent outage. The failure should reach a person with the cause attached: which store, which read, what error.
  4. Run the fallback on purpose. Corrupt the ledger in a test and confirm nothing goes out. A fail-closed branch that has never executed is a guess about what it will do.

None of this is sophisticated. It is a few lines per decision point. The work is finding every place where the system consults a record before acting and asking what happens when the read fails.

Test it by breaking it

The check is cheap, and it beats reading the code. Take any store your automation consults before it acts: a permission list, a sent-message record, a deduplication table, a rate counter. Then break it:

  • Rename the file.
  • Truncate it to zero bytes.
  • Revoke the credential that reads it.
  • Block the network path to it.

Watch what the system does next. If it keeps going, its permissions depend on its storage staying healthy, and a power cut is enough to widen them.

Pay attention to the second and third cases in particular. A zero-byte file and a revoked credential are the situations most likely to be handled by a broad except that returns an empty value, because they fail in ways the author did not picture when writing the happy path.

Documentation will not answer this for you. It describes the intended behaviour, and the ledger incident happened in the gap between the intended behaviour and the error branch. Read the branch itself.

Questions for your own systems

  • When one of your systems cannot tell whether something already happened, which way does it guess?
  • Which stores does it consult before acting, and what does the code do for each when the read fails?
  • Has anyone ever deliberately broken one of those stores to see what happens?
  • If the answer is “it carries on,” who decided that, and when?

The first question is the one to start with. A system’s guess under uncertainty is a design decision, and it is made whether or not anyone made it deliberately. The only choice is whether it was made on purpose.

If you run agents, schedulers or any automation that acts without a person approving each step, spend an hour breaking its ledgers. What the system does with an unreadable record tells you more about its real permission model than the permission list does.