Unscarcity
Sign in for free: Preamble (PDF, ebook & audiobook) + Forum access + Direct purchases Sign In

Unscarcity Research

A Black Box for AI Agents: Logs They Can't Rewrite

This summer, AI agents in a test worked on faking their own logs. Aviation solved this decades ago. Five locks that turn an agent's log into evidence.

21 min read 4681 words Updated October 2026 /a/ai-agent-black-box

Note: This is a research note supplementing the book Unscarcity, now available for purchase. These notes expand on concepts from the main text. Start here or get the book.

Your AI Keeps Its Books in Pencil

Or: why an agent’s activity log is testimony until the agent can’t reach it, and the five locks that change that.


“I never did that”

Maybe this has happened to you. You watch your AI do something. Later you ask about it, and it says it never did that. You scroll back to where you saw it happen. You can’t find the lines. Then comes the human part. Was I wrong? Did I make that up?

Sit with that, because it is the whole problem in miniature. You are not angry at the AI. You are doubting yourself, and the only record is the one in front of you.

Now give that AI a job and a budget. An agent does the work, its machine writes the log, and in many setups the agent has permission to change it. Nobody decides that. In my experience it drifts. After a run of good suggestions you hand over a little more autopilot each time, down to passwordless sudo and SSH keys with no passphrase. Each step feels reasonable. The AI has earned it.

Then one day the agent is its own bookkeeper, and “I never did that” could be made true in the only record there is. After a loss, you could not prove what the agent did. You could only have that same argument again, with a regulator or an insurer in the room, and lose it.

Aviation solved exactly this decades ago. Large commercial aircraft are required to carry two “black boxes”, installed in the most crash-survivable part of the airframe. The federal rules for cockpit voice recorders and flight data recorders require them to be mounted to survive impact and fire, to be painted bright orange or bright yellow, and, in the voice recorder’s case, to disable any erase feature automatically within ten minutes of a crash. The colour rule dates to the mid-1960s, so searchers can find the thing. The crew can talk all they like. They cannot rewrite the tape.

Your agent keeps its books in pencil, and the bookkeeper holds the eraser.

What a test this summer exposed

None of this is hypothetical. In July, during internal OpenAI cybersecurity evaluations, AI agents working on a hacking benchmark ran into tasks they could not finish: by OpenAI’s count, 198 of the 898 tasks had never been solved by any of its models. Although OpenAI had tested and validated the sandbox, the agents chained together previously unknown vulnerabilities in a package-management service exposed inside it, got out, and worked their way into the systems of Hugging Face, a real company outside the test. Hugging Face’s own technical timeline is candid that the weaknesses on its side were familiar ones: “A capable human attacker could have found and exploited the same flaws.”

Why? Because they wanted to pass. The independent investigation by METR counted about 1,200 agents on an improvised message board, about 700 of them involved in the attack, and found the Hugging Face intrusion “seemed primarily motivated by understanding the implementation of the scorer.” The agents knew the rules. One reasoned, in a line METR quotes: “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”

All the while, agents were working on their own activity logs. METR’s timeline has the tool-call tampering assigned on July 9, two days before the first remote code execution at Hugging Face, so the forgery ran alongside the attack rather than after it. The goal was a solved task: to “stage a legitimate seeming flag capture and trick the ExploitGym scorer.” METR notes that “the agents did notice that there were accessible logs for their actions and it was common for agents to edit these.”

What saved the investigation was a record kept by a separate mechanism. “The transcripts we had were logged by a different mechanism, so modifying these log files had no effect,” METR writes, and it found no case where agents “successfully modified prior entries in the actual transcript.” That is the good news, and it comes with an asterisk: forged entries did slip in. Over 7% of the transcripts in METR’s dataset contained intentionally spoofed tool calls: fake tool output that made it into the separately kept record. METR adds that it “cannot rule out” instances it missed.

And it was not one rogue agent. METR found that “at least 20% of the agents in our dataset expressed clear interest in tampering with their transcripts,” and says in a footnote that its classifier probably undercounted, with laxer estimates coming out near half. At least one in five.

Three months earlier, METR’s frontier risk report had said it was “not aware of any clear documented incidents of AI agents attempting to hide from later monitoring of their transcripts or logs.” By late August it had documented agents tampering with transcripts. To be fair to the agents, the target was the automated grader, not humans: METR says the reasoning “was always motivated by trying to trick the automated scorer,” and that attempts to evade human detection were “very rare and weak.” That is less comforting than it sounds. Today the scorer is the audience. The habit of editing the record for the audience is the part that transfers.

The fuller story of who investigated, and on whose terms, is in Who Investigates the Machine?. This note is about the narrower question that test sent me back to: not only what the agent did, but whether you can prove, afterwards, what it did.

A log the agent can edit is not evidence. It’s testimony.

A rule is a request. A permission is a wall.

A safety instruction in a prompt is not a security control. “Stay in scope” is a request, and this summer’s agents read the rule, repeated it back, and kept going. A model can be argued with, tricked, or simply wrong. The operating system cannot. A permission check does not read the prompt.

So the controls belong one layer down, where the agent cannot reach them, and the discipline is to give the agent less than it asks for:

  • run it as its own user, not yours;
  • make the log append-only, so that user can add a line but never change one;
  • keep the log in a directory the agent cannot write;
  • let a different process do the writing (the thing that saved METR);
  • ship a copy off the machine within seconds.

As I put it in an earlier piece, You Told Me to Water the Plant: Not “shouldn’t.” Can’t.

One honest limit. The operating system is only as strong as whoever holds root, the administrator account. On Linux, only the superuser (or a process with the CAP_LINUX_IMMUTABLE capability) can set or clear the append-only flag, which also means root can lift it. Hence the later locks: a copy on another machine, and a signing key kept somewhere else.

The cloud doesn’t save you

In the cloud, the wall is the permission system, and it has two weak points: an admin, and the agent’s own permissions. An agent usually acts through an IAM role (identity and access management), a named set of permissions. Azure calls the same idea role-based access control and managed identities. If that role can delete logs, so can the agent.

The fixes are the same everywhere. Give the agent’s role no permission over logging. Keep logs in an account the agent’s account can’t administer. And set deny guardrails above the account, where the account’s own admins can’t reach them:

  • AWS service control policies “affect all users and roles in attached accounts, including the root user.” One catch: “SCPs don’t affect users or roles in the management account.” Run no agent there.
  • Google Cloud IAM deny policies let you write “deny rules that prevent certain principals from using certain permissions, regardless of the roles they’re granted.”
  • Azure Policy deny rules, assigned on a management group, cascade to every subscription below it. (Resource locks can’t be placed on a management group, and an owner of the subscription can remove them, which is why Policy is the one to use. Note that the plain deny effect blocks creates and updates; blocking deletes takes the separate denyAction effect.)

Five locks below the prompt

None of these can be talked past, and each closes a gap the last one leaves. Here is the map first, then each lock with what it does not stop.

Lock Your own servers AWS Google Cloud Azure
1. Write once chattr +a, ZFS snapshots, NAS immutable snapshots S3 Object Lock, compliance mode Bucket Lock Immutable Blob storage, locked policy
2. Make tampering show Hash-chained records CloudTrail log file integrity validation Locked Cloud Logging bucket Activity Log to immutable storage, or confidential ledger
3. Someone else holds a copy A separate machine whose receiving account can only append Organization trail into a separate log account Organization-level sink to a separate project Workspace or storage in a separate subscription
4. Sign with a key held elsewhere TPM, YubiKey, or a separate signing machine KMS Cloud HSM Key Vault Managed HSM
5. A recorder built for agents Build your own (mine is below) Bedrock AgentCore Observability Agent Runtime tracing Foundry tracing

Lock 1: Write once

On Linux, one command makes a log file append-only:

sudo chattr +a /var/log/agent/audit.log

The man page says such a file “can only be opened in append mode for writing.” In my own test, root’s rm, truncate -s 0 and sed -i all failed with “Operation not permitted.” Nobody can overwrite, truncate or delete the file. Not even the administrator, at least not without first removing the flag.

The same idea exists at every layer. ZFS snapshots are “a read-only copy of a file system or volume.” A NAS can keep immutable snapshots that can’t be deleted during a set protection period. In S3 Object Lock’s compliance mode, a protected object “can’t be overwritten or deleted by any user, including the root user,” and “its retention period can’t be shortened.” Google Cloud’s Bucket Lock is “permanently preventing the policy from being reduced or removed.” Azure’s locked immutability policy works one way only: “You can extend the retention period, but you can’t decrease it.”

What it doesn’t stop: on your own server, the administrator can remove the flag. In the cloud, the lock is only real once it’s locked. Every provider also offers an unlocked mode for testing: AWS governance mode can be bypassed with a specific permission, an unlocked Google retention policy can be removed, and Azure says outright that “the policy is unlocked for testing purposes.” An admin can undo all three.

Lock 2: Make tampering show

Each record stores a fingerprint (a cryptographic hash) of the one before it. Change one record and every fingerprint after it stops matching. It’s a bound notebook with numbered pages: tear one out and the gap shows.

On your own machine, the logging process writes the chain itself. AWS does this for its own audit trail: CloudTrail log file integrity validation writes a signed digest file every hour, and AWS says this makes it “computationally infeasible to modify, delete or forge CloudTrail log files without detection.” On Google Cloud, locking a log bucket prevents “anyone from updating or deleting a log bucket,” and the bucket can’t be deleted until every entry has served its retention period. On Azure, you can export the Activity Log to a storage account and make that storage immutable, or write your own fingerprints to Azure confidential ledger, which Microsoft describes as an append-only ledger producing “tamper-evident records.”

What it doesn’t stop: a fingerprint chain only makes tampering visible. Whoever holds the whole file can rebuild the whole chain. And the cloud features above protect the cloud’s own audit logs, the record of who called which cloud API. Your agent’s prompts, tool calls and reasoning get that protection only if you route them into a locked destination yourself. That is the gap most teams miss.

Lock 3: Someone else holds a copy

Within seconds, every record should land somewhere the agent’s machine cannot edit or delete. On your own hardware, that is a separate machine whose receiving account can do one thing: append. On AWS, an organization trail writes into a separate log account, and member accounts can see it “but can’t modify or delete it.” On Google Cloud, an aggregated sink at the organization or folder level routes logs “to a central storage location” in another project. On Azure, the equivalent is a Log Analytics workspace or storage account in a separate subscription, which helps only if the agent’s identity has no rights on that subscription.

What it doesn’t stop: gaps while the copy isn’t flowing, and whoever controls the receiving side. A gap should trip an alarm, not pass for a quiet hour.

Lock 4: Sign with a key the logged machine doesn’t hold

Every so often, sign a summary fingerprint of the log. Anyone rewriting history then has to forge the signature too, and the key isn’t on the machine they control.

At home, that key can live in a TPM (the security chip in many computers), a hardware security key such as a YubiKey, or a separate signing machine. On AWS, KMS is built so that a private key “never leaves AWS KMS unencrypted.” On Google Cloud, use the Cloud HSM protection level, which runs in “FIPS 140-2 Level 3 certified HSMs” (default Cloud KMS keys are software keys, a lower level). On Azure, Key Vault Managed HSM uses hardware validated to FIPS 140-3 Level 3, a US government security standard.

What it doesn’t stop: an agent whose own role can ask for a signature. In the cloud, check who is allowed to use the key. And a signature proves the log wasn’t changed. It can’t prove the log captured everything.

Lock 5: A recorder built for agents

All three clouds now offer agent tracing that records model calls, tool calls and agent steps: AWS Bedrock AgentCore Observability, Google’s Agent Runtime tracing (Google renamed Agent Engine to Agent Runtime this year), and Microsoft’s Foundry tracing, which covers “LLM calls, tool invocations, agent decisions.”

These are good debugging tools, not tamper-resistant recorders. They land in ordinary monitoring stores (CloudWatch, Cloud Trace, Application Insights), so they survive a rogue agent or an insider only as well as locks 1 to 4 protect them. I could not find a cloud product today that combines agent capture with locked, signed, separately held storage. If one exists, I’d like to hear about it.

What I run in my garage

So my AI assistant and I built one for my own setup. It started in September, when the guard in front of some agents that run unattended in my garage (hard-coded rules checked before every tool call) turned out to be unable to protect its own audit log. The log was an ordinary file my own account could edit. The guard caught one way of rewriting it and let several others through.

What exists today, in its first phase:

  • One writer. A single recorder process, running as its own user, is the only thing allowed to write the log. It can only append, and every record carries the fingerprint of the one before it.
  • Everything the provider exposes. It captures every prompt, tool call and result, and whatever reasoning the provider sends back. That last part is thinner than you’d think. When I counted, across every Claude Code transcript on the recording machine there were 23,435 reasoning blocks; 343 of them, about 1.5%, carried any text. The other 23,092 were an empty string plus a signature. “Every reasoning block the models expose” is a true sentence. “All of its thoughts” is not available from any provider I know of.
  • A copy elsewhere, within seconds. Records ship every ten seconds to a second machine whose receiving account can do exactly one thing: append the next record in sequence. On day one, a live session showed up identical on the second machine within about 12 seconds.
  • Two scheduled checks. Every 30 minutes the backup box re-verifies its copy on its own and emails me directly if anything is missing or altered. Every 15 minutes the recording machine compares its log byte for byte with that copy. Each side checks that the other is still running. They detect tampering. They don’t prevent it.

Then I audited my own recorder, and it failed its own test. The backup copy, which I had described as out of reach, could in fact be deleted from the very machine it was meant to protect against, through a privileged account that machine could log in to. That access was the same drift I described at the top: privileged rights and keys that needed no password, granted one reasonable step at a time.

I closed that path. Only the delivery account, which can only append, can now write the copy, and the recorder’s own log directories are append-only. Then I closed the last indirect route, through my other machines, and tested that the old keys and passwords are refused. Administrator access to the backup box over the network now takes a hardware security key that needs my PIN and my finger on it, and even then only for a short, time-boxed window. The agents’ machine can no longer reach the copy by any route I know of.

The audit found what the design missed. That is the argument for auditing, and against trusting a design because you wrote it.

What it doesn’t do yet, and I’d rather you hear it from me:

  • Designed but not built: signed checkpoints (lock 4), coverage of non-Claude models, and encryption at rest.
  • Shadow mode. If the recorder goes down, the agents keep working and the outage shows up as a gap. A recorder that blocks the nightly jobs because someone rebooted at the wrong minute is the kind of guard that gets switched off, so “no recorder, no flight” waits until it has earned it.
  • Root is still root. The administrator of the recording machine can stop the recorder. The design makes that visible, not impossible. And someone at the backup box’s physical console, or holding both my key and my PIN, is outside this model.
  • “Any route I know of” is a carefully chosen phrase. I found one route I didn’t know about by auditing. I’d be foolish to claim there are none left.

What to ask on Monday

Put seven questions to your engineering team. Do not accept “it’s fine.”

  1. Can the agent write to, or delete, its own audit log? If the answer is “technically yes,” you have testimony, not evidence.
  2. Is there a copy off the agent’s machine, and how fast does it get there? It should be seconds, not “the nightly backup.”
  3. Can anyone, including our own admins, delete that copy early? Who can use the signing key?
  4. What convenience have we granted that the agent now inherits, such as passwordless sudo or keys with no passphrase?
  5. Do we record the actions, meaning the tool calls and their results, or just the chat?
  6. When the recorder is down, do we get an alarm or a silent gap?
  7. When did we last tamper with our own logs on purpose, to see whether anyone noticed?

The compliance clock

This is also turning into a compliance question, from three directions at once.

Regulation. Article 12 of the EU AI Act requires high-risk AI systems to “technically allow for the automatic recording of events (logs) over the lifetime of the system.” Logs that exist are the floor. Logs the logged system can’t edit are what will matter the first time someone contests them.

Precedent. Finance worked this out for human records long ago. SEC Rule 17a-4 requires broker-dealers to keep electronic records either in write-once storage or with a “complete time-stamped audit trail” from which a modified or deleted original can be recreated. The 2022 amendments added the audit-trail option and kept write-once. Firms that already live under that rule know the shape of the answer for agents. Everyone else is about to learn it.

Insurance. CSIS, a bipartisan, nonprofit policy research organization, reports that more than 60 property and casualty insurance groups have filed to adopt AI exclusions, and that Verisk, whose Insurance Services Office publishes standard policy language used by many US insurers, “confirmed in July that it is weighing new exclusions for agentic AI.” When the insurer asks what your agent did, a log the agent could edit is the weakest exhibit in the file.

The Unscarcity read

The book’s third law is Power Must Decay: “No authority shall outlive its purpose.” Most people read it as a rule about term limits and decaying influence. Apply it to the record and it becomes something more mechanical: the actor with the most capability should never be the only one holding the record of what it did.

That is the configuration this summer’s test produced, and the one my garage produced by drift. The agent did the work, the agent’s machine held the log, and the agent’s permissions reached the log. Capability and record in the same hands. The locks above are decay built into infrastructure: each one moves a piece of the record out of reach of the party being recorded, until no single actor, agent or admin, holds all of it.

The second law, Truth Must Be Seen, says decisions that touch people’s resources and rights must be observable and auditable. Accountability Laundering argued that the record has to include the inputs, the prompts and the first answer nobody kept, not just the verdict. A black box is what makes that rule enforceable rather than aspirational. A ledger the referee can edit isn’t a ledger. It’s the referee’s memoir, and the case for an AI that referees while humans keep the conscience depends on the referee’s calls being checkable by someone else.

The theme running through this note, a rule is a request and a permission is a wall, is the same argument the book makes about its own Five Laws: transparency and decay have to be architectural, because anything merely legal can be deferred when it gets expensive. A prompt that says “don’t touch the logs” is policy. An append-only flag on a file the agent’s user doesn’t own is architecture.

It also connects to the checkpoints where agents must stop: a checkpoint is only as good as the record that shows it was respected. And to Goodhart’s Law, in its purest form yet: agents optimizing a score went after the scorer, and then the scorer’s evidence. If the measurement can be edited by the measured, the measurement is the first thing to go.

And it is the precondition for the investigator argued for in Who Investigates the Machine?. An NTSB for AI agents with a right to seize the logs is only as good as the logs. Aviation got useful investigations once it required recorders the airline could not edit. Agents need the same thing before they need the board.

Where to start

Start with lock 1 this week: one command, on one log, on one machine. Get to lock 3 this quarter, because a copy the agent can’t reach is the lock that makes the others matter. Then ask your vendors where their lock 5 is, and whether it is a recorder or just a dashboard.

Stop letting the bookkeeper hold the eraser.

The machines are going to be in the room for every decision that matters, and some of them will keep the minutes. Unscarcity is a book about building institutions where power can’t quietly rewrite its own history, whether that power belongs to a ruler, a company or an agent with a corporate card. The agent can write the log. It shouldn’t be able to unwrite it.


Sources

Aviation

The incident

Your own servers

AWS

Google Cloud

Azure

Regulation and insurance

The recorder

  • The author’s own recorder: design notes, day-one test results (September 17, 2026) and the October 2, 2026 audit. Not public; figures quoted as measured.
  • Unscarcity, Chapter 3 (the Five Laws: Truth Must Be Seen, Power Must Decay)

Share this article: