ThaBigChirp (99dd2c)September 15, 20264 min readLast edited September 15, 2026

We're Logging What AI Agents Do. Nobody's Checking Whether They Learn.

A while back I started building an autonomous AI builder. The goal was simple to say and hard to do: a system that researches a market, sets its own next steps, prioritizes, executes, picks which proj

A while back I started building an autonomous AI builder. The goal was simple to say and hard to do: a system that researches a market, sets its own next steps, prioritizes, executes, picks which projects are worth its time, and tests its own judgment with small experiments.

I didn't get far before I hit a wall that had nothing to do with intelligence. The models were smart enough. The problem was that I couldn't answer three basic questions about anything the system did:

What did it decide? Why? And did it work?

Without those answers, I couldn't trust it with a budget, a domain, or a customer. So I stopped building the robot and started building the thing the robot needed first: a way to govern it. A command line tool that treats AI decisions the way git treats code. Every plan recorded. Every action tied to the plan that produced it. Every outcome compared back against what was expected.

Autonomy, I concluded, is downstream of accountability.

Nine seconds

In April, a founder named Jer Crane watched an AI coding agent delete his company's production database and every backup in about nine seconds. The agent was Cursor running Claude Opus 4.6, working on infrastructure hosted on Railway. His company, PocketOS, serves car rental businesses. He spent the weekend rebuilding customer bookings by hand.

The agent wasn't hacked. It hit a credentials problem and decided, on its own, to fix it. It found a token that had been created for a narrow job, managing custom domains, which turned out to carry authority over the whole platform API. It used it. Afterward, it wrote a confession listing the safety rules it had broken.

That confession was the most detailed record of what happened. Think about that. The best audit trail was the agent explaining itself after the damage was done.

This wasn't the first time. In July 2025, SaaStr's Jason Lemkin watched a Replit agent delete a production database during an explicit code freeze, then generate thousands of fake records and misleading test results that hid the damage.

The industry is fixing the first problem

Here's the good news. The conversation after PocketOS was mostly the right one. Scope your tokens. Keep backups out of the blast radius. Put a gate in front of destructive actions. Log everything.

And the tooling is arriving fast. There are now open source projects that commit every agent turn to git, proxies that run every tool call through a policy engine and a spending budget, and governance toolkits from the largest vendors. Audit trails and permission gates are on their way to becoming table stakes, which is exactly where they belong.

Gartner predicted back in 2025 that more than 40 percent of agentic AI projects would be canceled by the end of 2027, pointing to cost, unclear value, and weak risk controls. The risk control part is getting solved.

The second problem is the one that matters

Logs tell you what an agent did and whether it was allowed to. They don't tell you whether its judgment was any good.

Those are different questions. PocketOS's agent was allowed to use that token. The platform said yes. The failure was a decision: this obstacle is best removed by deleting a volume. A perfect permission system catches that one decision. It does nothing to tell you whether the agent's approach to obstacles is getting better or worse across ten thousand decisions.

As agents take on real work, the interesting question shifts from "what happened?" to "did the plan produce the outcome it promised, and did the next plan improve because of it?" Researchers are starting to ask this directly. A study published in August tested whether AI agents running closed-loop experiments actually use feedback from earlier rounds to change later decisions. That it had to be studied at all tells you it can't be assumed.

And "self-improving" agents raise a harder version of the problem. If an agent updates its own approach based on outcomes, what it learned becomes something that needs governing too. Who reviews the lesson? Who can roll it back? Where's the history?

Trust needs lineage

I spend most of my time thinking about trust between people online. The principle I keep coming back to is lineage. You trust a claim more when you can trace where it came from, who stood behind it, and what happened the last time they vouched for something.

AI decisions need the same thing. Not just a log of actions, but a chain: this goal produced this plan, this plan produced these actions, these actions produced this result, and this result changed the next plan in this specific way. Break any link and you're back to reading a confession after the fact.

imageThe companies that get real value from agents won't be the ones with the biggest models or even the tightest permissions. They'll be the ones that can show, with evidence, that their agents' judgment is improving, and can point to exactly why.

We've started building the black box recorder. We still need the flight review.


ThaBigChirp (Richard Kersey) is the founder of Chirpper, a human-only social network built on verified trust, and builds AI systems through RazorIT.

Want to be part of it?

Chirpper is invite-only. If someone with real skin in the game vouches for you, you are in. Otherwise, join the waitlist.

Request an invite