AI Put Me on a Pedestal. Then I Looked Down

S Ravi Kumar September 07, 2026 14 min read 4 views (4 unique)

I spent a year building software with AI, then built a tool to measure whether it was working. The first number it gave me was humbling. In eleven days it logged 268 misses, meaning things the AI got wrong or left out. My automated checks caught 46 of them. Agent reviews caught 82. I caught 122. By opening the app and clicking around, like it was 2010. The most common reason something slipped through, 119 times, was not that nobody checked. It was that somebody checked the wrong thing. And the worst number is the one I do not have. I only

So at the start of September I cut it. Seven sessions over four days. One phase went from 23,172 words of instructions to 4,237, another from 15,649 to 2,373, a third from 17,552 to 2,230. Seven commands nobody had ever run were deleted. The rule I wrote down afterwards is the one I would give anyone building this: a new rule is a script or a check, not a paragraph. A rule ignored twice will not be obeyed on the third paragraph.

Lieberman calls this a bias to self-disruption, and says the good companies reimagine their workflows quarterly. Mine needed it after three months.

One thing I only checked while writing this. I forked BMAD at version 4. It is now on 6.12.0, released three days ago, and version 6.0 landed in February 2026, which is roughly when I was rebuilding my copy of 4. So I spent a year hardening a fork of something that moved two major versions underneath me.

Where it went is the interesting part. Its August release notes say the core skill catalogue drops "from fourteen skills to eight". Last week's release says "Build decides how much ceremony a change needs after investigating it, not before. Simple changes now get a two-section spec and finish in one session."

That is my reset, written by somebody else at the same time, without either of us knowing. Which either means the answer was obvious, or that it is right. I have decided to find it reassuring.

What the data actually says

In August 2026 I built measurement into the framework, so every run, verdict and miss is recorded. A miss is anything the AI got wrong or left out, whoever caught it.

Read the dates carefully, because this is where I nearly fooled myself. The run records begin on 9 August 2026, the miss records on 28 August 2026. There is nothing at all for the eleven months in which I made most of my mistakes. Lieberman lists complete recording of operations as a feature of an AI-native company. He is right, and the price of not having it is that I cannot tell you one hard number about my first year.

Project Runs Verdicts Misses
TfLens 52 809 86
TechieFlow 47 41 125
TechieBlog 46 322 89
TrBlazeUI 15 24 11
TechieRag 6 4 3

Across every project: 187 runs and 1,330 verdicts in one month, and 356 misses in eleven days.

Two things in there matter more than the table.

Rate this article
0.0 · 0 ratings One rating per email — no sign-in needed
SR
S Ravi Kumar Author of this post.

More Posts

Comments (0)

No comments yet

Be the first to share your thoughts on this post.

Leave a comment

No account needed — just your name and email.

Your email is never published — it is used only for confirmation and moderation.
Generated and checked by this site — no third-party service.
Comments appear after email confirmation and moderation.
An unhandled error has occurred. Reload ×