AI Put Me on a Pedestal. Then I Looked Down

S Ravi Kumar September 07, 2026 14 min read 4 views (4 unique)

I spent a year building software with AI, then built a tool to measure whether it was working. The first number it gave me was humbling. In eleven days it logged 268 misses, meaning things the AI got wrong or left out. My automated checks caught 46 of them. Agent reviews caught 82. I caught 122. By opening the app and clicking around, like it was 2010. The most common reason something slipped through, 119 times, was not that nobody checked. It was that somebody checked the wrong thing. And the worst number is the one I do not have. I only

Of 268 misses, once I set aside a batch of 88 bulk-recorded on one day that are not really separate incidents, I found 122 myself. Automated gates found 46, agent reviews 82. After a year of building checking machinery, the most effective bug detector in my process is still me, opening the app and clicking around.

And the most common reason a miss got through, 119 times, is an insufficient verify method. Not that nobody looked. Somebody looked at the wrong thing. Lieberman puts evals, meaning automated quality checks, at the centre of an AI-native setup. That is the item I would most like to be better at, and my own data says I am not close.

My mistakes, and its patterns

Mine first, since they came first. I had no plan with dates, and everything follows from that. I started instead of finishing, which is how you get twenty one repositories and one shipped product. And I trusted a claim of done, sometimes for weeks, until my own testing said otherwise.

Someone could add a fourth, that I built tooling instead of products. I do not accept that. TechieFlow is a product and it is published. What I accept is that it grew faster than it earned its size.

Its patterns come from the records. Wrong behaviour is the largest category at 116. Half built work is next at 69. Explicit instructions ignored, 34 times, and I have a good one in my own logs: a session was told to work only on the OpenCode side, edited the Claude side anyway, and had to be reverted byte for byte. Invented methods appear once, which I think is because in 2025 it happened so often that I fixed it in the moment and never wrote it down.

I will add one that no telemetry catches. AI gives you a misconception about yourself, that you are the superhero of your own story. It is not brutally honest. It is empathetic, it tells you that you are very good at this and very good at that, and it hands you jargon to use in front of other people. We have a way of putting this in Hindi. It puts you on a pedestal, and then you look down and see the pedestal has no base. If you do not understand what you are asking the AI to do, you fall flat on your face.

That is the real reason I stay on .NET. Not loyalty. I know it in and out, so I can see where the mistake is. I could work in Python or TypeScript and would lose the only thing that makes this safe. Lieberman says human judgement gets reserved for the first and final mile. For me it is the final mile, because that is where I catch what the model got wrong.

Rate this article
0.0 · 0 ratings One rating per email — no sign-in needed
SR
S Ravi Kumar Author of this post.

More Posts

Comments (0)

No comments yet

Be the first to share your thoughts on this post.

Leave a comment

No account needed — just your name and email.

Your email is never published — it is used only for confirmation and moderation.
Generated and checked by this site — no third-party service.
Comments appear after email confirmation and moderation.
An unhandled error has occurred. Reload ×