AI Put Me on a Pedestal. Then I Looked Down

S Ravi Kumar September 07, 2026 14 min read 4 views (4 unique)

I spent a year building software with AI, then built a tool to measure whether it was working. The first number it gave me was humbling. In eleven days it logged 268 misses, meaning things the AI got wrong or left out. My automated checks caught 46 of them. Agent reviews caught 82. I caught 122. By opening the app and clicking around, like it was 2010. The most common reason something slipped through, 119 times, was not that nobody checked. It was that somebody checked the wrong thing. And the worst number is the one I do not have. I only

Project What it is Where it stands
TechieBlog My blog engine and site (where you are reading this blog) Shipped, August 2026
AppManager Identity, licences and subscriptions for my own apps In production, for me
TrBlazeUI My Blazor component library Published on NuGet
TechieRag My own retrieval library, which finds my documents and hands them to a model Published on NuGet, unfinished
TechieFlow My development framework Published on npm, in daily use
AI-First Playbook Team edition of that framework On npm, needs its own rebuild
TfLens Reads my development data and shows what AI cost and saved Ready for testing, data still wrong
Xpenser Expense tracker, built live on stream Parked, coming back
TrStudio Open source tool for talking head video Half built, parked
TrSetup A setup helper nobody asked for Being archived
Around fifteen others An IDE, a database tool, utilities Parked or half built

What kept going wrong

"You can't transform what you haven't mapped," Lieberman writes. That is the line that stung, because for most of this year I had no map. No plan with dates. No list of what I would finish and by when. I had a subscription, free evenings, and an enormous appetite for starting things.

I did have a process eventually. From mid September 2025 I used BMAD, then at version 4, where you talk to an analyst persona and it writes your requirements documents for you. Spec-driven development was everywhere that autumn, and BMAD's documents impressed me, which is why I stayed with it for months.

I also barely used it. BMAD is built on scrum, so it comes with sprints and stories and a lot of ceremony. What I did was let the analyst write the requirements and architecture, then tell Claude Code to turn that into a task list and build the lot, and test at the end. That is not spec-driven development. That is a wish list with extra steps.

Three failures came back on project after project.

The first is that it tells you it is done when it is not done. Every checklist line complete, then I open the app and the UI is broken and links go nowhere. This is not a 2025 problem I can look back on fondly. Last month I took TfLens into my own testing believing it was ready and found requirement after requirement only half built. My miss log has a run of them recorded on one day, 29 August 2026, against eight different screens. Most give the same reason: the check was too weak.

Rate this article
0.0 · 0 ratings One rating per email — no sign-in needed
SR
S Ravi Kumar Author of this post.

More Posts

Comments (0)

No comments yet

Be the first to share your thoughts on this post.

Leave a comment

No account needed — just your name and email.

Your email is never published — it is used only for confirmation and moderation.
Generated and checked by this site — no third-party service.
Comments appear after email confirmation and moderation.
An unhandled error has occurred. Reload ×