AI Put Me on a Pedestal. Then I Looked Down
I spent a year building software with AI, then built a tool to measure whether it was working. The first number it gave me was humbling. In eleven days it logged 268 misses, meaning things the AI got wrong or left out. My automated checks caught 46 of them. Agent reviews caught 82. I caught 122. By opening the app and clicking around, like it was 2010. The most common reason something slipped through, 119 times, was not that nobody checked. It was that somebody checked the wrong thing. And the worst number is the one I do not have. I only
| Project | What it is | Where it stands |
|---|---|---|
| TechieBlog | My blog engine and site (where you are reading this blog) | Shipped, August 2026 |
| AppManager | Identity, licences and subscriptions for my own apps | In production, for me |
| TrBlazeUI | My Blazor component library | Published on NuGet |
| TechieRag | My own retrieval library, which finds my documents and hands them to a model | Published on NuGet, unfinished |
| TechieFlow | My development framework | Published on npm, in daily use |
| AI-First Playbook | Team edition of that framework | On npm, needs its own rebuild |
| TfLens | Reads my development data and shows what AI cost and saved | Ready for testing, data still wrong |
| Xpenser | Expense tracker, built live on stream | Parked, coming back |
| TrStudio | Open source tool for talking head video | Half built, parked |
| TrSetup | A setup helper nobody asked for | Being archived |
| Around fifteen others | An IDE, a database tool, utilities | Parked or half built |
What kept going wrong
"You can't transform what you haven't mapped," Lieberman writes. That is the line that stung, because for most of this year I had no map. No plan with dates. No list of what I would finish and by when. I had a subscription, free evenings, and an enormous appetite for starting things.
I did have a process eventually. From mid September 2025 I used BMAD, then at version 4, where you talk to an analyst persona and it writes your requirements documents for you. Spec-driven development was everywhere that autumn, and BMAD's documents impressed me, which is why I stayed with it for months.
I also barely used it. BMAD is built on scrum, so it comes with sprints and stories and a lot of ceremony. What I did was let the analyst write the requirements and architecture, then tell Claude Code to turn that into a task list and build the lot, and test at the end. That is not spec-driven development. That is a wish list with extra steps.
Three failures came back on project after project.
The first is that it tells you it is done when it is not done. Every checklist line complete, then I open the app and the UI is broken and links go nowhere. This is not a 2025 problem I can look back on fondly. Last month I took TfLens into my own testing believing it was ready and found requirement after requirement only half built. My miss log has a run of them recorded on one day, 29 August 2026, against eight different screens. Most give the same reason: the check was too weak.
More Posts
Comments (0)
No comments yet
Be the first to share your thoughts on this post.
Leave a comment
No account needed — just your name and email.
