The Simulator Said 301. The Phone Said 155

S Ravi Kumar October 06, 2026 5 min read 7 views (7 unique)

Last month I counted 119 mistakes that got through because somebody checked the wrong thing. This is one more, caught in time. AI agents built the part of my library that runs a language model on the device itself, and they tested it on an iPhone simulator: 301 tokens a second. The real phone gave 155. The pretend Android phone was wrong in the other direction. Four real devices, one table of measured numbers, and what it taught me about the word "tested".

In my last post I counted my mistakes. One number from it has stayed with me. The most common reason a mistake got through, 119 times, was not that nobody checked. Somebody checked the wrong thing.

This is a small story about that same problem a month later. It has two numbers in it, and they are in the title.

What was unfinished

The table in that post had a line that read: TechieRag, published, unfinished.

TechieRag is the library my own applications are built on. Its job is to let an application answer from your own documents, which people call retrieval-augmented generation, without every team rebuilding the same plumbing.

It reads thirteen kinds of document, cuts them into passages, indexes them, and finds the right passages for a question. Then it hands them to a language model and gives you the answer with its sources.

Around that it carries what an AI application needs. One way of talking to all the main model providers, so you change provider in a settings file and not in code. Agents that can call tools and follow a flow of steps. Memory of the conversation. A budget for tokens and cost. Retries, and a second provider to fall back on when the first one fails.

And it has a fully offline path, where the search and the model both run on the device and your data never leaves it.

That last part was what was unfinished: running the model itself on the device, on four kinds of device. A Windows laptop, a Mac, an Android phone and an iPhone.

I did not write that code. AI agents did, from requirements I wrote. My job was the one I described last time: check before believing.

Checking the wrong thing, again

For the iPhone, the agents built the app and ran it on the iPhone simulator, the pretend iPhone that lives on a Mac. It worked. It reported 301 tokens a second, which is a lovely number.

A simulator is a small pedestal. It is not lying, exactly. It is a Mac wearing an iPhone's clothes, and it gives you the Mac's speed.

On Sunday I plugged in the real phone. 155. About half.

The pretend Android phone had been wrong in the other direction. It reported between 22 and 43. The real one gave 112. So one pretend device flattered me and the other insulted me, and from the pretend numbers alone I could not have told you which was which.

What saved me this time was dull. The record for this library names the device every number came from. The iPhone line said "simulator". It did not say "iPhone". A few months ago that line would have said "tested" and I would have believed it.

The numbers

Every row is a real device on my desk.

Rate this article
0.0 · 0 ratings One rating per email — no sign-in needed
SR
S Ravi Kumar Author of this post.

More Posts

Comments (0)

No comments yet

Be the first to share your thoughts on this post.

Leave a comment

No account needed — just your name and email.

Your email is never published — it is used only for confirmation and moderation.
Generated and checked by this site — no third-party service.
Comments appear after email confirmation and moderation.
An unhandled error has occurred. Reload ×