The Simulator Said 301. The Phone Said 155

S Ravi Kumar October 06, 2026 5 min read 7 views (7 unique)

Last month I counted 119 mistakes that got through because somebody checked the wrong thing. This is one more, caught in time. AI agents built the part of my library that runs a language model on the device itself, and they tested it on an iPhone simulator: 301 tokens a second. The real phone gave 155. The pretend Android phone was wrong in the other direction. Four real devices, one table of measured numbers, and what it taught me about the word "tested".

Device Model it ran First word Tokens a second Memory used
iPhone 17 Pro Qwen2.5 0.5B (small) 0.07 to 0.08 s 155 485 MB
Samsung Galaxy S23 Qwen2.5 0.5B (small) 0.06 to 0.10 s 112 to 114 about 750 MB
Mac, M4 Max Phi-3 mini (larger) 0.11 to 0.52 s 65 to 73 about 2 GB
Windows laptop, Core i5, 16 GB Phi-3 mini (larger) 0.26 to 0.32 s 7 to 13 3.4 GB

Two things in that table surprised me. The slowest device is the laptop, not a phone. It runs the larger model, and the size of the model decides the speed far more than the machine does. And memory is the real limit on a phone. 485 MB is fine. The 3.4 GB the larger model takes on the laptop would not be.

What this does not tell you

It tells you how fast and how heavy. It does not tell you how good the answers are. I have not measured that, so I will not claim it.

A model this small will not write your code. It is the right size for small jobs: sorting, labelling, a short answer from a note you hand it. What you get in return is that the note never leaves the device, there is no bill per question, and it works with the network switched off.

The part an agent could not do

One step needed a person. Someone had to plug the phone into the Mac, unlock it and press "Trust". The agents had done the rest: the build, the install, collecting the result.

Last time I said the best bug detector in my process was still me, clicking around. This was a better version of that. Not me finding what was broken afterwards. Me doing the one thing only a person could do, which was small, and known in advance.

What you can take from this

Do not accept "tested". Ask on what, and ask to see the run.

Do not trust a simulator for speed or memory. Mine was wrong in both directions.

Write down what was checked, not only that it was checked. That one habit is the difference between this story and the 119.

Where this leaves me

One line in last month's table can change. Not to "finished". To "measured on four devices", which is a smaller claim and a truer one.

The 301 is still in my records, next to the word "simulator". I am leaving it there.

Rate this article
0.0 · 0 ratings One rating per email — no sign-in needed
SR
S Ravi Kumar Author of this post.

More Posts

Comments (0)

No comments yet

Be the first to share your thoughts on this post.

Leave a comment

No account needed — just your name and email.

Your email is never published — it is used only for confirmation and moderation.
Generated and checked by this site — no third-party service.
Comments appear after email confirmation and moderation.
An unhandled error has occurred. Reload ×