The Simulator Said 301. The Phone Said 155
Last month I counted 119 mistakes that got through because somebody checked the wrong thing. This is one more, caught in time. AI agents built the part of my library that runs a language model on the device itself, and they tested it on an iPhone simulator: 301 tokens a second. The real phone gave 155. The pretend Android phone was wrong in the other direction. Four real devices, one table of measured numbers, and what it taught me about the word "tested".
| Device | Model it ran | First word | Tokens a second | Memory used |
|---|---|---|---|---|
| iPhone 17 Pro | Qwen2.5 0.5B (small) | 0.07 to 0.08 s | 155 | 485 MB |
| Samsung Galaxy S23 | Qwen2.5 0.5B (small) | 0.06 to 0.10 s | 112 to 114 | about 750 MB |
| Mac, M4 Max | Phi-3 mini (larger) | 0.11 to 0.52 s | 65 to 73 | about 2 GB |
| Windows laptop, Core i5, 16 GB | Phi-3 mini (larger) | 0.26 to 0.32 s | 7 to 13 | 3.4 GB |
Two things in that table surprised me. The slowest device is the laptop, not a phone. It runs the larger model, and the size of the model decides the speed far more than the machine does. And memory is the real limit on a phone. 485 MB is fine. The 3.4 GB the larger model takes on the laptop would not be.
What this does not tell you
It tells you how fast and how heavy. It does not tell you how good the answers are. I have not measured that, so I will not claim it.
A model this small will not write your code. It is the right size for small jobs: sorting, labelling, a short answer from a note you hand it. What you get in return is that the note never leaves the device, there is no bill per question, and it works with the network switched off.
The part an agent could not do
One step needed a person. Someone had to plug the phone into the Mac, unlock it and press "Trust". The agents had done the rest: the build, the install, collecting the result.
Last time I said the best bug detector in my process was still me, clicking around. This was a better version of that. Not me finding what was broken afterwards. Me doing the one thing only a person could do, which was small, and known in advance.
What you can take from this
Do not accept "tested". Ask on what, and ask to see the run.
Do not trust a simulator for speed or memory. Mine was wrong in both directions.
Write down what was checked, not only that it was checked. That one habit is the difference between this story and the 119.
Where this leaves me
One line in last month's table can change. Not to "finished". To "measured on four devices", which is a smaller claim and a truer one.
The 301 is still in my records, next to the word "simulator". I am leaving it there.
More Posts
Comments (0)
No comments yet
Be the first to share your thoughts on this post.
Leave a comment
No account needed — just your name and email.
