We have been checking the model leaderboards every day since the first complete OnDevice.AI shortlist went live. Most of those checks are ordinary watchlist work: scan the top rows, look for new open-weight releases, check whether licenses or mobile runtimes have changed, and leave the recommendations alone unless a candidate clears the full phone-ready bar. August 4, 2026 was different. The public leaderboards had moved enough in speech and image generation that the unchanged shortlist deserved a deeper written explanation.

The short version is that the recommendations stay in place. The longer, more useful version is that several public leaderboards have moved in interesting ways, especially for speech and image generation, but most of the movement is above the layer this site is trying to recommend. OnDevice.AI is not asking, "what is the best model anywhere?" It is asking, "what is the best open or open-weight model that is realistic for private, offline use on a modern phone with 8 GB of RAM?"

That difference matters. The current Arena leaderboard is full of excellent frontier systems across text, web development, vision, image generation, and video. That is useful signal, but the top rows are dominated by proprietary APIs, very large models, or hosted systems whose model weights and mobile runtimes are not available to an app developer building an offline phone experience. For OnDevice.AI, those entries are still important context, not direct replacements.

Text, coding, and reasoning are the least changed categories for our purposes. Frontier leaderboard positions have advanced, and there are larger Qwen, Kimi, Claude, Gemini, GLM, and GPT-class systems that score far ahead of small local models. But those are not 4B-class downloads that can be quantized into a comfortable phone footprint. Qwen3.5 4B remains the practical pick because it is current, Apache-2.0 licensed, strong across text, coding, and reasoning for its size, and available in usable GGUF formats. A bigger model might win a benchmark. A bigger model also changes the product: shorter battery life, tighter memory, slower replies, hotter devices, and a worse default experience.

Speech-to-text is more competitive than it was. Artificial Analysis now shows several proprietary or hosted models ahead of Whisper on word error rate, and it lists Voxtral Small as the leading open-weight speech-to-text model in its current comparison. That is exactly the kind of candidate we should keep watching. For the current shortlist, though, Whisper Large V3 Turbo still makes the cleanest on-device recommendation. It has a permissive MIT license, broad multilingual coverage, mature whisper.cpp support, Core ML deployment paths, and a known operating profile on consumer hardware. A narrower or less proven ASR model may be more accurate in one benchmark slice, but the default OnDevice.AI pick needs to be easy to ship privately.

Text-to-speech has also moved. The current Artificial Analysis TTS leaderboard is led by newer systems such as Qwen-Audio TTS and Speechify's Simba family, with other strong API-first voices close behind. Some open weights are now listed above Kokoro as well. That does not make Kokoro obsolete. It makes the tradeoff clearer. Kokoro 82M is tiny, Apache-2.0 licensed, widely mirrored into ONNX and mobile runtimes, and fast enough to feel native on modest hardware. A TTS model can sound better and still be the wrong recommendation if it is API-only, awkward to package, heavy for a phone, or ambiguous for commercial reuse.

Image generation is the category with the widest gap between leaderboard quality and phone-ready practicality. Arena and Artificial Analysis image rankings show major gains from GPT Image 2, Reve, Microsoft, Google, Ideogram, Cosmos, FLUX, and other newer systems. Several open-weight models now rank much higher than older SD 1.5-class pipelines. The constraint is that many of those models are large, non-trivial to run locally on phones, not clearly MIT or Apache-2.0, or better understood as hosted inference products than offline mobile defaults. LCM Dreamshaper v7 stays because it is MIT licensed, about 0.9B parameters, designed for very few inference steps, and realistic for ONNX, Core ML, or dedicated mobile diffusion runtimes.

Vision and photo analysis are similar. General vision leaderboards are now dominated by frontier multimodal systems, but most of those systems are not open mobile downloads. Gemma 4 E4B Instruct remains the best overall and photo-analysis pick because it gives us a compact, Apache-2.0, multimodal model with a plausible phone deployment story. It is not the strongest vision model in the world. It is the strongest model in this shortlist that can cover text, images, and general assistant behavior while still respecting the 8 GB target.

So the August 4 review changes the confidence note, not the shortlist. The frontier has moved, and the candidate watchlist is more interesting than it was in June. Voxtral-style ASR, newer open TTS systems, and modern open image generators all deserve continued testing. The bar for replacing a pick, however, is not just a better leaderboard row. A replacement needs a usable license, original downloads, proven mobile formats, realistic memory use, and enough community runtime support that a developer can build with it without turning the phone into a research workstation. For now, the models on the OnDevice.AI homepage still clear that full bar best.