You don’t always need a frontier model. A look at when a small, local model beats a large API call on latency, cost and privacy.

It’s easy to assume every AI feature needs the biggest, newest model available through an API. For a growing share of real products, a much smaller model running directly on the device turns out to be the better call — not the compromise.

The default assumption worth questioning

Reaching for the largest available model by default is understandable — it’s usually the most capable option on paper. But every API call to a large model costs money, takes a network round-trip, and depends on your user having a working internet connection at that exact moment.

For a lot of tasks, that’s a lot of overhead for a job a much smaller model can do just fine.

Where small models win

Small, efficient models — some just a few gigabytes, able to run on a phone or laptop — genuinely shine at narrow, well-defined tasks: cleaning up text, classifying a support ticket, transcribing speech, summarizing a short document.

Running locally also means no data ever leaves the device, which matters enormously for anything sensitive, plus answers arrive instantly with zero network latency.

What you give up

Small models are noticeably weaker at broad, open-ended reasoning, unfamiliar or unusual topics, and tasks that need a huge amount of background knowledge to get right.

Asking a small on-device model to do the job of a large frontier model is where teams get burned — the right move is matching model size to task difficulty, not picking one model size for everything.

A simple way to decide

Ask two questions: does this task need broad, general knowledge, or is it narrow and well-defined? And does privacy, latency, or working offline actually matter for this feature?

Two “narrow and yes” answers point toward a small local model; a “broad” task usually still needs a bigger model behind an API.