On-device AI vs cloud AI: what actually changes for you
The difference between on-device AI and cloud AI is one fact: where the model runs. Everything else, privacy, price, speed, what the model can do, follows from it. This is the comparison from the user's side of the screen, with the actual numbers where we have them.
The two designs
With cloud AI, your message travels to a data center, a very large model answers there, and the answer travels back. With on-device AI, your message stays on the phone and a smaller model answers there. Orvena is the second kind: it runs Qwen 3.5 4B, a one-time 3.1 GB download, on the iPhone's GPU, and on iOS 27 it can also answer with Apple Intelligence, Apple's on-device model, with nothing to download.
On-device AI vs cloud AI, side by side
| On-device (Orvena) | Cloud (ChatGPT and similar) | |
|---|---|---|
| Where your words go | Nowhere. The model is on the phone. | To the provider's servers. |
| Account | None. | Usually, for history, billing and limits; ChatGPT has a limited logged-out mode. |
| Works offline | Yes, including the phone tools. | No. |
| Latency | No network round trip. The answer streams as it is written. | Network plus queue; varies with load. |
| Model size | About 4 billion parameters for Orvena's own model. | Hundreds of billions. |
| Conversation memory | 16,384 tokens, about 12,000 words, with Orvena's own model. On iOS 27 Apple Intelligence reports 8,192, although Apple's documentation still says 4,096. | Much larger. |
| Price | The assistant is free, including Apple Intelligence on iOS 27. One $14.99 purchase for voice, connections, scheduled tasks, Private Cloud Compute on iOS 27 and cloud models with your key. | $0 to $200 a month for ChatGPT's tiers; $20 for Plus. |
| Phone needed | iPhone 15 Pro or newer, 8 GB of memory. | Any phone. |
| When the policy changes | Nothing about your past conversations changes. | Whatever the new policy says about stored conversations. |
What on-device does well
The requests people actually make to a phone. "Reschedule my 3pm to Thursday morning." "Set an alarm so I'm at the airport by 2pm tomorrow, counting travel time." "Tonight at 9, check tomorrow's weather against my first meeting." "What do you see in this photo?" A phone-sized model, Orvena's own or Apple Intelligence, handles the language; the phone's own tools do the calendar and the alarm, and Apple Maps supplies the travel time; the calendar entry and the photo are understood on the device and are not sent anywhere to be read.
Where the cloud is better
Scale. A data-center model knows more, reasons longer, writes better, and can hold a book-length document in its context. For long research, a specialized field, or writing that has to be excellent rather than good, you will feel the difference, and no amount of engineering on the phone closes that gap today. Cloud services also add features that need infrastructure, such as agents that browse for minutes at a time.
A hybrid, done carefully
You do not have to choose once. Orvena's Models page lists every engine it can use: under Built in, Apple Intelligence on iOS 27; under On-device models, Orvena's own downloaded model; and under Cloud, Apple's Private Cloud Compute on iOS 27, with no key to set up and a daily limit set by Apple, or any OpenRouter model on your own key. The default engine setting, Automatic (in Settings, Advanced), uses your downloaded model, or Apple Intelligence while none is ready, and never switches to the cloud on its own. With OpenRouter, Orvena asks before the first cloud request in a conversation whether to send that conversation and its tool results to the provider. Private Cloud Compute asks once, when you choose it. With either, health or financial data asks again every time. The voice screen shows which mode you are in: "on this iPhone" in green, "cloud" in amber.
How to decide
Write down the last ten things you asked an assistant. If most were about your own life, your schedule, your photos, a message you were writing, somewhere to go nearby, on-device is the right default: the data is personal and the tasks are well within a phone-sized model's reach. If most were research or long writing, you want cloud access some of the time, ideally as a choice you make per request rather than the only mode there is.
Common questions
Is on-device AI slower than cloud AI?
Per token, a phone is slower than a data-center GPU, but there is no network wait and the answer streams as it is written, so short answers feel immediate. What takes longer on a phone is reading a long conversation before the answer starts.
Is Private Cloud Compute as private as on-device AI?
Not in the same way. It runs on Apple's servers, so the request leaves the phone and needs a connection. Apple says "The data being processed in Private Cloud Compute is not stored or made accessible to Apple" and that "independent privacy and security researchers can verify this privacy promise at any time." On-device, there is no server in the conversation to make that promise about.
Does running AI on the phone drain the battery?
A short question costs very little. A long voice conversation warms the phone the way a game does and uses battery accordingly. Nothing runs in the background except a reply you already started.
Can on-device AI still search the web?
Yes. When you ask for current information the query goes to the search provider, as it would from a browser, and the reading and summarizing happen on the phone.
Orvena runs its model on the iPhone and lets you choose Apple's Private Cloud Compute, or a cloud model with your own key, when a request calls for it.
Download on the App Store