How to run an AI assistant offline on your iPhone: requirements and setup
Running an AI assistant on the phone itself, with no server behind it, is possible on recent iPhones. On iOS 27 with Apple Intelligence turned on it can work right away, and Orvena's own model takes about ten minutes to set up, most of it a download. This is the practical guide: which phones qualify, how much space it takes, how to set it up, what to expect, and exactly which features still need the internet.
Requirements for running AI offline on iPhone
| Phone | iPhone 15 Pro or newer: 15 Pro and 15 Pro Max; any iPhone 16, including the 16e; iPhone 17, 17e, Air, 17 Pro or 17 Pro Max; and iPhone 18 Pro, 18 Pro Max and iPhone Duo, which Apple lists for Apple Intelligence. Orvena checks for 8 GB of memory during setup. |
| Not supported | iPhone 15 and 15 Plus, iPhone 14 and earlier, and every iPhone SE. They have 6 GB or less. |
| iOS | iOS 26. Apple Intelligence and Private Cloud Compute in Orvena need iOS 27. |
| Storage | 3.1 GB for Orvena's own model, or no download in Orvena with Apple Intelligence on iOS 27 (Apple says Apple Intelligence itself takes up to 8 GB, or up to 14 GB on iPhone 17 Pro, 17 Pro Max and Air). Add 331 MB and 400 MB for the two voice packs if you use voice. Keep about 4 GB free before you start. |
| Connection | Only for the downloads. Wi-Fi is sensible for 3.1 GB, but cellular works. After that, none. |
The memory requirement is the real one. The model takes a little over 3 GB of memory while it runs, and iOS needs the rest. On a 6 GB phone the system would end the app before the model finished loading, so Orvena checks the phone's memory during onboarding and tells you plainly, rather than letting the download run and fail later.
Setting it up
- Install Orvena from the App Store. The assistant is free, and there is no account to create.
- Go through the onboarding pages. On the model page, if the iPhone runs iOS 27 with Apple Intelligence turned on and set to a language Apple Intelligence supports, an Apple Intelligence card reads "Built into iOS. Ready now, with nothing to download." You can start with it straight away; tap Use if it is not already checked. For Orvena's own model (Qwen 3.5 4B, under Apache 2.0), tap Get to start the download. It is 3.1 GB in ten files, each checked against a published hash when it arrives.
- Leave the app if you like. The download continues in the background, and you can explore the app while it runs. If it is interrupted, it resumes where it stopped when you tap it.
- Ask something. The first answer arrives as soon as the model is loaded. iOS asks for each permission, calendar, reminders, photos and so on, the first time a request needs it, with the reason shown.
- Optional: unlock voice. The two voice packs download in the background, and the first time you speak a language iOS fetches its on-device recognition model, which shows once as "Getting ready…".
What works with no connection
Everything that needs only the model and the phone:
- Conversation, drafting, summarizing text you paste, and reading a photo you attach. Orvena's own model has a vision encoder, so the photo is understood on the phone.
- Calendar: reading your events, adding one, moving one. Reminders: adding, listing, completing. Each change leaves a card with Undo.
- Alarms and timers, which ring even if you close the app, through silent mode and Focus.
- Searching your photos by date, album, place or kind, such as screenshots or selfies.
- Your Apple Health summary: steps, sleep and workouts.
- Looking up a contact, reading files you have imported, controlling Music, running a Shortcut.
- Voice conversations, once the voice packs are on the phone.
What still needs the internet, and where it goes
| Feature | Goes to |
|---|---|
| Web search | DuckDuckGo, when you ask for current information. |
| Opening a web page | That site, like a browser. |
| Weather | Apple's weather service. |
| Places, travel times, your current address | Apple Maps and geocoding. |
| Connected services | Notion, Linear, GitHub, or any MCP server you add, with sign-in to your own account. |
| Private Cloud Compute | Apple's servers, on iOS 27, after you choose it on the Models page. |
| Cloud models | OpenRouter, with your own key, after you consent. |
| Live cloud voice | OpenAI, with your own key, after you consent. |
In airplane mode these simply report that they cannot reach the service. If you chose Private Cloud Compute, Orvena answers with Apple Intelligence on the iPhone instead when it supports your language, and adds a note that says so: "This answer comes from this iPhone." Nothing else in the app has anywhere to go: there is no analytics, no account server, and no Orvena server between you and these services.
What to expect
Answers stream as they are written, and short ones start within a second or two. A long conversation takes longer before the answer begins, because reading the history is the expensive part on a phone. Orvena's own model reasons before answering by default; on the Models page you can turn that off for faster replies, and Settings, Advanced, lets you set your own budget. Apple Intelligence has no thinking setting. With Orvena's own model the conversation window is 16,384 tokens, about 12,000 words. Orvena reads Apple Intelligence's window from iOS, and on iOS 27 it reported 8,192 tokens, so long conversations run out sooner there. When a conversation outgrows the window, Orvena keeps the most recent whole turns and leaves out the oldest; it does not summarize them. If even that does not fit on Apple Intelligence, it says "This conversation is too long for Apple Intelligence. Start a new chat, or switch to Orvena's downloaded model in Models."
A short question costs very little battery. A long voice conversation warms the phone the way a game does. If you switch apps, a reply from Apple Intelligence or a cloud engine keeps going for a short while, and a reply from Orvena's own model pauses and continues when you come back; alarms and timers ring on their own; a scheduled task notifies you at the time you asked for and runs when you open it.
Keeping it lean
The Models page lets you switch to a mixed-precision variant of the same model (3.57 GB, more accurate tool calls) or remove the model to free space. Each voice pack can be removed separately. Conversation history is a file on the phone, encrypted with iOS complete file protection and excluded from backups. Deleting the app removes everything: the model, the voices and the history.
If something goes wrong
- The download stopped. Tap it. It resumes from where it was, and every file is re-checked against its hash.
- "This device can't run Orvena." The phone has less than 8 GB of memory. There is no setting that changes this.
- "Preparing the speech model took too long." The first voice session in a new language downloads that language's recognizer from Apple and needs a connection once. Try again in a moment.
- Answers feel slow. Turn thinking off on the Models page, or ask shorter questions in a fresh conversation. On iOS 27 you can also switch to Apple Intelligence under Built in on the Models page. Low Power Mode also halves the thinking budget automatically.
- "Apple Intelligence doesn't support this language yet." Apple's model covers fewer languages than Orvena's own. Download Orvena's model on the Models page to chat in that language.
Common questions
Does the model download every time I open the app?
No. The 3.1 GB download happens once. After that the model loads from the phone's storage and runs with no network at all.
Does it really work in airplane mode?
Yes. Conversation, images and every phone tool work with the radios off. Only web search, opening web pages, weather, places, connected services, Private Cloud Compute, cloud models and live cloud voice need a connection, and they say so when they cannot reach it.
How much storage does it take?
No download in Orvena if you use Apple Intelligence on iOS 27. Orvena's own model is 3.1 GB, plus 331 MB and 400 MB for the two voice packs if you use voice. Keep about 4 GB free before you start.
Can I run it on an iPhone 15 or 14 Pro?
No. Those phones have 6 GB of memory, and the model needs a little over 3 GB of it while iOS needs the rest. The iPhone 15 Pro and everything newer qualify.
Does Apple Intelligence work offline?
The on-device model does. Apple says apps can use it and "the features you build work offline." In Orvena, Apple Intelligence answers and uses the phone tools with the radios off. Private Cloud Compute, Apple's larger model on its servers, needs a connection.
Orvena is free on the App Store for iPhone 15 Pro and newer. On iOS 27 it can start on Apple Intelligence, and its own model downloads once and runs on the phone.
Download on the App Store