On-device

Running Qwen 3.5 4B on an iPhone with MLX

UPDATED 29 SEPTEMBER 2026 · 6 MIN READ · ORVENA LABS

If you run Qwen on a laptop, you already know the family. This is the concrete setup that runs it on an iPhone: which variant, how it is quantized, how much memory and storage it needs, what the runtime does between turns, and what a four-billion-parameter model can and cannot do once it is connected to your calendar.

Orvena is free for iPhone 15 Pro and newer. The model runs on the phone itself. Download

The model

Qwen 3.5 4B is one of the small Qwen 3.5 models from Alibaba's Qwen team (the family also has 0.8B, 2B and 9B sizes), published in early 2026 under the Apache 2.0 license. It has about four billion parameters, a vision encoder so it can read images, a thinking mode that is on by default, a native context of 262,144 tokens, and training for tool calling. Alibaba lists 201 languages and dialects. On a phone, the two properties that matter most are the tool calling and the size: this is roughly the largest model that runs with room to spare in 8 GB of memory.

Orvena downloads mlx-community/Qwen3.5-4B-4bit, pinned to a single revision. The download is ten files totalling 3.06 GB, of which the weights are 3.03 GB; the app rounds this to "3.1 GB". A second option, "Qwen 3.5 4B Mixed", keeps some of the most sensitive layers (the embeddings, half the MLP down projections and some attention value projections) at eight bits and the rest at four. It is 3.57 GB and produces more accurate tool-call arguments, which is where four-bit quantization shows first. You can switch between the two on the Models page.

Why four bits

At sixteen bits per weight, four billion parameters are about 8 GB, which is the entire memory of the phone. At four bits the weights are about 3 GB on disk and a little more in memory, leaving room for iOS, the app and the working memory the model needs while it answers. That is the reason for the hardware requirement: Orvena checks the phone's physical memory on first launch and does not offer the download below roughly 8 GB. The phones that have it are the iPhone 15 Pro and 15 Pro Max, every iPhone 16 including the 16e, and the iPhone 17 and 17e. The iPhone Air and the 17 Pro models have 12 GB. Apple also lists the iPhone 18 Pro, 18 Pro Max and iPhone Duo for Apple Intelligence; on any phone, the memory check during setup decides.

The runtime

Inference runs through MLX, Apple's array framework for its own silicon, using the Swift bindings. The model runs on the GPU with the phone's unified memory, so there is no copying of weights between processors. Once loaded, the weights stay in memory between turns; a memory warning from iOS clears the caches but not the model.

Two caches make the phone feel faster than its raw throughput would suggest. When the model loads, Orvena computes the fixed part of the prompt, the system prompt and the tool descriptions, and keeps that state, so a new conversation does not pay for those tokens again. Within a conversation the key-value cache is carried from turn to turn, so each reply only has to read your latest message and the tool results, not the whole history.

The context window is capped at 16,384 tokens on the phone, rather than the model's full 262k. That is about 12,000 words of conversation, and the cap exists because reading a long prompt is the slow part on a phone and the cache for a long context is what runs the memory out. When a conversation grows past the window, Orvena keeps the most recent whole turns and leaves out the oldest ones rather than summarizing them. If a single tool result is too large to fit, it says "The latest tool result is too large for this model's context window. Try a narrower request."

Sampling is temperature 1.0 with top-p 0.95 for the whole reply when thinking is on, and temperature 0.7 with top-p 0.8 when thinking is off.

Thinking mode

Qwen 3.5 reasons before it answers unless told not to. Orvena gives that reasoning a token budget: 288 tokens on an 8 GB phone and 576 on a 12 GB one, half that in Low Power Mode. When the budget runs out the app forces the model to close its thinking and start the answer, so a hard question cannot stall indefinitely. The Models page has two settings, Automatic and Off, and Settings, Advanced, adds a custom budget between 128 and 2,048 tokens. Thinking helps most on multi-step tool tasks, an alarm that depends on travel time, for example, and costs a few seconds; for a quick factual question, Off is noticeably faster.

Tool calling

Tools are described to the model as JSON function schemas and rendered into the prompt by the model's own chat template, so the model emits calls in the format it was trained on. Orvena keeps the prompt small by disclosing tools progressively: the model starts with a few core tools and short summaries, then loads the skill it needs, calendar, reminders, maps, health and so on, and only then sees the full schemas. A request may take up to 32 model rounds. Every call is checked against the declared schema; a call to an undeclared name or with malformed arguments is returned to the model as an error it can read and correct, rather than silently dropped.

Arithmetic that matters is done by tools rather than by the model. When you ask for an alarm that gets you somewhere on time, the travel tool computes the leave-by time and the model only reports it.

Speed

There is no published benchmark for Orvena yet. Text streams as it is written, and the first words of a short answer appear within a second or two. What you will notice is that a long conversation takes longer before the answer starts, because reading the prompt is the expensive part on a phone, and that thinking adds a visible pause.

Languages and images

The model answers in whatever language you write. To make that reliable for short messages, Orvena detects the language of each message with Apple's on-device language recognizer and appends a runtime note stating it; the note is stripped from anything the model writes back, so you never see it. Attached photos go through the vision encoder on the phone.

Compared with running Qwen on a Mac

On a Mac you can pick the larger Qwen 3.5 sizes, and they are stronger models. On a phone, 4B is the practical ceiling today. What the phone adds is everything around the model: your calendar, reminders, alarms, photos, health summaries, maps, and a record of every action the model takes that you can read and undo. If you want a bigger model from your phone, Premium adds Apple's Private Cloud Compute on iOS 27, with no key to set up, or any OpenRouter model with your own key, and Orvena asks for consent before any conversation leaves the device.

Common questions

Which Qwen variant does Orvena download?

By default mlx-community/Qwen3.5-4B-4bit at a pinned revision, 3.06 GB in ten files. A mixed four- and eight-bit variant, 3.57 GB, is available on the Models page for more accurate tool calls.

Can I load my own GGUF or MLX model?

No. The download catalog has two entries, and each is tested with tool-loop tests on physical devices before it is listed. On iOS 27 the Models page also offers Apple Intelligence, which is built into iOS. Orvena is an assistant, not a general model runner.

Where do the weights come from, and are they verified?

From Orvena's mirror, with Hugging Face as the fallback. Every file is checked against the pinned size and SHA-256 hash before the model is activated. The download runs in the background, resumes if interrupted, and is excluded from iCloud backup.

Is my data used to train Qwen?

No. The model runs on your phone. Nothing you type or say is sent to Alibaba or to Orvena. A tool sends what it needs to its own service, a web search query to DuckDuckGo for example, and your conversation leaves the phone only if you choose Private Cloud Compute or connect a cloud model yourself.

Qwen 3.5 4B, on the phone, with your calendar.

Orvena is free on the App Store for iPhone 15 Pro and newer. The 3.1 GB download happens once and runs on the phone.

Download on the App Store