Skip to content

Local Models

Grafida can run an AI model on your device itself — no account, no API key, no internet connection once the model is downloaded, and nothing about your article ever leaving the device.

This is different from the On-Device (Apple Intelligence) option described in AI Services. That one uses Apple's own model, which is built into iOS and iPadOS on eligible devices. This one uses an open-weights model that you download and manage yourself, and it works on devices where Apple Intelligence is unavailable.

What you need

Local models are demanding, and Grafida will only offer one where it can realistically run:

  • Enough memory. This is the requirement that decides whether your device qualifies at all — your device needs to report at least 6.5 GiB, which in practice means an 8 GB-class device; a nominally 6 GB device reports about 5.8 GiB and does not qualify. If your device does not clear this, the Models tab is not shown, and nothing else about Grafida changes — every other feature works exactly the same, including the AI assistant with any hosted provider you configure.
  • Enough free storage for the model you choose. Grafida tells you how much it needs and how much you have. Unlike the memory requirement, this resolves the moment you free up space — the Models tab stays visible and offers the download again.

If your device qualifies on memory but has an A-series chip rather than an M-series one (M1 or later), Grafida still offers the model — it will simply run noticeably slower, and the model card says so before you download anything. A-series devices can technically run these models; they are just slow enough that sustained use is unlikely to be worth it. If a specific model also recommends more memory than your device has, its card says that too — the model is offered either way, since it may well work fine for shorter conversations.

The models on offer

Open AI in the sidebar, then the Models tab.

Model Size What it does
Qwen 3.5 2B about 1.7 GB Text, and it can see images
Ternary Bonsai 8B about 2.3 GB Text only

Both are free and openly licensed. Their licences are shown on the model card before you download anything, and again in the About screen.

Grafida does not include either model and does not distribute them. When you tap Download, the files come directly from Hugging Face, a public repository for machine-learning models, onto your device.

Downloading

Tap Download on the model you want. You will see how much has arrived, and an estimate of how long is left once there is enough evidence to give a sensible one.

IMPORTANT Keep Grafida open and your device awake while a model downloads. If you switch to another app or the screen turns off, the download pauses.

Pausing costs you nothing — nothing already downloaded is lost. Tap Download again and it carries on from where it stopped, whether it paused a minute ago or you came back the next day. The same is true if you tap Cancel Download, or if your connection drops partway through.

Downloads use Wi-Fi only. Grafida will not spend a couple of gigabytes of your mobile data without being asked.

Using a model once it is downloaded

Downloading a model adds it to the provider list. To use it:

  1. Go to the Services tab and add a service.
  2. Choose the model from the provider list — it appears near the top, above the hosted providers.
  3. Give it a name and save.

There is no endpoint to enter and no API key, because there is no service to connect to. Whether the model can see images is decided by the model itself, so there is no setting for it either.

From then on it behaves like any other AI service: pick it in the editor's AI panel, or set it as your default.

Deleting a model

Tap Delete Model on its card. The files are removed from your device and the model disappears from the provider list.

Any AI service you set up for that model is kept. Deleting a model to reclaim storage should not throw away settings you spent time on. If you use that service before downloading the model again, Grafida tells you the model is missing rather than failing in some vaguer way — download it again and the service works as before.

How it compares to a hosted provider

A local model is slower than a good hosted one, and less capable — these are small models, chosen to fit on a device. What they give you instead:

  • Nothing leaves your device. Not your article, not your images, not your prompt.
  • No account, no API key, no bill.
  • It works offline, once downloaded.

Whether that trade is worth it depends on what you are writing. Nothing stops you keeping both: a hosted service for heavy work and a local one for everything else, switching between them per tool or per conversation.