Ultra TechArabic edition

Ollama

Ollama 0.40 runs supported models on MLX by default on Apple Silicon

Checked by machine against Ollama’s page on .

Ollama 0.40.0 changes the engine that runs your models on an Apple Silicon Mac: model architectures that the MLX runtime supports now run on MLX automatically. The v0.40.1 release, the one an update installs, is dated 7 October, 23:22. The v0.40.0 page carries an earlier date, 25 September, 03:31 — GitHub dates a release by the day its page was opened, and this one opened while it was still a pre-release. An engine here is the program that loads a model's file into memory and computes its answers. On an Apple Silicon Mac, the engine under several common models changes without you asking; on any other machine, the headline does not apply to you.

What changed

The release note says it in its heading: models run on MLX on Apple Silicon by default. What decides it is the model's architecture — the internal design a family of models shares — not its name: any architecture the MLX runtime supports moves over. The page's own example is qwen3.8, pulled and run with the usual two commands,ollama pull qwen3.8 and ollama run qwen3.8. It names gemma4, qwen3.6 and qwen3.5 as further models.

Two other kinds of model come with them. The decision models — Nimble, tev1, clef and clef-flash — are now available on MLX as well ; these are models that return choices, probabilities and scores instead of text. And MLX now supports an embedding model, embeddinggemma-2 — the kind of model that turns a passage into numbers, so a search over your documents can find passages close in meaning. The page ends the list with an intention rather than a date: the project will continue testing and enabling additional models.

v0.40.1 is a short list of changes on top of it. For a Mac, its one MLX line drops a Metal residency patch the project had been carrying, now that the change is upstream. The rest concern other machines and other parts of Ollama: a fix for clef head reads past 2GiB on Windows , the manifest avoids symlinks on Windows , the account step is gone from the command line's onboarding , and the server now proxies the cloud usage and balance APIs.

What it was before

MLX is not new to Ollama. The page lists its changes as those since v0.35.1 , and that release updated both llama.cpp and the MLX engine — so Ollama already carried two engines, and what 0.40.0 changes is the one a supported model gets by default on Apple Silicon. Two of the decision models arrived just before: 0.35.1 added Clef and Clef Flash, Cloudflare's open-source decision models, through/v1/systemone. The pages do not say which engine each of these models ran on before 0.40.0.

What it means for you

If you have a Mac with Apple Silicon, updating is the whole step: there is nothing to switch on. A model whose architecture MLX supports moves to MLX by itself. If you run qwen3.8 , gemma4, qwen3.6 or qwen3.5 , one of the decision models , or embeddinggemma-2 , you are on the list the page names. If your model is not named, the page gives no full list of architectures — only that more will be enabled after testing.

Update straight to 0.40.1 rather than stopping at 0.40.0: it is the later of the two, and its MLX line is the one change in it that touches MLX.

To check the change, start with what the page does not give you: it names no command that shows which engine a loaded model is running on. So check what you can see. Before you update, run the model you use most — ollama run qwen3.8, if you follow the page's example — with one prompt you keep, and note how fast the answer comes and how much memory the machine is using; after the update, run the same prompt again. The page gives no speed figures; your two runs are the measurement. What decides speed on a Mac — how much memory it has and how fast that memory can be read — is explained in our guideMac or PC for Running Models Locally, and Which Machine Runs a Model Locally shows you how to read those figures on the machine you already have.

If something gets worse, neither page names a setting that takes a model back off MLX, or a way to choose the engine yourself; they say only that testing continues. If a model you rely on is slower after the update, or answers differently, keep your before-and-after notes and report them to the project with the model's name and the version you updated to.

If you use embeddinggemma-2 to search your own documents and keep an index built before the update, the page says nothing about whether the numbers it produces stay the same on MLX. Embed a few passages again and compare them with the stored ones before you mix old and new in one index.

The source

The release pages on GitHub: Ollama v0.40.0 and Ollama v0.40.1; for what came before, Ollama v0.35.1 and Ollama v0.35.0.