Ollama
Ollama 0.40.2 upgrades your models in the background and keeps the originals as backups
Checked by machine against Ollama’s page on .
Updated .
Ollama 0.40.2 upgrades the models you downloaded with earlier versions in the background, the first time you run each one, for better performance and compatibility when they run on llama.cpp. The release page is dated 8 October, 16:53. The upgrade has a cost you will see on your disk: to make going back to an older version safe, Ollama keeps the original copy of each upgraded model as a backup, so for now those models take up space twice. A backup here is simply the model as it was before the upgrade, set aside so that an older Ollama can still use it. This release tells you how to clear those backups, and the one case in which you should keep them.
What changed
The release note puts it under one heading, model upgrades. A model you pulled with an earlier version is not converted the moment you install 0.40.2: it is upgraded in the background the first time you run it, and the reason given is better performance and compatibility on llama.cpp, the engine that loads the model and computes its answers. So the change reaches your models one at a time, as you use them.
The original copy stays behind as a backup, so that downgrading is safe, and while it does the upgraded model sits on disk beside it. The page calls this temporary: a future release will remove these backups automatically. It does not say which release.
If you do not want to wait, the page gives a command that removes the backups now, and it needs jq. jq is a small command-line program that reads JSON, the text format Ollama's server answers in. The command is four lines long; copy it from the release page linked under "The source" below, exactly as it is printed there.
Here is what it does, a step at a time. It first asks your local Ollama server for the list of your models and takes each model's name. For each name, it then asks the server to describe that model and look at its manifests — a manifest is the short file that lists what a stored model is made of. If the model has a manifest for the llama.cpp runner, which means it has been upgraded, the command picks the manifest it still keeps for the old ggml runner and prints that manifest's digest: a fingerprint that names one stored copy exactly. Finally it drops repeated digests and hands each one to ollama rm. The page is explicit about what it touches: this only deletes the backups. And it states what that costs you: if you later downgrade to a version older than 0.40, you will need to pull those models again.
Two smaller fixes come with it. ollama listno longer shows duplicate entries for upgraded models , so if a model you had run appeared twice in that list, this is the fix. Andollama launch claudenow gives the model all of its context — everything the model can take in at once.
What it was before
The page lists its changes as those since v0.40.1 , the release dated 7 October, 23:22. That release, and the switch in 0.40 to running supported models on MLX on Apple Silicon, are in our itemOllama 0.40 runs supported models on MLX by default on Apple Silicon.
What it means for you
Start with disk space, because that is the part you will notice. Every model you run after updating gets an upgraded copy while its original stays as a backup , so for each model you use, you need room for both copies until the backups go. The growth does not come all at once at the update: it arrives model by model, the first time you run each one. If your models are large and your disk is close to full, check your free space before you update, and again after you have run the models you use most.
Then decide about the backups. Their one job is to make a downgrade safe. If you might go back to a version older than 0.40 — because a model behaves differently after the upgrade, or a tool you rely on needs the older version — keep them: without them, you would have to pull those models again. If you are staying on 0.40 or later, removing them takes nothing you need , and you can either wait for the release that removes them by itself or run the command now.
Before you run it, check two things. jq must be installed ;jq --version answers if it is. And the Ollama server must be running on this machine on its default port, the address written into the first two lines, because those are the lines that ask it for your models; if you run the server on another address, change it in both places. If you want to see what would be removed before anything is, drop | xargs -n1 ollama rmfrom the last line: the command then only prints the digests. A model you have not run since updating has no llama.cpp manifest yet, so the command leaves it alone — it has no backup to remove.
The two smaller fixes ask nothing of you. If you use ollama launch claude, the model gets all of its context by itself after the update , andollama listshows each upgraded model once.
The source
The release pages on GitHub: Ollama v0.40.2 and, for the release before it, Ollama v0.40.1.