Ultra TechArabic edition

Ollama

Ollama 0.40.3 pauses the background model upgrade that 0.40.2 brought in

Checked by machine against Ollama’s page on .

Ollama 0.40.3 pauses the background model upgrade that arrived with 0.40.2, and the reason given is the speed of embedding and other requests. The release page is dated 11 October, 01:39. The change behind it says more: it temporarily removes the background migration until a future release, because the migration added overhead to every request and put pressure on garbage collection when a server handles a high volume of requests, embedding models being the example named. If you run Ollama as a server for many requests, this is the release to install; if you use it now and then on your own machine, it changes less for you, and below is what it changes and what it leaves unsaid.

What changed

The release lists three changes, and one of them matters. The background model upgrade is paused. The pull request that made the change — the place on GitHub where the change was proposed and merged — describes it as a temporary removal: the migration is taken out until a future release. "Migration" here means the same thing as the upgrade: converting a model you downloaded earlier into the form the new version prefers.

The reason has two parts. The first is overhead per request: while the migration was in place, every request to the server carried a little extra work. The second is garbage-collection pressure. Garbage collection is the clean-up a program does to free memory it no longer uses; when it has to run often, the program spends time on it that it could have spent answering. Both costs are small for one request and grow with many, which is why the page names high-volume use and gives embedding models as the example. An embedding model turns a piece of text into a list of numbers that captures its meaning, for search or for matching similar documents, and such a model is usually sent a long run of short requests rather than one long conversation.

The two other changes are about ollama launch, the command that starts a coding tool wired to a local model. ollama launch claudeno longer shows a warning about Claude connectors , andollama launch codexno longer warns about service tiers. Neither changes what the tools do; the warnings simply stop appearing.

What it was before

The page lists its changes as those since v0.40.2. In 0.40.2, models downloaded with earlier versions of Ollama were upgraded in the background the first time you ran them, for better performance and compatibility on llama.cpp. To keep going back to an older version safe, Ollama kept the original copy as a backup, so upgraded models were kept on disk twice for a time , and the page said a future release would remove these backups automatically. Our itemOllama 0.40.2 upgrades your models in the background and keeps the originals as backups explains that release and the command it gives for removing the backups.

What it means for you

If you serve many requests — an embedding model behind a search tool, or an application that sends Ollama a steady stream of short calls — update to 0.40.3. The per-request overhead and the garbage-collection pressure were named for exactly that kind of use , so it is where you are most likely to see the difference.

If you have not updated past 0.40.1 yet, know what you get: 0.40.3 does not run the background upgrade , so the models you run are not converted on their first run as 0.40.2 described. And since the backups came from that upgrade , this version has no reason to make new ones.

If you are on 0.40.2 and some models have already been upgraded, the 0.40.3 page says nothing about them: not whether they stay in their upgraded form, and not what happens to the backups made beside them. The only word on the backups is still the one in 0.40.2 — that a future release will remove them automatically. So nothing on your disk changes by itself when you install 0.40.3. If you need the space now, the command in our 0.40.2 item still applies; read its conditions there before you run it, especially the one about going back to a version older than 0.40.

The migration is paused, not dropped: the change says it is removed until a future release. Expect it to come back, and when a release brings it back, check its notes for how it avoids the cost that stopped it this time.

If you use ollama launch claude or ollama launch codex, there is nothing to do: the warnings about Claude connectors and about service tiers no longer appear.

The source

The release page on GitHub, Ollama v0.40.3; the change behind the pause, pull request #18908; and the release before it, Ollama v0.40.2.