Ollama
Ollama 0.35 adds decision models that answer with choices, probabilities and scores instead of text
Checked by machine against Ollama’s page on .
Ollama 0.35.0 adds decision models, served on your own machine through /v1/systemone, an endpoint the release page says is based on TypeSafe's Jev API. A decision model returns choices, probabilities and scores instead of text , and the page suggests it for ticket triage, model routing and content classification. The release page dates 0.35.0 to 28 September, 21:23 ; 0.35.1, dated 29 September, 20:14, adds a second family of these models that can also look at images.
What changed
0.35.0 ships two decision models: Nimble from Bespoke Labs and Tev1 from Together AI. You fetch one the usual way, withollama pull nimble, and the page's example sends its request tohttp://localhost:11434/v1/systemone.
A request carries context and one or more questions. In the page's example, the model isnimble and the context — the statefield — is a support ticket: "Our checkout has returned 500 errors since 9am." The one question, namedlabel, is of type choice. It asks "Which label fits this ticket?" and lists three options, each with a short description: billing for payments and refunds, bug for software errors, accountfor login and account access.
The answer is not a sentence. It is the option the model chose, bug, with a probability for every option — 0.9781 for bug, 0.0125 for billing, 0.0093 for account — and a confidenceof 0.8906.
choiceis one of three question types the API supports. Achoicequestion selects an option and returns a probability for each ; anoulquestion returns the probability that a condition is true ; ascorequestion returns a score across an ordered set of criteria.
0.35.1 adds Clef and Clef Flash, which the page calls Cloudflare's new open-source decision models, on the same endpoint. Clef is 27B and Clef Flash 9B, and both are multimodal: a request can include images alongside the text state, shared by all the questions and scored together with it. The page's example sendsclef-flash the state "The user took this screenshot." with one base64-encoded image in an imageslist , and asks a singlenoulquestion: "Does this image contain Ollama?" The answer is one number, the probability: 0.958.
0.35.1 also changes how these models present themselves: ollama show and the model list now report decisionas their only capability, so clients no longer offer them for general chat, tools or thinking. Modelfiles gainCAPABILITYdeclarations, so whoever builds a model can state what it can do. And llama.cpp and the MLX engine are updated.
Elsewhere in 0.35.0, Settings opens without waiting for model discovery , stalled MLX downloads no longer hang , and the deprecatedtypical_pparameter logs a warning instead of failing the request.
What it was before
The release page lists its changes against v0.34.4. The nearest thing in that release's notes is structured output: on thinking models it now applies in a single pass, which the page says makes it faster and more reliable.
What it means for you
If you pull labels out of a local model today, you probably ask a chat model for JSON and hold it to a schema — structured output, which our guide What Structured Output Iswalks through. That is still the right tool when the answer is text: a summary, an extracted name, fields whose values you cannot list in advance. A decision model returns choices, probabilities and scores instead of text , so it fits the other case: the answer is one of a set you wrote down yourself.
What it adds is the number next to the choice. A schema makes a reply the right shape, which, as the guide warns, says nothing about whether the value is right. With a choicequestion you get a probability for each option , so you can route a ticket automatically when the top option is far ahead and hand it to a person when two are close. Withnoul, a yes-or-no check becomes one probability you set your own threshold on. Those are the jobs the page names: triage, routing, classification.
To try it, update to 0.35.0, run ollama pull nimble, and send your context and questions to/v1/systemonerather than to the chat endpoint. If the context is a screenshot or a photo, update to 0.35.1 and use Clef or Clef Flash. The page gives no pull command for Clef, and gives the sizes, not the memory they need; our guideWhich Machine Runs a Model Locally explains what decides whether a model fits.
After 0.35.1, don't look for these models in your chat app's model picker: clients no longer offer them for chat. If you run llama.cpp rather than Ollama, or would rather not run the model yourself, see our items onllama.cpp 0.6.0 and OpenAI's Decisions API.
The source
The release pages on GitHub, in the ollama/ollama repository: Ollama v0.35.0 and Ollama v0.35.1, and for the previous release, Ollama v0.34.4.