Ultra TechArabic edition

llama.cpp

llama.cpp 0.6.0 adds GLM-5.3-Flash and Clef, a decision-model API in the server, and new Metal kernels

Checked by machine against llama.cpp’s page on .

llama.cpp 0.6.0 introduces a new extended batch API, adds GLM-5.3-Flash, the Clef decision model and MTP speculative decoding for Qwen4Exp, gives the server a /v1/systemone API for decision models, rebuilds the Web UI around the Hugging Face Hub, adds flash-attention kernels for Metal and Vulkan, and moves ggml to v0.26.0. The release page dates it 5 October, 16:56. If you run llama-server on a Mac, two new Metal kernels change what happens under it; on an NVIDIA card, the parts to look at are a new path for two quantization formats and a faster way to decode Qwen4Exp.

What changed

The release adds an extended batch API, llama_batch_ext, with llama_process(): one batch can now mix tokens and embeddings. The examples, speculative decoding, mtmd (the multimodal part) and the server have all moved to it. The formats for saved state moved with it:LLAMA_SESSION_VERSION is now 11 and LLAMA_STATE_SEQ_VERSION4.

On models, GLM-5.3-Flash (GLM5-Next) arrives, a 320B hybrid model that reads text and images and is built as a mixture of experts. The Clef decision model is fully supported, with text and vision. Qwen4Exp gets what the page calls high-quality support, with MTP speculative decoding — about 1.5x faster decoding on DGX Spark — and correctness fixes.

The server gains a /v1/systemone API for five decision models: laya, julia-1, lev, openjev (with vision) and kev ; a sixth, nimble, joins them in the same release.GET /v1/models and GET /models now report each model's input and output modalities, in a new architectureobject./v1/embeddingsaccepts image, audio and video input, and answers an invalid request with HTTP 400. In router mode, the logs now keep child commands separate, and a preset can set the log file. The built-in Web UI gets a model download pipeline, a Hugging Face Hub data layer and an estimate of whether a model fits in memory.

Under the models, Metal gets a flash-attention kernel for an F16 KV cache , and few-row mat-mul kernels for speculative and batched decoding that the page puts at up to about 3x faster mat-mul on Apple GPUs. ggml moves to v0.26.0 ; in it, CUDA gains a W4A4 mul_mat path for NVFP4 and MXFP4 models, plus a fusion for shared experts , and model loading is faster, with stricter checks on GGUF sizes.

What it was before

The page lists all of the above as the changes since v0.5.0. That release took ggml to v0.25.0 ; this one takes it to v0.26.0. For Qwen4Exp, 0.5.0 added its hyper-connection operations and sparse flash attention ; 0.6.0 is the release that calls its support high-quality and adds MTP.

What it means for you

On a Mac, the two Metal kernels each cover a particular case. The flash-attention kernel is for an F16 KV cache ; if you have set the cache to a quantized type, it is not the case this kernel was written for. The mat-mul kernels are for speculative and batched decoding — so the gain should show when you run a draft model or serve several requests at once, and "up to about 3x" describes the mat-mul alone, not the words per second you will see. Measure on your own machine, before and after.

On an NVIDIA card, the W4A4 path concerns models in NVFP4 or MXFP4 ; if your files use another quantization, it is not the change to look for. If you run Qwen4Exp, update: MTP speculative decoding is now there, and the page's 1.5x was measured on DGX Spark, the one machine it names — your card will give its own figure.

On either machine, if you fetch models through the Web UI, it can now estimate whether one fits before you download it. What decides that answer — how much memory the model needs and how fast that memory is read — is explained in our guideWhich Machine Runs a Model Locally, and how a Mac and a PC differ on both in Mac or PC for Running Models Locally.

If you save state to a file and load it later, the format versions moved ; the page does not say whether files saved by 0.5.0 still load, so try one before you rely on it.

If your own code talks to the server, it can now read from /v1/modelswhether a model takes images or audio before it sends any , and code that requests embeddings should expect HTTP 400 for an invalid request. If you serve several models through the router, its logs now keep child commands apart , which helps when one model fails to start. And the decision models are new ground: the page names them and their endpoint, not what a request to it looks like , so read the server's documentation before you build on it.

The source

The release page on GitHub: llama.cpp v0.6.0, under the ggml-org organisation. For the version before it: llama.cpp v0.5.0.