OpenAI
OpenAI adds Ultrafast mode to GPT-6.1 Sol, at six times the Standard price and with EU data residency
Checked by machine against OpenAI’s page on .
Updated .
OpenAI added Ultrafast mode for GPT-6.1 Sol to the Responses API on 8 October 2026. You switch it on per request: sendgpt-6.1-sol with service_tier: "ultrafast", and the time between generated output tokens shrinks. It is open to all API users, subject to rate limits, with global processing and with both US and EU data residency — the EU part is what GPT-6 Astra's Ultrafast mode does not offer. The price is six times Standard.
What changed
A service tier is how OpenAI processes and bills one request: the same model, under the same id, can come at a different speed and price depending on the tier you ask for. OpenAI's guide calls Ultrafast mode the fastest service tier in its API , broadly available for GPT-6 Astra and GPT-6.1 Sol. The GPT-6.1 Sol model page now says the same: for the fastest response speeds, use Ultrafast mode withmodel: "gpt-6.1-sol" and service_tier: "ultrafast"in the Responses API.
The price is a multiplier, not a column of figures. The model page says Ultrafast mode prices are 6x Standard , beside Fast mode at 2x Standard and Batch and Flex at 50% lower than Standard. The Standard figures being multiplied are on the same page: per 1M tokens, $2.00 input, $0.10 cached input, $2.50 cache writes and $10.00 output. The Ultrafast figures themselves sit in their own table, which the guide points to for input, cached input, cache write and output prices.
Rate limits — the ceiling on how many tokens your organization may send a model per minute — are counted apart: Ultrafast has separate rate limits from Standard and Fast modes. For GPT-6.1 Sol the defaults are 1,000,000 tokens per minute at Build, 4,000,000 at Launch and 40,000,000 at Grow. Those are the three usage tiers OpenAI moved to on 6 October ;our item on the change explains how an organization moves between them.
Data residency means your requests are processed inside a region you choose rather than wherever capacity happens to be. The guide says Ultrafast mode for GPT-6.1 Sol supports US and EU data residency and global processing , and the model page says the model keeps US and EU data residency with Fast and Ultrafast modes alike.
What it was before
What is new is Ultrafast on this model; the mode itself is not new at OpenAI. The changelog first announced Ultrafast mode on 13 August as a new API service tier for GPT-5.6 Sol , available in limited preview to select customers. On 29 September it added Ultrafast mode for GPT-6 Astra in the Responses API , with global processing and US data residency, and with EU and other regional inference residency not supported.
For GPT-6.1 Sol, which we covered in OpenAI releases GPT-6.1 Sol, the faster option until now was Fast mode, at 2x Standard.
What it means for you
Decide whether the speed is worth six times the price.The guide's own advice is to use it when speed justifies the higher cost. The model id staysgpt-6.1-sol; only the service tier changes. That suits a person watching an answer appear, or an agent that chains step after step and waits on each one. A job nobody is waiting on belongs at the other end of the list, where Batch and Flex cost 50% less than Standard. Before you commit traffic, read the actual figures in the Ultrafast pricing table.
If you run an agent, use a WebSocket.A WebSocket is one connection kept open between your application and the API, so each new request does not pay to open a fresh one. OpenAI strongly recommends WebSockets, especially for agentic applications that make many tool calls in quick succession , because without a persistent connection, network overhead can reduce the latency gains. Over a WebSocket you set the model togpt-6.1-sol and service_tier to ultrafast in each response.createevent. Plain HTTP requests through the SDK work too : in Python, a single call isclient.responses.create(model="gpt-6.1-sol", service_tier="ultrafast", input=...).
Check your limits before you shift traffic.Because Ultrafast's limits are separate , the headroom you have on Standard tells you nothing here. The guide says to check your organization’s limits before increasing traffic ; at Build, the default for GPT-6.1 Sol is 1,000,000 tokens per minute. If your organization works with an OpenAI account team, that is who to ask for higher limits.
If your data must stay in the EU, this is the Ultrafast model to use.GPT-6.1 Sol's Ultrafast mode supports EU data residency ; GPT-6 Astra's does not. Budget for the regional premium too: the pricing page charges data residency endpoints a 10% uplift for models released on or after 5 March 2026 , and the model page states the same 10% premium for regional processing.
The source
OpenAI's API changelog, the entry of 8 October 2026, with the Ultrafast mode guide, the GPT-6.1 Sol model page and the pricing page.