“Up to 8× faster” describes a ceiling for token generation, not the time to fix a bug and pass its tests. With GPT-6.1 Sol now offered in three speed tiers, this guide gives developers the official rate comparison, a worked cost example, and a repeatable pilot for routing only the tasks that benefit.
What changed with GPT-6.1 Sol Ultrafast
OpenAI's October 8 announcement adds an Ultrafast service tier for GPT-6.1 Sol across the API, Codex, and ChatGPT Work. In the API it is still gpt-6.1-sol: set service_tier: "ultrafast". Standard and Fast remain options. API access and the subscription-based Codex/Work entitlement follow different rules.

The launch image establishes the service tier; pricing, eligibility, and speed claims need separate checks.
This is a speed tier for the same model, not a new model ID or an automatic increase in reasoning quality. Its value depends on whether waiting for generated output lies on your task's critical path. A faster stream does little for a task dominated by package installation, browser loads, test execution, or human review.
The DevDay stage demonstration makes the difference visually obvious. It does not measure how long your own end-to-end workflow will take.

Use the stage demo to understand the positioning, then benchmark your own full task.
Translate “up to 8×” into task time
In its DevDay recap, OpenAI says Ultrafast can deliver up to 8× faster token generation in Codex and up to 6× in the API. Neither number is a guarantee that an entire task completes eight or six times faster. Queueing, network hops, prompt processing, tool calls, tests, and review remain in the workflow.
For example, if a 100-second job spends 40 seconds generating output and 60 seconds elsewhere, a 6× improvement in generation reduces the total to about 66.7 seconds, or roughly 1.5× end-to-end. If generation takes only 10 seconds and the rest takes 90, the total becomes about 91.7 seconds. Measure where your own time goes before paying for a speed tier.
One developer's public test used four runs of the same prompt and reported a median near 304 versus 42 tokens per second, with total response times near 4.3 versus 17 seconds. This is a useful example of what to record, not a service-level promise or an independent cross-task benchmark.

Track token speed, time to first output, and total response time separately; agent tasks also include tools and tests.
Log request start, first visible output, model completion, and accepted task completion. Hold prompt, reasoning effort, tools, and cache conditions fixed across tiers. Compare medians and p95, not a single promotional clip.
The three API prices, calculated on the same workload
The official pricing page lists the following USD prices per million tokens for short-context GPT-6.1 Sol requests. Fast costs twice Standard; Ultrafast costs six times Standard. That multiplier buys a service tier, not six times the allowance or six times the task success rate.
| Tier | Input | Cached input | Output | Relative to Standard |
|---|---|---|---|---|
| Standard | $2.00 | $0.10 | $10.00 | 1× |
| Fast | $4.00 | $0.20 | $20.00 | 2× |
| Ultrafast | $12.00 | $0.60 | $60.00 | 6× |

The tier table is a starting point. Cache writes, tool fees, long context, and regional processing can change the invoice.
Take a request with 20,000 uncached input tokens and 5,000 output tokens, excluding tools. Standard costs 0.02×$2 + 0.005×$10 = $0.09; Fast costs $0.18; Ultrafast costs $0.54. On the same token count, Astra Standard's input and output alone would cost $0.45. Sol Ultrafast can therefore cost more per request than Astra Standard. Which is better still depends on acceptance rate and latency.
The following individual developer screenshot compares a real Standard charge with the estimated Ultrafast charge for the same token usage. It demonstrates the price multiplier, not a measured reduction in work time at equal quality.

An estimate based on one task is not a representative average.
Another comparison puts Sol Standard, estimated Sol Ultrafast, and Astra Standard together. The apparent bargain changes when you change the baseline: one-fifth the price of Astra Ultrafast does not mean cheaper than every Astra configuration.

Recalculate this single-user example using your own task set and the official rate card.
How much waiting must the premium actually remove?
For the 20,000-input/5,000-output-token example above, Ultrafast costs $0.45 more per call than Standard. Five sequential calls of the same size add $2.25. If an engineer is truly blocked by those calls and you value that time at an illustrative $60 per hour, the tier needs to save at least 2 minutes 15 seconds of active waiting to cover the premium. A background job finishing sooner does not automatically free paid time; this is a decision calculation, not a measured speedup or a quoted labor rate.
Suppose those five calls shorten accepted completion by three minutes while the engineer waits: the assumed time value is $3, above the $2.25 premium. If the stream starts earlier but tests and review leave the accepted result only 20 seconds sooner, the premium does not pay for itself on that measure. Log incremental API spend, blocked waiting, and accepted completion time separately. Tokens per second alone cannot answer the purchase question.
Cache and long context
The model page says requests exceeding 272K input tokens use long-context prices for the whole request. Ultrafast long-context prices are $24/M input and $90/M output. At 300,000 uncached input tokens and 20,000 output tokens, the text-token subtotal is 0.3×24 + 0.02×90 = $9.00. A short-context unit-price example should not be extrapolated across this boundary.
Cached input may lower input charges; cache writes, retries, tool calls, and regional processing can raise the total. The pricing page lists a possible 10% regional-processing uplift. Treat these examples as decision aids, and confirm your actual invoice and current rates before procurement.
API eligibility is different from Codex and Work eligibility
The Ultrafast API guide says GPT-6.1 Sol Ultrafast is available to API users, has rate limits separate from Standard and Fast, and supports global processing plus eligible US and EU data residency. API usage is billed under API rates; a ChatGPT subscription is not a free API allocation.
For Codex and ChatGPT Work, OpenAI's rollout notice names Pro 500, eligible usage-based Enterprise, and credit-based Edu. Enterprise administrators must enable it. GPT-6.1 Sol is a Work/Codex model rather than the ordinary Chat model. Do not infer that everyone seeing GPT-6 in Chat has Sol Ultrafast. Plan allowances and credit consumption should be checked in the account's current usage view; multipliers shown for other models or plan configurations may not transfer to every Sol account.

The Pro 500 slide describes a subscription route, not API entitlement or pricing.
A user also reported rapid weekly-allowance consumption. It is a prompt to inspect your own quota before a pilot, not a universal percentage-per-minute rule.

Quota experiences vary with model, effort, context length, and the actual path through a task.
Run a fair API pilot
For a functional smoke test, the Responses API HTTP example accepts the same model ID with an explicit service tier. The Python snippet below is illustrative and needs an API key and an installed OpenAI SDK; it is not executed here. Compare default, fast, and ultrafast with identical prompts, effort, tools, and data, and record the tier and usage returned. If your project changes default routing, confirm what default actually maps to.
from time import perf_counter
from openai import OpenAI
client = OpenAI() # Reads OPENAI_API_KEY from the environment
started = perf_counter()
response = client.responses.create(
model="gpt-6.1-sol",
service_tier="ultrafast",
reasoning={"effort": "medium"},
input="Explain empty-value handling in a Python form and propose two reproducible tests.",
)
print(response.output_text)
print({
"elapsed_seconds": round(perf_counter() - started, 2),
"actual_tier": response.service_tier,
"usage": response.usage,
})
This non-streaming example measures the completed response only; it cannot measure time to first token. To measure first output, enable streaming and timestamp the first content event.
For agents that call tools repeatedly in one job, OpenAI recommends a persistent WebSocket connection. Recreating the connection on every turn adds network overhead that can erase part of the inference gain. HTTP remains supported; use it to validate basic behavior, then measure whether a persistent connection improves total task time. Ultrafast has separate rate limits, so test peak traffic and retry handling as well.
Useful columns are task_id, tier, reasoning_effort, input_tokens, cached_input_tokens, output_tokens, time_to_first_token, model_duration, tool_duration, end_to_end_duration, accepted, rework_minutes, cost_usd. Run the same representative tasks across tiers at least twice. Record p50 and p95, quality failures, and manual rework. A successful HTTP response is not the same as an accepted outcome.
Turn the pilot into a decision record
Run the same acceptance script for each task_id under all three tiers and keep failed attempts. For a task such as “fix a date-field crash on an empty value,” acceptance means the specified test passes, existing form submission still works, and a reviewer accepts the diff. Interleave tier runs to reduce time-of-day bias. Keep cache state, concurrency, and retry policy as similar as possible; if you cannot, record those differences rather than hiding them in an average.
| Task ID | Tier | First token / full response | Tools and tests | Accepted completion | Charge and rework | Outcome |
|---|---|---|---|---|---|---|
bug-017 |
One row per tier | Measured seconds | Measured seconds | Measured seconds | Actual dollars / minutes | Accepted, failed, or retried |
Compare acceptance rate, p50/p95, and cost per accepted task within each workload category. Enable Ultrafast only where it shortens a blocked critical path within budget. If the separate Ultrafast limit is hit or the gain disappears, fall back to a tier already tested on that workload. Avoid switching all traffic at once.
Faster steering is a separate product change
The ChatGPT release notes also describe faster steering in Codex desktop: send a follow-up while a task runs so it can incorporate the change sooner. In Settings → General → Follow-up behavior, you can choose whether the default is to steer the current run or queue the message for the next one. The Codex help article explains Steer versus Queue. This is not an Ultrafast-only switch.

Steering changes collaboration during a running task; model speed tier is a separate choice.
In the demo, the request changes from a cat to a dog and then to three ducks. The next two frames show an intermediate correction and the result. The demo interface displays GPT-6 Astra, so it cannot prove Sol Ultrafast timing or success on the same task.

First frame: a new goal arrives during the current run.

Second frame: the run follows the new goal; the Astra model shown here is not a Sol Ultrafast test.
Steering is useful for correcting a premise, adding a missing file, or narrowing scope. It does not automatically undo file edits or external actions already completed. Review actual artifacts and tests after a direction change.
A task-based decision rule
| Workload | Starting tier | What to verify |
|---|---|---|
| Overnight summaries and queueable reports | Standard; evaluate lower-cost asynchronous options | No person waits for each response. |
| Interactive reviews and short coding loops | Compare Fast with Ultrafast on the same set | Does accepted task time and rework fall? |
| Incident response, live UX, repeated tool calls | Trial Ultrafast with a fallback | Is model waiting on the critical path? |
| Long context and repeated cached prompts | Recalculate cache and long-context costs first | Token composition may dominate tier selection. |
Use cost per accepted task = API charge + retries + human rework cost, alongside p50/p95 time to accepted completion. If only the text stream starts sooner, the sixfold API premium may be hard to defend. If model waiting holds up an incident or a high-value interactive workflow, the premium may be rational.
Turn a Code0 latency test into routing rules
Separate tasks into short requests with a person waiting, multi-turn tool jobs, and queueable offline work. Give each group a tolerable delay and measure acceptance rate per tier. A code fix must include tests and review, not merely a fast patch. If Standard produces the same accepted outcome and is only eight seconds slower when no one is blocked, keep Standard. If each minute of incident recovery has a clear operational cost, route that narrow class to Ultrafast. Put the policy into configuration and monitoring rather than switching every request globally.
In a Code0 task dashboard, attach latency targets and spend caps to task types rather than the whole model. Review actual service_tier, accepted completion, and fallback counts weekly. If the existing route meets its p95 target, that category has no reason to default to Ultrafast.
Pilot checklist
- Select 20–30 historical tasks spanning short answers, tool-heavy work, and long context. Hold prompts, tools, and effort fixed.
- Run each tier at least twice; log first visible output, complete response, and accepted task time.
- Save returned
usage, actualservice_tier, failures, retries, and cost; separate requests above 272K input tokens. - Review output by the same acceptance rules and count rework minutes; do not equate HTTP 200 with completion.
- Define task-level routing, a spend cap, and a Standard/Fast fallback when Ultrafast adds little value or hits a limit.
Sources
- OpenAI GPT-6.1 Sol model page and official pricing: model ID, rates, and long-context treatment.
- OpenAI Ultrafast guide and rollout notice: parameters, WebSockets, rate limits, and access.
- ChatGPT release notes and Codex help: steering, queueing, and product entry points.



