Updated 22 Aug 2026: OpenAI cut standard GPT-5.6 Sol API pricing on Aug 21 to $4/M input and $20/M output, down 20% and 33% respectively, with promotional pricing guaranteed at least through Nov 21. This materially changes the cost baseline against which the dossier evaluates the still-unpriced Ultrafast tier.

Key details

  1. OpenAI announced GPT-5.6 Sol Ultrafast on August 13, 2026.
  2. OpenAI says Ultrafast reaches up to 750 output tokens per second and up to 14× Standard processing speed.
  3. The service is powered by Cerebras and remains a limited preview.
  4. On August 21, standard GPT-5.6 Sol pricing fell to $4/M input and $20/M output tokens, down 20% and 33% respectively.
  5. OpenAI says the promotional Standard pricing will be available at least through November 21, 2026.
  6. OpenAI still has not published Ultrafast pricing or broad availability dates.

What builders should take away

  1. Recalculate Sol unit economics using $4/M input and $20/M output rather than the launch-era $5/$30 rates; output-heavy agent paths see the largest percentage reduction.
  2. Benchmark end-to-end task latency, not only tokens per second; tool calls, retrieval and network round trips can erase headline generation gains.
  3. If you build voice, coding, support or interactive research products, compare Standard, Fast and Ultrafast on completed-task cost and latency rather than treating faster tiers as automatically better.
  4. Keep Standard or Fast fallback routing because Ultrafast remains capacity-limited and preview-only.
  5. Do not make Ultrafast margin assumptions until OpenAI publishes its actual rates; a cheaper Standard tier increases the premium Ultrafast must justify.

What changed

OpenAI announced on August 13, 2026 that GPT-5.6 Sol is available in a limited-preview Ultrafast service tier powered by Cerebras, with claimed generation of up to 750 output tokens per second and up to 14× Standard speed. On August 21, OpenAI also cut standard GPT-5.6 Sol API pricing to $4 per million input tokens and $20 per million output tokens, reductions of 20% and 33% respectively. OpenAI says the promotional pricing will remain available at least through November 21, 2026. Ultrafast pricing itself remains undisclosed.

Why it matters

The latency argument for Ultrafast now sits against a cheaper Standard baseline. Builders evaluating low-latency coding, voice, support, research or agent workflows need to compare not only quality and time-to-first/last-token but the incremental price of Ultrafast versus a standard Sol request whose output cost has fallen by one-third. The standard price cut can materially improve unit economics even for teams that never receive Ultrafast access, while the lack of Ultrafast pricing still prevents a complete cost-per-completed-task comparison.

This is a service-tier change, not a new model

Ultrafast runs GPT-5.6 Sol rather than introducing a separate distilled or smaller model. OpenAI says the tier is powered by Cerebras and reaches up to 750 output tokens per second. That distinction matters because builders can evaluate whether latency-sensitive workloads can keep Sol-level capability instead of routing to a faster but weaker model.

Standard Sol just became materially cheaper

On August 21, OpenAI reduced GPT-5.6 Sol standard API pricing from $5 to $4 per million input tokens and from $30 to $20 per million output tokens. OpenAI describes the rates as promotional and guarantees them at least through November 21, 2026. For output-heavy workloads, the one-third reduction is the larger unit-economics shift.

The comparison point now has three dimensions

OpenAI already offers Fast mode for GPT-5.6 Sol, advertised at up to 2.5× Standard speed at a premium, while Ultrafast is advertised at up to 14× Standard. The August 21 Standard price cut means teams should compare Standard, Fast and Ultrafast on completed-task cost, latency distribution and quality rather than extrapolating from older token rates. OpenAI still has not published Ultrafast pricing.

Interactive workloads remain the clearest fit

OpenAI highlights coding, voice, financial research, support and incident response as early Ultrafast use cases. These are workloads where human-perceived latency is often part of product quality. The benefit will be smaller for jobs dominated by external tool latency, batch processing or long-running reasoning, while the cheaper Standard tier may now be economically preferable for many asynchronous paths.

Capacity and pricing uncertainty remain

Ultrafast is available only to a limited group of customers during preview. OpenAI says access will expand as capacity grows. Builders evaluating it should retain Standard or Fast fallback routing and should not commit unit economics until Ultrafast rates, quotas and operational guarantees are published.

What to watch next

  • Public pricing for the Ultrafast service tier.
  • Whether the $4/$20 promotional Sol rates are extended or changed after November 21, 2026.
  • Broader Ultrafast API availability, quotas and regional coverage.
  • Independent latency distributions beyond peak output-token throughput.
  • How Ultrafast behaves with long context, tool calling and high reasoning effort.

Still unclear

  • The 750 tokens-per-second and 14× figures are OpenAI’s stated maxima; independent production latency distributions are not yet available.
  • Ultrafast pricing, service-level guarantees and general-availability timing have not been published.
  • The new Standard Sol rates are explicitly promotional, although OpenAI says they last at least through November 21, 2026.

Sources

Direct reading behind this dossier.

4 sources
Changelog
OpenAI API primary

Primary changelog announcing per-request regional processing with region-prefixed domains for Global-geography projects.

Pricing
OpenAI API primary

Current API pricing reference for Sol and service-tier comparisons.