HOME>Blog>OpenAI previews GPT-5.6 Sol Ultrafast mode at up to 14x the speed
AI News5 min read

OpenAI previews GPT-5.6 Sol Ultrafast mode at up to 14x the speed

OpenAI is previewing Ultrafast, a new API service tier for GPT-5.6 Sol that can generate up to 750 output tokens per second. Here is what changes for teams building latency-sensitive workflows.

Kylon TeamProduct

The short version

OpenAI is previewing Ultrafast, a new service tier for GPT-5.6 Sol in the API. OpenAI says it can run up to 14 times faster than Standard processing and generate up to 750 output tokens per second. The preview is limited to a select group of customers today. OpenAI's announcement is the source for these figures and availability details.

This is a speed-tier announcement rather than a new model launch. The change matters because it places frontier-level reasoning closer to workflows where a person, customer, or live system is waiting for the next step.

What Ultrafast changes

OpenAI describes Ultrafast as a new service tier powered by Cerebras. It runs GPT-5.6 Sol with a focus on lower latency, rather than asking teams to choose a smaller model whenever response time matters. OpenAI says that the tier is intended to make more useful work possible per second.

The announcement highlights several possible workflows:

  • Incident response, where logs, traces, recent code changes, and engineer reports need to be reviewed while an outage is still unfolding. Source
  • Financial research and security review, where signals and transactions change during the analysis. Source
  • Customer support and voice experiences that need multi-step answers without breaking the conversation. Source
  • Commerce, including product questions, inventory checks, recommendations, and checkout support. Source

These are examples from OpenAI's announcement, not a guarantee of performance for every production system.

How it compares with other current directions

The current model market is separating two decisions that used to be bundled together: how much reasoning a workflow needs and how quickly the response must arrive.

OpenAI's Ultrafast preview keeps GPT-5.6 Sol as the underlying model and changes the processing tier. Google’s Gemini 3.7 Flash announcement takes a different route: it focuses on software engineering, knowledge work, and web development improvements, while offering an introductory API price of $0.75 per million input tokens and $3.75 per million output tokens through the end of 2026. Google's official announcement provides those pricing and capability details.

In practical terms, teams should not compare speed claims in isolation. A latency-sensitive support interaction may value Ultrafast's response time. A high-volume agent that can tolerate more waiting may prioritize Gemini 3.7 Flash's introductory token price and workflow improvements. The right comparison is task-level cost, latency, review burden, and output quality together.

Availability and what is not yet public

GPT-5.6 Sol on Ultrafast mode is available in a limited preview to a select group of customers. OpenAI says access will expand as capacity grows and provides a sign-up form for updates. OpenAI's availability section does not publish a general launch date or a public price for Ultrafast in this announcement.

That means teams should treat the preview as an evaluation opportunity, not as a generally available production option. Before changing an architecture, measure the current workflow with representative prompts, tool calls, retries, and human review.

What teams should measure first

A useful test should capture more than tokens per second:

  1. Time to a usable answer, including tool calls and application overhead.
  2. The number of retries or manual corrections required.
  3. Cost per completed task, not only cost per model request.
  4. Whether a human can still inspect the evidence and approve the final action.
  5. How the workflow behaves when a connected system is slow or unavailable.

This is especially important for agents. Faster generation does not automatically make an end-to-end workflow faster if the agent spends most of its time waiting on external systems or asking for repeated approvals.

What this means for AI-native workspaces

A speed tier becomes more useful when the surrounding work is already structured. The request, source context, permissions, connected tools, and approval history need to stay visible to the people responsible for the result.

That is the operational problem Kylon is designed around. Teams can keep their channels, files, agents, connections, and recurring workflows in one workspace, then decide which model or processing tier belongs at each step. The model can change without losing the work context and review path around it.

Bottom line

  • OpenAI is previewing Ultrafast for GPT-5.6 Sol in the API.
  • OpenAI says it can run up to 14 times faster than Standard processing and produce up to 750 output tokens per second. Official source
  • Availability is limited today, and public pricing is not included in the announcement. Official source
  • Teams should evaluate latency, completed-task cost, retries, and human review together.

For the full technical and availability details, read OpenAI's official Ultrafast announcement.

Hire the AI agent team that runs your entire business.