AWS lists a 6× Ultrafast lane on Bedrock — speed is the cash print
AWS today listed Ultrafast for GPT-6 Astra on Bedrock with a published 6× price card. The cash print is a scarce inference lane — not a new model.
WIRE · HYPERSCALERS · US · Data as at 30 Sep 2026
Amazon Web Services today listed Ultrafast as a paid speed tier for OpenAI GPT-6 Astra on Amazon Bedrock. The model has been generally available since 8 September. What printed on 30 September is a 6× price card on a scarce inference lane — not a new campus. On the hyperscalers layer the stuck step is who can sell priority tokens while the fleet stays full.
What happened
On 30 September 2026 AWS said Ultrafast mode for GPT-6 Astra is available on Amazon Bedrock — a premium speed tier for real-time coding assistants, interactive agents, and customer-facing apps. Speed is attributed to OpenAI: up to 6× faster inference in the API, up to 300 tokens per second (AWS What's New, 30 Sep 2026 — company).
The Bedrock model card is the cash document. Ultrafast is six times the corresponding Standard prices, requested with "service_tier": "ultrafast". GPT-6 Astra launched on Bedrock on 8 September 2026 (1.05M context, 128k max output, knowledge cutoff 30 April 2026). Priority and Flex are not supported for this model (Amazon Bedrock, GPT-6 Astra model card — company).
OpenAI opened the same lane on its own API the day before, at DevDay on 29 September 2026 — GPT-6 Astra Ultrafast in the API, and in ChatGPT Work and Codex on Pro 500 and Enterprise (OpenAI, DevDay 2026 Recap, 29 Sep 2026 — company). AWS’s 30 September note is the hyperscaler attach, billed through AWS.
| Figure | Date | |
|---|---|---|
| Event | Ultrafast on Bedrock for GPT-6 Astra | AWS, 30 Sep 2026 |
| Claimed speed (OpenAI, via AWS) | Up to 6× API; up to 300 tok/s | Same |
| Model GA on Bedrock | 8 Sep 2026 | Model card |
| Standard Global, short (≤272k in) | $10 in · $50 out per 1M tokens | Same |
| Ultrafast Global, short | $60 in · $300 out (6×) | Same |
| Ultrafast In-Region / US Geo, short | $66 in · $330 out (10% over OpenAI) | Same |
| Long-context gate | Whole request reprices above 272k input | Same |
| Ultrafast Global, long | $120 in · $450 out | Same |
All Bedrock figures are USD per 1 million tokens. Short-context rates apply at 272,000 input tokens or fewer; above that, long-context rates apply to the entire request.
Why cash cares
Ultrafast does not add a megawatt. It reprices the same model so a customer who needs the next token now pays six times Standard. Who gets paid is the priority lane — AWS on the Bedrock invoice, OpenAI billed through AWS. Who pays is the app that cannot wait in the Standard queue. See What Is GPU as a Service?; this print is the token-meter version of the same seat.
Our arithmetic, from the model card, not from a run:
- Six times the token. Ultrafast Global short output is $300 per 1M versus Standard $50 — a $250 delta per million output tokens. If the OpenAI-attributed 6× token rate actually prints, a 1M-output job finishes in one-sixth the wall-clock and still costs six times Standard. You buy wait-time, not a discount.
- The 272k cliff reprices the whole book. A 272,000-input request at Ultrafast Global short is 272,000 × $60 / 1,000,000 = $16.32. One extra token trips long-context rates on the entire request: 272,001 × $120 / 1,000,000 = $32.64. Output also steps from $300 to $450 per million on that whole request.
- Residency is a 10% add, already in the card. AWS says commercial In-Region and Geo CRIS prices include a 10% premium over the corresponding OpenAI rates; you do not add it again. Ultrafast short output is $330 In-Region / US Geo versus $300 Global — $30 extra per million output tokens to keep the job on the US geographic path.
OpenAI’s price list matches Global Ultrafast short at $60 / $300, and says OpenAI models on Bedrock and Azure are billed through those clouds (OpenAI API pricing — company). The Bedrock card is the invoice AWS will send.
What a speed tier is
Ultrafast is a service tier, not a second model. The model ID stays openai.gpt-6-astra. Standard remains pay-per-token with no commitment. Priority and Flex are off for Astra — no Flex discount, no Priority SKU beside Ultrafast.
Two clocks sit under that flag. Route: Mantle Ultrafast is listed in N. Virginia (us-east-1) only; Oregon Mantle is Standard only. Runtime Ultrafast is US geographic CRIS and Global CRIS, not in-Region. Quota: Runtime TPM uses a 10× burndown. OpenAI’s Ultrafast guide lists default Ultrafast TPM of 500k / 1M / 5M by usage tier, and US-or-global processing only — not EU regional endpoints (OpenAI, Ultrafast mode — company). Treat that TPM table as OpenAI’s API, not AWS’s Bedrock quota.
The second-order read
| Who | Seat | What this means |
|---|---|---|
| AWS Bedrock | Invoice owner | Collects Standard or 6× Ultrafast tokens, plus 10% on In-Region / Geo CRIS |
| OpenAI | Model provider | Paid through AWS on Bedrock; listed Ultrafast on its own API 29 Sep |
| Agent / coding apps | Who pays | Buy wait-time at 6× per token if the Standard queue is too slow |
| Standard-tier jobs | Same model, other lane | Still clear at $10 / $50 Global short — they wait |
| GPU / HBM / CoWoS | Upstream | The fast lane is still a seat on a scarce rack, not a new wafer start |
| Azure OpenAI path | Peer hyperscaler | OpenAI says Azure also bills its models; no Azure Ultrafast card in today’s AWS note |
Who gets paid. AWS on the Bedrock meter at 6× (and 10% more In-Region / US Geo). OpenAI is paid as the model behind that meter. Who pays. Interactive agents, coding assistants, and any loop that cannot sit in Standard. Upstream. The lane still needs attached GPUs, HBM, and power — the same delivery bind the hyperscalers stack already maps. Ultrafast does not print a new hall. Downstream. Token-by-token UX clears the 6× line first; batch jobs should not.
By clock. Model: Astra since 8 September. Lane: 30 September’s Ultrafast flag and 6× card. Route: N. Virginia Mantle and US/Global CRIS, not Oregon Mantle. Price: 272,000 input tokens is a hard step. Our read: the stuck step is priority inference seats. Ready is not paid. Fast is paid at six times the token.
Two cautions
Speed is “up to.” AWS attributes “up to 6×” and “up to 300 tokens per second” to OpenAI — not a Bedrock SLA. OpenAI’s Ultrafast guide currently heads the tier as “up to 8× faster than Standard,” and recommends WebSockets so network overhead does not eat the gain. Do not treat 6×, 8×, or 300 tok/s as a guaranteed Bedrock rate.
Do not mix cards. Bedrock’s In-Region / Geo rows already include the 10% OpenAI premium. Priority and Flex are unsupported here, so an Azure- or OpenAI-direct “Fast” number is not a Bedrock Astra invoice. OpenAI’s Pro 500 plan is a ChatGPT subscription, not the Bedrock API price. Long-context rates apply to the entire request once input exceeds 272k.
What would change the view
- AWS publishing a dated, measured tokens-per-second print for Ultrafast on Bedrock, with a region and model ID — replacing the OpenAI-attributed “up to.”
- A price-card revision that breaks the 6× multiple, or that turns on Priority / Flex for Astra.
- Ultrafast landing on Oregon Mantle, EU regional processing, or in-Region Runtime.
- Azure listing an equivalent Ultrafast card for GPT-6 Astra.
- Evidence that Standard-tier Astra latency has caught the Ultrafast claim, collapsing the reason to pay 6×.
What to watch
- Price card: next GPT-6 Astra model-card edit — Ultrafast rows, the 272k gate, the 10% In-Region / Geo line.
- Route: whether
us-west-2Mantle gains Ultrafast, and whether Runtime offers in-Region Ultrafast. - Quota: any published Bedrock Ultrafast TPM distinct from OpenAI’s 500k / 1M / 5M table.
- Peer: Azure’s GPT-6 Astra service-tier list after 30 Sep 2026.
- Fleet: the next AWS utilisation print — attaching seats, or only a SKU.
Read: Hyperscalers — Capex Pull and the Power Bind · What Is GPU as a Service?
Get the Wire. A free brief every day, here and in your inbox. Subscribe free →
Sources: AWS What's New, 30 Sep 2026 — company; Amazon Bedrock, GPT-6 Astra model card — company (prices, routes, 272k gate); OpenAI, DevDay 2026 Recap, 29 Sep 2026 — company; OpenAI, Ultrafast mode — company; OpenAI API pricing — company; AWS What's New, 8 Sep 2026 — company (Astra GA). Not investment advice.