> ## Content Index
> Fetch the complete content index at: https://www.sotkn.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# AWS lists a 6× Ultrafast lane on Bedrock — speed is the cash print
- URL: https://www.sotkn.com/wire/aws-bedrock-gpt6-astra-ultrafast-wire/
- Published: 2026-09-30T23:51:20.000Z
- Updated: 2026-09-30T23:51:20.000Z
- Description: AWS today listed Ultrafast for GPT-6 Astra on Bedrock with a published 6× price card. The cash print is a scarce inference lane — not a new model.
- Author: Sean
- Tags: #desk-wire

`WIRE` · `HYPERSCALERS` · `US` · Data as at 30 Sep 2026

Amazon Web Services today listed **Ultrafast** as a paid speed tier for OpenAI **GPT-6 Astra** on Amazon Bedrock. The model has been generally available since 8 September. What printed on 30 September is a **6× price card** on a scarce inference lane — not a new campus. On the [hyperscalers layer](https://www.sotkn.com/stack/hyperscalers/) the stuck step is who can sell **priority tokens** while the fleet stays full.

## What happened

On 30 September 2026 AWS said Ultrafast mode for GPT-6 Astra is available on Amazon Bedrock — a **premium speed tier** for real-time coding assistants, interactive agents, and customer-facing apps. Speed is attributed to OpenAI: **up to 6×** faster inference in the API, **up to 300 tokens per second** ([AWS What's New, 30 Sep 2026](https://aws.amazon.com/about-aws/whats-new/2026/09/openai-gpt-6-astra-ultrafast-on-amazon-bedrock/?ref=sotkn.com) — company).

The Bedrock model card is the cash document. Ultrafast is **six times** the corresponding Standard prices, requested with `"service_tier": "ultrafast"`. GPT-6 Astra launched on Bedrock on **8 September 2026** (1.05M context, 128k max output, knowledge cutoff **30 April 2026**). Priority and Flex are **not supported** for this model ([Amazon Bedrock, GPT-6 Astra model card](https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-openai-gpt-6-astra.html?ref=sotkn.com) — company).

OpenAI opened the same lane on its own API the day before, at DevDay on **29 September 2026** — GPT-6 Astra Ultrafast in the API, and in ChatGPT Work and Codex on Pro 500 and Enterprise ([OpenAI, DevDay 2026 Recap, 29 Sep 2026](https://openai.com/index/devday-2026-recap?ref=sotkn.com) — company). AWS’s 30 September note is the hyperscaler attach, billed through AWS.

| Print                               | Figure                                  | Date             |
| ----------------------------------- | --------------------------------------- | ---------------- |
| Event                               | Ultrafast on Bedrock for GPT-6 Astra    | AWS, 30 Sep 2026 |
| Claimed speed (OpenAI, via AWS)     | Up to 6× API; up to 300 tok/s           | Same             |
| Model GA on Bedrock                 | 8 Sep 2026                              | Model card       |
| Standard Global, short (≤272k in)   | $10 in · $50 out per 1M tokens          | Same             |
| Ultrafast Global, short             | $60 in · $300 out (6×)                  | Same             |
| Ultrafast In-Region / US Geo, short | $66 in · $330 out (10% over OpenAI)     | Same             |
| Long-context gate                   | Whole request reprices above 272k input | Same             |
| Ultrafast Global, long              | $120 in · $450 out                      | Same             |

All Bedrock figures are USD per **1 million tokens**. Short-context rates apply at **272,000 input tokens or fewer**; above that, long-context rates apply to the **entire request**.

## Why cash cares

Ultrafast does not add a megawatt. It **reprices the same model** so a customer who needs the next token *now* pays six times Standard. Who gets paid is the **priority lane** — AWS on the Bedrock invoice, OpenAI billed through AWS. Who pays is the app that cannot wait in the Standard queue. See [What Is GPU as a Service?](https://www.sotkn.com/explained/gpu-as-a-service-explained/); this print is the token-meter version of the same seat.

Our arithmetic, from the model card, not from a run:

- **Six times the token.** Ultrafast Global short output is $300 per 1M versus Standard $50 — a **$250** delta per million output tokens. If the OpenAI-attributed **6×** token rate actually prints, a 1M-output job finishes in one-sixth the wall-clock and still costs six times Standard. You buy wait-time, not a discount.
- **The 272k cliff reprices the whole book.** A 272,000-input request at Ultrafast Global short is 272,000 × $60 / 1,000,000 = **$16.32**. One extra token trips long-context rates on the entire request: 272,001 × $120 / 1,000,000 = **$32.64**. Output also steps from $300 to $450 per million on that whole request.
- **Residency is a 10% add, already in the card.** AWS says commercial In-Region and Geo CRIS prices **include a 10% premium** over the corresponding OpenAI rates; you do not add it again. Ultrafast short output is $330 In-Region / US Geo versus $300 Global — **$30** extra per million output tokens to keep the job on the US geographic path.

OpenAI’s price list matches Global Ultrafast short at $60 / $300, and says OpenAI models on Bedrock and Azure are billed through those clouds ([OpenAI API pricing](https://developers.openai.com/api/docs/pricing?ref=sotkn.com) — company). The Bedrock card is the invoice AWS will send.

## What a speed tier is

Ultrafast is a **service tier**, not a second model. The model ID stays `openai.gpt-6-astra`. Standard remains pay-per-token with no commitment. Priority and Flex are **off** for Astra — no Flex discount, no Priority SKU beside Ultrafast.

Two clocks sit under that flag. **Route:** Mantle Ultrafast is listed in **N. Virginia** (`us-east-1`) only; Oregon Mantle is Standard only. Runtime Ultrafast is US geographic CRIS and Global CRIS, not in-Region. **Quota:** Runtime TPM uses a **10× burndown**. OpenAI’s Ultrafast guide lists default Ultrafast TPM of 500k / 1M / 5M by usage tier, and US-or-global processing only — not EU regional endpoints ([OpenAI, Ultrafast mode](https://developers.openai.com/api/docs/guides/ultrafast-mode?ref=sotkn.com) — company). Treat that TPM table as OpenAI’s API, not AWS’s Bedrock quota.

## The second-order read

| Who                 | Seat                   | What this means                                                                      |
| ------------------- | ---------------------- | ------------------------------------------------------------------------------------ |
| AWS Bedrock         | Invoice owner          | Collects Standard or 6× Ultrafast tokens, plus 10% on In-Region / Geo CRIS           |
| OpenAI              | Model provider         | Paid through AWS on Bedrock; listed Ultrafast on its own API 29 Sep                  |
| Agent / coding apps | Who pays               | Buy wait-time at 6× per token if the Standard queue is too slow                      |
| Standard-tier jobs  | Same model, other lane | Still clear at $10 / $50 Global short — they wait                                    |
| GPU / HBM / CoWoS   | Upstream               | The fast lane is still a seat on a scarce rack, not a new wafer start                |
| Azure OpenAI path   | Peer hyperscaler       | OpenAI says Azure also bills its models; no Azure Ultrafast card in today’s AWS note |

**Who gets paid.** AWS on the Bedrock meter at 6× (and 10% more In-Region / US Geo). OpenAI is paid as the model behind that meter. **Who pays.** Interactive agents, coding assistants, and any loop that cannot sit in Standard. **Upstream.** The lane still needs attached GPUs, HBM, and power — the same delivery bind the [hyperscalers stack](https://www.sotkn.com/stack/hyperscalers/) already maps. Ultrafast does not print a new hall. **Downstream.** Token-by-token UX clears the 6× line first; batch jobs should not.

**By clock.** Model: Astra since 8 September. Lane: 30 September’s Ultrafast flag and 6× card. Route: N. Virginia Mantle and US/Global CRIS, not Oregon Mantle. Price: 272,000 input tokens is a hard step. Our read: the stuck step is **priority inference seats**. Ready is not paid. Fast is paid at six times the token.

## Two cautions

Speed is **“up to.”** AWS attributes “up to 6×” and “up to 300 tokens per second” to OpenAI — not a Bedrock SLA. OpenAI’s Ultrafast guide currently heads the tier as “up to 8× faster than Standard,” and recommends WebSockets so network overhead does not eat the gain. Do not treat 6×, 8×, or 300 tok/s as a guaranteed Bedrock rate.

Do not mix cards. Bedrock’s In-Region / Geo rows already include the 10% OpenAI premium. Priority and Flex are **unsupported** here, so an Azure- or OpenAI-direct “Fast” number is not a Bedrock Astra invoice. OpenAI’s Pro 500 plan is a ChatGPT subscription, not the Bedrock API price. Long-context rates apply to the **entire** request once input exceeds 272k.

## What would change the view

- AWS publishing a dated, measured tokens-per-second print for Ultrafast on Bedrock, with a region and model ID — replacing the OpenAI-attributed “up to.”
- A price-card revision that breaks the 6× multiple, or that turns on Priority / Flex for Astra.
- Ultrafast landing on Oregon Mantle, EU regional processing, or in-Region Runtime.
- Azure listing an equivalent Ultrafast card for GPT-6 Astra.
- Evidence that Standard-tier Astra latency has caught the Ultrafast claim, collapsing the reason to pay 6×.

## What to watch

- **Price card:** next GPT-6 Astra model-card edit — Ultrafast rows, the 272k gate, the 10% In-Region / Geo line.
- **Route:** whether `us-west-2` Mantle gains Ultrafast, and whether Runtime offers in-Region Ultrafast.
- **Quota:** any published Bedrock Ultrafast TPM distinct from OpenAI’s 500k / 1M / 5M table.
- **Peer:** Azure’s GPT-6 Astra service-tier list after 30 Sep 2026.
- **Fleet:** the next AWS utilisation print — attaching seats, or only a SKU.

Read: [Hyperscalers — Capex Pull and the Power Bind](https://www.sotkn.com/stack/hyperscalers/) · [What Is GPU as a Service?](https://www.sotkn.com/explained/gpu-as-a-service-explained/)

---

**Get the Wire.** A free brief every day, here and in your inbox. **[Subscribe free →](https://www.sotkn.com/#/portal/signup)**

Sources: [AWS What's New, 30 Sep 2026](https://aws.amazon.com/about-aws/whats-new/2026/09/openai-gpt-6-astra-ultrafast-on-amazon-bedrock/?ref=sotkn.com) — company; [Amazon Bedrock, GPT-6 Astra model card](https://docs.aws.amazon.com/bedrock/latest/userguide/model-card-openai-gpt-6-astra.html?ref=sotkn.com) — company (prices, routes, 272k gate); [OpenAI, DevDay 2026 Recap, 29 Sep 2026](https://openai.com/index/devday-2026-recap?ref=sotkn.com) — company; [OpenAI, Ultrafast mode](https://developers.openai.com/api/docs/guides/ultrafast-mode?ref=sotkn.com) — company; [OpenAI API pricing](https://developers.openai.com/api/docs/pricing?ref=sotkn.com) — company; [AWS What's New, 8 Sep 2026](https://aws.amazon.com/about-aws/whats-new/2026/09/openai-gpt-6-astra-on-amazon-bedrock/?ref=sotkn.com) — company (Astra GA). Not investment advice.