What Is GPU as a Service? Billed Accelerator Hours for AI Work
GPU as a service is billed accelerator time: a cloud operator puts GPUs in a power-ready hall and sells hours — on-demand, reserved, or as a dated capacity block — instead of the customer buying the cards.
EXPLAINED · Mechanism · Last reviewed: 24 Sep 2026 · Next review: after next hyperscaler AI-cloud earnings or a GPU instance / Capacity Blocks price-list update
GPU as a service is billed accelerator time: a cloud operator puts GPUs in a power-ready hall and sells hours — on-demand, reserved, or as a dated capacity block — instead of the customer buying the cards.
In short. GPU as a service is how the hyperscalers layer turns CapEx into hours. The stuck step is not the brochure SKU. It is utilisation — whether the card is earning — and whether the hall is energised. List price is not realised price.
Who gets paid when AI teams buy hours instead of silicon, and which clock you are watching — list rate, reserved block, or megawatts that can light up. Compare the specialist seller in what a neocloud is and the demand dial in what a hyperscaler is.
What is GPU as a service?
GPU as a service (GaaS) is billed accelerator capacity. The operator owns or leases the servers, places them in a hall with power, cooling, and fabric, and bills the customer for time on those GPUs. The customer gets accelerator hours without buying the cards, the CDU, or the transformer.
The product is the hour, not the brand on the tray. A hyperscaler sells that hour as one line among storage, databases, and general compute. A neocloud sells little else. Same accelerators can appear in both stories. Who gets paid, and what a demand dip hits, are different — see neocloud vs hyperscaler.
| Stage | Role | Read |
|---|---|---|
| Install | Finished GPUs in a hall | Hardware from packaging + HBM |
| Energise | Watts + cooling | Hours are not billable until the rack is live |
| Offer | On-demand, reserved, or a block | Access shape, not just the SKU |
| Earn | Billed hours | Utilisation × realised price |
Oracle’s Q1 FY27 print is the dated hyperscaler example. On 10 September 2026 it said Cloud Infrastructure revenue rose 121% to US$7.4 billion, that it delivered 850 MW of additional datacenter capacity and more than 300,000 GPUs since the end of Q4, booked more than US$30 billion of new AI cloud contracts, and lifted remaining performance obligations to US$664 billion (Oracle / PR Newswire, 10 Sep 2026). Capacity delivered is not hours billed. The slides add the utilisation clock: 97.9% in the quarter, and GPUs up for renewal repriced about 20% above prior contracts (Oracle Q1 FY27 slides, 10 Sep 2026 — company-reported). A fleet that stays nearly full after a record add means the short clock is still delivery. Wire: Oracle 850 MW — fleet still 97.9% full.
How do you buy a GPU hour?
Three commercial shapes matter. They are not the same product just because they share a GPU name.
On-demand. Start and stop when you need the card. The posted rate is a starting point. Large buyers negotiate off the page.
Reserved / committed. Book a block for months. You pay for access even if you do not use every hour. The operator gets a more predictable cash line; you get supply certainty. This is how most scaled fleets fill.
Capacity blocks. A dated reservation for a cluster that starts on a future day and ends on a known day. Amazon EC2 Capacity Blocks for ML let you reserve accelerated instances in UltraClusters for up to six months, in sizes from one to 64 instances (512 GPUs or 1,024 Trainium chips), up to eight weeks ahead (AWS product page, accessed 24 Sep 2026). You pay a reservation fee up front; the price is set at purchase. AWS says the current reservation rates are scheduled to update next in October 2026 (AWS pricing, accessed 24 Sep 2026).
| Shape | What you buy | Cash clock |
|---|---|---|
| On-demand | Hours you start and stop | Rate × used hours |
| Reserved / committed | Access for a term | Term fee whether or not every hour runs |
| Capacity block | A dated cluster window | Up-front reservation + OS while running |
Microsoft’s ND H100 v5 series is the hyperscaler instance shape: one VM, eight NVIDIA H100 80 GB GPUs (Microsoft Learn, accessed 24 Sep 2026). That is a SKU definition, not a price. You buy the instance hour, not the tray.
What does a GPU hour cost — and why is list not the earn?
There is no single GPU-hour price. The same card carries different effective rates by commitment, region, and whether the hall is live today.
| Snapshot (list / reservation table) | Rate | As at | Source type |
|---|---|---|---|
| AWS Capacity Blocks p5.48xlarge (8× H100), US East (N. Virginia) | US$5.191 per accelerator-hour | Accessed 24 Sep 2026 | AWS pricing table |
| AWS Capacity Blocks p6-b300.48xlarge (8× B300), US East (N. Virginia) | US$14.04 per accelerator-hour | Accessed 24 Sep 2026 | AWS pricing table |
| Nebius NVIDIA HGX H100 on-demand | US$3.85 now → US$4.50 from 1 Oct 2026 (+17%) | Price list checked 23 Sep 2026 | Company list |
| Nebius NVIDIA HGX B300 on-demand | US$7.85 now → US$9.50 from 1 Oct 2026 (+21%) | Same | Company list |
AWS’s table is a reservation rate that updates with supply and demand; the next scheduled update is October 2026. Nebius’s October column is a posted on-demand hike — the second in about three months (Reuters, 17 Sep 2026, reported). Commitment discounts still apply; Nebius’s own list still offers up to 35% off for reserved clusters (Nebius price list, checked 23 Sep 2026).
List price ≠ realised price. A hyperscaler can discount inside a multi-year cloud deal. A neocloud can post a rise and still fill most of the book on older committed rates. The earn is billed hours × the rate after discounts, not the landing-page number.
A second unit has shown up next to the hour: revenue per megawatt. Power, not chip count, is the scarce envelope once the cards exist. Nebius’s Q2 2026 letter said its four largest new deals averaged more than US$1 billion each at US$20–25 million of annual revenue per megawatt, with a short-term opportunity of US$40–50 million per megawatt (shareholder letter, 12 Aug 2026). CoreWeave said it signed short-dated Q3 deals of about three to six months at approximately US$40 million per megawatt of annualised revenue (Business Wire, 17 Sep 2026). At US$20 million per megawatt, a 100 MW site is about US$2 billion a year; at US$40 million, about US$4 billion — our arithmetic on those company figures. The chips barely change. The price of the hour does.
Why does utilisation bind?
Utilisation is the share of the installed fleet that is actually earning a customer bill. A GPU that sits idle still burns depreciation or lease cost, hall space, and a share of power. High utilisation spreads that carry across more revenue hours. Low utilisation turns a fleet into a waiting room.
Oracle’s 97.9% utilisation after adding 850 MW is the hyperscaler version of that fact (earnings slides, 10 Sep 2026). The new capacity was absorbed, not parked. The opposite print — capacity arriving while utilisation falls — would mean hours are no longer scarce.
Upstream, CoWoS and HBM decide when the next tray can ship. Downstream, speed-to-power decides when a delivered tray becomes a billable hour. GaaS sits in the middle.
Who gets paid — and where is the stuck step?
On the hyperscalers layer, who gets paid tracks which step clears first when AI teams want hours.
| Step | Who | Why it matters |
|---|---|---|
| Demand dial | Hyperscalers (and specialist neoclouds) | CapEx and fleet adds set how many hours can exist |
| Stuck step | Utilisation + energised MW | Idle cards still cost; contracted MW is not live MW |
| Upstream feed | Packaging, HBM, servers | No finished GPU, no hour to sell |
| Power + cooling | Halls, interconnect, CDUs | Watts make hours billable |
| Customers | AI labs and platforms | Pay for access certainty and speed |
Cash sticks at the operator when the fleet is full and realised prices hold — Oracle’s 97.9% utilisation and ~20% renewal premium on older GPUs is one dated print of that (company-reported, 10 Sep 2026). Cash sticks at the customer when they prepaid for a block that has not energised yet. Cash sticks upstream when CoWoS slots or HBM stacks bind and the next tray slips.
The US market door is where most public GaaS exposure sits: US speed-to-power when halls wait on watts, and CoreWeave neocloud contracts when the specialist book is the story. China is the other hyperscaler-layer door when domestic hours cannot use the same stack — China HBM access.
A useful test: if GPU contracts dominate growth, the GaaS line behaves like a neocloud even inside a hyperscaler brand. If accelerator revenue is a slice of a broad book, a soft hour is absorbed elsewhere. That is a concentration test, not a ticker call.
Common misconceptions
- People assume GPU as a service is a neocloud-only product. Actually hyperscalers sell the same hour as one line among many. The mechanism is billed time. The risk shape changes with revenue concentration — see neocloud vs hyperscaler.
- People assume the list GPU-hour is what the operator earns. Actually realised price after discounts, reserved mix, and utilisation decide the earn. A posted hike can sit next to older committed rates that still fill most of the book.
- People assume more delivered megawatts means more profit this quarter. Actually installs and billed hours run on different clocks. A hall can be contracted and still dark.
- People assume a capacity block is the same as on-demand. Actually a block is a dated window you pay for up front. AWS sets the reservation price at purchase; the public table can move later without changing a block you already bought.
Not investment advice. Do your own research.
FAQ
What is GPU as a service?
GPU as a service is billed accelerator capacity. A cloud operator installs GPUs in a power-ready hall and bills customers for hours — on-demand, reserved, or as a dated capacity block — instead of the customer buying the cards.
How is GPU as a service different from owning GPUs?
Owning puts hardware carry, power, cooling, and obsolescence on the buyer’s book. GaaS moves those lines to the operator. The operator gets paid only when hours bill at a realised rate that covers carry.
What is the difference between on-demand, reserved, and a capacity block?
On-demand is start-and-stop. Reserved capacity is a term of access you pay for even if you skip hours. A capacity block is a dated cluster window — AWS Capacity Blocks can be reserved up to eight weeks ahead for up to six months (accessed 24 Sep 2026).
What does a GPU hour cost?
There is no single rate. As at 24 September 2026, AWS listed Capacity Blocks H100 at US$5.191 per accelerator-hour in US East (N. Virginia). Nebius’s public on-demand H100 moves from US$3.85 to US$4.50 on 1 October 2026. List is not realised price.
Why does utilisation matter?
It is the share of the fleet earning a bill. Idle GPUs still cost power and capital. Oracle reported 97.9% utilisation in Q1 FY27 after delivering 850 MW (company materials, 10 Sep 2026).
Who gets paid when GPU hours bind?
Operators with energised, uncommitted hours get paid first. Upstream, packaging and HBM get paid when the next tray cannot ship. The stuck step is utilisation plus live megawatts, not the brochure SKU.
Is GPU as a service only something neoclouds sell?
No. Hyperscalers sell GPU instances as one line in a broad cloud. Neoclouds concentrate on the same hour. The useful test is revenue concentration, not the company label.
Related terms
- Neocloud — specialist seller of accelerator hours.
- Hyperscaler — the CapEx dial that sets how many hours can exist.
- Neocloud vs hyperscaler — same GPUs, different who gets paid.
- HBM — memory that must match the GPU before an hour can be sold.
- CoWoS — packaging attach that can hold the next tray.
Markets and stack doors
Layer: hyperscalers. Sibling layer: neoclouds. Markets: US speed-to-power · CoreWeave neocloud contracts · China HBM access.
Get the Wire — short notes when GPU hours or hall clocks move. Wire: Oracle 850 MW · Nebius prices. Analysis: where the stuck step moves.
How we check this
Figures come from company filings and government releases, each dated in the list below. Capacity and timing from press reports are labelled as reported, not guided. Where we work something out ourselves, the arithmetic is shown in full.
Last reviewed: 24 Sep 2026 · Next review: after next hyperscaler AI-cloud earnings or a GPU instance / Capacity Blocks price-list update
Spot an error? Tell us and we will correct it and note the change here.
Sources
- Oracle / PR Newswire — Q1 FY27; IaaS US$7.4B (+121%); 850 MW; >300,000 GPUs; RPO US$664B (10 Sep 2026)
- Oracle Q1 FY27 slides — 97.9% utilisation; ~20% GPU renewal premium (10 Sep 2026)
- AWS — EC2 Capacity Blocks for ML (accessed 24 Sep 2026)
- AWS — Capacity Blocks pricing; H100 US$5.191/h (N. Virginia); B300 US$14.04/h; next table update Oct 2026 (accessed 24 Sep 2026)
- Microsoft Learn — Azure ND H100 v5; 8× H100 80 GB (accessed 24 Sep 2026)
- Reuters — Nebius on-demand GPU prices +17–21% from 1 Oct 2026 (17 Sep 2026, reported)
- Nebius — GPU price list; H100 US$3.85 → US$4.50; B300 US$7.85 → US$9.50 (checked 23 Sep 2026)
- Nebius Q2 2026 letter — deals >US$1B TCV at US$20–25m/MW (12 Aug 2026)
- CoreWeave via Business Wire — ~US$40m/MW short-dated Q3 deals (17 Sep 2026)