What Is GPU as a Service? Billed Accelerator Hours for AI Work

GPU as a service is billed accelerator time: a cloud operator puts GPUs in a power-ready hall and sells hours — on-demand, reserved, or as a dated capacity block — instead of the customer buying the cards.

Share
What Is GPU as a Service? Billed Accelerator Hours for AI Work

EXPLAINED · Mechanism · Last reviewed: 24 Sep 2026 · Next review: after next hyperscaler AI-cloud earnings or a GPU instance / Capacity Blocks price-list update

GPU as a service is billed accelerator time: a cloud operator puts GPUs in a power-ready hall and sells hours — on-demand, reserved, or as a dated capacity block — instead of the customer buying the cards.

In short. GPU as a service is how the hyperscalers layer turns CapEx into hours. The stuck step is not the brochure SKU. It is utilisation — whether the card is earning — and whether the hall is energised. List price is not realised price.

Who gets paid when AI teams buy hours instead of silicon, and which clock you are watching — list rate, reserved block, or megawatts that can light up. Compare the specialist seller in what a neocloud is and the demand dial in what a hyperscaler is.

What is GPU as a service?

GPU as a service (GaaS) is billed accelerator capacity. The operator owns or leases the servers, places them in a hall with power, cooling, and fabric, and bills the customer for time on those GPUs. The customer gets accelerator hours without buying the cards, the CDU, or the transformer.

The product is the hour, not the brand on the tray. A hyperscaler sells that hour as one line among storage, databases, and general compute. A neocloud sells little else. Same accelerators can appear in both stories. Who gets paid, and what a demand dip hits, are different — see neocloud vs hyperscaler.

Stage Role Read
Install Finished GPUs in a hall Hardware from packaging + HBM
Energise Watts + cooling Hours are not billable until the rack is live
Offer On-demand, reserved, or a block Access shape, not just the SKU
Earn Billed hours Utilisation × realised price

Oracle’s Q1 FY27 print is the dated hyperscaler example. On 10 September 2026 it said Cloud Infrastructure revenue rose 121% to US$7.4 billion, that it delivered 850 MW of additional datacenter capacity and more than 300,000 GPUs since the end of Q4, booked more than US$30 billion of new AI cloud contracts, and lifted remaining performance obligations to US$664 billion (Oracle / PR Newswire, 10 Sep 2026). Capacity delivered is not hours billed. The slides add the utilisation clock: 97.9% in the quarter, and GPUs up for renewal repriced about 20% above prior contracts (Oracle Q1 FY27 slides, 10 Sep 2026 — company-reported). A fleet that stays nearly full after a record add means the short clock is still delivery. Wire: Oracle 850 MW — fleet still 97.9% full.

How do you buy a GPU hour?

Three commercial shapes matter. They are not the same product just because they share a GPU name.

On-demand. Start and stop when you need the card. The posted rate is a starting point. Large buyers negotiate off the page.

Reserved / committed. Book a block for months. You pay for access even if you do not use every hour. The operator gets a more predictable cash line; you get supply certainty. This is how most scaled fleets fill.

Capacity blocks. A dated reservation for a cluster that starts on a future day and ends on a known day. Amazon EC2 Capacity Blocks for ML let you reserve accelerated instances in UltraClusters for up to six months, in sizes from one to 64 instances (512 GPUs or 1,024 Trainium chips), up to eight weeks ahead (AWS product page, accessed 24 Sep 2026). You pay a reservation fee up front; the price is set at purchase. AWS says the current reservation rates are scheduled to update next in October 2026 (AWS pricing, accessed 24 Sep 2026).

Shape What you buy Cash clock
On-demand Hours you start and stop Rate × used hours
Reserved / committed Access for a term Term fee whether or not every hour runs
Capacity block A dated cluster window Up-front reservation + OS while running

Microsoft’s ND H100 v5 series is the hyperscaler instance shape: one VM, eight NVIDIA H100 80 GB GPUs (Microsoft Learn, accessed 24 Sep 2026). That is a SKU definition, not a price. You buy the instance hour, not the tray.

What does a GPU hour cost — and why is list not the earn?

There is no single GPU-hour price. The same card carries different effective rates by commitment, region, and whether the hall is live today.

Snapshot (list / reservation table) Rate As at Source type
AWS Capacity Blocks p5.48xlarge (8× H100), US East (N. Virginia) US$5.191 per accelerator-hour Accessed 24 Sep 2026 AWS pricing table
AWS Capacity Blocks p6-b300.48xlarge (8× B300), US East (N. Virginia) US$14.04 per accelerator-hour Accessed 24 Sep 2026 AWS pricing table
Nebius NVIDIA HGX H100 on-demand US$3.85 now → US$4.50 from 1 Oct 2026 (+17%) Price list checked 23 Sep 2026 Company list
Nebius NVIDIA HGX B300 on-demand US$7.85 now → US$9.50 from 1 Oct 2026 (+21%) Same Company list

AWS’s table is a reservation rate that updates with supply and demand; the next scheduled update is October 2026. Nebius’s October column is a posted on-demand hike — the second in about three months (Reuters, 17 Sep 2026, reported). Commitment discounts still apply; Nebius’s own list still offers up to 35% off for reserved clusters (Nebius price list, checked 23 Sep 2026).

List price ≠ realised price. A hyperscaler can discount inside a multi-year cloud deal. A neocloud can post a rise and still fill most of the book on older committed rates. The earn is billed hours × the rate after discounts, not the landing-page number.

A second unit has shown up next to the hour: revenue per megawatt. Power, not chip count, is the scarce envelope once the cards exist. Nebius’s Q2 2026 letter said its four largest new deals averaged more than US$1 billion each at US$20–25 million of annual revenue per megawatt, with a short-term opportunity of US$40–50 million per megawatt (shareholder letter, 12 Aug 2026). CoreWeave said it signed short-dated Q3 deals of about three to six months at approximately US$40 million per megawatt of annualised revenue (Business Wire, 17 Sep 2026). At US$20 million per megawatt, a 100 MW site is about US$2 billion a year; at US$40 million, about US$4 billion — our arithmetic on those company figures. The chips barely change. The price of the hour does.

Why does utilisation bind?

Utilisation is the share of the installed fleet that is actually earning a customer bill. A GPU that sits idle still burns depreciation or lease cost, hall space, and a share of power. High utilisation spreads that carry across more revenue hours. Low utilisation turns a fleet into a waiting room.

Oracle’s 97.9% utilisation after adding 850 MW is the hyperscaler version of that fact (earnings slides, 10 Sep 2026). The new capacity was absorbed, not parked. The opposite print — capacity arriving while utilisation falls — would mean hours are no longer scarce.

Upstream, CoWoS and HBM decide when the next tray can ship. Downstream, speed-to-power decides when a delivered tray becomes a billable hour. GaaS sits in the middle.

Who gets paid — and where is the stuck step?

On the hyperscalers layer, who gets paid tracks which step clears first when AI teams want hours.

Step Who Why it matters
Demand dial Hyperscalers (and specialist neoclouds) CapEx and fleet adds set how many hours can exist
Stuck step Utilisation + energised MW Idle cards still cost; contracted MW is not live MW
Upstream feed Packaging, HBM, servers No finished GPU, no hour to sell
Power + cooling Halls, interconnect, CDUs Watts make hours billable
Customers AI labs and platforms Pay for access certainty and speed

Cash sticks at the operator when the fleet is full and realised prices hold — Oracle’s 97.9% utilisation and ~20% renewal premium on older GPUs is one dated print of that (company-reported, 10 Sep 2026). Cash sticks at the customer when they prepaid for a block that has not energised yet. Cash sticks upstream when CoWoS slots or HBM stacks bind and the next tray slips.

The US market door is where most public GaaS exposure sits: US speed-to-power when halls wait on watts, and CoreWeave neocloud contracts when the specialist book is the story. China is the other hyperscaler-layer door when domestic hours cannot use the same stack — China HBM access.

A useful test: if GPU contracts dominate growth, the GaaS line behaves like a neocloud even inside a hyperscaler brand. If accelerator revenue is a slice of a broad book, a soft hour is absorbed elsewhere. That is a concentration test, not a ticker call.

Common misconceptions

  • People assume GPU as a service is a neocloud-only product. Actually hyperscalers sell the same hour as one line among many. The mechanism is billed time. The risk shape changes with revenue concentration — see neocloud vs hyperscaler.
  • People assume the list GPU-hour is what the operator earns. Actually realised price after discounts, reserved mix, and utilisation decide the earn. A posted hike can sit next to older committed rates that still fill most of the book.
  • People assume more delivered megawatts means more profit this quarter. Actually installs and billed hours run on different clocks. A hall can be contracted and still dark.
  • People assume a capacity block is the same as on-demand. Actually a block is a dated window you pay for up front. AWS sets the reservation price at purchase; the public table can move later without changing a block you already bought.

Not investment advice. Do your own research.

FAQ

What is GPU as a service?
GPU as a service is billed accelerator capacity. A cloud operator installs GPUs in a power-ready hall and bills customers for hours — on-demand, reserved, or as a dated capacity block — instead of the customer buying the cards.

How is GPU as a service different from owning GPUs?
Owning puts hardware carry, power, cooling, and obsolescence on the buyer’s book. GaaS moves those lines to the operator. The operator gets paid only when hours bill at a realised rate that covers carry.

What is the difference between on-demand, reserved, and a capacity block?
On-demand is start-and-stop. Reserved capacity is a term of access you pay for even if you skip hours. A capacity block is a dated cluster window — AWS Capacity Blocks can be reserved up to eight weeks ahead for up to six months (accessed 24 Sep 2026).

What does a GPU hour cost?
There is no single rate. As at 24 September 2026, AWS listed Capacity Blocks H100 at US$5.191 per accelerator-hour in US East (N. Virginia). Nebius’s public on-demand H100 moves from US$3.85 to US$4.50 on 1 October 2026. List is not realised price.

Why does utilisation matter?
It is the share of the fleet earning a bill. Idle GPUs still cost power and capital. Oracle reported 97.9% utilisation in Q1 FY27 after delivering 850 MW (company materials, 10 Sep 2026).

Who gets paid when GPU hours bind?
Operators with energised, uncommitted hours get paid first. Upstream, packaging and HBM get paid when the next tray cannot ship. The stuck step is utilisation plus live megawatts, not the brochure SKU.

Is GPU as a service only something neoclouds sell?
No. Hyperscalers sell GPU instances as one line in a broad cloud. Neoclouds concentrate on the same hour. The useful test is revenue concentration, not the company label.

  • Neocloud — specialist seller of accelerator hours.
  • Hyperscaler — the CapEx dial that sets how many hours can exist.
  • Neocloud vs hyperscaler — same GPUs, different who gets paid.
  • HBM — memory that must match the GPU before an hour can be sold.
  • CoWoS — packaging attach that can hold the next tray.

Markets and stack doors

Layer: hyperscalers. Sibling layer: neoclouds. Markets: US speed-to-power · CoreWeave neocloud contracts · China HBM access.

Get the Wire — short notes when GPU hours or hall clocks move. Wire: Oracle 850 MW · Nebius prices. Analysis: where the stuck step moves.

How we check this

Figures come from company filings and government releases, each dated in the list below. Capacity and timing from press reports are labelled as reported, not guided. Where we work something out ourselves, the arithmetic is shown in full.

Last reviewed: 24 Sep 2026 · Next review: after next hyperscaler AI-cloud earnings or a GPU instance / Capacity Blocks price-list update

Spot an error? Tell us and we will correct it and note the change here.

Our editorial standards →

Sources

Sean

Writes Second Order by TKN — a plain-English map of AI infrastructure and semiconductors for non-specialist investors. Focus: who gets paid, where money sticks, and which step is stuck.

Author page → · Editorial standards →