> ## Content Index
> Fetch the complete content index at: https://www.sotkn.com/llms.txt
> Use this file to discover other available public pages before exploring further.

# Neocloud, Explained: Who Rents GPUs When AI Demand Spikes
- URL: https://www.sotkn.com/explained/neocloud-explained/
- Published: 2026-09-12T07:56:39.000Z
- Updated: 2026-09-15T07:03:13.000Z
- Description: A neocloud is a specialist cloud that leases GPU clusters for AI work. What it is, how it earns, and why utilisation decides everything.
- Author: Sean
- Tags: Explained

A neocloud is a specialist cloud that rents GPU capacity to AI teams. It earns by billing accelerator time faster than hardware, power, and operations eat the margin.

**In short:** neoclouds buy or lease AI servers, put them in a hall with power and cooling, and bill customers for GPU time. Profit turns on two dials — utilisation, meaning how busy the fleet is, and realised price, meaning what customers actually pay after discounts.

## What is a neocloud?

A neocloud concentrates on accelerators rather than selling a broad menu of cloud services. Customers rent GPU time by the hour, by the instance, or under longer reserved contracts.

The core loop is six steps:

1. Buy or lease AI servers with GPUs
2. Place them in a hall with power, cooling, and network
3. Offer instances to customers
4. Bill for time, reservations, or committed use
5. Pay for hardware depreciation, power, space, staff, and software
6. Keep the gap if utilisation and pricing cover costs

That gap is the earn. Everything else is detail around those six steps.

## What are the key differences between a hyperscaler and a neocloud?

A hyperscaler sells hundreds of services and rents GPUs as one line among many. A neocloud sells accelerator time and very little else. That single difference drives everything else — revenue concentration, pricing power, balance-sheet shape, and what happens to each when AI demand wobbles.

|                             | Hyperscaler                                 | Neocloud                                                |
| --------------------------- | ------------------------------------------- | ------------------------------------------------------- |
| Service range               | Storage, databases, general compute, GPUs   | Accelerators, plus the minimum around them              |
| Revenue concentration       | GPUs are a slice                            | GPUs are effectively all of it                          |
| What funds expansion        | Cash from unrelated businesses              | Debt, leases, equity, prepayments                       |
| Pricing behaviour           | Anchors the market, discounts strategically | Prices against the anchor, competes on access and speed |
| Customer lock-in            | Data gravity, service breadth               | Supply certainty and price                              |
| Sensitivity to a demand dip | Absorbed by other lines                     | Hits the whole business                                 |
| Time to deploy new capacity | Slower, process-bound                       | Faster, which is the core advantage                     |

The practical reading: hyperscalers set the alternative supply. When they expand AI instances, neocloud pricing feels pressure. When their queues are long, neoclouds absorb overflow. Customers frequently use both — training on one, serving inference on the other.

Concentration cuts both ways. It is why neoclouds can move faster on a scarce GPU generation, and why a demand dip lands on them first.

## What are the top neocloud companies?

The category splits into three groups, and they earn differently enough that lumping them together produces bad reading.

| Group                          | Examples                                             | What defines them                                                      |
| ------------------------------ | ---------------------------------------------------- | ---------------------------------------------------------------------- |
| Scaled, capital-markets funded | CoreWeave, Nebius, Crusoe                            | Large committed fleets, long contracts, heavy debt or lease structures |
| Developer-first rental         | Lambda, RunPod, Vast.ai, Together AI                 | Self-serve, hourly, smaller commitments, price-competitive             |
| Sovereign and regional         | Operators built around national or regional capacity | Policy-driven demand, local data requirements                          |

The first group is where the capital-markets story lives: large contracted backlogs, and financing structures that behave like infrastructure rather than software. The second competes on price and immediacy and is closest to a commodity market. The third is growing because data-residency rules and national AI programmes create demand that cannot legally move.

Second Order does not publish buy lists. The useful question is not *which name*, but *which group*, because the three respond differently to the same scarcity. When GPUs are scarce, the first group's secured supply becomes a moat and the second group's pricing gets squeezed.

## Is Oracle a neocloud?

Not by the definition used here, though it is the closest large incumbent to one. Oracle sells a full cloud platform, so it is structurally a hyperscaler. But its AI infrastructure business behaves like a neocloud: large committed GPU contracts with a small number of AI labs, concentrated revenue, and heavy capex against future contracted demand.

That hybrid position is worth naming because it breaks the neat two-category split. A useful test is where the revenue concentration sits rather than what the company calls itself:

- If GPU contracts dominate the growth and a handful of customers dominate those contracts, it behaves like a neocloud regardless of the label
- If accelerator revenue is a slice of a broad services business, it behaves like a hyperscaler

By that test, Oracle's AI infrastructure line reads as neocloud-like inside a hyperscaler.

## What does it cost to rent a GPU?

GPU rental prices are set by scarcity of the specific accelerator, length of commitment, and whether the hall is power-ready — not by a published list rate. The same card can carry very different effective prices depending on contract length and who is asking.

Four things move a quoted GPU rental price:

**Which accelerator.** The newest generation carries a premium while supply is allocated. Previous generations fall in price as they become plentiful, which is why a fleet's average realised rate drifts down unless it keeps refreshing.

**Commitment length.** On-demand hourly is the highest rate and the least certain. Multi-month reserved commitments trade rate for predictability in both directions.

**Scale.** A large committed customer negotiates a discount ladder that never appears on a pricing page. This is why list price is a poor guide to what a provider actually earns.

**Where the capacity sits.** Power-ready halls with network in place command more than capacity that exists on a slide.

The reading rule: **separate list price from realised price.** A provider advertising a headline hourly rate while discounting heavily for the customers who fill the fleet is earning far less than the page suggests. "Sold out" without a discount ladder tells you nothing about margin.

## How does the revenue actually come in?

**On-demand rental.** Customers start and stop GPUs as needed at a posted rate. Revenue rises when many users run jobs and falls when machines sit idle.

**Reserved and committed capacity.** Customers book blocks for months, paying for access even if they do not use every hour. The neocloud gets predictable cash; the customer gets supply certainty.

**Managed extras.** Storage, networking, orchestration help, and support add fees. The base earn still rides on GPU time.

**Prepaid credits.** These pull cash forward but still convert into usage later. Accounting can differ from cash timing; the economic engine remains billed accelerator time.

## Why is utilisation the quiet boss?

Utilisation is the share of the fleet actively earning customer bills. A GPU that sits idle still burns rent, power, and capital cost.

Think of taxis. Owning cars is not the business. Busy cars with paying riders is the business.

High utilisation spreads fixed hardware cost across more revenue hours. Low utilisation turns a shiny fleet into a waiting room. What drives it:

- Demand from AI training and inference
- Mix of reserved versus on-demand use
- Reliability and support quality
- Ability to place small jobs into gaps
- Scarcity of competing capacity elsewhere

## How does pricing work, and why can it diverge from utilisation?

Prices move with how scarce the GPU type is, how long the customer commits, how ready the hall is, and what hyperscalers charge for similar instances.

The key insight is that **volume and rate are separate**. A fleet can be fully booked at soft prices. Another can run half empty at premium rates for a scarce part. Earnings depend on both.

Spot-like prices fall when fleets look empty. Reserved prices hold when scarce parts stay booked. Separate list price from realised price after discounts — a "sold out" claim means little without the discount ladder.

Contract mix matters too. Short on-demand jobs fill gaps but leave weekends soft. Multi-month reserved deals stabilise utilisation but lock prices in.

## What does it cost to run?

| Cost                   | Why it bites                                                                                   |
| ---------------------- | ---------------------------------------------------------------------------------------------- |
| Hardware               | GPUs and servers are expensive; cost spreads over useful life as depreciation or lease expense |
| Power and cooling      | Dense racks draw heavy power and make heavy heat; bills scale with use                         |
| Space and connectivity | Colocation rent, build cost, and network transit                                               |
| People and software    | Operators, support, security, billing, and the scheduler that raises utilisation               |

Think per GPU-month. Revenue is billed hours times effective rate. Cost is hardware carry plus power plus cooling share plus space plus operations share. If revenue stays above cost for long stretches, the neocloud earns.

Newer accelerator generations can age older fleets for some workloads, which is a depreciation risk as much as a competitive one.

## How does scarcity upstream affect the earn?

When finished GPUs are scarce because of memory allocation or packaging queues, neoclouds with secured supply keep fleets full and rental rates firm.

The same scarcity raises input cost and lead times, so buying the next tranche is harder. **Earn on the existing fleet can look strong while expansion stays constrained** — scarcity is simultaneously a revenue tailwind for owners of live GPUs and a brake on growth.

See [HBM explained](https://www.sotkn.com/explained/hbm-explained/) and [CoWoS explained](https://www.sotkn.com/explained/cowos-explained/) for the upstream clocks, [speed-to-power](https://www.sotkn.com/explained/speed-to-power-explained-for-investors/) for the hall clock, and [the US market hub](https://www.sotkn.com/markets/us/) for where most listed neocloud exposure sits.

## What are the failure modes?

1. Idle fleet after a demand dip or a wave of new supply
2. Power or cooling limits that prevent full clock speeds
3. A generation shift that strands older cards for top workloads
4. Customer concentration if a few buyers dominate revenue
5. Upstream delays that stop expansion while rivals catch up

Financing shapes the risk. Debt or leases speed growth but add interest and covenants, and hardware that earns less than planned becomes balance-sheet stress. Equity buys runway but does not remove the need for utilisation.

## What would change this view?

The firm-pricing read weakens if utilisation falls while fleets grow, discounting deepens on realised prices, hyperscalers expand AI instance supply aggressively, or upstream scarcity eases enough that GPU access stops being a moat.

It strengthens if scarce parts stay booked, reserved mix rises, and power-ready capacity lands on schedule.

## Practical takeaways

- The earn is billed GPU time minus hardware carry, power, cooling, space, and operations
- Utilisation is the quiet boss of unit economics
- List price is not realised price; the discount ladder is the number that matters
- Scarcity can firm rental revenue and slow fleet growth at the same time
- Ask whether GPUs are idle for lack of customers or lack of power
- Judge the group — scaled, developer-first, or sovereign — before judging the name

## FAQ

**What is a neocloud?**  
A specialist cloud provider that rents GPU and AI accelerator capacity, usually by the hour or under reserved contracts, rather than offering a broad menu of cloud services.

**What are the key differences between a hyperscaler and a neocloud?**  
A hyperscaler earns across many services with GPUs as one line; a neocloud's revenue is concentrated on accelerator time. That drives different pricing power, different financing, and very different sensitivity to an AI demand dip.

**Is Oracle a neocloud?**  
Structurally it is a hyperscaler, because it sells a full cloud platform. But its AI infrastructure business behaves like a neocloud — concentrated GPU contracts with few customers, and heavy capex against contracted demand.

**What are the top neocloud companies?**  
They divide into scaled capital-markets-funded operators such as CoreWeave and Nebius, developer-first rental platforms such as Lambda and RunPod, and sovereign or regional operators. The three respond very differently to the same scarcity.

**How does a neocloud make money?**  
It bills accelerator time and aims for utilisation high enough to cover hardware carry, power, cooling, space, and operations.

**What is utilisation?**  
The share of the fleet actively earning customer bills. Idle GPUs still cost money, so low utilisation directly hurts the earn.

**What does it cost to rent a GPU?**  
There is no single rate. Price depends on the accelerator generation, commitment length, negotiated scale discount, and whether the capacity is power-ready today.

**Why do power and cooling matter to neocloud earnings?**  
Dense GPUs draw heavy power and generate heavy heat. Those are large operating and capital lines, and without them the hardware cannot run as rentable capacity.