← Back to blog

Let the Hyperscalers Buy the GPUs

October 5, 2026

Let the hyperscalers do that.

That was my first reaction to a discussion about GPU rental economics: aging hardware, falling rental prices, and the enormous capital required to keep buying the next generation. I never posted the reply. The conversation became a bigger question about where I would choose to build an AI business.

The money, in my eyes, is in making models useful with your data and turning that knowledge into a marketable solution. Open-weight models give us a way to own and operate more of that system. Buying a fleet of GPUs gives us a different business, with a different set of obligations.

For the domain AI I want to build, an unprepared corpus and a model endpoint are ingredients. The product comes when they solve a problem someone will pay to have solved.

The capital treadmill

GPU rental is a serious infrastructure business. Its economics reach far beyond the accelerator purchase price and the electricity bill.

Servers, networking, storage, cooling, utility interconnects, transformers, switchgear, redundant feeds, UPS systems, generators, staff, spares, and financing all belong in the calculation. Then come utilization and depreciation. Capacity that sits idle still costs money, while the hardware continues aging.

The next generation does not need to make the old one useless to change its economics. Better performance per watt, memory capacity, or throughput can make newer hardware more attractive for a customer's workload. Older equipment can keep doing valuable work while facing pressure on its rental price.

Buy capacity. Deploy it. Find customers. Recover the investment. Refresh the fleet. Repeat.

A smaller operator can succeed with the right customers or an unusual advantage. But hyperscalers bring purchasing leverage, financing, infrastructure, and a wide range of workloads to that cycle. I would rather use the capacity they build than make continuously financing hardware my primary business.

Amazon puts the question in focus

On October 2, 2026, Reuters reported, citing the Financial Times, that Amazon was discussing a structure to move roughly $8 billion of NVIDIA Grace Blackwell chips into a special-purpose vehicle funded by outside investors and lease them back. Amazon declined to comment to Reuters.

That is a reported proposal. It is not evidence of a completed transaction, a confirmed count of 100,000 GPUs, or a particular Bedrock arrangement. The eventual terms would determine the obligations and risks each party retains.

Separately, AWS and NVIDIA announced on August 26 that AWS plans to deploy two million additional NVIDIA GPUs in 2027–2028.

My interpretation is that even at Amazon's scale, how you finance compute matters alongside how you deploy it. The proposed deal does not prove the hardware is unprofitable. It does illustrate the scale of the capital question.

If that is the game, let the hyperscalers play it. I am interested in what we build with the compute.

Power is a system, not a price

Datacenters consumed substantial power before AI. Redundancy, backup generation, UPS capacity, and cooling exist because availability has to survive failures. Dense AI workloads add more demand to an already physical system.

Cheap electricity helps. It does not install another feed, increase transformer capacity, or remove heat from a rack.

Solar can contribute meaningfully, too. But annual energy production and continuous power availability are different requirements. If a workload must run through the night and bad weather, the design needs storage, grid supply, complementary generation, or some combination. The Department of Energy's explanation of solar and storage describes why storing energy changes when it can be used.

Take the idea further and imagine orbital datacenters harvesting sunlight. Even granting favorable solar exposure, the business still has launch, hardware, heat rejection, communications, maintenance, and replacement problems. Sunlight does not make the whole system free. Failed or obsolete equipment still represents invested capital.

More efficient accelerators and better energy systems can improve these economics. I welcome that progress. I do not need to wait for it to start solving a customer's knowledge problem today.

Engineering does not guarantee a return

The use case can be real while a particular investment has bad math.

A capable model and a working GPU cluster demonstrate engineering results. Whether their useful output pays for the capital, energy, financing, operations, and refresh cycle is another question.

Then there are human incentives. Investors expect returns. Optimistic assumptions about utilization, pricing, and residual hardware value can become painful when expectations meet actual cash flow. Greed is a human problem. Models, GPUs, cooling, and power are engineering problems. Better engineering does not automatically correct the incentives around it.

I do not need to claim every frontier model business is losing money to make this point. AI can be useful and transformative while people still lose substantial capital financing it poorly.

The corpus is where I would build

A proprietary corpus becomes valuable through the work around it: provenance, rights to use it, cleaning, metadata, retrieval, evaluations, human corrections, and workflows that turn knowledge into useful action.

Some facts belong in retrieval, where they can be updated and traced to a source. Some validated examples can support fine-tuning a model's behavior or skills. RAG supplies context at inference time; it does not update model weights. Saving feedback is not proof a model learned anything. A change has to be applied and measured.

My product test would be concrete: does the system perform better on the domain's actual tasks, at an acceptable cost and reliability level? Held-out evaluations and customer outcomes matter more than collecting impressive quantities of data.

The underlying model can then be a component I evaluate and replace. Migration still takes work, and open weights do not eliminate license restrictions or operating costs. But a better model or cheaper accelerator can improve the product without erasing the knowledge system around it.

That is the economic position I want: benefit when intelligence becomes cheaper.

Physics gets the final vote

There is another limit to centralizing every decision: the answer has to arrive in time.

Some healthcare applications can wait seconds. Certain medical devices, vehicles, aircraft, industrial machines, and rockets have control or safety functions with millisecond deadlines, depending on the function. Those functions need bounded responses and a safe outcome when connectivity disappears.

A remote request crosses networks, routing, queues, inference scheduling, and model execution before the response makes the return trip. More capacity can reduce queueing. Better networks can reduce overhead. Neither removes propagation delay or guarantees that every dependency stays available.

The speed of light is not getting a firmware upgrade.

An overloaded endpoint cannot become the reason a machine misses its braking, stabilization, or protection deadline. NASA's description of spacecraft onboard systems explains why spacecraft need autonomy and onboard fault protection when operating far from Earth.

My expectation is a hierarchy: sensors and validated control close to the event; specialized local intelligence where appropriate; regional and frontier models for work that can tolerate the trip. These are roles, not a requirement to send every decision through every layer.

Moving an LLM onto the device does not make it deterministic or safe. Critical control still needs validated behavior, bounded execution, and appropriate assurance. Local models can support the system within those constraints.

That creates room for specialized, deployable intelligence. It need not know everything. It needs to do its defined job well, on the hardware available, within the time allowed.

Capital, power, depreciation, and physics all shape where AI belongs. The opportunity I see today is preparing the corpus, proving the behavior, and delivering useful intelligence where the customer needs it.

Let Amazon worry about financing the next two million GPUs. I'll worry about what we teach them.