What determines the total cost of running AI GPUs?
By Nathan Benaich (Air Street Capital)
Cost framework · reviewed October 11, 2026
The GPU purchase price or hourly rental rate is only part of the cost. An owned deployment also needs servers, networking, storage, power, cooling, maintenance, software, and staff. Financing and the hardware’s eventual resale value affect the economics. For inference, compare cost per completed request or per token at the same model quality, context length, throughput, and latency requirements. A cheaper GPU-hour can cost more per useful result if the job takes longer or requires more machines.
How to read this finding
Choose a consistent comparison period and workload. For rented capacity, identify what the service price includes and add charges such as storage or data transfer where applicable. For owned capacity, distinguish purchase cash spending from depreciation expense; do not count both as separate cash costs. AWS’s Inference Recommender reports cost per hour and per inference alongside throughput and latency, illustrating why one price metric is insufficient. This index does not contain a matched cost benchmark and cannot name the cheapest GPU or cloud provider.
Company filings and hardware documentation reviewed October 11, 2026. Accounting examples use 2025 annual reports; vendor claims are attributed. Research citations do not measure profitability, and this index has no rental-price or resale-value series.
Benaich, Nathan. “What determines the total cost of running AI GPUs?” State of AI Report Compute Index. Web page updated 2026-10-11; data periods as specified above.