AI Capex vs Opex Calculator: Owned GPUs vs API Subscriptions Over 3 Years
This free AI capex vs opex calculator compares the 3-year total cost of owning GPU hardware and operations staff against paying for the equivalent workload through an API or SaaS subscription. Enter your GPU count, hardware price, infrastructure overhead, operations staffing, and the API cost of an equivalent workload, and the tool returns a clear 3-year savings or premium figure for owning versus subscribing. This is the decision most CIOs actually face once a pilot proves value: keep paying a growing API bill indefinitely, or make a capital commitment that only pays off at real, sustained utilization.
Your numbers
H100 80GB street price runs roughly $25,000-$32,000 in 2026.
Additional capex on top of raw GPU cost for servers, networking, storage, and rack space.
What the same volume of inference would cost today through a hosted API or SaaS AI subscription.
Your results
Planning estimate only, excluding depreciation tax treatment and hardware residual value. Real GPU utilization below 50-60% shifts the economics meaningfully toward the API option.
Get your full capex vs opex model
We will email you a personalized 3-year cost comparison built from your real workload volume and utilization forecast, and a Netray infrastructure architect will follow up with a 30-minute review.
No spam. Your results stay private. Unsubscribe anytime.
The real cost of owning is more than the GPU price
The GPU purchase price is only the starting point of capex, not the total. Networking, storage, and rack or facility overhead typically add 20-40% on top of raw hardware cost, and that number climbs toward the high end for a new buildout rather than spare capacity in an existing datacenter. On the opex side, dedicated operations headcount and power and cooling are recurring costs that persist for the entire life of the hardware, not one-time expenses. Compare the full 3-year picture, not the sticker price of the GPUs alone, or the capex option will look artificially attractive.
- Infrastructure overhead (networking, storage, facilities) commonly adds 20-40% on top of raw GPU cost.
- Operations staffing is a recurring cost for the full hardware life, not a one-time line item.
- Power and cooling for a multi-GPU cluster typically run tens of thousands of dollars per year, not a rounding error.
- Compare 3-year totals, not year-one capex against year-one API spend.
Why utilization decides this comparison, not sticker price
Owned hardware cost is fixed once purchased regardless of how much it is used, while API or SaaS cost scales directly with actual consumption. That means the owning option only wins economically at sustained high utilization, typically above 50-60% average across a full week including nights and weekends, not just business-hours peaks. A cluster sized for peak demand but averaging 20% utilization has effectively doubled or tripled its real cost per unit of work, and the API alternative frequently wins at that utilization level even when the raw hardware math looked favorable on paper.
When the capex path makes sense regardless of the number
Some organizations should own hardware even when the pure cost comparison is close, because data residency, export control, or contractual requirements rule out sending data to a third-party API entirely. For aerospace, defense, and electronics manufacturers under ITAR or CMMC obligations, the capex decision is frequently a compliance requirement first and a cost optimization second, and the 3-year savings figure becomes a planning input rather than the deciding factor.
- Data residency and export control requirements can make owning the only viable option, regardless of the cost comparison.
- A close or slightly unfavorable 3-year number is still worth owning when compliance rules out the API alternative.
- Multi-tenant SaaS AI tools are frequently disqualified outright for ITAR-controlled data regardless of price.
- Document the compliance rationale alongside the cost comparison for the board record.
Reading your result
A large positive 3-year savings favoring ownership only holds if your realistic utilization forecast supports it; re-run this calculator with a conservative utilization assumption before presenting the number to finance. A negative result favoring the API option is common and not a failure, it simply means your workload volume has not yet reached the scale where dedicated hardware pays for itself. Revisit the comparison annually as workload volume grows, since the crossover point where owning becomes cheaper shifts every year model prices and hardware prices both change.
How Netray helps make this decision and execute on it
Netray sizes and deploys on-prem GPU infrastructure for manufacturers who need AI inside their network boundary, and we build the utilization discipline into every deployment, workload consolidation, capacity planning against real traffic, that determines whether the capex bet actually pays off. We also run the reverse analysis when the numbers favor staying on API or SaaS, since our incentive is the right decision, not a hardware sale. Engagements start with a two-week workload and utilization modeling phase using your real traffic data.
Frequently Asked Questions
What utilization level makes owning GPUs cheaper than API pricing?
As a general guide, owning tends to win once sustained average utilization exceeds roughly 50-60% across a full week, not just business-hours peaks. Below that, the fixed hardware and operations cost is spread across too little actual work, and the API option, which scales directly with consumption, usually comes out ahead. Model your realistic utilization curve, including nights and weekends, before committing capital.
Does this calculator include power and cooling costs?
Yes, as an annual recurring input alongside operations staffing. Power and cooling for a multi-GPU cluster commonly add tens of thousands of dollars per year and are frequently underestimated in a first capex proposal. It does not include depreciation tax treatment or hardware residual value, which a full financial model should add separately.
Should compliance requirements override a negative cost comparison?
Often yes. If data residency, ITAR, or CMMC obligations rule out sending your data to a third-party API entirely, the capex decision becomes a compliance requirement rather than a pure cost optimization, and a modestly unfavorable 3-year number is still the correct choice. Document that rationale explicitly for the board rather than presenting only the cost figure.
How often should this comparison be re-run?
At least annually, since both GPU prices and API or SaaS pricing change meaningfully year to year, and your own workload volume typically grows as well. A comparison run once at project kickoff can be stale within twelve months, especially as new GPU generations shift the price-to-performance ratio on the owned side of the ledger.
Get a 3-year capex versus opex model built from your real workload volume and compliance requirements.
Related Tools
Private AI Total Cost of Ownership Calculator
Model the full 3-year cost of an on-prem AI deployment, including amortized hardware, power, staff time, and support, against comparable API spend.
On-Prem AIOn-Prem LLM Total Cost of Ownership Calculator
Model the full multi-year cost of running LLMs on your own hardware, including GPU capex, power, cooling, support contracts, and operations staffing.
On-Prem AIAI Investment Payback Calculator
Turn upfront AI investment, ongoing run cost, and expected savings or revenue into a ramp-adjusted payback period and 3-year net value figure.
Go Deeper
Budgeting an On-Prem AI Project: A Line-Item Guide
Budgeting an on-prem AI project: realistic 2026 line items for GPU hardware, licensing, integration engineering, and the change management costs teams skip.
On-Prem AI Cost Benchmarks for 2026
On-prem AI cost benchmarks for 2026: four realistic project tiers from a $50k scoped pilot to $2M-plus enterprise programs, and what moves you between them.
On-Prem AI Managed Services: What Good SLAs Look Like
On-prem AI managed services: what should be in scope, which SLA metrics actually matter, typical pricing models, and questions to ask before signing.