Your API bill grows in perfect lockstep with your success. Every new customer adds tokens, every token adds cost, and the whole amount lands in your cost of goods sold every single month. At some volume, that metered bill crosses a line: buying the GPUs outright and running inference yourself becomes cheaper than renting someone else's. The question is where that line sits for you — because crossing it is not just an engineering decision. It rewrites your balance sheet, your gross margin, and your tax return.
This guide walks through the breakeven math, why the vLLM serving stack moved the line, and what changes on your books the day you stop expensing API calls and start capitalizing a GPU cluster.
Two Ways to Pay for Inference
Every token your product serves is paid for one of two ways, and they hit your financials completely differently.
The API path is pure operating expense. You pay per million tokens, the bill scales with usage, and the full amount is a period cost — most naturally booked as cost of goods sold, since inference is directly tied to delivering your product to customers. Zero upfront, zero assets, zero depreciation. Your gross margin takes the hit every month at a constant rate no matter how big you get.
The self-hosted path is mostly capital expenditure. You buy GPU servers (or sign a long reservation), put a fixed asset on the balance sheet, and depreciate it over its useful life. Your monthly P&L then shows depreciation plus operating costs — electricity, colocation or rack space, monitoring, and the engineering hours that keep the cluster healthy — instead of a per-token bill.
That structural difference is the whole game. API costs scale linearly with volume forever. Self-hosted costs are front-loaded and mostly fixed, so the effective cost per token falls as utilization rises. Somewhere those two curves cross. Everything below is about finding where.
The Breakeven Math
A widely cited 2026 analysis of more than 50 production deployments put the rule of thumb at roughly $20,000 per month in API spend. Below that, the engineering and operations cost of running your own cluster almost always exceeds the savings. Above $50,000 per month, self-hosting the bulk of traffic typically wins by 50 to 70 percent. Between those lines sits a gray zone where the details — your models, your traffic shape, your team's GPU experience — decide.
The same analysis sketched the investment behind those numbers: $50,000 to $500,000 upfront for GPU hardware depending on model size, plus $3,000 to $15,000 per month in ongoing operations. That ranges from a single used A100-class box serving a 70-billion-parameter model to a multi-node H100 cluster.
It matters enormously which API you are replacing
The breakeven volume shifts by roughly 100x depending on what you pay per token today:
| Monthly volume | Frontier API cost (approx.) | Self-hosted cost (owned hardware) | Monthly savings |
|---|---|---|---|
| 100M tokens | $1,750 | $5,500 | -$3,750 (loss) |
| 500M tokens | $8,750 | $7,500 | $1,250 (20-month payback) |
| 1B tokens | $17,500 | $10,000 | $7,500 (7-month payback) |
| 5B tokens | $87,500 | $25,000 | $62,500 (1-month payback) |
Against a premium frontier API charging on the order of $1,750 per 100 million tokens, the crossover lands around half a billion to a billion tokens a month, and the payback accelerates hard from there.
But replace a budget API charging closer to $80 per 100 million tokens and the math inverts: you would need 50 billion-plus tokens a month to break even — a volume that demands a serious multi-GPU cluster and a dedicated infrastructure team. If a cheap API meets your quality bar, self-hosting rarely makes financial sense at any volume a small business will reach.
Utilization is everything
Other cost breakdowns land the breakeven much lower — one detailed build-vs-rent model puts it near $4,200 a month in API spend — and the gap between estimates is itself the lesson: there is no universal breakeven point. The entire calculation hinges on utilization. A GPU serving steady traffic around the clock spreads its fixed cost over billions of tokens. The same GPU serving spiky daytime traffic sits idle all night, depreciating with nothing to show for it, and the idle hours quietly double or triple your true cost per token.
Before trusting anyone's breakeven figure, including this article's, measure your own traffic shape: sustained tokens per second, peak-to-average ratio, and growth slope. Flat, predictable, growing volume favors self-hosting. Spiky or low volume favors the API's zero idle cost.
Why vLLM Moved the Line
The serving stack you choose is a financial variable, not just a technical one. Throughput per GPU decides how many GPUs you must buy, and the spread between a naive serving loop and a modern inference engine is enormous.
vLLM has become the default production choice through two techniques that ship enabled out of the box:
- Continuous batching keeps every GPU slot filled by mixing new and in-flight requests instead of waiting for a whole batch to finish, lifting throughput roughly 2 to 3x over static batching.
- PagedAttention manages the key-value cache in blocks rather than pre-allocating contiguous memory, cutting fragmentation waste and supporting 2 to 4x more concurrent requests on the same card.
Combined with tensor parallelism across GPUs and support for quantized weights, vLLM delivers on the order of 24x the throughput of a naive transformers serving loop — which means roughly 24x fewer GPUs to buy for the same traffic. A cluster sized without these techniques is not just slower; it is a capital-expenditure mistake.
Two more properties matter for the business case. First, vLLM exposes OpenAI-compatible endpoints, so migrating bulk traffic off an API is largely a URL and key change rather than a rewrite — the switching cost stays low. Second, the constraint: you can only self-host open-weight models such as Llama, Qwen, DeepSeek, and Mistral. Frontier models from the big labs stay API-only, which is why the common end state is a hybrid: self-host an open model for the 80 percent of traffic that is routine, and keep routing the 20 percent that needs frontier reasoning through an API.
What Changes on Your Books When You Self-Host
The day the GPU servers arrive, your accounting changes in five places. Get these right and the breakeven math you ran actually shows up in your financials.
1. The hardware becomes a fixed asset
Purchased GPU servers are capital assets, not supplies. You record the purchase price plus the direct costs of getting the cluster into service — freight, rack installation, initial configuration labor — as a fixed asset on the balance sheet, then depreciate it. Computers and related equipment are generally 5-year MACRS property, and depreciation begins when the asset is placed in service (ready and available for its intended use), not when you pay the invoice.
The $2,500 de minimis safe harbor that lets you expense small purchases will not cover a GPU server. Do not run a $60,000 cluster through office supplies.
2. You get a genuine tax-timing choice in 2026
For federal taxes, current law gives you three speeds for the same hardware:
- Section 179 expensing: deduct up to $2,560,000 of qualifying equipment placed in service in 2026, with the benefit phasing out dollar-for-dollar once total qualifying purchases exceed $4,090,000. The deduction cannot exceed your taxable business income, but unused amounts carry forward.
- 100 percent bonus depreciation: permanently restored for qualifying property acquired after January 19, 2025. Unlike Section 179, it can create a loss and has no dollar cap.
- Regular MACRS: spread the deduction over 5 years.
A profitable small business buying its first cluster will often expense the whole thing in year one. A pre-profit startup may prefer MACRS, preserving deductions for years when there is income to offset. Either way, track book and tax depreciation separately — your financial statements should reflect economic reality over the asset's useful life even when the tax return takes it all at once.
3. Rented GPUs stay operating expense
None of the above applies if you rent cloud GPUs by the hour. Hourly rentals, reserved instances, and GPU cloud subscriptions are period operating costs — closer to the API path than to ownership. That is a legitimate middle ground (no upfront capital, still metered), but do not expect a balance-sheet asset or a depreciation deduction from a rental bill. The CapEx-vs-OpEx decision and the build-vs-buy decision are two separate axes.
4. Ongoing cluster costs split between COGS and OpEx
Once running, classify costs by what they support:
- Into COGS (they scale with serving customers): electricity for production nodes, colocation and bandwidth for the serving cluster, monitoring and logging for production, and the depreciation on production GPUs.
- Into operating expense: the prototype box your engineers experiment on, staging environments, and general R&D infrastructure.
The same split applies to labor. Engineering time spent keeping production inference up supports delivery and can sit in COGS; time spent evaluating next quarter's model is R&D. SaaS gross margins are typically benchmarked at 70 to 85 percent, and misclassifying either direction makes your margin incomparable — put API and hosting spend in overhead and your margin looks artificially wonderful while your OpEx looks bloated.
5. Your gross margin should expand past breakeven
Watch this identity monthly after migration: the API line in COGS should fall toward zero (or toward the 20 percent frontier remainder in a hybrid setup), replaced by a smaller combined depreciation-plus-hosting figure. If gross margin does not improve within a quarter or two of steady-state operation, either utilization is lower than modeled or hidden operations costs ate the savings — which is your cue to revisit the decision rather than defend it.
In a plain-text ledger, the purchase itself is one balanced entry — an asset swap from cash to equipment, with depreciation entries recognizing the cost month by month. If you have never modeled fixed assets this way, the Beancount documentation walks through accounts, depreciation bookings, and reports step by step.
Mistakes That Erase the Savings
Most failed self-hosting migrations do not fail on GPU benchmarks. They fail on costs the spreadsheet omitted:
Sizing for the peak, paying for the average. A cluster provisioned for your busiest hour runs half-empty the rest of the day. Every idle hour is depreciation with no tokens. Autoscaling helps on rented GPUs; on owned hardware, the only fix is enough baseline volume.
Forgetting the operations tail. One detailed teardown puts hidden monthly costs at $4,700 to $7,900 on top of hardware for a modest 4-GPU setup: 8 to 20 engineering hours a month ($2,500 to $5,000 at loaded salaries), electricity ($400 to $600), networking and storage, monitoring, and a redundancy spare. None of it appears in a GPU-price comparison.
Ignoring the uptime gap. API providers typically back 99.9 percent SLA uptime. A self-run cluster without serious redundancy investment realistically lands at 95 to 99 percent — failed cards, out-of-memory crashes, a CUDA driver update that breaks deployment at 2 a.m. On $100,000 a month of AI-driven revenue, one extra point of downtime costs $1,000 a month, before counting the incident response.
Booking API spend as overhead. If inference costs sit in general expenses instead of COGS, your gross margin is fiction in both directions: overstated before migration, and the post-migration improvement invisible. Fix the classification before you compare.
Useful-life optimism. Some hyperscalers depreciate GPUs over 5 to 6 years while analysts argue the economic life is closer to 2 to 3 given how fast each generation obsoletes the last. If your cluster's resale value collapses when the next architecture ships, book an impairment rather than carrying a fantasy asset.
Commingling prototype spend with production. The GPU you bought to evaluate fine-tuning is R&D. The cluster serving customer traffic is production. Mixing them in one account corrupts both your gross margin and your R&D credit support.
A Decision Checklist Before You Buy
Run through these six questions with real numbers, not vibes:
- Volume: Is sustained monthly API spend above roughly $20,000, or credibly headed there within two quarters?
- Which API are you replacing? Premium frontier pricing breaks even near a billion tokens a month; budget API pricing may never break even.
- Traffic shape: Is load flat enough that GPUs stay busy, or will peak provisioning strand capacity?
- Expertise: Does someone on the team already speak CUDA, quantization, and vLLM tuning — or are you budgeting a hire into the payback math?
- Cash and taxes: Can you fund the upfront capital, and do you have the income to use a Section 179 or bonus deduction — or would MACRS serve you better?
- Non-financial drivers: Do privacy, compliance, or sub-100ms latency requirements force self-hosting regardless of cost?
Three or more weak answers means staying on the API (or the hybrid split) for now. The line will still be there when your volume grows into it — and by then, the next GPU generation will have moved it in your favor.
Keep Your Inference Spend Legible
Whether you pay per token or per depreciation schedule, inference is now one of your largest cost lines, and it deserves better than a single mystery total on the P&L. Separating API spend from hosting, production GPUs from prototype boxes, and book depreciation from tax depreciation is what turns the breakeven line from a one-time calculation into a number you can watch every month.
Beancount.io provides plain-text accounting that gives you complete transparency and control over your financial data — no black boxes, no vendor lock-in. Track the cluster as a fixed asset, book depreciation on schedule, and watch it flow through your reports in Fava. Get started for free and keep your AI infrastructure spend as legible as your infrastructure code.





