Self-Hosted LLM: What It Costs and When It Pays Off
Every self-hosted LLM guide reprints the same breakeven number, and their own cost tables contradict it. Here is the honest math on GPUs, throughput, and the engineer time nobody counts, plus the cheaper lever almost nobody pulls.

Key takeaways
- The 2 million tokens per day breakeven printed across this topic is wrong. The source table it comes from shows the API winning at that volume.
- Against the cheapest hosted models, a fully loaded self-hosted box does not break even until roughly 100 million tokens a day.
- Against frontier-model pricing the crossover arrives near 8 million tokens a day, which is the comparison most guides quietly make.
- Engineer time is the largest recurring line item in self-hosting and appears in none of the published cost tables.
- Self-hosting changes the price of a token. It changes nothing about how many tokens your system spends.
For most teams, self-hosting an LLM is worth it for data sovereignty and regulatory constraint, and it is not yet worth it for cost.
One more thing to know before you read any further. The cost numbers circulating on this topic are a single recycled dataset. The Prem AI guide published a cost table in February 2026, and the Alpacked guide reproduced it in May 2026 nearly line for line, down to the same four volume rows, the same $850 monthly figure, and the same two unattributed customer cases. Neither derives the breakeven it states. This article derives its own, using their published figures as inputs so you can check the working against their pages.
What a self-hosted LLM actually is
A self-hosted LLM is an open-weight model running on infrastructure your organization controls, whether that is a rack in your building or GPU instances inside your own cloud account. You hold the weights, you operate the inference server, and prompts never cross into a vendor's network. You pay for capacity rather than per token.
Local inference is the single-machine case of that. One workstation, one GPU, usually Ollama or llama.cpp, usually one user at a time. It shares the sovereignty property and none of the operational burden of running a service. Most people who say they have self-hosted an LLM have run a local one, and the two behave very differently once a second concurrent user shows up.
Managed or dedicated inference is the third thing, and it gets conflated with self-hosting constantly. A provider runs the GPUs and gives you an isolated tenancy, so your traffic does not share a pool with other customers. Modular's handbook draws the line usefully: with a dedicated or serverless endpoint you are still handing the vendor "deployment, monitoring, storage and transfer, redundancy, GPU operations," and you get back control over data location without owning the operations (Modular). That is a real category and it is often the right answer for a team that wants the compliance property without hiring for it.
Why teams self-host
Data sovereignty is the first reason and the strongest. If your prompts contain patient records, trade positions, or unreleased source code, the question of whether inference happens inside your network boundary is a compliance question with a binary answer. Kong surveyed 550 IT leaders and developers between February and March 2025 and found 44% naming governance and security as one of their greatest barriers to LLM adoption (Kong). That is the barrier self-hosting actually removes.
Model control is the second. A hosted model can be deprecated, silently updated, or have its safety behaviour retuned underneath you. An open-weight checkpoint on your own disk does none of those things, which matters if you have evaluation suites pinned to a specific version.
Latency and offline availability are third. Local inference removes a network round trip and keeps working when the WAN does not.
Cost is fourth, and it is the weakest of the four for almost every team. The rest of this article is about why.
What it actually costs
Hardware
VRAM is the binding constraint, and the rule of thumb is about 0.5 GB per billion parameters at 4-bit quantization, roughly double that at FP16. A 70B model needs about 35 to 40 GB quantized and around 140 GB at full precision.
| Model size | VRAM at Q4 | Example GPU | Approximate tokens/sec | | --- | --- | --- | --- | | 7-8B | 4-6 GB | RTX 4060, RTX 5060 Ti | 30-50 | | 13B | 8-10 GB | RTX 4060 Ti 16GB, RTX 5070 | 20-35 | | 32-34B | 16-20 GB | RTX 4090, RTX 5070 Ti | 12-20 | | 70B | 35-40 GB | RTX 5090, dual RTX 4090 | 7-12 |
Figures from Alpacked. Read the throughput column as single-stream, not served capacity.
Acquisition price is where published guides go quiet. Alpacked is the exception and reports RTX 5090 MSRP at $1,999 against a street price of $3,500 to $4,000 in early 2026, with used RTX 4090s at $1,600 to $2,000. Take a single RTX 5090 at $3,750 street, add $1,500 for a chassis, CPU, 64 GB of system RAM, and the 1,200W supply the card wants, and the box is about $5,250. Amortised over 36 months, that is $146 a month.
Power and throughput
The RTX 5090 draws 575W and the 4090 draws 450W. At $0.16/kWh that is roughly $65 and $50 a month respectively for the card alone (Alpacked). Add the rest of the machine and call the box $85 a month.
The serving stack matters more than the card. Alpacked measures Ollama at about 41 tokens/sec under load against vLLM at about 793 on the same RTX 4090 baseline, roughly a 19x gap, with vLLM holding sub-100ms P99 at 128 concurrent users where Ollama reaches 673ms. If you are evaluating self-hosting on numbers you got from Ollama, you are evaluating the wrong configuration.
The line item nobody prints
Somebody has to keep this alive. Driver upgrades, CUDA version drift, a model swap when a better checkpoint lands, capacity planning, an on-call rotation for the inference tier, and the first two weeks of getting vLLM tuned. No published guide on this topic prices it.
Here is my assumption, stated plainly so you can substitute your own. A platform engineer at a fully loaded cost of $120 an hour, spending four hours a month on steady-state care of a single box. That is $480 a month, and it is the largest recurring number in the model. It is an assumption I am declaring, not a survey figure.
Single box, fully loaded: $146 hardware plus $85 power plus $480 labour equals $711 a month.
The breakeven math, corrected
Start with the error. Both guides state that self-hosting becomes competitive above roughly 2 million tokens a day. Both publish this table:
| Daily volume | GPT-4o mini monthly | Self-hosted monthly | Winner | | --- | --- | --- | --- | | 500K | ~$15 | ~$850 | API | | 2M | ~$60 | ~$850 | Roughly even | | 10M | ~$300 | ~$850 | Self-hosted | | 50M | ~$1,500 | ~$850 | Self-hosted (significantly) |
At 2 million tokens a day their own numbers read $60 against $850. That is not roughly even. Their API column scales at $30 per month per million daily tokens, so $850 is crossed at about 28 million tokens a day, fourteen times the figure they print.
The API column has its own problem. GPT-4o mini is published at $0.15 per million input tokens and $0.60 per million output (OpenAI). Even at 100% output tokens, 500K a day is 15 million a month, which is $9, not $15. A realistic agent mix of 80% input and 20% output blends to $0.24 per million, putting 500K a day at $3.60. The column cannot be reproduced from the vendor's published prices under any input/output split.
Redo it with the real rate of $7.20 per month per million daily tokens, and with a fully loaded self-hosted cost:
| Daily token volume | Hosted API monthly cost | Self-hosted monthly cost (fully loaded) | Which wins | | --- | --- | --- | --- | | 500K | $4 | $711 | API, by a wide margin | | 2M | $14 | $711 | API | | 10M | $72 | $711 | API | | 50M | $360 | $711 | API | | 100M | $720 | $711, above a single box's realistic ceiling | Even on paper, API in practice | | 250M | $1,800 | $1,736 on a four-GPU node | Self-hosted, if utilisation stays above ~55% |
The corrected crossover is about 100 million tokens a day, not 2 million. Working: $711 monthly cost divided by $7.20 per million daily tokens equals 98.75 million.
There is a capacity trap underneath that number. Take Alpacked's vLLM figure of 793 tokens/sec on a 4090 and their claim that the 5090 runs 67% faster, and a single 5090 tops out near 1,320 tokens/sec, or 114 million tokens a day at 100% utilisation. Nobody runs at 100%. At a realistic 40%, the box serves about 46 million a day, which is less than half the volume it needs to break even. You reach the crossover only by adding hardware, and adding hardware moves the crossover. A four-GPU node at $1,736 a month needs about 241 million tokens a day, and can only serve that if sustained utilisation clears roughly 55%. Sizing for peak and running hot are opposing requirements.
One honest qualification, and it is the one that flips the answer. All of the above compares against a cheap hosted model. Compare instead against frontier pricing at $1.25 input and $10.00 output per million, and the same 80/20 mix blends to $3.00 per million, or $90 a month per million daily tokens. Now the single box crosses at about 8 million tokens a day. That is a real and reachable number.
The catch is that a 4-bit 32B open-weight model is a substitute for the cheap tier, not the frontier tier. If your workload genuinely needs frontier reasoning, the model you can fit on your GPU probably will not do it, and the favourable comparison is the wrong one. Which hosted model you are actually replacing decides the entire calculation, and not one circulating table says so.
How to decide
- Measure your real token volume over 30 days, split into input and output. Not your projection. Your logs.
- Name the specific hosted model you would be replacing, and pull its current published price.
- Compute blended monthly API cost at your actual input/output mix.
- Price the self-hosted configuration for a model that can genuinely do your job: hardware amortised over 36 months, power at your utility rate, and a stated number of engineer hours per month.
- Compute your capacity ceiling. Served tokens per second times 86,400 times your honest utilisation assumption.
- Find the crossover by dividing monthly self-hosted cost by monthly API cost per million daily tokens.
- Self-host on cost only when 30-day sustained volume exceeds that crossover by 2x and projected GPU utilisation clears 50%. Below either threshold, self-host for sovereignty or not at all.
Common misconceptions
Self-hosting is cheaper by default. It is cheaper per token at high sustained utilisation and more expensive at everything else, because you pay for idle capacity and for the person maintaining it.
A 24GB consumer card runs a frontier-class model. It runs a good 32B model at 4-bit. The gap between that and a 400B-parameter open-weight model is not closed by quantization; Llama 4 Maverick needs about 206 GB at INT4 (Onyx).
Network isolation is the same as governance. Self-hosting changes where your data goes, which is a specific and real benefit. It gives you no access control, no permissions model, no audit trail, and no way to answer who asked the model what in March. Those are properties of the application around the model, which is what enterprise-grade actually means, and they need control at the point of action rather than at the network edge.
Cutting per-token price is the main cost lever. Your bill is price multiplied by quantity. Self-hosting attacks price. Almost nobody attacks quantity, and quantity is where the waste is. An agent that reasons its way through the same invoice reconciliation from scratch on every run is paying full price for a conclusion it already reached last Tuesday. Routing between models trims the price side further, and coordinating models across an app layer is how you start trimming the other side.
What we're building at Major in response
My position is that self-hosting is a legitimate lever aimed at the smaller of the two variables. Price per token is bounded below by physics and above by competition, and it has fallen steadily without any of us doing anything. Quantity of inference is unbounded, is set entirely by how your system is built, and in most deployments is inflated by the same structural habit: an agent that re-derives a workflow it has already solved, every single run, forever.
On Major, agents build the software they run on. When an agent works out how to reconcile a payment or triage a ticket, it writes an app for the repeatable part of that job and then runs the app instead of reasoning through the step again. The app holds its own state in a managed database, keeps its files, writes its own logs, and enforces permissions on every action, so the work is inspectable after the fact rather than reconstructed from a prompt. The model still gets called, for the judgment that actually needs judgment. It just stops getting called for the parts that were settled months ago. Cost is front-loaded and then flat instead of climbing with every run. Reason once, run forever.
This argument is orthogonal to hosting, and I want to be clear that it is. The same structure holds whether the weights are Anthropic's, OpenAI's, or a Llama running on a rack you bought. Self-hosting plus a smaller reasoning volume is a better position than either one alone, and if you have regulated data, you should self-host regardless of what any breakeven table says. We are not an inference host and we do not sell GPUs, so I have nothing to sell you on the hosting decision. What I would push back on is spending a quarter negotiating the unit price of work your system should not be repeating, which is how agentic workflows get structured when nobody asks how many times the model is being asked the same question.
If you are cost-motivated enough to be pricing GPUs, the same instinct is worth pointing at your reasoning volume first. You can see how Major's agents turn repeated work into deterministic apps and measure the difference against your own traffic.
Related articles
Frequently asked questions
- Is it worth self-hosting an LLM?
- For data sovereignty or regulated workloads where prompts cannot leave your network, yes, and the cost calculation is beside the point. For cost alone, usually not yet. Against cheap hosted models a fully loaded self-hosted box does not break even until roughly 100 million tokens a day. Against frontier-model pricing the crossover falls near 8 million a day, but the open-weight model you can fit on your GPU may not be a fair substitute for a frontier model.
- What is the best self-hosted LLM model?
- Pick by VRAM budget rather than by leaderboard, since the honest answer is the largest model that fits your card at acceptable throughput. A 16GB card runs a 13B model at Q4 comfortably. A 24GB RTX 4090 runs a 32B model at 4-bit, which is the sweet spot for most single-box deployments. Reaching 70B means about 35 to 40GB, so an RTX 5090 or dual 4090s.
- Can I run an LLM on my own computer?
- Yes, with VRAM as the ceiling. Budget roughly 0.5GB per billion parameters at 4-bit quantization, so a 24GB card handles a 32B model and stops well short of frontier-scale open weights. Falling back to CPU or system RAM works but costs you 10 to 100 times the speed, dropping from tens of tokens per second to single digits.
- How much does it cost to self-host an LLM?
- Around $700 to $1,800 a month fully loaded, depending on scale. For a single RTX 5090 box: roughly $146 in hardware amortised over 36 months on a $5,250 build, about $85 in power at $0.16/kWh, and $480 in engineer time assuming four hours a month at $120 an hour. That labour line is the largest of the three and appears in almost no published cost table.
- Is self-hosting an LLM more secure?
- It changes where your data goes, which is a specific and real benefit for regulated workloads. It does not give you access control, a permissions model, or an audit trail. Knowing that inference happened inside your network does not tell you who asked the model what, under whose credentials, or what it did afterward. Those properties come from the application around the model, not from the network boundary.