Local LLM Tokenomics: Self-Hosted Cost Per Token (2026)
What a self-hosted LLM token really costs in 2026: cost per token across owned hardware, why memory bandwidth sets speed, and where buying beats the API.
The price of intelligence
Token prices are the sticker. Context, retries, subscriptions and unfinished work decide the bill.
Topic archive · page 8
Analysis of AI pricing, API bills, subscriptions and the real cost of completing useful work.
What a self-hosted LLM token really costs in 2026: cost per token across owned hardware, why memory bandwidth sets speed, and where buying beats the API.
Google ended free Gemini CLI access on June 18, 2026. The best alternatives (Claude Code, Aider, OpenCode, Antigravity) ranked by real cost per task.
Claude Fable 5 costs exactly 2x Opus 4.8 per token: $10/$50 vs $5/$25. Whether it is cheaper per task depends on loop count, not the sticker rate.
Gemini 3.5 Flash lists at $1.50/$9.00 per million tokens: about 2x Haiku per token, yet roughly 3x cheaper per task than GPT-5.5 on agentic coding work.
A cost-first guide to building an AI agent harness: loop, context, tools, verification, and guardrails, and the token cost each layer controls.
Microsoft calls MAI-Code-1-Flash its cheapest coding model. In Copilot it bills like Claude Haiku 4.5 (0.33x); its token edge holds only on easy benchmarks.
Qwen 3.7 Max lists at half Claude Opus 4.8 and runs the same eval suite for a third of the cost. The catch is not hidden cost. It is what the price buys.
Codex leads Claude Code 83.4% to 78.9% on Terminal-Bench 2.1, but the prices match and the cheaper agent flips with your harness. The honest scorecard.
Composer 2.5 finishes a coding task for about $0.07, 10-60x under Claude Opus and GPT-5.5 at near-equal benchmark scores. What the cheap headline leaves out.