Why Local LLMs Got Good in 2026: Capability & Cost
Open-weights LLMs crossed from toy to useful in 2026. What actually changed, and the cost math for when running a model yourself beats paying an API.
Memory capacity sets what a machine can load, bandwidth sets how fast it answers, and utilization decides whether owning hardware beats renting it. With DRAM in a shortage cycle the arithmetic keeps moving: a build that paid for itself last quarter may not this one, and the break-even turns on how many hours the box actually runs.
This guide frames the system before you move into the newest signals.
Open-weights LLMs crossed from toy to useful in 2026. What actually changed, and the cost math for when running a model yourself beats paying an API.
The featured guide stays above; the stream below moves as new analysis is published.
Apple refreshed the Mac mini and Mac Studio on August 25 2026. Every memory ceiling stayed flat, and bandwidth became a paid upgrade.
The Plaud Note costs $159 plus a subscription to do anything useful. What it really costs per transcribed hour, who pays too much, and when to skip it.
Apache 2.0 weights for Qwen3.8-27B landed on August 14, 2026. Strong vendor scores, hosted tokens at 80 percent off the flagship, and a catch in the local math.
YMTC now ships more NAND bits than Micron or Kioxia but ranks only fifth by revenue. What that gap means for flash prices in 2026 and 2027.
NAND contract prices ran 70 to 75 percent in a single quarter and a 2TB drive doubled. What drove it, and why the surge is already slowing.
Samsung, SK hynix and Micron reversed course and committed about 2 trillion dollars to new memory fabs. The first wafers arrive in December 2028.
Identical open weights score 67.8 or 90.0 on SWE-bench depending on the harness. What the research says about where local coding agents break.
Nvidia RTX cards are up 20 to 30 percent in a third 2026 hike. Here is the memory-cost chain behind it, and whether to buy a GPU now or wait.
CXMT closed 466 percent above its offer price on debut, worth 3.28 trillion yuan. What that valuation implies about the DRAM shortage.
SK hynix says 2027 will be the worst supply year ever for memory and ADATA expects a decade-long shortage. Here is how long the RAM squeeze may last.
CXMT is the largest DRAM maker in China, now the number four global supplier, closing in on Micron on capacity while its HBM chips stay years behind.
The best open-weight AI models in 2026, ranked by use case: coding, long context, multimodal, on-device, and the real cost per finished task.
RAM prices roughly tripled in 2026. Here are 7 ways to save money buying memory: right-size capacity, buy used, find bundle deals, time the market, and more.
OpenAI is spending nearly 6.5 billion dollars to build a screenless AI smart speaker. The business logic, the running costs, and the Apple lawsuit risk.
Consumer DDR5 has gone flat while contract and HBM prices keep climbing. Here are the signals that show whether the 2026 memory cycle is peaking.
The 2026 memory shortage explained: why AI and HBM demand starved consumer DRAM, how high prices have climbed, and when the RAM shortage is likely to end.
RAM prices roughly tripled in 2026 as AI memory demand starved consumer DRAM. Here is where the trend is heading and whether to buy now or wait.
Apple raised Mac and iPad prices on June 25 2026 as the memory shortage bit. Here is what the unified-memory tax means for running local AI.
How much RAM you need to run a local LLM in 2026: what models 8GB to 512GB can run, the per-billion-parameter math, and the device for each tier.
RAM keeps getting pricier in 2026 because the memory giants profit more from scarcity than supply. Inside the most lucrative shortage in chip history.
Seven Chinese firms now ship AI accelerators, the best near NVIDIA H100 class. A fact-checked 2026 map of who makes China's GPUs and what is real.
Cohere North Mini Code is free on the API and open-weight. Here is what it really costs per task once you self-host it on a single H100.
What a self-hosted LLM token really costs in 2026: cost per token across owned hardware, why memory bandwidth sets speed, and where buying beats the API.
The economics of local inference, RAM, memory bandwidth, self-hosting and the hardware behind AI.
This topic collects 24 analyses, and they are written to be read together rather than one at a time: the featured guide sets out the shape of the problem, and the pieces below work through the individual numbers, tradeoffs and edge cases behind it. Every figure is attributed to a primary source at the point it is used and carries the date it was verified, because prices and benchmark results in this area go stale in weeks rather than years. Where a number is modeled rather than measured, the assumptions are stated so the arithmetic can be checked.
The topics on this site overlap by design. This one runs into Models & benchmarks and Compute infrastructure, and a question that starts in one usually ends in another: a pricing decision turns into a hardware decision, a benchmark result turns into a cost question. Follow the links inside the posts rather than treating these archives as separate shelves.
Check what your hardware can runWatch the RAM price indexCompare GPU street prices