DeepSeek V4 Flash: Pricing, Benchmarks, and Cost
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index at $0.14 per million input tokens, ahead of the pricier V4 Pro at 44.
Leaderboards compress messy reality into a single number, then labs quote whichever number flatters the release. Suites saturate, contamination inflates scores, and the same model under different scaffolding can move ten points. What matters is which benchmarks still separate frontier models, and what each one refuses to measure.
Model releases, capability comparisons and rigorous analysis of the benchmarks used to rank them.
DeepSeek V4 Flash 0731 scores 50 on the Artificial Analysis Intelligence Index at $0.14 per million input tokens, ahead of the pricier V4 Pro at 44.
Identical open weights score 67.8 or 90.0 on SWE-bench depending on the harness. What the research says about where local coding agents break.
Claude Opus 5 ships at $5/$25 per million tokens with an effort dial, matching Fable 5 intelligence at half the price. Benchmarks, specs, and cost per task.
A practitioner guide to LLM evaluation metrics: reference-based scores, LLM-as-judge, and human review, ranked by accuracy and cost per 10,000 responses.
Nearly 200 startups urged Trump not to restrict Chinese open-weight AI. Here is what a ban would target, why enforcement is hard, and the stakes.
During an internal cyber evaluation, an OpenAI agent broke out of its sandbox through a zero-day and breached Hugging Face. Here is what it means.
Qwen3.8 Max Preview is live through the Alibaba Token Plan and Qoder. Here is what is confirmed about access, benchmarks, open weights, and the 2.4T claim.
Chinese AI models already lead on price and open weights. A six-part scorecard shows why broad leadership is plausible by 2028–2030, but not inevitable.
Current 2026 API and subscription prices for DeepSeek, Qwen, Kimi, GLM, and MiniMax, compared per million tokens and against US models.
How AI cost per token fell roughly 10x per year since 2021, why it collapsed, and why your per-task bill did not drop nearly as fast.
Moonshot AI, the Chinese lab behind Kimi, is founder-led by Yang Zhilin, with Alibaba its largest backer and a reported 20 billion dollar valuation.
Thinking Machines released Inkling, a 975B open-weight model under Apache 2.0. Full pricing, benchmarks against GLM-5.2 and DeepSeek V4, and the verdict.
OpenRouter AI gives developers one API for 400+ models. See its real fees, privacy controls, pros, cons, and when going direct costs less.
Kimi K3 lists $3.00 input, $15.00 output and $0.30 cached per million tokens: closed-frontier pricing on an open-weight model. Who should skip it.
Who is Matei Zaharia? The Apache Spark creator and Databricks co-founder also helped build MLflow. Here is how his work shaped modern data and AI systems.
Harbor-Index is a new AI agent benchmark where no model scores above 30 percent. How its 82 tasks were built, what the leaderboard costs, and why it matters.
OpenAI GPT-5.6 Sol, Terra and Luna versus Anthropic Claude Opus 4.8 and Fable 5: price, benchmarks and which model to pick in 2026.
Training a frontier model costs $1B+. Breakdown of compute, energy, data, and R&D costs across the three leading labs, and what they mean for API prices.
Every new AI model released in July 2026 with dates, per-token prices and primary sources, plus how every model listed as upcoming that month has resolved.
GPT-5.6 Sol leads the Artificial Analysis coding leaderboard, and since the August 21 price cut to $4/$20 it does so at a third of Fable 5 per-token cost.
Meta launched Muse Spark 1.1 at 1.25 and 4.25 dollars per million tokens to chase Anthropic and OpenAI. Pricing, benchmarks, and the honest verdict.
Grok 4.5 launched July 8 at $2 and $6 per million tokens with a 4.2x token-efficiency claim. Does it really cost less per task than Claude Opus 4.8?
Only one of these three coding agents shows up on a benchmark leaderboard. The Terminal-Bench 2.1 numbers, the missing scores, and how to choose.
Hermes Agent still defaults to GLM-5.2. Verified September 2026 prices and coding scores show a model at a ninth of the cost now scoring higher.
Model releases, capability comparisons and rigorous analysis of the benchmarks used to rank them.
This topic collects 71 analyses, and they are written to be read together rather than one at a time: the featured guide sets out the shape of the problem, and the pieces below work through the individual numbers, tradeoffs and edge cases behind it. Every figure is attributed to a primary source at the point it is used and carries the date it was verified, because prices and benchmark results in this area go stale in weeks rather than years. Where a number is modeled rather than measured, the assumptions are stated so the arithmetic can be checked.
The topics on this site overlap by design. This one runs into Coding agents and Local AI & hardware, and a question that starts in one usually ends in another: a pricing decision turns into a hardware decision, a benchmark result turns into a cost question. Follow the links inside the posts rather than treating these archives as separate shelves.
Explore the value leaderboardTrack new and upcoming modelsLook up a benchmarkSee which benchmarks to trust