← Blog
AUG 24, 2026 · MAPLEHILL LABS · UPDATES

ModelIndex updates — August 2026

The first of a monthly series: what changed in the data and the site, and why. August was the month ModelIndex went from an overnight build to a measured, audited index at modelindex.ai.

Methodology v1.1

Our first throughput probe used repetitive filler text. Speculative-decoding stacks inflated some providers' tok/s by up to 97% on it, and one model's safety layer refused it outright. v1.1 uses varied deterministic filler, and aggregation only publishes runs matching the current methodology version — old and new numbers never blend into one median. Every re-measured pair is live; the worst corrections are documented in the findings post.

Service tiers are now a real dimension

Flex, batch, and priority endpoints were being treated as list prices, which understated OpenAI and Google costs by 2× on affected models. Tiers are now badged on every endpoint row and excluded from best-price figures, the "Standard" pricing variant, and the calculator. Spot-checked against official pricing pages after two external audits: every checked model now matches.

Quality benchmarks, with provenance

Model pages gained a curated benchmark panel — SWE-bench Verified, GPQA Diamond, LMArena Elo, Humanity's Last Exam, and Terminal-Bench 2.1 — across 32 models and 69 sourced scores. Every cell carries its source, an as-of date, and an INDEPENDENT vs SELF-REPORTED badge; disagreements between vendor and independent figures are shown, not averaged; unsourceable scores are simply absent. These are third-party results and labeled as such — distinct from the TTFT/tok/s and capability screens ModelIndex measures itself.

Daily measurement, including first-party APIs

Probes now run daily and automatically: catalog prices, a starter benchmark set, direct probes against Anthropic, OpenAI, Google, Groq, Cerebras, Fireworks, DeepSeek, and Perplexity's own APIs (marked distinctly from OpenRouter-routed measurements), and a retry lane that re-attempts failed endpoints oldest-first. Capability screens found providers disagreeing on 30 of 39 multi-provider models — the per-probe breakdown is the SMOKE column on every model page.

Pricing coverage beyond the aggregator

Human-verified reference pricing (authority-ranked below a keyed official API, above aggregator data) now covers the current Anthropic, Google, OpenAI, and DeepSeek lineups — including DeepSeek's 50%-off-peak windows, Gemini's >200K long-context tier, and corrections where aggregator rows turned out to be mislabeled service tiers. Conflicts are recorded and published per model, not silently overwritten.

Also this month

  • Subscriptions: 32 hand-verified consumer and coding plans across 11 providers, each with source and verification date.
  • modelindex.ai launched — apex domain, per-model RSS feeds, embeddable live badges, and this blog.
  • The model directory gained release dates, OpenRouter 7-day usage (the default sort is now "most used"), and a LAST CHANGE column that reports observed price moves with their age instead of a mostly-empty 30-day window.

Data corrections, methodology versions, and spend ledgers are all in the open. If a number looks wrong, it probably has a provenance trail — and if it doesn't, we want to know.