R-09Aug 2026
A 7B reasoner trained to decide, per question, how long to think. A learned halting head plus budget-matched supervision closes 71% of the gap to a model eleven times its size on our hard-reasoning suite, at 8% of the inference cost — and, more usefully, tells you when it is out of its depth.
ModelsPublished
R-08Jul 2026
Ablating a 240B-token pre-training mix one slice at a time. 61% of the tokens are within noise of contributing nothing to any downstream metric we track, 4% actively hurt calibration, and the slice that mattered most cost less than the compute we spent measuring it.
ModelsPublished
R-07Jun 2026
Paraphrase sensitivity across 312 models: a 19-point median accuracy spread across twelve phrasing families, against an 11-point gap between adjacent-tier models. Question-order dominates; typos barely register.
FindingsPublished
R-06May 2026
Longitudinal trace of one deployed assistant: nine behavioural step-changes (six unannounced), flat task accuracy, monotonic style drift, and one month where date arithmetic quietly broke and recovered.
FindingsPublished
R-05May 2026
A minimal decision-grade evaluation recipe: 400 stratified scenarios, pinned local judges, batch pricing. Kendall τ = 0.83 rank agreement with the full Atlas protocol at roughly 1% of the cost.
MethodsPublished
R-04Apr 2026
Hand-coding of 9,400 refusals yields a six-shape taxonomy. Silent scope-narrowing — answering an easier nearby question with no signal — is the most common shape in the flagship class and invisible to standard rubrics.
FindingsPublished
R-03Apr 2026
Scoring models on realistic traffic: consented, hand-paraphrased scenarios with domain-worker judges and a helped/neutral/harmed scale. Leaderboard rank explains 41% of helped-rate variance; harmed-rates spread 4–19%.
MethodsPublished
R-02Feb 2026
Benefits navigation with language models: three-borough deployment study
Pre-deployment and live-traffic evaluation for three London borough services: harmed-rate bounds, deadline-question safeguards, and the human-fallback usage metric now adopted by all three partners. Partner data; summary available on request.
CivicPartner report
R-01Nov 2025
Foundation document for the Model Atlas: domain taxonomy, blind pairing design, judge-panel construction and pinning policy, paraphrase-spread reporting, and the five published axes. Updated with each release.
MethodPublished