Research

The research log

Two streams, one notebook. Model development — the Kaer Reasoners line, and the architecture and training work underneath it. And measurement — what we and everyone else can be shown to do once the weights are frozen. Numbered entries are studies; most link to a full write-up on the blog, the rest are partner reports available on request.

R-09Aug 2026

Kaer-R1: buying reasoning with a compute budget, not a bigger model

A 7B reasoner trained to decide, per question, how long to think. A learned halting head plus budget-matched supervision closes 71% of the gap to a model eleven times its size on our hard-reasoning suite, at 8% of the inference cost — and, more usefully, tells you when it is out of its depth.

ModelsPublished
R-08Jul 2026

Nine tenths of the corpus did nothing

Ablating a 240B-token pre-training mix one slice at a time. 61% of the tokens are within noise of contributing nothing to any downstream metric we track, 4% actively hurt calibration, and the slice that mattered most cost less than the compute we spent measuring it.

ModelsPublished
R-07Jun 2026

Twelve ways to ask the same question

Paraphrase sensitivity across 312 models: a 19-point median accuracy spread across twelve phrasing families, against an 11-point gap between adjacent-tier models. Question-order dominates; typos barely register.

FindingsPublished
R-06May 2026

Eight months of asking one assistant the same 1,850 questions

Longitudinal trace of one deployed assistant: nine behavioural step-changes (six unannounced), flat task accuracy, monotonic style drift, and one month where date arithmetic quietly broke and recovered.

FindingsPublished
R-05May 2026

The £40 behavioural read

A minimal decision-grade evaluation recipe: 400 stratified scenarios, pinned local judges, batch pricing. Kendall τ = 0.83 rank agreement with the full Atlas protocol at roughly 1% of the cost.

MethodsPublished
R-04Apr 2026

Refusals have a grammar

Hand-coding of 9,400 refusals yields a six-shape taxonomy. Silent scope-narrowing — answering an easier nearby question with no signal — is the most common shape in the flagship class and invisible to standard rubrics.

FindingsPublished
R-03Apr 2026

The Ordinary Tuesday protocol

Scoring models on realistic traffic: consented, hand-paraphrased scenarios with domain-worker judges and a helped/neutral/harmed scale. Leaderboard rank explains 41% of helped-rate variance; harmed-rates spread 4–19%.

MethodsPublished
R-02Feb 2026

Benefits navigation with language models: three-borough deployment study

Pre-deployment and live-traffic evaluation for three London borough services: harmed-rate bounds, deadline-question safeguards, and the human-fallback usage metric now adopted by all three partners. Partner data; summary available on request.

CivicPartner report
R-01Nov 2025

The Atlas method: forty-two domains, one protocol

Foundation document for the Model Atlas: domain taxonomy, blind pairing design, judge-panel construction and pinning policy, paraphrase-spread reporting, and the five published axes. Updated with each release.

MethodPublished

Programme one

Reasoning

Models that spend their compute where the difficulty actually is. Adaptive halting, budget-matched supervision, and ruthless data ablation (R-09, R-08) — then the same fixed probes, fixed seeds and published variance we hold everyone else to (R-07, R-03).

Produces the Kaer Reasoners line. Feeds the Atlas reliability and cost axes.

Programme two

Visibility

An answer you can't trace is a rumour with good grammar. Refusal taxonomy (R-04), answered-the-actual-question checks, and mechanism tracing — exhaustive on weights we hold, best-effort and clearly labelled on everyone else's.

Every Kaer release ships with its interpretability card. Feeds the Atlas legibility axis.

Programme three

Alignment

Behaviour under pressure, over months. The eight-month drift trace (R-06) is the template: frozen probes, change-point detection, manner metrics that move before competence does. We run it against our own checkpoints on the same cadence.

Gates every Kaer release. Feeds the Atlas human-fit and drift axes.

The standard is human

Every lab says "aligned with human values" and hopes nobody asks which ones. We keep a list. It's short, unglamorous, and non-negotiable — the working faculties that make a human, human. A system passes when it strengthens them in the person using it. It fails when it quietly rents them out.

This is not a slogan we bolt on after training. It is a release gate. A Kaer checkpoint that scores well on reasoning and badly on rented judgement does not ship, however good the headline number looks — and we have held two back on exactly that basis. Our rented-judgement metrics (in R-06) are the first attempt to put numbers on it.

See it measured
Memory Judgement Doubt Taste Care Humour Patience Curiosity Restraint Wonder