Kaer Labs vs Epoch AI vs METR: who measures what in AI evaluation

The independent AI evaluation field has split into distinct specialisms — trends and compute, dangerous-capability testing, and deployed behaviour. A guide to who measures what, and which one to call.

Renaissance scholar peering into a brass microscope

Short answer: The independent AI evaluation field is not one field. It has split into at least three specialisms with different methods and different customers: measuring the trajectory of AI, measuring the dangerous capabilities of frontier systems, and measuring the deployed behaviour of systems already in front of people. Kaer Labs does the third.

Three distinct jobs

SpecialismCore questionTypical methodWho needs it
Trends & computeWhere is the field going, and how fast?Public data, databases, trend estimationPolicymakers, funders, strategists
Dangerous capabilityCan this frontier model do something it must not?Pre-release access, threat modelling, task suitesDevelopers, regulators, safety institutes
Deployed behaviourWhat is this live system doing to real users?Realistic traffic, domain judges, longitudinal tracesInstitutions deploying, the public

These are not ranked. They answer questions that do not substitute for one another, and an organisation good at one is not automatically good at another — the staffing alone differs. Trend work needs economists and data engineers. Dangerous-capability work needs threat modellers and pre-release agreements with developers. Deployed-behaviour work needs caseworkers, nurses and advisers on the payroll as judges, which is a payroll problem more than a research problem.

What Kaer Labs measures, precisely

We work on the failures that never make the news. Not a model producing a bioweapon protocol — a welfare assistant that gets a deadline slightly wrong, at 11pm, for someone with no case worker to catch it. Individually small, and they accumulate across millions of interactions in exactly the populations with the least slack to absorb a mistake.

The awkward part: we also train models

Most independent evaluators do not ship models. We do — the Kaer Reasoners line — and that is a genuine conflict rather than a synergy to be spun. Here is the arrangement, in full, so it can be checked:

  1. Kaer checkpoints are entered into the Atlas blind, under the same identifiers as everyone else's.
  2. The unblinding key is held by one person outside the model team, released only after scores are locked.
  3. We do not enter a checkpoint we have not already frozen, so there is no tuning against a score we have seen.
  4. Our entries are flagged in the published release with raw traces attached, so anyone can replay the arithmetic.

This constrains the conflict; it does not remove it. The real safeguard is that the protocol, rubric and archive are open enough for someone else to score us and publish a disagreement. We would rather ask for that scrutiny than pretend it is unnecessary.

What building buys, and the reason we do both: we would not trust our own evaluation numbers if we had never had to hit them ourselves. Every methodological complaint we make about other people's evaluations is one we have had to answer for our own releases first.

Where the field is converging

Regulation is pushing all three specialisms toward the same grammar. The EU AI Act's post-market monitoring duties for high-risk systems treat oversight as a continuing obligation rather than a launch-day signature. That is the deployed-behaviour question becoming a legal requirement rather than a courtesy — and UK procurement, which still tends to validate once and rarely obliges a vendor to flag a behavioural change, has some catching up to do.

For eight months, on one assistant, we gave a partner the change notices their contract never obliged the vendor to send. Nobody else was in a position to. That gap is the whole reason this specialism exists.

← All posts Best small reasoning models in 2026 →

Common
questions

What does an independent AI evaluation organisation actually do?

Broadly one of three things. Track the trajectory of the field — compute, cost, dataset and capability trends over time. Test frontier models for dangerous capabilities before or around release, usually under agreement with the developer. Or measure what already-deployed systems do to the people using them. These need different methods, different staff and different funding, which is why organisations specialise.

Why do you not test frontier models for dangerous capabilities?

Because it is a distinct discipline requiring pre-release access agreements, specialist threat modelling and a very different staffing profile, and because organisations already do it better than we would. We work on deployed behaviour in ordinary domains — the failures that never make the news but accumulate across millions of interactions.

Is Kaer Labs independent if it also trains models?

It is a real tension and we do not claim it away. Our safeguards: Kaer checkpoints are entered blind under the same identifiers as everyone else's, the unblinding key sits with one person outside the model team and is released only after scores are locked, and our own entries are flagged in the published release with raw traces attached. If a Kaer model ever tops a domain in our Atlas, treat it with more suspicion than the rest of the table and check it.

Who should I contact for what?

For trend and compute questions, the organisations that specialise in tracking the field. For pre-deployment dangerous-capability testing of a frontier system, the organisations built for that. For what your deployed assistant is doing to your users this quarter, us.