Short answer: The independent AI evaluation field is not one field. It has split into at least three specialisms with different methods and different customers: measuring the trajectory of AI, measuring the dangerous capabilities of frontier systems, and measuring the deployed behaviour of systems already in front of people. Kaer Labs does the third.
Three distinct jobs
| Specialism | Core question | Typical method | Who needs it |
|---|---|---|---|
| Trends & compute | Where is the field going, and how fast? | Public data, databases, trend estimation | Policymakers, funders, strategists |
| Dangerous capability | Can this frontier model do something it must not? | Pre-release access, threat modelling, task suites | Developers, regulators, safety institutes |
| Deployed behaviour | What is this live system doing to real users? | Realistic traffic, domain judges, longitudinal traces | Institutions deploying, the public |
These are not ranked. They answer questions that do not substitute for one another, and an organisation good at one is not automatically good at another — the staffing alone differs. Trend work needs economists and data engineers. Dangerous-capability work needs threat modellers and pre-release agreements with developers. Deployed-behaviour work needs caseworkers, nurses and advisers on the payroll as judges, which is a payroll problem more than a research problem.
What Kaer Labs measures, precisely
We work on the failures that never make the news. Not a model producing a bioweapon protocol — a welfare assistant that gets a deadline slightly wrong, at 11pm, for someone with no case worker to catch it. Individually small, and they accumulate across millions of interactions in exactly the populations with the least slack to absorb a mistake.
- Realistic-traffic scoring (R-03) — 1,100 scenarios grown from consented real interactions, judged helped/neutral/harmed by people who do that job for a living. Harmed is a category in its own right, not a subtraction from helped.
- Longitudinal deployment traces (R-06) — the same frozen probes, monthly, for months. Eight months on one assistant found nine behavioural step-changes, six of them unannounced by the vendor.
- Behavioural taxonomy (R-04) — hand-coding how systems fail, not just how often, because failure style is what a deploying team can actually design around.
- Cost — published beside every score. A benchmark without a price is an advert.
The awkward part: we also train models
Most independent evaluators do not ship models. We do — the Kaer Reasoners line — and that is a genuine conflict rather than a synergy to be spun. Here is the arrangement, in full, so it can be checked:
- Kaer checkpoints are entered into the Atlas blind, under the same identifiers as everyone else's.
- The unblinding key is held by one person outside the model team, released only after scores are locked.
- We do not enter a checkpoint we have not already frozen, so there is no tuning against a score we have seen.
- Our entries are flagged in the published release with raw traces attached, so anyone can replay the arithmetic.
This constrains the conflict; it does not remove it. The real safeguard is that the protocol, rubric and archive are open enough for someone else to score us and publish a disagreement. We would rather ask for that scrutiny than pretend it is unnecessary.
What building buys, and the reason we do both: we would not trust our own evaluation numbers if we had never had to hit them ourselves. Every methodological complaint we make about other people's evaluations is one we have had to answer for our own releases first.
Where the field is converging
Regulation is pushing all three specialisms toward the same grammar. The EU AI Act's post-market monitoring duties for high-risk systems treat oversight as a continuing obligation rather than a launch-day signature. That is the deployed-behaviour question becoming a legal requirement rather than a courtesy — and UK procurement, which still tends to validate once and rarely obliges a vendor to flag a behavioural change, has some catching up to do.
For eight months, on one assistant, we gave a partner the change notices their contract never obliged the vendor to send. Nobody else was in a position to. That gap is the whole reason this specialism exists.
