Euler CenterMeasurement & EvaluationWrite to the lab
Euler Center · AI evaluation

How to evaluate an AI agent at work.

Most AI evaluation measures the model. The question that decides a rollout is whether the work performs, which is a different measurement.

the instrument on screen, ASSAY on cubelet.ai
I / What benchmarks miss

The model is not the unit

A benchmark tells you how a model scores on tasks someone else chose. It does not tell you what happens when your analyst has to check the output, when an exception arrives that nobody specified, or when the person who has to answer for the result cannot explain how it was produced. Those are properties of the arrangement, not of the model, and they are where pilots stall.

MeasuredThe arrangement, not the model
UnitOne bounded workflow
ComparedCurrent practice against one proposal
Ends inAn operating decision, including Stop
II / What we measure instead

Evaluate the arrangement.

Quality, not just speed

Speed is one measure. Quality, review effort, exception handling, authority, and recovery are measured in the context where the work actually has to perform.

The whole arrangement

People, AI, process, and governance are evaluated together, because that is the unit that either works or does not.

Against a real baseline

The current arrangement is measured before the proposed one, so the comparison is against practice rather than against a hope.

With the protocol frozen

Criteria and vocabulary are agreed and locked before testing, so a result cannot be reinterpreted after it arrives.

III / The instrument

The Work Performance Review

The instrument for this is the Work Performance Review: one bounded workflow, one accountable owner, one decision, and a brief that says what the evidence supports.

One critical workflow. One operating decision. A bounded evaluation of how the work performs.

Start with the decision

What do you need to know?

Discuss a workflow

Tell us about the work and the decision you need to make. A paragraph is enough to tell whether this is a fit.

Mail to hello@euler.center is read by Sravan Ankaraju, who founded the lab. Expect a reply within 48 hours, and a straight answer about fit rather than a discovery call.