Most AI evaluation measures the model. The question that decides a rollout is whether the work performs, which is a different measurement.

A benchmark tells you how a model scores on tasks someone else chose. It does not tell you what happens when your analyst has to check the output, when an exception arrives that nobody specified, or when the person who has to answer for the result cannot explain how it was produced. Those are properties of the arrangement, not of the model, and they are where pilots stall.
Speed is one measure. Quality, review effort, exception handling, authority, and recovery are measured in the context where the work actually has to perform.
People, AI, process, and governance are evaluated together, because that is the unit that either works or does not.
The current arrangement is measured before the proposed one, so the comparison is against practice rather than against a hope.
Criteria and vocabulary are agreed and locked before testing, so a result cannot be reinterpreted after it arrives.
The instrument for this is the Work Performance Review: one bounded workflow, one accountable owner, one decision, and a brief that says what the evidence supports.
One critical workflow. One operating decision. A bounded evaluation of how the work performs.
Tell us about the work and the decision you need to make. A paragraph is enough to tell whether this is a fit.
Mail to hello@euler.center is read by Sravan Ankaraju, who founded the lab. Expect a reply within 48 hours, and a straight answer about fit rather than a discovery call.