Teams instrument their agents before they grade them, 89 percent run observability and only 52 percent run evals. Watching what an agent did is not the same as knowing whether it was any good.

As open coding models hit similar capability ceilings, the differentiator is internal evals tied to your product. Here is one you will actually run.