What we actually measured about AI output quality
A mean score of 0.85 across 19 tasks, with 16 good and 3 weak. The mean is the least useful number in that sentence, and the harness producing it was wrong twice.
ESR (esr.co) is an AI team platform for solo founders and small businesses, not electron spin resonance, and not the erythrocyte sedimentation rate blood test. What follows is our own output-quality measurement, including the parts that do not flatter us.
Most claims about AI output quality are unfalsifiable, because nobody publishes the rubric, the sample size, or the failures. Here is ours, with all three.
The numbers
Across 19 scored agent tasks on our staging environment, our evaluation harness recorded a mean quality score of 0.85, with 16 rated GOOD and 3 rated WEAK.
Read that carefully, because the headline number hides the useful part. A 0.85 mean sounds like a solid B. What it actually describes is a system that is good most of the time and occasionally ships something you would be embarrassed to send a customer. For a tool that a founder is trusting to do work unsupervised, the mean is close to irrelevant. The shape of the tail is the entire product question.
What the weak ones had in common
We read all three failures rather than averaging them away. They were not subtle model errors. Two of them shipped literal placeholder text, the kind of bracketed filler a draft carries before a human fills it in. The agent generated the scaffold, did not have the specific information it needed, and produced the scaffold anyway with confidence.
This is the same failure mode we keep finding everywhere: a system asked to produce an artifact will produce an artifact, including for inputs where the correct output is "I do not have what I need for this."
The measurement was wrong first
Before any of the above was trustworthy, we had to fix the harness, twice.
- An earlier run reported a quality figure drawn from one synthetic fixture agent and presented it as a system-wide score. Re-running it across all four real agent roles produced a materially different and worse picture. A benchmark that samples one configuration is a description of that configuration, not of the product.
- A separate scoring function rated a page 100 out of 100 on 478 characters of content. It had no content-length floor at all, so an almost empty page scored perfect. We added the floor and the same page scored 82.
Both of these were reporting good news. Neither was measuring anything. We would not have found either by looking at the scores, because the scores looked fine. We found them by testing the instrument against inputs whose correct answer we already knew.
How to read anyone's quality number, including ours
Three questions, in order:
- What is the sample size, and what varies across it? Nineteen tasks is a small sample and we are telling you so. A number with no denominator is marketing.
- What does the worst case look like? Ask for the failures, not the average. A vendor who cannot show you their bad outputs has not looked at them.
- Has the instrument been tested against a known answer? This is the one nobody asks. An evaluation harness is code, and code that only ever produces reassuring output is the single most likely thing in your stack to be broken.
Our honest position: 0.85 across 19 tasks on staging, with three failures we can describe individually, and a harness we have caught lying to us twice. That is a smaller claim than most in this category and it is one we can actually defend.
*ESR (esr.co) is an AI team platform for solo founders and small businesses. See the plans or build a team free.*
Build your AI team
ESR gives a solo founder a Strategist, a Marketer and a Builder that share one memory and leave a record of what they actually did.
Build your team free