Synthetic sample scorecard
The cheaper candidate lost.
A transparent demonstration of the measurement and rejection process using six invented scenarios and a local model. This is not a customer result or a savings claim.
Synthetic proof only. Six cases demonstrate the harness plumbing and acceptance logic. They do not establish production reliability, customer savings, or cloud-provider cost.
| Strategy | Quality gates | Total tokens | Average latency | Eligible? |
|---|---|---|---|---|
| Verbose baseline | 6 / 6 | 1,196 | 1,844 ms | Yes |
| Shorter candidate | 5 / 6 | 904 | 1,566 ms | No |
Decision
Reject the shorter strategy.
The candidate used 292 fewer tokens and reduced average latency by 278 milliseconds, but it incorrectly excluded an eligible provider in one comparison. The declared threshold was 100%, so the efficiency gain could not be recommended.
What the harness demonstrated
- Token and latency measurement on identical cases.
- Automated checks for JSON structure, required keys, and scenario criteria.
- Retention of failed cases instead of removing them from the report.
- A lower-resource candidate can—and should—lose when quality regresses.
Limitation found
The rubric itself needed correction.
The first run used overly literal keyword requirements and rejected semantically correct answers. Those initial results were preserved. The rubric was corrected to test the actual decision requirements, then rerun.
Before real customer proof
The evaluation set should expand to 20–30 representative cases, add current provider pricing, define business-specific scoring, and compare at least one multi-provider workflow using approved sanitized data.
Residual Forge / founding cohort
Know which AI costs are buying quality—and which are just waste.
Tell us which AI workflow you run repeatedly and where cost, latency, or quality is creating friction.