Validation
Measured, not magical.
OpAgent publishes measured results with the methodology attached — sample size, ground truth, and date. A number without a method is marketing; these are measurements.
Weighted combined accuracy across 359 held-out surgical cases spanning 5 procedure families. 97.6% on non-ambiguous outputs; 4.2% of cases self-flagged as ambiguous and routed to human review — the engine tells you when it is guessing.
n=359 · locked test set · April 2026 · scoped to CodeIQ OR, not implied across agents
Live production count
How this was measured
The validation dataset is 3,951 de-identified operative cases (2018–2026) from a single high-volume arthroplasty and trauma practice, held as a Limited Data Set. Cases were split with a fixed seed into few-shot, evaluation, and locked-test partitions; the locked test set was physically quarantined and never viewed during prompt iteration. Ground truth is the operating surgeon's own billing-record CPT assignments — the surgeon performed and coded every case. Prompt changes were accepted only when the evaluation set improved; any change that regressed it was reverted, not shipped.
| Procedure family | Test n | Combined accuracy |
|---|---|---|
| Hip arthroplasty | 196 | 98.5% |
| Knee arthroplasty | 140 | 96.4% |
| Hip trauma | 11 | 90.9% |
| Knee non-arthroplasty | 10 | 100.0% |
| Hip infection | 2 | 100.0% |
| Weighted | 359 | 97.5% |
Scope: this validation covers CodeIQ OR's surgical coding only. It is not a claim about other OpAgent agents, other specialties, or national representativeness. Sample size, methodology, and date update together when the validation is re-run.
In production
RadiologiQ has drafted 6,026 imaging reports in production since April 2026 — every one reviewed and signed by a physician. The count above is live from the production system, not a marketing snapshot.