How Redback is measured
Whether Redback's figures are right is measured, not assumed. Figures we already hold with a checked quote are hidden, Redback researches each one as it would any gap, and code scores what comes back. This page describes the design, the scoring and the results so far.
The answer key
The answer key is the research ledger: every figure in it was found on a page and written with the words that state it, re-read on that page when it was written, and the record still holds it. It is a stronger key than the records' own confidence scores, which are our judgement and were, in part, built from figures later removed as invented.
The sample is drawn by a seeded shuffle and frozen before any research runs, and records worked in the same week are left out, because their searches and pages are already remembered. The first measurement used 56 figures on 25 records, every record the key held at the time.
The blindfold
For each figure the research runs as it would on a real gap: the field is hidden, the page it came from is withheld, and the record's own citations are not handed over, so the answer has to be found rather than re-read. Redback's run is its real one: the same models, search chain, readers, checks and evidence rules.
The scoring
Code, not a model, scores. A written figure is right if it is within the tolerance of a rounding somebody did when converting: 2% or 1 mm for a length, 3% or 5 g for a mass. Every hidden figure ends in one of: written and right, written and wrong, filed as a doubt (right or wrong), refused, not found, or nothing to research.
Two numbers matter, and they are reported separately: accuracy, the share of written figures that are right, and coverage, the share of hidden figures written right. A system can buy accuracy with coverage by refusing more; reporting both keeps that trade visible.
Results
Redback was run on the same 56 hidden figures after each round of fixes the runs themselves exposed. Every row is the same key, so the change between rows is the change in Redback.
| Run | Written, right | Written, wrong | Doubts (right / wrong) | Discarded |
|---|---|---|---|---|
| First measured run | 23 | 4 | 0 / 1 | 14 |
| After the source and variant rules | 18 | 1 | 8 / 5 | 11 |
| After the table, identity and number fixes | 23 | 1 | 8 / 5 | 5 |
The one wrong figure written in the last run is a genuine disagreement: a reference work's weight for a rifle's military version against its maker's weight for the civilian one. Five wrong figures that earlier runs wrote now reach a person as doubts instead.
Each round was driven by reading what went wrong, not by tuning a number: a pounds-and-ounces parse that turned 771 g into 454 g, a page's title taken from a product code, an unclassified airsoft site, a review of a Full-Size model for a Compact record.
What the measurement cannot tell
A held figure can itself be wrong, so a written wrong is a disagreement to read, occasionally in Redback's favour. The web moves, so a run is a measurement of its day. And 56 figures is a wide interval, not a rate: no wrong figure through in 56 is a strong result, not a proof of none.