GunSpec

How Redback is measured

Whether Redback's figures are right is measured, not assumed. Figures we already hold with a checked quote are hidden, Redback researches each one as it would any gap, and code scores what comes back. This page describes the design, the scoring and the results so far.

The answer key is the research ledger: every figure in it was found on a page and written with the words that state it, re-read on that page when it was written, and the record still holds it. It is a stronger key than the records' own confidence scores, which are our judgement and were, in part, built from figures later removed as invented.

The sample is drawn by a seeded shuffle and frozen before any research runs, and records worked in the same week are left out, because their searches and pages are already remembered. The first measurement used 56 figures on 25 records, every record the key held at the time.

DiagramHow Redback's accuracy is measured
100%
How Redback's accuracy is measuredFigures already held with a checked quote form the answer key. A sample drawn from a fixed starting value is frozen before the run. Each figure is hidden and its page withheld, and the real research runs on it. What the evidence rules hold back is counted as held back; what is written is compared with the held figure and counted right or wrong.

For each figure the research runs as it would on a real gap: the field is hidden, the page it came from is withheld, and the record's own citations are not handed over, so the answer has to be found rather than re-read. Redback's run is its real one: the same models, search chain, readers, checks and evidence rules.

Code, not a model, scores. A written figure is right if it is within the tolerance of a rounding somebody did when converting: 2% or 1 mm for a length, 3% or 5 g for a mass. Every hidden figure ends in one of: written and right, written and wrong, filed as a doubt (right or wrong), refused, not found, or nothing to research.

Two numbers matter, and they are reported separately: accuracy, the share of written figures that are right, and coverage, the share of hidden figures written right. A system can buy accuracy with coverage by refusing more; reporting both keeps that trade visible.

Redback was run on the same 56 hidden figures after each round of fixes the runs themselves exposed. Every row is the same key, so the change between rows is the change in Redback.

RunWritten, rightWritten, wrongDoubts (right / wrong)Discarded
First measured run2340 / 114
After the source and variant rules1818 / 511
After the table, identity and number fixes2318 / 55

The one wrong figure written in the last run is a genuine disagreement: a reference work's weight for a rifle's military version against its maker's weight for the civilian one. Five wrong figures that earlier runs wrote now reach a person as doubts instead.

Each round was driven by reading what went wrong, not by tuning a number: a pounds-and-ounces parse that turned 771 g into 454 g, a page's title taken from a product code, an unclassified airsoft site, a review of a Full-Size model for a Compact record.

A held figure can itself be wrong, so a written wrong is a disagreement to read, occasionally in Redback's favour. The web moves, so a run is a measurement of its day. And 56 figures is a wide interval, not a rate: no wrong figure through in 56 is a strong result, not a proof of none.