GunSpec

What Redback learns, and who decides

Redback keeps what it finds out between runs, and the rule for what it may keep is the distinction between a measurement and a claim. A measurement is kept as it is taken; a claim waits for a person. This page describes each kind of memory and the way both reach the reviewed catalogue.

Redback keeps two kinds of thing between runs, and they are treated differently on purpose. A measurement (how a site responded when it was read) is kept as it is taken, because it is not a claim about any gun and there is nothing to review. A claim (that a maker says two variants share a figure) changes what a figure may be written from, so it waits for a person.

DiagramWhat Redback keeps, and who may change it
100%
What Redback keeps, and who may change itEvery page read updates how each site reads, a measurement that is not reviewed. A maker's statement, with its page and words, is proposed as a maker fact, and a person accepts it (then the evidence rules use it) or rejects it (kept, never used). Both live in the platform's database, are proposed to the catalogue, are merged there by a person, and are loaded by the next deployment.

Every page Redback reads updates its site's counts, and the team's own measurements add to them: how many reads got the page as served, how many needed a browser (which only the team's measurements can tell), how many met a verification wall, how many failed, and how many pages stated figures at all. From those counts each site is filed under one word: reads plainly, drawn in the browser, a verification wall, gone, no figures, or mixed.

Research uses it before spending anything: a maker's site that answers scripts with a wall is not searched, and the model is told to look for the maker's PDF, a review or an archived copy. A person can add notes to a site (Tech Specs print barrel length without a unit), and no measurement overwrites them.

A maker fact is a statement, with its page and its words, that a figure is the same across a model's chamberings. It is proposed by a person or by the team's research tools (which check its words are on the page before filing it), and used only once a person accepts it in the staff console. It confirms a chambering doubt for the fields it names and nothing else: Walther's CCP M2 in .380 shares the 9 mm model's barrel and length, and says nothing of its weight, which differs.

An accepted fact can be re-read on its page at any time, and one whose words have left its page no longer confirms anything.

Search answers and pages are kept for a day (see How Redback chooses its models and tools). That is a saving, not learning: nothing about them outlives the day, and nothing in them changes what counts as evidence.

A person's rejection is memory too. The rejected answer and the reviewer's note lead the next attempt's question, and a result filed again carries the earlier attempt with it, so a reviewer sees both.

Each proposal carries its full trace: the question each model was asked, the excerpts it was shown, every tool call with what it returned, the provider and model of each step, the tokens and cost, and what each check said. A discarded answer keeps its trace too, with the rule that refused it.

Each model call is also logged on its own, with its outcome, so the staff console can show which provider served, which key rested and why, and what a day of research cost.

The live memory is in the platform's database; the reviewed record is in the catalogue, the repository every figure lives in. A deployment is loaded from the catalogue without undoing what reads taught the database; what the database learned (new site measurements, accepted facts) is proposed back to the catalogue in the same pull request that carries accepted results, and a person merges it. A database not yet loaded never proposes an empty memory.

So nothing in the catalogue is written by the platform: it learns in the database, a person reviews, and a pull request carries the reviewed state home.

Redback does not retrain or tune itself. It improves in two ways, and both can be seen. What it measures (how sites read) and what a person decides (maker facts, rejections) change the next run automatically. Everything else changes when a person reads what went wrong and changes a rule.

  1. Measure blind against the answer key (see Measuring Redback).
  2. Read the trace of every wrong or discarded figure, to find the step that failed: the search, the excerpt, the reading, or a rule.
  3. Change that step, with a test built from the case, never by adjusting the number that scored it.
  4. Measure again on the same key, so the difference between the runs is the change.

Each round of fixes in the results on the measurement page came from this loop.