GunSpec

One Redback task, end to end

The other Redback pages each describe one part. This one puts them together: every stage a task passes through, who or what does it, what each stage produces and where that output goes, followed through a real run from the blind measurement.

The record is the Glock 21 SF, and the field is its empty weight. The run is taken from the blind measurement: the figure we hold was hidden, the page it came from withheld and the record's own citations left out, so Redback met it as an empty field with nothing to start from but the record's name and maker. Everything below is what that run did and produced.

MeasureThis run
Searches1
Pages read3
Excerpts shown to the model3
Models asked2
Model tokens2,253
Time6.2 s

A task can end in three places: discarded with the rule that refused it, filed as a doubt for a person to settle, or filed as a proposal. Only a person turns a proposal into a change to the catalogue, and a rejection sends the task back with the reviewer's note.

DiagramOne task from the finding to the catalogue
100%
One task from the finding to the catalogueA check finds a fault and a task is queued. A run claims it with its record's other tasks. Code searches, reads pages and cuts excerpts; a model reads the excerpts; code finds the quote on the page. Another site, a second model and the evidence rules then decide: refused is discarded with its rule, a doubt goes to a stronger model, and a pass becomes a proposal with its full trace. A person reviews it: accepted goes to the catalogue by pull request, rejected returns to the queue with a note.

The table follows a gap task. A verification skips the search and reads only the maker's page; a contradiction and a vocabulary spelling skip the reading and ask a model to judge what the record already holds. Every task ends in the same filing and the same review.

StageDone byTakesProduces
1. FindingA data quality checkThe catalogueA task: the record, the field, and the check's own line
2. ClaimThe runThe queueThe task and its record's other open tasks, as one dossier: the record, its cited pages and their rank, its maker's site, its findings, and what a person last decided about it
3. SearchCodeThe record's name and the field in wordsResults, the maker's site first; off-topic ones set aside
4. Read pagesCodeThe strongest results, up to sixEach page's own stated facts and its text; a measurement of how the site read
5. Cut excerptsCodeThe pagesUp to sixteen short windows that name the field with a number in them
6. ReadA reader modelThe record without the figure, and the excerptsWhich excerpt, the quote, the figure as stated, a confidence
7. HoldCodeThe readingThe quote found on the page as fetched, the number found in the quote, the figure in the stored unit
8. CorroborateCodeThe figureAny other site read that states the same figure
9. CheckA second model, another providerThe excerpt and the figureYes, no or unsure
10. Evidence rulesCodeThe figure, quote, page and recordPassed, a doubt with its reason, or refused with its rule
11. EscalateA stronger model, on doubt onlyThe same excerptsA second, independent reading
12. FileCodeEverything aboveA proposal: the figure, its words, its sources, confidence, checks, reasoning and full trace
13. ReviewA personThe proposalAccepted, or rejected with a note
14. CatalogueA personThe accepted resultsA pull request, merged by a person, loaded by the next deployment

The same stages, with what each one produced for the Glock 21 SF.

  1. Search. One query for the model's name and the field, with the maker's site first. Among the results was Glock's own page for the model.
  2. Read pages. Three pages were read. Glock's page is classed as the maker's own, the strongest kind of source, and its technical data table was read as stated facts rather than as prose.
  3. Cut excerpts. Three excerpts were cut. Glock's technical data table lists three weights one after another (740 g without magazine, 825 g with an empty magazine, 1100 g with a loaded one), and the excerpt from its page was cut around them.
  4. Read. The reader model named that excerpt and read the figure as 740 g, from the line labelled without magazine.
  5. Hold. Code found the quote on Glock's page as it was fetched, and 740 in the quote. The page states grams, so nothing was converted.
  6. Corroborate. No other site read in this run stated the figure, so the answer rests on the maker's page alone.
  7. Check. A second model, from a different provider, was shown the excerpt and the figure and answered yes.
  8. Evidence rules. The source is the maker, so it is citable. The quote is on the page. The page is titled for the Glock 21 SF, so it names this model. The magazine rule reads only up to the next label, so the 825 g on the next line, labelled with empty magazine, does not count against a figure labelled without one. The figure is within bounds. It passed every rule.
  9. Escalate. Nothing was in doubt, so no stronger model was asked.
  10. File. The result was 740 g at a confidence of 0.95, with the reasoning: read off the maker's page, quote present as fetched, confirmed by a second model. In the measurement it was then scored against the figure we hold, 740 g, and counted right.

One run leaves more than its answer. Each output has one place, and the ones that change what the next run does are either measurements or a person's decisions.

DiagramEvery output of one run, and where it goes
100%
Every output of one run, and where it goesOne run produces a proposal and its trace, which a person accepts (it goes to the catalogue by pull request) or rejects with a note (it heads the next attempt's question). It also leaves a log line per model call, search answers and page text kept for a day, and a measurement of how each site read, which is proposed to the catalogue too.
OutputWhere it goesWho reads itHow long
The proposalThe task, in the review queueA reviewer in the staff consoleKept after review, with the decision
The traceWith the proposalA reviewer, and anyone measuring a runAs long as the proposal
A discarded answerThe task, with the rule that refused itA reviewer, and anyone measuring a runAs long as the task
A log line per model callThe call logOperators, as usage per provider and keyA bounded window
Search answersThe shared search memoryEvery callerOne day
Page textThe shared page memoryEvery callerOne day
How each site readThe site memoryEvery later run, then the catalogue by pull requestKept
A reviewer's rejectionThe taskThe next attempt, at the top of its questionUntil the task is settled
An accepted figureThe catalogue record, with its page and quoteEveryone who reads the recordKept, with its history

A proposal is written to be read top down. Its first line is the reasoning, and a doubt, when there is one, comes first in it: what the rule saw and why it held the figure back. Then the figure as the page stated it and in the stored unit, the quote, and each page that stated it, every one a link to open. Then the checks: whether a second site agreed, what the second model said, whether a stronger model was asked, and what the evidence rules decided.

The trace is underneath, for when the answer is not enough: every question put to a model, the excerpts it was shown, every page and search with what came back, and the provider and model of each step. A figure a reviewer doubts can be followed back to the exact words a model was given.