One Redback task, end to end
The other Redback pages each describe one part. This one puts them together: every stage a task passes through, who or what does it, what each stage produces and where that output goes, followed through a real run from the blind measurement.
The case
The record is the Glock 21 SF, and the field is its empty weight. The run is taken from the blind measurement: the figure we hold was hidden, the page it came from withheld and the record's own citations left out, so Redback met it as an empty field with nothing to start from but the record's name and maker. Everything below is what that run did and produced.
| Measure | This run |
|---|---|
| Searches | 1 |
| Pages read | 3 |
| Excerpts shown to the model | 3 |
| Models asked | 2 |
| Model tokens | 2,253 |
| Time | 6.2 s |
The whole path
A task can end in three places: discarded with the rule that refused it, filed as a doubt for a person to settle, or filed as a proposal. Only a person turns a proposal into a change to the catalogue, and a rejection sends the task back with the reviewer's note.
Every stage, and what it produces
The table follows a gap task. A verification skips the search and reads only the maker's page; a contradiction and a vocabulary spelling skip the reading and ask a model to judge what the record already holds. Every task ends in the same filing and the same review.
| Stage | Done by | Takes | Produces |
|---|---|---|---|
| 1. Finding | A data quality check | The catalogue | A task: the record, the field, and the check's own line |
| 2. Claim | The run | The queue | The task and its record's other open tasks, as one dossier: the record, its cited pages and their rank, its maker's site, its findings, and what a person last decided about it |
| 3. Search | Code | The record's name and the field in words | Results, the maker's site first; off-topic ones set aside |
| 4. Read pages | Code | The strongest results, up to six | Each page's own stated facts and its text; a measurement of how the site read |
| 5. Cut excerpts | Code | The pages | Up to sixteen short windows that name the field with a number in them |
| 6. Read | A reader model | The record without the figure, and the excerpts | Which excerpt, the quote, the figure as stated, a confidence |
| 7. Hold | Code | The reading | The quote found on the page as fetched, the number found in the quote, the figure in the stored unit |
| 8. Corroborate | Code | The figure | Any other site read that states the same figure |
| 9. Check | A second model, another provider | The excerpt and the figure | Yes, no or unsure |
| 10. Evidence rules | Code | The figure, quote, page and record | Passed, a doubt with its reason, or refused with its rule |
| 11. Escalate | A stronger model, on doubt only | The same excerpts | A second, independent reading |
| 12. File | Code | Everything above | A proposal: the figure, its words, its sources, confidence, checks, reasoning and full trace |
| 13. Review | A person | The proposal | Accepted, or rejected with a note |
| 14. Catalogue | A person | The accepted results | A pull request, merged by a person, loaded by the next deployment |
The run, stage by stage
The same stages, with what each one produced for the Glock 21 SF.
- Search. One query for the model's name and the field, with the maker's site first. Among the results was Glock's own page for the model.
- Read pages. Three pages were read. Glock's page is classed as the maker's own, the strongest kind of source, and its technical data table was read as stated facts rather than as prose.
- Cut excerpts. Three excerpts were cut. Glock's technical data table lists three weights one after another (740 g without magazine, 825 g with an empty magazine, 1100 g with a loaded one), and the excerpt from its page was cut around them.
- Read. The reader model named that excerpt and read the figure as 740 g, from the line labelled without magazine.
- Hold. Code found the quote on Glock's page as it was fetched, and 740 in the quote. The page states grams, so nothing was converted.
- Corroborate. No other site read in this run stated the figure, so the answer rests on the maker's page alone.
- Check. A second model, from a different provider, was shown the excerpt and the figure and answered yes.
- Evidence rules. The source is the maker, so it is citable. The quote is on the page. The page is titled for the Glock 21 SF, so it names this model. The magazine rule reads only up to the next label, so the 825 g on the next line, labelled with empty magazine, does not count against a figure labelled without one. The figure is within bounds. It passed every rule.
- Escalate. Nothing was in doubt, so no stronger model was asked.
- File. The result was 740 g at a confidence of 0.95, with the reasoning: read off the maker's page, quote present as fetched, confirmed by a second model. In the measurement it was then scored against the figure we hold, 740 g, and counted right.
Where every output goes
One run leaves more than its answer. Each output has one place, and the ones that change what the next run does are either measurements or a person's decisions.
| Output | Where it goes | Who reads it | How long |
|---|---|---|---|
| The proposal | The task, in the review queue | A reviewer in the staff console | Kept after review, with the decision |
| The trace | With the proposal | A reviewer, and anyone measuring a run | As long as the proposal |
| A discarded answer | The task, with the rule that refused it | A reviewer, and anyone measuring a run | As long as the task |
| A log line per model call | The call log | Operators, as usage per provider and key | A bounded window |
| Search answers | The shared search memory | Every caller | One day |
| Page text | The shared page memory | Every caller | One day |
| How each site read | The site memory | Every later run, then the catalogue by pull request | Kept |
| A reviewer's rejection | The task | The next attempt, at the top of its question | Until the task is settled |
| An accepted figure | The catalogue record, with its page and quote | Everyone who reads the record | Kept, with its history |
Reading a proposal
A proposal is written to be read top down. Its first line is the reasoning, and a doubt, when there is one, comes first in it: what the rule saw and why it held the figure back. Then the figure as the page stated it and in the stored unit, the quote, and each page that stated it, every one a link to open. Then the checks: whether a second site agreed, what the second model said, whether a stronger model was asked, and what the evidence rules decided.
The trace is underneath, for when the answer is not enough: every question put to a model, the excerpts it was shown, every page and search with what came back, and the provider and model of each step. A figure a reviewer doubts can be followed back to the exact words a model was given.