How Redback works a task
Redback is the agent that works the data quality queue. This page follows one task from the check that raised it to the record a person may change, and says at each step what is done by code, what is asked of a model, and what is left to a person.
What Redback is, and what it is not
The data quality checks find faults: a figure invented and removed, two records that contradict each other, a spelling no vocabulary knows, a figure no page supports. Each finding becomes a task. Redback is the agent that works those tasks, and it works four kinds: a gap (a figure to find), a contradiction (a disagreement to judge), a verification (a figure to confirm on its maker's page), and a vocabulary spelling (a value to map to a known term).
Redback never writes to the catalogue. Its work ends in a proposal: a value, the page it came from, the words on that page that state it, and its reasoning. A person reads the proposal and accepts or rejects it, and only an accepted proposal can reach a record, through a pull request a person merges. A task Redback closes as resolved means a proposal was filed, never that the record changed.
The life of a finding
The diagram follows one finding. What is automatic and what is judged are kept apart: everything up to the proposal is code and models, everything after it is a person.
A run
Redback works in runs. A run is started by the schedule, by a person pressing run in the staff console, or by a task whose wait has expired. It claims a batch of tasks in the order its assignment sets, and each claim is conditional, so two runs can never work the same task.
When a run takes a task on a record, it also takes the record's other open tasks (up to six) and works them as one dossier. The record is read once: its cited pages with their standing, its maker's page, its open findings, and what a person last confirmed about it. Every page and search in the run is kept for the rest of the run, so the maker's page read for a length is not fetched again for a weight. Each task still gets its own answer and its own review, because a reviewer judges one figure at a time.
A task a provider could not serve (a busy host, keys all resting, a spent allowance) is deferred rather than failed: it is released with a time to try again, from the provider's own retry time or a doubling wait up to an hour, and after six deferrals it fails. A run that finds a task it cannot work as generated skips it; one bad task never stops the queue behind it.
Research is code
For a gap, research is done by code and not by a model. A model given a search tool and told to find a figure was measured to plan poorly, search the same thing three times and read whole pages it could not hold; the division of labour fixes each of those.
- Start from the record's own cited pages, strongest source first: a maker's page before a reference work before the press before a shop.
- Search, with the record's name and the field in words, and the maker's own site first.
- Read the strongest results, up to six pages a task, preferring a maker, a standards body, a government or a reference work over the press, shops and forums. A reference work's own citations are followed to the pages it cites.
- Cut each page to the few excerpts that name the field: windows around the words a page uses for it, with a number in them, up to sixteen excerpts a task. Each excerpt is a slice of the page with only whitespace collapsed, so a quote copied from it can be found on the page again.
The model reads
The model is shown the record (without the figure it is looking for), what the check found, and the excerpts. It is asked for one thing: which excerpt states the figure, the words that state it, and the value as written. It is not given tools, it does not search, and it does not convert units; code does the arithmetic.
Only when research found nothing to read at all does a tool-using fallback run, in which a model may search and fetch within a fixed budget of twenty tool calls. Contradictions work that way from the start, because judging which of two records is wrong needs looking around; a verification reads only the maker's page, chosen by code.
Checked before it is filed
A reading is held by code before anything else: the excerpt it names must be one of those shown, the quote must be on the page as it was fetched, and the number the model says it read must be in the quote. A real quote beside an invented number is still an invented number.
Then three independent checks. Code looks for the same figure on another site among the pages read, which raises the confidence and adds evidence. A second model, on a different provider, is shown the excerpt and asked only whether it states the figure: yes, no or unsure. And the evidence rules (the next page) decide whether the figure can be written at all.
A reading in doubt goes to a stronger model: a no from the second model, an unsure nothing corroborates, a confidence under 0.6, a size read off a shipping box, a reading that was discarded or could not be parsed. The stronger model reads the same excerpts and answers independently. A plain none of these state it is not escalated; it is the commonest honest answer, and escalating it would double the cost of honesty.
Filed for a person
What is filed carries its evidence: the value in the stored unit, the words as the page wrote them, the quote, every page that stated it, the confidence, what each check said, and the full trace of what was read and asked. A doubtful figure is filed with its doubt first and its confidence capped at 0.3.
A person reviews it in the staff console. An accepted result is proposed to the catalogue in the next pull request; a rejected one is kept, and the rejection and its note lead the question the next attempt is asked, with the instruction not to give the same answer again unless it answers the objection.