How Redback chooses its models and tools
Redback is not one model. Each job inside a task (reading excerpts, checking a reading, reading again when in doubt, judging a contradiction, placing a spelling) is its own request, with its own chain of models chosen for what that job needs. Code does the searching and fetching, and keeps what it has paid for. This page describes how a job gets its model, how a call finds a working key, how search works, and what is kept and for how long.
Jobs, not one agent
Redback works four kinds of task: a gap (a figure to find), a contradiction (a conflict to judge), a verification (a figure to confirm on its maker's page) and a vocabulary spelling (a value to place in a known term). Each kind is split into jobs, and each job is one request to a model with one question. A model is never asked to plan a task: code decides what to read and in what order, and asks a model only the part that needs reading or judgement.
| Job | What it is asked | What it needs from a model | Its default chain starts with |
|---|---|---|---|
| Reading a gap | Which excerpt states the figure, in what words | Quotes verbatim, and answers none rather than guess | A large open model on a fast host |
| Checking a reading | Does this excerpt state this figure: yes, no or unsure | One word, from a different provider than the reader | A mid-size model on another host |
| Escalating a doubt | The same excerpts, answered again | The strongest model available, since it runs only on the hard cases | The largest model in the chain |
| Verifying on the maker's page | What the page states for each checked figure, with the sentence | Never shown our value, so it cannot simply agree with it | A large open model on a fast host |
| Judging a contradiction | Is the check right that something is wrong | Reasoning over the record, and saying cannot tell | A reasoning model |
| Placing a spelling | Second spelling, new value, or wrong question | One of three words, from a list given in full | A fast host |
Tools are functions a model may call, each described in words it chooses by (the staff console shows them verbatim), and each states what it reaches: the web, our own catalogue, or nothing outside the platform. Only a contradiction, and a gap with nothing to read, are given tools. A tool never fails a run: an error comes back as a sentence the model can work with, or one dead site would end a run.
The tools: a web search and a search within one site; a page reader, a PDF reader, an encyclopaedia search and article reader; a structured data query; three catalogue readers (a record, a search, a record's relations); a unit converter; a source's standing; and a notepad.
How a job gets its model
Each job resolves its chain in a fixed order, and the staff console shows which step answered. First, a chain an operator assigned to that job. Otherwise, every model in the pool tagged with the capability the job declares (extraction, classification, reasoning and so on), in the operator's pool order. Otherwise, the job's own default chain, and only then the platform's single default model.
The order is always the operator's, never derived. Providers report their limits differently or not at all, and ranking by whichever looks most available would route on a number we made up. Tagging a model once for a capability configures every job that declares it, including jobs written later.
The default chains follow three rules. No router that picks a model per call, because then the model behind an answer is unknowable. No small model, because the failure that matters is a plausible invented number. And the hops alternate between providers and the hosts behind them, because consecutive hops on one overloaded host fail together.
The check on a reading avoids the provider that produced it: its chain is filtered to other providers, and only when none is left does it fall back to the same one. A model is better at catching another model's misreading than its own.
How a call finds a working key
A model is named together with its provider, because the same open model is served by several providers with different limits and different bills. Each provider holds a ring of keys, sealed in the staff console and opened only by the platform; no caller ever holds a key.
- Keys are tried least recently used first, so five keys share the load rather than one working while four wait.
- Only a key's own fault moves the call to the next key. A rate limit rests the key, an empty balance marks it spent, and a refused key is switched off. A timeout or an unreadable answer would fail the same on every key, so it is an answer, not a reason to spend four more requests arriving at it.
- A rate limit that names one model locks only that model on that key, until the moment the provider says it lifts; a daily limit rests the whole key. When every key is locked, the call reports the wait, and the task is deferred until then rather than failed.
- A provider with no usable key is skipped for the next in the chain. A busy host or a retired model moves to the next hop. A timeout or an unreadable answer stops the chain, because the next model would be sent the same thing.
- Usage is counted per provider and reserved once per answer, not per attempt: a key that was refused used nothing, and one provider reaching its limit never locks out the providers behind it in the chain.
Every call is logged with its provider, key, model, tokens, cost, time and outcome, and every step of a task's trace names the provider and model that answered it, so a figure can always be followed back to the model that read it.
The search chain
Search runs through its own chain of providers, in an order the operator sets, with key rings on the same rules as the models.
Each provider's allowance is read the way that provider publishes it. Where the allowance belongs to each key, the keys are used one after another until each is spent. Where it belongs to the account, only the first key is sent, since a second key on the same account buys nothing more. A key that hits a rate limit or runs out of credit rests and the next is tried; a refused key is switched off.
A provider with no usable key passes the query on. An answer about something else (results that never name the query's defining words) passes it on too, and is kept in case nothing better comes. No search is refused to save allowance.
What is kept, and for how long
Every search and page Redback fetches is kept long enough to be reused, and no longer. A kept copy only ever saves a repeat: it never becomes evidence on its own, and what failed is never kept.
| What | Kept for | Shared with | Why |
|---|---|---|---|
| A search answer | One day | Every caller | The same question, in any word order, costs nothing the second time. A failure or an off-topic answer is not kept, so it is asked again. |
| A page, as extracted text | One day | Every caller | A batch on one maker reads the same page many times. Only a page that was read is kept, never an error, and never the raw markup. |
| A site's robots rules | Checked before any kept page is used | Every caller | A publisher who says no is not served from our own copy. |
| The pages and searches of a run | The run | Every task on the record | A maker's page read for a length is not fetched again for a weight. |
| Site memory and accepted maker facts | A few minutes | Every run | A run reads them dozens of times; a fact accepted a minute ago reaches the next run. |
| A model's rate limit on a key | Until the stated reset | Every job | The ring skips that model on that key, and nothing else. |
A kept page cannot produce a citation that was never there: the quote is checked against the text that run was handed, so a kept page is checked against itself, and the proposal carries the address it was read from for a person to open.
Reading a page
A page is fetched politely: its site's robots rules are honoured, requests to one site are spaced, and what was extracted is kept for a day. The reader separates what a page states in its own tables and structured data from its prose, and gives the facts first. A table with a header row is read one row at a time with that row's labels, so a spec table's barrel keeps the caliber it belongs to; a table without headers is left alone, because flattening a comparison invents relationships.
PDFs are read as documents, with their tables rebuilt (Reading a PDF, below). Pages drawn only in a browser, and verification walls, cannot be read by Redback: a wall is recognised and never tried past, and such sites are marked in the memory (What it learns) so research looks for a maker's PDF or an archived copy instead.
Reading a PDF
Makers publish their specifications as PDFs more often than as pages: spec sheets, owner's manuals, datasheets and product catalogues. Redback reads them as documents, in our own Worker, rather than as a page, because a PDF read as a page comes back as noise that a model would treat as read.
The hard part is the tables. A PDF has no table in it, only words placed at points on a page, so reading it in the order the file stores them runs a table's labels into one line and its figures into another, and which figure belongs to which label is lost. Redback rebuilds the table from where each word sits instead.
- Words on one baseline become a line, and a wide gap within a line starts a new column.
- A run of label lines above a line of figures is read as the table's header.
- Each figure is placed under the label whose column it sits in, and the row is written out as a fact: "Barrel length: 4.7 in, Weight: 710 g", the same shape a table on a web page gives.
- This is done on every page of the document, not only the first few, because a catalogue's specification table is often forty pages in. The document's prose is read separately, in order, so a two-column paragraph still reads the way it was written.
| What the document is | What Redback does |
|---|---|
| A spec sheet or manual with tables | Its tables come first as facts, row by row with their labels, followed by the prose |
| A long catalogue or manual | The tables are gathered from every page; the prose follows from the start, as far as the reading budget for one document allows |
| A scanned PDF with no text in it | Said plainly: the document has pages but no text to read. Nothing is guessed from it, and Redback looks for another source |
| A file over twelve megabytes, or one that will not open | Refused with the reason, so the task does not treat it as read |
| A site whose robots rules forbid the path | Not fetched, and the refusal is said |
A figure Redback takes from a PDF is held to the same rule as one from a page: it must be quoted, and before it can reach a record the quote is checked again against the document, read the same way. A quote taken from a rebuilt table is checked against that table, so the figure and its label have to agree a second time.
Reference works
An encyclopedia article is a pointer, never a citation. Its infobox is read as fields, grouped as the article groups them, and its references are ranked by the source hierarchy and followed to the pages they cite. A structured-data lookup gives a second opinion in the same way: a lead to a page, not a value.
The catalogue itself
A contradiction is a disagreement between two records, and a task carries one of them; judging whether a firearm predates its maker needs the maker. The catalogue tools read a record, search the catalogue and follow a record's relations, each against an allow-list, because a table name cannot be a bound parameter. The standing of a source tells a model a site's kind, rank and how often the catalogue cites it, and how that site reads to a script, before it quotes it.
The notepad
A model in the tool loop has a scratchpad of twelve notes of 240 characters, re-shown on every turn from outside the conversation, so pruning the history cannot remove a conclusion. It is for what was ruled out: the measured failure was three near-identical searches, because nothing recorded that the first two had found nothing. A note is not evidence; a figure in a note has no source behind it.