GunSpec

How Redback chooses its models and tools

Redback is not one model. Each job inside a task (reading excerpts, checking a reading, reading again when in doubt, judging a contradiction, placing a spelling) is its own request, with its own chain of models chosen for what that job needs. Code does the searching and fetching, and keeps what it has paid for. This page describes how a job gets its model, how a call finds a working key, how search works, and what is kept and for how long.

Redback works four kinds of task: a gap (a figure to find), a contradiction (a conflict to judge), a verification (a figure to confirm on its maker's page) and a vocabulary spelling (a value to place in a known term). Each kind is split into jobs, and each job is one request to a model with one question. A model is never asked to plan a task: code decides what to read and in what order, and asks a model only the part that needs reading or judgement.

JobWhat it is askedWhat it needs from a modelIts default chain starts with
Reading a gapWhich excerpt states the figure, in what wordsQuotes verbatim, and answers none rather than guessA large open model on a fast host
Checking a readingDoes this excerpt state this figure: yes, no or unsureOne word, from a different provider than the readerA mid-size model on another host
Escalating a doubtThe same excerpts, answered againThe strongest model available, since it runs only on the hard casesThe largest model in the chain
Verifying on the maker's pageWhat the page states for each checked figure, with the sentenceNever shown our value, so it cannot simply agree with itA large open model on a fast host
Judging a contradictionIs the check right that something is wrongReasoning over the record, and saying cannot tellA reasoning model
Placing a spellingSecond spelling, new value, or wrong questionOne of three words, from a list given in fullA fast host

Tools are functions a model may call, each described in words it chooses by (the staff console shows them verbatim), and each states what it reaches: the web, our own catalogue, or nothing outside the platform. Only a contradiction, and a gap with nothing to read, are given tools. A tool never fails a run: an error comes back as a sentence the model can work with, or one dead site would end a run.

The tools: a web search and a search within one site; a page reader, a PDF reader, an encyclopaedia search and article reader; a structured data query; three catalogue readers (a record, a search, a record's relations); a unit converter; a source's standing; and a notepad.

Each job resolves its chain in a fixed order, and the staff console shows which step answered. First, a chain an operator assigned to that job. Otherwise, every model in the pool tagged with the capability the job declares (extraction, classification, reasoning and so on), in the operator's pool order. Otherwise, the job's own default chain, and only then the platform's single default model.

The order is always the operator's, never derived. Providers report their limits differently or not at all, and ranking by whichever looks most available would route on a number we made up. Tagging a model once for a capability configures every job that declares it, including jobs written later.

The default chains follow three rules. No router that picks a model per call, because then the model behind an answer is unknowable. No small model, because the failure that matters is a plausible invented number. And the hops alternate between providers and the hosts behind them, because consecutive hops on one overloaded host fail together.

The check on a reading avoids the provider that produced it: its chain is filtered to other providers, and only when none is left does it fall back to the same one. A model is better at catching another model's misreading than its own.

DiagramHow a job gets its model, and a call its key
100%
How a job gets its model, and a call its keyA job uses the chain an operator assigned to it; otherwise every model tagged for what the job needs; otherwise its own default chain. The chain is tried in order. A model with no available key on its provider is skipped. A busy host or a retired model moves to the next model. A timeout or an unreadable answer stops the chain, because the next model would be sent the same thing. An answer is returned with its usage recorded.

A model is named together with its provider, because the same open model is served by several providers with different limits and different bills. Each provider holds a ring of keys, sealed in the staff console and opened only by the platform; no caller ever holds a key.

  1. Keys are tried least recently used first, so five keys share the load rather than one working while four wait.
  2. Only a key's own fault moves the call to the next key. A rate limit rests the key, an empty balance marks it spent, and a refused key is switched off. A timeout or an unreadable answer would fail the same on every key, so it is an answer, not a reason to spend four more requests arriving at it.
  3. A rate limit that names one model locks only that model on that key, until the moment the provider says it lifts; a daily limit rests the whole key. When every key is locked, the call reports the wait, and the task is deferred until then rather than failed.
  4. A provider with no usable key is skipped for the next in the chain. A busy host or a retired model moves to the next hop. A timeout or an unreadable answer stops the chain, because the next model would be sent the same thing.
  5. Usage is counted per provider and reserved once per answer, not per attempt: a key that was refused used nothing, and one provider reaching its limit never locks out the providers behind it in the chain.

Every call is logged with its provider, key, model, tokens, cost, time and outcome, and every step of a task's trace names the provider and model that answered it, so a figure can always be followed back to the model that read it.

Search runs through its own chain of providers, in an order the operator sets, with key rings on the same rules as the models.

Each provider's allowance is read the way that provider publishes it. Where the allowance belongs to each key, the keys are used one after another until each is spent. Where it belongs to the account, only the first key is sent, since a second key on the same account buys nothing more. A key that hits a rate limit or runs out of credit rests and the next is tried; a refused key is switched off.

A provider with no usable key passes the query on. An answer about something else (results that never name the query's defining words) passes it on too, and is kept in case nothing better comes. No search is refused to save allowance.

DiagramHow a search finds a key and a provider
100%
How a search finds a key and a providerA query asked already that day is answered from the saved answer without a new request. Otherwise the next provider in the chain is tried, with a key that has allowance left and is not resting. A key's own fault, a rate limit or no credit, rests that key and the ring moves on; a provider with no usable key hands the query to the next. An answer about something else also moves on. An on-topic answer is returned and its spend recorded.

Every search and page Redback fetches is kept long enough to be reused, and no longer. A kept copy only ever saves a repeat: it never becomes evidence on its own, and what failed is never kept.

WhatKept forShared withWhy
A search answerOne dayEvery callerThe same question, in any word order, costs nothing the second time. A failure or an off-topic answer is not kept, so it is asked again.
A page, as extracted textOne dayEvery callerA batch on one maker reads the same page many times. Only a page that was read is kept, never an error, and never the raw markup.
A site's robots rulesChecked before any kept page is usedEvery callerA publisher who says no is not served from our own copy.
The pages and searches of a runThe runEvery task on the recordA maker's page read for a length is not fetched again for a weight.
Site memory and accepted maker factsA few minutesEvery runA run reads them dozens of times; a fact accepted a minute ago reaches the next run.
A model's rate limit on a keyUntil the stated resetEvery jobThe ring skips that model on that key, and nothing else.

A kept page cannot produce a citation that was never there: the quote is checked against the text that run was handed, so a kept page is checked against itself, and the proposal carries the address it was read from for a person to open.

A page is fetched politely: its site's robots rules are honoured, requests to one site are spaced, and what was extracted is kept for a day. The reader separates what a page states in its own tables and structured data from its prose, and gives the facts first. A table with a header row is read one row at a time with that row's labels, so a spec table's barrel keeps the caliber it belongs to; a table without headers is left alone, because flattening a comparison invents relationships.

PDFs are read as documents, with their tables rebuilt (Reading a PDF, below). Pages drawn only in a browser, and verification walls, cannot be read by Redback: a wall is recognised and never tried past, and such sites are marked in the memory (What it learns) so research looks for a maker's PDF or an archived copy instead.

Makers publish their specifications as PDFs more often than as pages: spec sheets, owner's manuals, datasheets and product catalogues. Redback reads them as documents, in our own Worker, rather than as a page, because a PDF read as a page comes back as noise that a model would treat as read.

The hard part is the tables. A PDF has no table in it, only words placed at points on a page, so reading it in the order the file stores them runs a table's labels into one line and its figures into another, and which figure belongs to which label is lost. Redback rebuilds the table from where each word sits instead.

  1. Words on one baseline become a line, and a wide gap within a line starts a new column.
  2. A run of label lines above a line of figures is read as the table's header.
  3. Each figure is placed under the label whose column it sits in, and the row is written out as a fact: "Barrel length: 4.7 in, Weight: 710 g", the same shape a table on a web page gives.
  4. This is done on every page of the document, not only the first few, because a catalogue's specification table is often forty pages in. The document's prose is read separately, in order, so a two-column paragraph still reads the way it was written.
What the document isWhat Redback does
A spec sheet or manual with tablesIts tables come first as facts, row by row with their labels, followed by the prose
A long catalogue or manualThe tables are gathered from every page; the prose follows from the start, as far as the reading budget for one document allows
A scanned PDF with no text in itSaid plainly: the document has pages but no text to read. Nothing is guessed from it, and Redback looks for another source
A file over twelve megabytes, or one that will not openRefused with the reason, so the task does not treat it as read
A site whose robots rules forbid the pathNot fetched, and the refusal is said

A figure Redback takes from a PDF is held to the same rule as one from a page: it must be quoted, and before it can reach a record the quote is checked again against the document, read the same way. A quote taken from a rebuilt table is checked against that table, so the figure and its label have to agree a second time.

An encyclopedia article is a pointer, never a citation. Its infobox is read as fields, grouped as the article groups them, and its references are ranked by the source hierarchy and followed to the pages they cite. A structured-data lookup gives a second opinion in the same way: a lead to a page, not a value.

A contradiction is a disagreement between two records, and a task carries one of them; judging whether a firearm predates its maker needs the maker. The catalogue tools read a record, search the catalogue and follow a record's relations, each against an allow-list, because a table name cannot be a bound parameter. The standing of a source tells a model a site's kind, rank and how often the catalogue cites it, and how that site reads to a script, before it quotes it.

A model in the tool loop has a scratchpad of twelve notes of 240 characters, re-shown on every turn from outside the conversation, so pruning the history cannot remove a conclusion. It is for what was ruled out: the measured failure was three near-identical searches, because nothing recorded that the first two had found nothing. A note is not evidence; a figure in a note has no source behind it.