Metadata-Version: 2.4
Name: scibraid
Version: 0.1.2
Summary: Deterministic tools for building and pooling evidence subgraphs of the scientific literature
Keywords: literature-review,evidence-synthesis,knowledge-graph,provenance,openalex,claude-code
Author: Steve Phelps
Author-email: Steve Phelps <phelps.sg@gmail.com>
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Science/Research
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Scientific/Engineering :: Information Analysis
Requires-Dist: httpx>=0.28.1
Requires-Dist: pydantic>=2.13.5
Requires-Dist: fastembed>=0.8.0 ; extra == 'embeddings'
Requires-Dist: pypdf ; extra == 'fulltext'
Requires-Python: >=3.13
Project-URL: Homepage, https://github.com/phelps-sg/scibraid
Project-URL: Repository, https://github.com/phelps-sg/scibraid
Project-URL: Issues, https://github.com/phelps-sg/scibraid/issues
Provides-Extra: embeddings
Provides-Extra: fulltext
Description-Content-Type: text/markdown

# scibraid

scibraid is a multi-agentic approach to literature review: agents build traceable evidence graphs from the scientific literature. Given a research question, they identify hypotheses, experiments, conditions and results, grounding each relationship in evidence from the original papers. Independent reviews can then be braided together to discover connections and contradictions that emerge only when different questions bring different parts of the literature into contact.

![A subgraph in the viewer, with the evidence for and against each hypothesis listed beside it](https://raw.githubusercontent.com/phelps-sg/scibraid/main/docs/viewer-overview.png)

## Why?

A researcher reviewing a literature builds a mental model of it: which experiments tested which hypotheses, under what conditions, with what results, and which results conflict. That model is the valuable product of the review. It is also private, and usually disappears when the project ends. The next researcher asking a nearby question starts again from the papers.

**scibraid keeps the model.**

The strands of a braid stay distinct, which is how scibraid combines graphs without pretending that they are one canonical representation. Given a research question, an agent reads the relevant papers and records what it finds as a graph of hypotheses, experiments, conditions, observations and interpretations.

Every link carries:
- a confidence level
- whether the relationship was stated by the paper's authors or inferred by the agent
- a verbatim quotation from the source

A deterministic check verifies each quotation against the paper's text, where that text is held, and rejects quotations that cannot be found in it. This means the graph can be audited back to the evidence it claims to represent.

Failed replications and null results are recorded alongside positive results. They matter because they mark the parts of hypothesis space that have already been searched.

Evidence is counted by source, not by paper. Papers with an author in common (by OpenAlex author id, or by name where there is none) are one group, and the dossier gives, beside the number of papers supporting a hypothesis, the number of independent groups. A result reported four times by one laboratory is one source reporting four times. This is the one credibility measure the graph computes; citation counts, venues and sample sizes are recorded and shown, and weighed by whoever writes, not folded into a score.

Results age, fastest in the fields that publish fastest, and a result is more often overturned by a later paper that reran it on something newer than by one that calls itself a replication. `scibraid sweep` searches among the works that cite the papers a subgraph rests on most, newest first and for the vocabulary of failure, and `lint` asks for the sweep until it has been done. The graph records what each result was obtained on (the model, the population, the year) as conditions, so a later result under a newer condition sits beside the older one rather than replacing it.

## The first review is already useful

scibraid does not require a corpus-wide knowledge graph to exist before it can answer a question.

A single graph is useful on its own. It is already a structured literature review whose claims can be inspected and traced to quoted passages: the hypotheses in contention, the evidence for and against them, conflicting results, and the conditions under which those results were obtained.

`scibraid lint` identifies structural gaps, such as hypotheses without direct evidence, experiments without conditions, missing failures and unverified quotations. `scibraid view` lets a reader browse the resulting review.

Pooling comes later. It is not a prerequisite.

## From separate reviews to new connections

Graphs built for different questions are pooled, but never merged. Each remains an independent representation of the question that generated it.

A second agent pass judges which nodes in different graphs refer to the same thing, recording each judgement as a link with its own rationale. The resulting structure can expose things that no individual review was looking for:

- a condition shared by failures in apparently unrelated lines of work
- a contradiction that the differing conditions of the two results may explain
- an experiment relevant to a hypothesis from another literature
- a hypothesis that has been tested under conditions relevant to another question
- experiments that have not yet been run

The key idea is that the system does not need to construct a universal ontology of science in advance. **The questions determine what gets represented; the points of contact between questions determine where further reasoning is worth doing.**

A third pass checks promising leads against the sources and wider literature. Each lead records what was checked, whether the connection held, whether it was already known or fell apart, and what question should be investigated next.

The process therefore forms a loop:

**question → agent-built review → alignment → candidate connections → verification → new question**

When a question reaches the point where the existing literature cannot settle it, scibraid records the open research question and the experiment that would resolve it.

## How it works

There are no language-model calls and no model API keys in the Python code itself. The only model it runs is the optional local embedding model. The agent harness, currently Claude Code, does the reading, extraction and judgement through four skills and five subagents, each with a fixed model. The `scibraid` command-line tool handles the deterministic parts:

- literature search and retrieval of open full text
- schema validation
- quote verification
- storage
- pooling
- candidate ranking
- structural queries

This separation is deliberate. The agent is used where interpretation is required; deterministic code checks and preserves the resulting structure.

## Related work

Most scientific graphs take the paper as the node. [OpenAlex](https://openalex.org/) and [Semantic Scholar](https://www.semanticscholar.org/) connect papers through citations and metadata. [scite](https://scite.ai/) goes a step further by classifying citations as supporting, mentioning or contrasting the cited work. These systems tell us which papers are connected, rather than what experiment was run, under what conditions, or what it found.

Extraction pipelines go below the paper to triples of entities. [SemMedDB](https://lhncbc.nlm.nih.gov/research/informatics/semmed/) contains large numbers of subject-predicate-object statements mined from PubMed. [GraphRAG](https://github.com/microsoft/graphrag) and related systems use language models to extract similar structures from arbitrary corpora. [MR-KG](https://www.medrxiv.org/content/10.64898/2025.12.14.25342218) extracts structured evidence from Mendelian randomisation studies.

These approaches generally process a corpus before a specific research question is asked. That makes them expensive to build and means the schema must anticipate the questions researchers will eventually ask. A simple triple also loses the experimental context behind a claim: which experiment produced it and under what conditions. A null result survives at best as a negated predicate, with nothing to say where the effect was absent.

Curated approaches preserve more of this structure. The [Open Research Knowledge Graph](https://orkg.org/) represents papers through structured properties so that contributions can be compared. [Nanopublications](https://nanopub.net/) package assertions together with provenance. [Discourse graphs](https://discoursegraphs.com) represent questions, claims and evidence in researchers' notes.

scibraid shares their emphasis on provenance, but uses agents to construct the structure from the literature rather than requiring people to do the extraction by hand.

The closest recent system is [ASKS](https://arxiv.org/abs/2608.29612), which uses an agent to compile papers into a persistent graph with deterministic checks and links back to source material. scibraid differs in making the **research question**, rather than a fixed corpus, the unit around which the graph is built. Separate question-driven graphs remain distinct and are aligned only when there is a reason to compare them.

The broader idea of discovering connections by combining otherwise separate literatures goes back to Don Swanson's work on literature-based discovery. scibraid follows that tradition but looks for a more specific kind of connection: **shared experimental conditions**, rather than merely shared terms. Conditions are often where apparently conflicting results separate.

## Example

Two questions were put to the agent separately:

```text
/evidence-subgraph Are emergent abilities in large language models real, or an artefact of how performance is measured?
/evidence-subgraph Does chain-of-thought prompting only help above a certain model scale, and why?
```

The first produced a graph of 96 nodes and 158 links from 12 papers. The second produced 63 nodes and 103 links from 10 papers. Every quoted passage was verified against the source text.

The two graphs shared one paper and three node IDs. Other overlaps were hidden behind different names, such as `c:model-family-palm` in one graph and `c:palm-models` in the other.

```text
/align-subgraphs
```

The cheap candidate stage proposed 56 cross-graph pairs. (At the time it proposed every pair of hypotheses; it no longer does, for the reason given under Pooling and alignment.) The agent judged 10 to be the same thing, 8 narrower or broader, 17 related and 21 different. The last category included plausible-looking lexical matches that turned out to be conceptually different.

`scibraid observe` then surfaced a failure shared by two lines of work that neither original question had asked about. Both results concerned base models without instruction tuning. One came from the emergent-abilities literature and one from the prompting literature.

The third pass checked the observation and related candidates against the sources. That connection turned out to be in the literature already: one of the two papers makes its remark "echoing Fu et al. (2022)", an essay that traces chain-of-thought ability to how a model was trained and not to its size. The pool had recovered it from two papers that do not frame it that way. Of the other candidates, one held and one fell apart under closer inspection. The system keeps those outcomes too, so a rejected lead is not rediscovered and pursued again.

The important result is not that every candidate is an insight. Most are not. The point is to create a **searchable space of cross-literature hypotheses**, then spend expensive agent reasoning and source checking on the connections where the structure suggests it may pay off.

This pool ships with scibraid, so you can look at it before building anything. It is the two reviews above and the third that one of their leads prompted, with the alignment judgements and the checked leads, one of each of `known`, `refuted` and `open`. It costs no model time and needs no OpenAlex key:

```bash
scibraid view --example
```

To run the other commands on it, write it to a data directory of its own and point scibraid there:

```bash
scibraid example ~/scibraid-example
export SCIBRAID_HOME=~/scibraid-example
scibraid observe --new
scibraid lead list
scibraid agenda
```

The example holds each paper's metadata and abstract and the passages the graphs quote, with the result of checking them. It does not hold the papers' full text, which is not ours to redistribute. `scripts/export_example.py` regenerates it from a working data directory.

## Install

Requires [uv](https://docs.astral.sh/uv/) and Claude Code. Install scibraid as a plugin:

```text
/plugin marketplace add phelps-sg/scibraid
/plugin install scibraid@scibraid
```

The plugin puts a `scibraid` command on the PATH of Claude Code's shell for as long as it is enabled. The command runs the copy of the package that ships with the plugin, in an environment that uv builds on first use and keeps outside the plugin directory, so nothing else needs installing. Outside a Claude Code session the command is not on your PATH. To use it there, install it from [PyPI](https://pypi.org/project/scibraid/):

```bash
uv tool install "scibraid[embeddings]"
```

To work on scibraid itself, install it from a clone instead:

```bash
git clone https://github.com/phelps-sg/scibraid && cd scibraid
uv tool install --editable ".[embeddings]"
```

Reading PDFs uses `pdftotext` from poppler if it is installed, and otherwise needs the `fulltext` extra. Compiling a write-up needs `pdflatex` and `bibtex`; without them the draft is checked but not compiled.

Each agent names the model it runs on. Where your plan does not include that model, Claude Code runs the agent on the newest available model of the same family, or on the session's model, and warns you which. See Which model does what.

Run `scibraid doctor` to see what is missing on your machine and how to fix it (`scibraid healthcheck` does the same). It checks that OpenAlex answers and that a key is set, and whether a PDF reader, LaTeX and the embedding model are present, and it exits non-zero only when something the tool cannot work without is missing. `--offline` skips the network check.

Literature search uses [OpenAlex](https://openalex.org/), which needs no account. Requests without a key share one free daily budget per IP address, and a day of building subgraphs can exhaust it, so get a free key and either set `OPENALEX_API_KEY` or put the key in `~/.openalex-tok` (another path can be named in `OPENALEX_API_KEY_FILE`). A file is the easier of the two, because every session and subagent on the machine finds it.

The `embeddings` extra adds a small local embedding model (fastembed, about 70 MB on first use, no API key) to help the alignment stage find matching conditions that share no words. Leave it off and the system still works using word overlap. The clone install above includes it. Under the plugin it is off unless you set `SCIBRAID_EXTRAS=embeddings` in the environment Claude Code starts from.

To develop the skills or agents without installing the plugin, start Claude Code inside a clone (it finds them through `.claude/skills` and `.claude/agents`), or point it at the clone with `claude --plugin-dir /path/to/scibraid`.

## Use

Installed as a plugin, the skills are namespaced: type `/scibraid:evidence-subgraph`, `/scibraid:align-subgraphs`, `/scibraid:pursue-leads` and `/scibraid:write-up`. Inside a clone they are `/evidence-subgraph` and so on, which is the form the examples below use.

Ask a question. The agent frames the hypotheses in contention, searches, fetches open full text, hands each paper to an extractor subagent, and hands the extracted graph to a synthesiser that draws the links spanning papers. It checks the result with `scibraid lint`, reports what the evidence shows, and can pool it with other reviews.

```text
/evidence-subgraph why is the ego-depletion effect still disputed?
```

Anything written after the question frames it: hypotheses you want tested, a distinction every paper should be read for, literatures that bear on it under other names. The agent records that framing on the subgraph (`scibraid show` prints it) and adds the rival hypotheses you did not give.

```text
/evidence-subgraph Does telling LLM agents that their partner is a copy of the same model change how they coordinate?
Test these: identity alone does the work; common knowledge of it is required; belief alone suffices.
Keep apart whether the partner was the same model and whether the agents were told so.
Cover superrationality and program equilibrium, not only papers about LLMs.
```

Browse what it built:

```bash
scibraid view
```

Once two or more graphs have been pooled:

```text
/align-subgraphs
/pursue-leads
```

You can then inspect what the aligned pool suggests:

```bash
scibraid observe --new
scibraid lead list
scibraid followups
scibraid agenda
```

A lead can seed the next review:

```bash
scibraid new graded-cot-metrics --lead cot-threshold-never-scored-with-graded-metric
```

When the pool has something to say on a topic, have it written up:

```text
/write-up what the literature shows about LLM agents told that their partner is a copy of themselves, for an economics audience
```

The text after the command is the steer: it chooses what the paper is about, and does not choose which evidence counts. The agent picks the pooled subgraphs the steer bears on and builds a dossier of their evidence. It then shows you a one-page spine (thesis, argument, what each section does) before writing anything. The draft is a single LaTeX file that arXiv accepts, with a bibliography generated from cached metadata. Every reference comes from a fetched record and none from the model's memory. `scibraid draft check` refuses a citation that is not in the bibliography and a quotation that is not verbatim in the source it cites, and it lists the phrases that make prose read as machine-written.

The paper is written in the voice of an exemplar you choose, usually a paper of your own:

```bash
scibraid voice set ~/papers/my-best-paper.pdf --name mine --default   # or a DOI or arXiv id with an open copy
```

The exemplar is kept in the data directory, not in any draft, and the agent takes its register and rhythm and none of its sentences.

The tools can also be driven by hand. A batch is a JSON file of nodes and links; the format is documented in `skills/evidence-subgraph/SKILL.md`.

```bash
scibraid search "ego depletion replication" --limit 10
scibraid search "replication" --citing W2499154041   # among the works that cite a paper
scibraid paper get 10.1177/1745691616652873          # a paper you already know of, by DOI or arXiv id
scibraid paper show W2499154041
scibraid fetch W2499154041                           # attach open full text, if there is any
scibraid new ego-depletion --question "Why is ego depletion still disputed?"
scibraid add ego-depletion batch.json
scibraid lint ego-depletion
scibraid pool ego-depletion
```

## Viewer

```bash
scibraid view
```

The viewer opens a local page at `http://127.0.0.1:8765/`. It shows hypotheses, experiments, conditions, observations and interpretations as distinct node types.

Every quoted passage links to its paper. The citation opens a paper panel with links out to the paper, its DOI, arXiv and OpenAlex, the abstract, a BibTeX entry with a copy button, and every relationship in the graph that rests on that paper, which are highlighted. The papers list has buttons to copy all the BibTeX or download it as a `.bib` file.

Selecting a link reveals the evidence behind it: its confidence, whether it was asserted by the paper or inferred by the agent, the agent's rationale, and the source passage used to support it.

Graphs can be filtered by node type, who asserted a relationship and confidence. Choosing "All subgraphs together" shows the local graphs side by side with the pool's alignment judgements between them and the rationale behind each. Leads show the evidence and checks on which they depend.

## Commands

| Command | Purpose |
|---|---|
| `search "<query>"` | Search OpenAlex by title and abstract, and cache the results with their abstracts and whether an open copy exists |
| `search --citing <id>` / `--references-of <id>` | Search among the works that cite a paper, or among its references |
| `paper get <identifier>` | Cache a paper's OpenAlex record by DOI, arXiv id or URL, PubMed id or OpenAlex id. An arXiv id is checked against arXiv's own title |
| `paper refresh <id ...> \| --all` | Re-fetch cached OpenAlex records, reporting ids that OpenAlex has merged away and papers newly marked retracted |
| `paper reid <old id> [identifier]` | Give a hand-added paper its OpenAlex record in every subgraph and lead that cites it, keeping the text its passages were checked against |
| `paper show\|add\|text` | Read a paper, add one OpenAlex lacks, or attach full text by hand |
| `fetch <id ...> \| --subgraph <slug>` | Find an open copy (arXiv HTML, Europe PMC, PDF) and attach its full text |
| `new <slug> --question "..." [--hypothesis ...] [--condition ...] [--brief ...]` | Start a subgraph, recording how the question is framed |
| `frame <slug> [--hypothesis ...] [--condition ...] [--brief ...]` | Add to the framing of an existing subgraph |
| `add <slug> batch.json` | Validate and apply a batch |
| `show <slug> [--format summary\|json\|mermaid]` | Inspect a subgraph |
| `sweep <slug> [--top N] [--limit K] [--paper id ...]` | The newest works citing the papers the subgraph rests on, and those reporting failures: what has been published since |
| `lint <slug>` | Find structural gaps and unverified evidence |
| `duplicates <slug>` | List conditions and hypotheses within a subgraph that may be one thing under two ids |
| `merge <slug> <keep> <drop>` | Fold one node into another, moving its links and passages |
| `relabel <slug> <node> --label "..." --reason "..."` | Reword a node's label so that it is true of what sits under it, keeping the old wording and why it went |
| `fit list <slug> [node] [--unexplained]` | For each condition several papers share, each paper's reason for being under it and a passage from it, side by side |
| `fit add <slug> <node> <paper> --why "..."` | Record, after the fact, why a paper belongs under an id |
| `retract <slug> --edge <source> <relation> <target> \| --node <id> --reason "..."` | Take out a wrong link, or a node with its links, keeping a record of what it was and why |
| `view [slug ...] [--example]` | Browse a graph in the browser, or the bundled example |
| `example [dir]` | Write the bundled example pool to a data directory of its own, to try the other commands on it |
| `doctor` (or `healthcheck`) `[--offline]` | Check that this machine is set up for scibraid, that cached OpenAlex ids still exist, and say how to fix what is not |
| `pool <slug>` | Add a subgraph to the pool |
| `candidates [--budget N] [--type T] [--lexical]` | Rank unjudged cross-graph pairs |
| `hypotheses` | Show hypothesis lists for pooled subgraphs |
| `align add\|list` | Record or inspect alignment judgements. `add` reports any verdict it records that clashes with those already held |
| `align check [slug]` | Verdicts that cannot all be right: a pair judged different that a chain of `same` verdicts joins, or two nodes of one subgraph that such a chain makes one |
| `observe [--new]` | Find candidate observations in the aligned pool |
| `lead add\|list` | Record or inspect checked leads |
| `followups` | Show follow-up questions raised by leads |
| `agenda` | Show open research questions, the experiment each needs, and whether the literature already poses it |
| `voice set\|list\|show` | Exemplar papers whose voice a write-up takes |
| `draft start <dir> <slug ...> [--focus ...] [--voice ...]` | Make or refresh a draft: the evidence dossier, `refs.bib` and a LaTeX skeleton arXiv accepts |
| `draft cite <dir> <identifier ...>` | Add a reference by DOI, arXiv id or paper id, and print its citation key |
| `draft check <dir>` | Citations resolve, quotations are verbatim, arXiv's requirements hold, prose tells, and it compiles |
| `bibtex [slug ...] [-o refs.bib]` | BibTeX for the papers the subgraphs cite, generated from cached metadata |
| `builder add <slug> --person --model ...` | Record who built a subgraph made before builders were recorded |
| `repair list\|resolve` | List and resolve faults in subgraphs or judgements found while checking leads (recorded on the lead) |

The commands that list things take `--format markdown` and print tables that paste into a README, an issue or a note: `list`, `pool --list`, `show`, `candidates`, `align list`, `lead list`, `followups`, `agenda`, `repair list` and `observe`. Most also take `--format json`.

```bash
scibraid agenda --format markdown
```

Data lives in `$SCIBRAID_HOME`, by default `~/.local/share/scibraid`. Several sessions can work against it at once. Each should build its own subgraph; the pool is SQLite and locks itself, files are written atomically, additions to one subgraph are serialised by a file lock, and a lead edited from a stale copy is refused. The file lock is advisory and local, so it does not protect a data directory shared over a network file system.

## Data model

The core graph contains five node types:

- `hypothesis`
- `experiment`
- `condition`
- `observation`
- `interpretation`

Links describe relationships such as `tests`, `performed_under`, `yields`, `observed_under`, `supports`, `contradicts`, `explains`, `proposes`, `competes_with` and `fails_to_replicate`.

Within a subgraph, two papers share a condition only because their readers chose the same id, so the choice is a judgement that the two are one thing. It is recorded as one. A batch that puts a paper under a condition already in the graph must say why (`reuses`), `add` refuses it otherwise, and the reason is kept on the node with who gave it. `fit list` sets the reasons and a passage from each paper side by side, which is where an id that lumps two different things shows. Across subgraphs the same judgement is an alignment verdict.

Each piece of evidence remains attached to the paper that provides it. If two papers support the same relationship, they create two links rather than one aggregated fact. This allows later evidence to strengthen, qualify or weaken an earlier claim.

Observations are kept separate from interpretations because papers often combine the two in a single sentence.

A quote must match the held source text. If the source text is unavailable, the quote is marked unverified rather than silently treated as checked.

## Pooling and alignment

The pool is a local SQLite database. Subgraphs remain namespaced and unchanged; alignment adds links between them.

Each alignment judgement is one of:

`same`, `broader`, `narrower`, `related`, `different`

and carries its own confidence and rationale.

Candidate matches are ranked using word overlap, shared identifiers, shared papers and, when enabled, embedding similarity. The agent makes the final judgement. This lets the cheap matching stage narrow the search space before expensive semantic judgement is applied.

The system does not rely on similarity to relate hypotheses. Similarity finds restatements of one claim. Whether two hypotheses bear on each other is a substantive research judgement, not a text-similarity problem, and on three pooled subgraphs neither word overlap nor embeddings ranked the related pairs better than chance. So `scibraid hypotheses` gives the agent both lists to read in context.

`observe` searches the aligned structure for patterns such as:

- conditions reached from different questions
- failures sharing a condition
- experiments that may bear on another question's hypothesis
- contradictions whose differing conditions may explain the disagreement
- hypotheses linked across questions
- experiments that have not yet been run

Most such patterns are not discoveries. A **lead** is a pattern that has been written down to be checked. Its status is `candidate` until it has been: the tool will not accept any other status without recorded checks. After checking it is `holds` (the premise is verified, no duller explanation was found, and a brief search did not find the connection stated), `known` (the literature already says it), `refuted`, or `open`. `open` is reserved for an open research question: the literature has been reviewed and does not settle the claim, so only new empirical work can. Open is a judgement made by this process, not a statement that the field regards the question as open, and it is not the same as new. Each open lead therefore records where the literature already poses the question (`posed_in`), or that a search found it posed nowhere, and `scibraid agenda` lists first the questions that nobody was found to have asked. A lead cannot be more confident than the weakest alignment judgement it rests on.

Each lead records the claim, the graph elements and alignment judgements it depends on, the checks performed, what would confirm or refute it, and a follow-up question.

The loop therefore preserves not only what the literature says, but also what the system considered, checked, rejected and left unresolved.

## Who did the work

Agreement between two subgraphs counts as confirmation only if they were read independently, and a lead checked by whoever built its evidence has not had a second reader. So every subgraph records its builders, every alignment judgement its judge and every lead its checker: the person, the agent harness, the model and the session. The tool takes the person, harness and session from the environment. It cannot see the model, so the agent passes `--model`.

Independence is reported as one of four levels, weakest first: `unknown` (nothing recorded, which counts as not independent), `same reader` (same person and model, even in a new session), `same model` (different people, one model, so shared blind spots), and `different model`. `observe` gives the level for each candidate that spans subgraphs, and `agenda` gives it for the checker of each open question against the builders of its evidence.

## Which model does what

The steps do not all need the same model. Reading one paper and recording what it did is the bulk of the tokens, and the tool checks that work: every quoted passage must appear in the source. Deciding what several papers mean together, how two hypotheses bear on each other, and whether a lead is real is a small share of the tokens, and nothing checks it.

So the work is divided among five subagents with a fixed model each, and the division holds whichever model the session itself runs on.

| Agent | Model | Does |
|---|---|---|
| `paper-extractor` | Sonnet | Extracts one paper, with a fresh context, and records its own model on what it adds |
| `evidence-synthesiser` | Fable | Merges ids that parallel extractors minted twice (`scibraid duplicates`, `scibraid merge`), draws the links that span papers, and re-reads the results the question turns on |
| `hypothesis-aligner` | Fable | Reads pairs of hypothesis lists in full and records how the claims bear on each other |
| `lead-checker` | Fable | Triages what `observe` returns, checks leads against the sources and the literature, and records them |
| `synthesis-writer` | Fable | Writes the spine and then the paper for a write-up, and edits it against the check |

The session frames the question, retrieves the papers, judges the ranked candidate pairs during alignment, and relays what the agents report. To use another model for a step, change the `model:` line in `agents/<name>.md`; where the top model is not available to you, put the strongest you have.

The extraction tier comes from one comparison, not a benchmark. Three models extracted the same five papers for the same question. Sonnet recorded 92 links to the top model's 36, five conditions per experiment to its three and eight failures to its one, and caught an error in the top model's reading of one paper. Haiku had 6 of its 13 batches rejected, recorded one failure, rated its own links at a mean confidence of 0.90, and lost the claim under test: each of its ten hypotheses restated one paper, and none was evidenced by more than one. Sonnet connected the papers as well as the top model had, so the synthesis pass sits with the top model for a structural reason, that extractors who each see one paper cannot link two, and as a second reading of what they extracted.

A useful side effect is that a subgraph extracted by one model and checked by another has had a second reader of a different kind, which `agenda` and `observe` report.

## Limitations

The current system is a research prototype.

- A check that a lead is already known, or that a question has already been asked, is a brief agent search, not a systematic literature review. `holds`, `open` and "not found posed anywhere" therefore mean that nothing was found, not that nothing exists.
- OpenAlex merges duplicate records without notice, and the old id then answers to nothing. `doctor` reports cached ids that have vanished and `paper reid` moves a paper's citations to the surviving record; sixteen of the first 582 records cached were merged away within weeks.
- Retraction is taken from OpenAlex's `is_retracted` flag, which lags the retraction notices and misses preprints; a paper withdrawn on arXiv is not flagged. A retracted paper's evidence is kept and marked, not removed.
- The first backfill of reasons, over three graphs and 349 paper-condition pairs, found the label fitting outright in a little over half of them (194); the rest were narrower (76), true only in part (45), not decidable from the text held (15) or, for thirteen links, false. Two labels written when the questions were posed fitted almost none of the papers under them and were reworded. One reader per paper judged each pair, against the paper's full text where it was held; nobody has measured whether a second reader would agree.
- A reason for reusing an id is required of conditions only, and only from now on. Hypotheses are reused without one, on the ground that a `supports` or `contradicts` link already carries its own passage; graphs built before reasons were required hold none, and `lint` counts the conditions affected. A reason is a sentence from the reader, not a check: it makes lumping visible to the next reader and does not prevent it.
- Whether two nodes are the same thing is a judgement, made pair by pair, and judgements made separately can disagree. `align check` and `lint` find the disagreements that follow from the verdicts alone (a pair judged different that `same` verdicts join; two nodes of one subgraph made one), with the chain and its weakest link. They cannot find a wrong `same` that nothing else contradicts.
- Quote verification establishes that a passage exists in the source. It does not establish that the passage actually supports the relationship the agent attached to it.
- The example graphs were built and aligned by the same agent in one session, so they are not independent in the way reviews produced by different researchers would be. The tool now reports this (`same reader`) instead of leaving it to be remembered.
- The model tiers rest on a single five-paper comparison on one question. Whether hypothesis alignment or the checking of leads can move to a cheaper model has not been tested, so both are pinned to the top model. Framing, retrieval and the judging of candidate pairs run on the session's model, and have only been tried with the top model as the session.
- Builders are recorded per subgraph, not per link, so a subgraph that two readers contributed to counts as the weaker of the two everywhere.
- No domain expert has audited the example extractions.
- The example comparison with related work is based on the author's knowledge and a brief search rather than a systematic survey.
- A write-up is a draft for its authors to check, not a finished paper. The check verifies that quotations are verbatim and that references exist; it cannot verify that a sentence characterises a cited result fairly, and a reference added with `draft cite` has not been read by anyone in the pipeline. The write-up skill has not yet produced a full paper.
- Framing steers a review towards what its author expected. The framing is recorded and the skill requires rival hypotheses, but nothing checks that the rivals were sought as hard as the favoured ones.
- Word search finds a minority of the relevant papers. Three hand-written queries per question, top 25 results each, returned 20 of the 62 papers the six example subgraphs cite; the rest were found through reference lists and the agent's own knowledge. Following citations (`--citing`, `--references-of`) is the remedy the skill prescribes, and its effect has not been measured.
- OpenAlex's default search also matches full text, which it holds only for open papers: on the same queries 2% of its results were closed, against 19% when matching on title and abstract, which is now the default. Recall was the same either way (18 and 20 of 62).
- OpenAlex files a few unrelated records under the arXiv DOIs of well-known papers (2 of the 22 arXiv ids tried, one of them Wei et al.'s chain-of-thought paper). `paper get` checks an arXiv id against arXiv's own title and then looks for the paper by title. A DOI that is not an arXiv DOI is not checked.
- Authors are not resolved to people. Each paper keeps OpenAlex's author ids and ORCIDs where OpenAlex has them (every author in 52 of the 66 papers the example subgraphs cite, some in 12, none in 2, the last being recent preprints), but nothing yet uses them, so two results from one group count as independent.
- OpenAlex metadata can contain errors. BibTeX is generated from that metadata: author lists are stored cut at eight names (the entry then ends "and others"), and venues and entry types should be checked before use.
- Full text comes only from open copies. A closed paper is found by search and can be extracted from its abstract, but its body is unread unless someone attaches the text by hand, so `open` and "not found posed anywhere" describe the literature that could be read. The share that is closed varies widely by field and rises with age; it was about two thirds for one materials-science query.
- Text taken from a PDF loses section headings and mangles mathematics. Some publishers refuse automated requests even for open papers, and `fetch` then reports the failure.
- An observation inherits every condition recorded for its experiment, including ones that do not apply to it, which produces spurious shared-condition patterns.
- A condition that a paper used but the agent did not record looks the same as a condition that was never tested, which produces spurious "experiments that have not been run".
- An observation's outcome is relative to what its own experiment was looking for, so two "negative" results can point in opposite directions. `observe` treats them as comparable.
- The candidate stage finds nodes that look alike. It does not solve the harder problem of discovering whether two hypotheses bear on each other. The current system asks the agent to reason over hypothesis lists instead.
- Scaling alignment across hundreds of subgraphs will require better strategies for deciding which graphs and which node pairs are worth comparing.

## Development

```bash
uv run pytest
```

## Design principle

scibraid is built around a simple inversion:

**Don't build a knowledge graph of the literature and then ask questions of it. Let research questions build the parts of the graph that matter, and connect those graphs when new questions make the connections useful.**

That makes the evidence graph a by-product of research rather than a prerequisite for it.
