- One-minute verdict
- What the essay argues
- Fact-check table
- The three substantive errors
- Diagram: Peirce's triad
- Diagram: rebutting vs undercutting
- Diagram: the Nixon diamond
- Diagram: inheritance with exceptions
- Diagram: Lakatos's belt
- What the essay gets right
- Where the argument strains
- Is this AI slop?
- Case study: a partial glass box (KISS Sorcar)
- Final scorecard
- Sources
One-minute verdict
This is a good essay built on a good idea: our AI agents are trained to be confidently right and have no machinery for being productively wrong — for holding a hypothesis, marking exactly what would kill it, and retracting cleanly when it dies. The scholarship it draws on is real, and most of the checkable claims hold up against the primary sources. But the fact-check surfaced three substantive problems: a citation dated twenty-one years early, a hyperlink that sends readers to the wrong Lakatos book, and — most seriously — a reconstruction of the Biggs & Wilson argument that gets their central criterion backwards. Several framing claims ("the newest paper is from 2004," "the oldest is from 1987") are contradicted by the essay's own citations.
There is an irony worth stating plainly: an essay arguing that agents should track their own defeaters ships with three unflagged defeaters of its own. None of them sink the thesis. Below, each checkable claim is verified against primary and scholarly sources, the errors are dissected, and five diagrams reconstruct the essay's key ideas so you can judge them directly.
What the essay actually argues
Stripped to its spine, the argument runs like this:
1. Real reasoning is defeasible. Most useful inference is not deduction from certainties; it is jumping to a good-enough conclusion that new evidence can later overturn. The essay leans on Peirce's triad (abduction, deduction, induction) and on the formal literature of non-monotonic logic.
2. Defeaters come in types. Following Pollock, a belief can be attacked head-on (a rebutting defeater says "your conclusion is false") or sideways (an undercutting defeater says "your reason no longer supports the conclusion"). A good reasoner knows which kind just hit it.
3. A provocative inversion. The essay claims abduction (hypothesis-forming) is the indefeasible arbiter while deduction is defeasible — borrowing a thesis from Biggs & Wilson.
4. The punchline. LLM agents have none of this. They can't introspect a forward pass, can't name the belief that took the hit, and can't say what would change their mind back. The essay proposes building agents as glass boxes: explicit hypotheses, typed defeaters, a tunable credulous-vs-skeptical "temperament," and a Lakatosian hard core / protective belt so anomalies get routed, not silently swallowed.
It is a builder's essay resting on a philosophical literature, and the fit is mostly good. Now the inspection.
Fact-check: the checkable claims
Each row was verified against the primary literature or authoritative references — journal metadata, the papers themselves where accessible, PhilPapers/PhilArchive, Crossref, the Stanford Encyclopedia of Philosophy, and the Natural History Museum (full source list in Sources). ✓ verified · ✗ wrong · ≈ imprecise or qualified.
| Claim in the essay | What the sources say | Verdict |
|---|---|---|
| Delrieux (2004) formalizes Peirce's triad for defeasible research programmes | Claudio Delrieux, "Abductive inference in defeasible reasoning: a model for research programmes," J. Applied Logic 2(4):409–437, 2004. Title, author, year all correct. Low citation count, so "not famous" is fair for this one. | ✓ |
| Pollock (1987): rebutting vs undercutting defeaters; the "looks red" example | John L. Pollock, "Defeasible Reasoning," Cognitive Science 11(4):481–518, 1987. Terminology and the red-illumination example are exactly Pollock's. | ✓ |
| Horty, Thomason & Touretzky (1990) on skeptical inheritance | "A Skeptical Theory of Inheritance in Nonmonotonic Semantic Networks," Artificial Intelligence 42(2–3):311–348, 1990. Exact match. Note: the credulous/skeptical distinction predates it (Reiter & Criscuolo 1981; Touretzky, Horty & Thomason, "A Clash of Intuitions," IJCAI 1987) — HTT 1990 argues for skepticism rather than introducing the fork. | ≈ |
| The Nixon Diamond (Quaker → pacifist; Republican → not) | Canonical non-monotonic puzzle; the credulous vs skeptical resolutions are standard semantics (truth in some extension vs truth in all extensions). | ✓ |
| Platypus: reached Britain in 1799, suspected a hoax; Shaw took scissors to the pelt hunting for stitches | The pelt and sketch were sent by Captain John Hunter around 1798; George Shaw described the animal in 1799 and did check the specimen for stitches (Robert Knox also suspected fraud). The essay compresses 1798–99 into "1799," which is close enough; the hoax story itself is well documented. | ✓ |
| Caldwell's 1884 telegram: "Monotremes oviparous, ovum meroblastic," ending an ~85-year controversy | William Hay Caldwell, telegram read to the British Association meeting in Montreal, September 1884. Wording exact; the Natural History Museum frames it as ending an "85-year" controversy (1799→1884). | ✓ |
| Bondarenko, Dung, Kowalski & Toni (1997): abstract argumentation-theoretic default reasoning | "An abstract, argumentation-theoretic approach to default reasoning," Artificial Intelligence 93(1–2):63–101, 1997. The foundational Assumption-Based Argumentation paper; the essay's four-part summary (monotonic base logic, assumptions, attack via contrary, admissible = conflict-free + self-defending) is faithful. | ✓ |
| Peirce triad; abduction gives "reason to suspect"; SEP files abduction under defeasible reasoning | The SEP "Abduction" entry supports Peirce's generative sense and notes that abduction violates monotonicity (i.e. is defeasible). It does not endorse the stronger Biggs–Wilson "indefeasible" thesis, which is a separate, contested claim. | ✓ |
| Sakana AI Scientist critique quote (arXiv 2502.14297) | Beel et al. 2025 verbatim: "It heavily depends on user input, struggles with methodological soundness, and lacks the ability to critically assess its own results." The second sentence the essay splices in ("AI models inherit biases from historical data and cannot independently distinguish between scientific quality and consensus") is also verbatim from the same paper. Quoted accurately. | ✓ |
| A 2026 case study found every autonomous-research framework produced "sophisticated hallucinations" that get structurally integrated into pipelines | Verified. Agrawal et al., "Can AI Conduct Autonomous Scientific Research? Case Studies on Two Real-World Tasks," bioRxiv, DOI 10.64898/2026.01.05.697809 (posted Jan 2026). The abstract states eight frameworks were tested on two tasks, none completed a full cycle, and "Every framework produced sophisticated hallucinations." The recommendation to separate speculative from computed statements is in the paper. Caveat: it is an un-peer-reviewed preprint testing eight open-source frameworks on two tasks — "every system" is scoped to that sample, not all AI scientists. | ✓ |
| "Frontier systems" reinvented the choreography of argumentation (propose, critique, debate, rank, evolve) then discarded the auditable bookkeeping | The linked source is Gottweis et al., "Accelerating scientific discovery with Co-Scientist," Nature 655:487–496 (2026). The generate/reflect/rank/evolve choreography is accurate. But the "discarded all bookkeeping" reading is too strong: the paper describes explicit hypothesis state, deep-verification reviews that decompose hypotheses into constituent assumptions, and evidence weighed for and against. What these systems lack is Pollock-style typed defeaters and retraction, not structured state as such. | ≈ |
| Pauli, per Peierls, said of a young physicist's paper "It is not even wrong" | The attribution to Rudolf Peierls's recollection of Wolfgang Pauli is the standard one. Correct. | ✓ |
| "The newest paper I'm going to lean on is from 2004. The oldest is from 1987." | Both false on the essay's own citations. The Biggs & Wilson chapter it leans on is 2025 (see below); it also cites a 2026 bioRxiv study and a 2026 Nature paper. And it links a 1983 inheritance paper (Etherington & Reiter) as its inheritance-with-exceptions source, older than 1987. | ✗ |
| Biggs & Wilson, "The Indefeasibility of Abduction," dated 2004 | Actually 2025 — a chapter in Inductive Metaphysics (eds. Hüttemann & Schurz, Routledge; DOI 10.4324/9781003514404-7; PhilPapers BIGTIO-4). Off by 21 years. | ✗ |
| "Delrieux models a theory the way Imre Lakatos did" — link points to Proofs and Refutations | The hard-core / protective-belt / heuristics apparatus is from Lakatos's methodology of scientific research programmes (the 1970 essay "Falsification and the Methodology of Scientific Research Programmes," collected 1978), not Proofs and Refutations (1976, about mathematics). Concept correctly attributed to Lakatos; link lands on the wrong work. | ✗ |
| How the essay reconstructs Biggs & Wilson's proof that "deduction is defeasible" (via reductio) | Misrepresents the argument. See the deep-dive below: the essay's reductio route defeats a premise, which is exactly what does not count as defeasibility under Biggs & Wilson's own definition. | ✗ |
Tally: 12 checkable statements verified as accurate (some with minor caveats), 3 clear errors, and a cluster of imprecisions flagged with ≈. This is not an exhaustive enumeration of every sentence in the essay — it covers the load-bearing factual and citation claims.
The three substantive errors, in detail
Error 1 — Biggs & Wilson is a 2025 chapter, not 2004 FACTUAL
The essay cites Stephen Biggs and Jessica Wilson's "The Indefeasibility of Abduction" as a "2004 chapter." It is a 2025 chapter in Inductive Metaphysics (Routledge, eds. Andreas Hüttemann & Gerhard Schurz; DOI 10.4324/9781003514404-7). The essay even links the PhilArchive copy, whose BibTeX record reads year = {2025}.
This matters because the essay builds a rhetorical frame around age — "the newest paper I'm going to lean on is from 2004" — and the load-bearing thesis of the whole piece (abduction is the indefeasible arbiter) comes from this 2025 source. The framing collapses on the essay's own citations, which also include 2026 material.
Error 2 — the Lakatos link points to the wrong book CITATION
Where the essay says Delrieux "models a theory the way Imre Lakatos did," the hyperlink lands on the Wikipedia article for Proofs and Refutations (Lakatos's 1976 book about how mathematical conjectures evolve). But every concept the essay then uses — the hard core, the protective belt, the negative and positive heuristics, progressive vs degenerating research programmes — comes from Lakatos's methodology of scientific research programmes (the 1970 essay, collected in the 1978 volume of that name). The attribution to Lakatos is right; the specific link sends a curious reader to the wrong work.
Error 3 — the essay misreconstructs Biggs & Wilson's argument for defeasible deduction FIDELITY
This is the most consequential slip, because it touches the essay's central inversion. Biggs & Wilson define reasoning as defeasible when a conclusion "can be defeated even when its premises remain undefeated." To show deduction is defeasible, they give the "Dina" example: two valid chains yield conflicting conclusions ("the wall is probably red" / "probably not red") from independently justified premises, so one conclusion must be given up without any premise being defeated.
The essay instead argues that deduction is defeasible via reductio ad absurdum: "if the chain is valid and the conclusion is false, a premise must go." But that is premise defeat — precisely the case Biggs & Wilson say does not establish defeasibility. In their framework, a conclusion defeated only by defeating a premise is the signature of indefeasibility (it's how they class abduction as indefeasible: any defeater hits the "catch-all premise"). So the essay reaches Biggs & Wilson's conclusion using the very mechanism their definition rules out. The right illustration is conflicting conclusions with undefeated premises, not modus tollens against a premise.
Two housekeeping notes for balance. First, none of these three damages the essay's central practical thesis. Second, the essay is honest that the indefeasibility thesis is contested ("Not every philosopher buys this"), which is the right instinct even though the supporting argument is mis-drawn.
Diagram — Peirce's triad, and where defeat enters
Diagram — rebutting vs undercutting (Pollock, 1987)
This is the essay's most useful distinction. A rebutting defeater attacks the conclusion directly. An undercutting defeater leaves the conclusion alone but severs the link from reason to conclusion — it says "your evidence no longer bears on this."
Diagram — the Nixon diamond
Two good defaults, one contradiction. Quakers are (defeasibly) pacifists; Republicans are (defeasibly) not. Nixon was both. In formal terms, a credulous reading accepts a conclusion true in at least one extension; a skeptical reading accepts only what holds in every extension — so here the skeptic withholds both "pacifist" and "not pacifist." The essay's "temperament" dial is exactly this choice, made explicit and tunable.
Diagram — inheritance with exceptions (the platypus)
The essay's best-told story. "Mammals bear live young" is a superb default and a false universal. The platypus lays eggs, so the specific fact overrides the inherited default — which is exactly what non-monotonic inheritance networks are built to handle, and what took science until 1884 to accept.
is-a links run upward from the specific class to the general one; the specific fact ("lays eggs") wins over the inherited default ("bears live young"). Verified: Shaw 1799 hoax suspicion, Caldwell 1884 telegram, ~85-year gap.Diagram — Lakatos's hard core & protective belt
The essay's proposed architecture: keep a small, protected hard core, surround it with a protective belt of auxiliary hypotheses, and route anomalies to the belt first. Lakatos's actual criterion for progressive vs degenerating is not "how deep the anomaly reaches" but whether the belt's revisions predict novel facts that then check out (progressive) or merely accommodate known anomalies after the fact (degenerating, ad hoc). Replacing the hard core is a change of research programme, not the definition of degeneration.
What the essay gets right (and it is a lot)
The two-questions framing is the real contribution STRONG
When an agent hits a defeater, the essay says it should answer two things: "Which of your beliefs just took the hit?" and "What would change your mind back?" That is a clean, testable specification for an introspective agent, and more concrete than most of the "AI should reason better" genre. It turns a philosophical point into an interface requirement.
The glass-box proposal is buildable
Explicit hypotheses, typed defeaters (rebutting vs undercutting), a tunable credulous/skeptical temperament, and an explicit hard-core/protective-belt structure are all inspectable objects. The essay names the data structures a better agent would expose rather than just complaining that LLMs are opaque. The Assumption-Based Argumentation framework it points to (Bondarenko et al. 1997) really does supply semantics and existence theorems for exactly this bookkeeping. Several practical approximations to that wishlist already run in an existing agent framework — see the KISS Sorcar case study below.
Honest about the contested parts
The load-bearing claim — that abduction is indefeasible — is a minority position, and the essay flags it ("not every philosopher buys this"). Signposting your own weakest link is exactly the habit the essay preaches.
The core diagnosis is also correct and important: a single forward pass through an LLM has no dedicated place to store "this belief is provisional and here is its defeater," and there is a real, growing literature showing that a model's narrated chain-of-thought is often an unfaithful, post-hoc rationalization rather than the mechanism that produced the answer. The essay is right that we mostly train and reward agents for confident correctness, not graceful retraction.
Where the argument strains
"LLMs cannot represent defeaters" overstates a real point NUANCE
Three things are worth separating. (1) A forward pass cannot give a faithful mechanistic account of itself — this is well supported. (2) But a model can perfectly well represent a defeater as text, and an agent can store one in external state — which is exactly what the author is building. (3) Chain-of-thought is often unfaithful, but it is not simply "generated after the fact" with no causal role; intermediate tokens can and do influence later ones. So the sharp claim should be "today's agents mostly don't track typed defeaters," not "LLMs cannot."
The "frontier systems keep only a score" reading is too strong
The linked Nature Co-Scientist paper actually describes explicit hypothesis state, deep-verification reviews that decompose a hypothesis into its constituent assumptions, and evidence weighed for and against. What these systems lack is Pollock-style typed defeaters and principled retraction — a narrower and more accurate charge than "they discarded the bookkeeping."
"Deduction is the fragile one" — heading vs argument
Deductive validity is not fragile; a valid argument stays valid. What Biggs & Wilson call defeasible is the conclusion. As noted in Error 3, the essay's reductio illustration does not actually establish their sense of defeasibility — so this section overreaches both in its heading and in its supporting argument.
"These papers aren't famous" undersells the canon
Pollock 1987 has well over a thousand citations; Horty-Thomason-Touretzky 1990 is on the field's list of classic AI papers; the 1997 ABA paper is foundational. To a general audience they are obscure; to anyone in knowledge representation they are standard references. Fair rhetoric for lay readers, imprecise as fact.
None of these sink the piece. They are the difference between a sharp essay and an airtight one.
Is this AI slop?
The task asked specifically to check for "AI slop." The honest answer has two parts.
On style: the essay does not carry the usual generic-LLM signature — it is not the smooth, hedged, symmetrical, list-of-three prose that reads as machine-default. The voice is distinctive, opinionated, and domain-literate, with idiosyncratic asides (Tony Stark, "reorganize the zoo," the Pauli anecdote). By the ordinary meaning of "AI slop" — low-effort, generic, filler text — this is not it.
On authorship: it is worth being careful here. The essay contains a crop of small errors a careful editor would catch:
probabilisitic · adductive (for "abductive") · you lawn is wet · refuses to be conclude · The paper argue · is throws out · does not collapses · Do no underestimate · a self-tangling phrase, "mammals really do, as a rule, are viviparous," and a garbled sentence, "most of them are a tree search is an explore-versus-exploit machine."
These read like fast, unedited human typing. But it must be said plainly: typos do not prove human authorship, and cannot. Models can produce typos and disfluencies, and any text may be human-drafted then machine-polished, or machine-drafted then hand-edited. Authorship is not reliably recoverable from the text alone. The defensible verdict is: no strong AI-slop signature, a distinctive human-sounding voice, and copy-editing that a final pass would have fixed. Whether a model helped draft it is not something this review can settle.
The genuine irony stands on its own: an essay demanding that agents surface their own defeaters ships with three unflagged defeaters (the 2004 date, the wrong Lakatos link, the mis-drawn deduction argument) and a handful of typos. That does not undermine the thesis — if anything it illustrates the need for the tooling it argues for.
Case study — a partial glass box: KISS Sorcar
The essay's proposal can feel ambitious: explicit hypotheses, typed defeaters, a temperament dial, and a hard core with a protective belt. A fair test is to compare that wishlist with an agent framework that exists today. This review was itself produced with KISS Sorcar, an open-source, local-first agent framework (Apache-2.0, about 2,900 lines of core-agent code) built for long-horizon tasks and AI discovery. Sorcar exposes practical analogues for several items on the list, largely through prompts, files, git worktrees, and a trajectory database. It does not implement a defeasible-logic engine. The comparison therefore separates operational bookkeeping available now from the formal epistemic machinery still missing.
How to read the status column: each label measures Sorcar against the essay's full specification, not merely whether Sorcar has a related feature.
| What the essay asks for | Nearest KISS Sorcar mechanism | Status |
|---|---|---|
| Explicit hypotheses as inspectable objects, not activations in a forward pass | The AI-discovery and optimization prompt templates instruct the agent to maintain a journal file of tried ideas, recording outcomes so failed approaches are not silently retried. For runs using those templates, this creates an inspectable prose ledger. It is not a typed hypothesis object, and compliance is prompt-driven rather than enforced by the framework. | Partial |
| Defeaters actively sought, not just absorbed — a reasoner that attacks its own conclusions | Sorcar supports dynamic model switching, and its optimization and bug-fixing workflows specify a cross-vendor reviewer instructed to hunt for missed claims, broken wiring, and newly introduced bugs. This is defeater-seeking review, implemented as a workflow rather than a logic rule. In this report, that pass undercut the draft's inference that typos prove human authorship and caught a third factual error. | Built (workflow) |
| Answer to "which of your beliefs just took the hit?" | Every tool call, model response, and cost is persisted to a replayable trajectory store (sorcar.db). A reviewer can inspect the event where a claim first appeared, while the mandatory research file keeps source-indexed notes. That is useful forensic provenance, but it is recorded per event, not per belief: no dependency graph links each conclusion to its live premises or automatically identifies every claim affected by a bad source. |
Partial |
| Answer to "what would change your mind back?" | The templates call for explicit numeric stopping conditions, held-out evaluation splits, and anti-reward-hacking instructions declared before search. These provide auditable acceptance criteria for candidates and outputs. They do not say what evidence would reinstate a particular defeated belief, and prompt clauses do not technically prevent the model from gaming a metric. | Partial analogue |
| Lakatosian hard core + protective belt, with anomalies routed to the belt | Branch-per-task git worktree isolation keeps speculative changes off the user's main working tree and makes a failed trial easy to discard. That resembles a protective belt operationally. It is not a Lakatosian hard core: a branch may revise any assumption or file, successful changes can later be merged, and no semantic rule routes anomalies to belt rather than core. | Partial analogue |
| Beliefs that survive outside the opaque forward pass | At a context-window boundary, the continuation protocol asks the model for a structured chronological summary that seeds the next sub-session. This externalizes working state instead of relying on hidden activations alone. The summary is still model-written and potentially lossy, however; it is not a complete belief store or a premise-to-conclusion graph. | Partial |
| Typed defeaters: rebutting vs undercutting, with defeat semantics (ABA-style) | Nothing. Defeats happen and are logged, but they are not typed, and no admissibility computation decides what survives. This is prompt discipline, not Pollock. | Missing |
| A tunable credulous-vs-skeptical temperament dial | Nothing explicit. The system prompt is uniformly skeptical-leaning (verify before finish, no fabricated source counts), but there is no per-task dial and no extension semantics behind it. | Missing |
What the comparison shows BOTH WAYS
For the essay: its wishlist is not purely speculative. Sorcar shows that hypothesis journals, cross-model challenge, replayable traces, isolated experiments, and pre-declared evaluation criteria can all be assembled from ordinary prompts, files, git, and a database. These are useful approximations to the bookkeeping layer, not a completed implementation of it.
For the framework (and similar general-purpose agents): the sharpest items remain absent: typed defeaters with explicit defeat semantics, belief-level dependency tracking, principled reinstatement, and a temperament dial grounded in extension semantics. Specialized argumentation systems supply parts of that formal layer; Sorcar does not yet connect them to its operational plumbing. The essay is valuable because it names that integration target.
One caveat for symmetry: this mapping is analogical, not causal. Sorcar's mechanisms were engineered for long-horizon reliability, not derived from defeasible-logic theory. Their overlap supports a narrower claim than the essay's full thesis: reliable long-running agents benefit from external state, challenge procedures, provenance, and rollback. It does not by itself show that Sorcar represents beliefs, defeaters, or argumentation extensions.
Final scorecard
| Dimension | Grade | Notes |
|---|---|---|
| Core thesis (agents lack defeater machinery) | A | Correct, important, freshly framed. |
| Technical accuracy of the KR / logic content | B+ | Pollock, HTT, ABA, Nixon, Peirce faithful; the Biggs & Wilson deduction argument is mis-drawn. |
| Citation accuracy | C | Biggs & Wilson mis-dated by 21 years; Lakatos link points to the wrong book. |
| Framing discipline ("newest is 2004," "oldest 1987") | C− | Contradicted by the essay's own 2025–2026 and 1983 citations. |
| Intellectual honesty | A− | Flags its contested premise openly; overstates the "cannot" claims. |
| Practical, buildable proposal | A− | Glass-box + two questions = a real spec, grounded in ABA. |
| Writing style (AI-slop test) | Pass | No generic-slop signature; distinctive voice; needs copy-editing. Authorship not determinable from text. |
Bottom line: a smart, well-read, genuinely useful essay that would be excellent after a fact-check pass of its own. Fix the Biggs & Wilson date (2025, not 2004), drop the "nothing newer than 2004 / older than 1987" framing, repoint the Lakatos link to the methodology of scientific research programmes, and redraw the deduction argument to match Biggs & Wilson's actual criterion (conflicting conclusions with undefeated premises, not a reductio against a premise). The central thesis survives all four corrections intact.
Sources
Primary papers and authoritative references consulted for this fact-check. Where the essay's own link is correct it is noted; where it is wrong the correct target is given.
- Delrieux, C. (2004). "Abductive inference in defeasible reasoning: a model for research programmes." J. Applied Logic 2(4):409–437. ScienceDirect
- Pollock, J. L. (1987). "Defeasible Reasoning." Cognitive Science 11(4):481–518. Wiley
- Horty, Thomason & Touretzky (1990). "A Skeptical Theory of Inheritance in Nonmonotonic Semantic Networks." Artificial Intelligence 42(2–3):311–348. ScienceDirect
- Bondarenko, Dung, Kowalski & Toni (1997). "An abstract, argumentation-theoretic approach to default reasoning." Artificial Intelligence 93(1–2):63–101. ScienceDirect
- Etherington & Reiter (1983). "On Inheritance Hierarchies With Exceptions." AAAI-83 (the "inheritance-with-exceptions" paper the essay links). AAAI PDF
- Biggs, S. & Wilson, J. (2025). "The Indefeasibility of Abduction." In Hüttemann & Schurz (eds.), Inductive Metaphysics, Routledge. DOI 10.4324/9781003514404-7. PhilPapers · full text (author copy)
- Lakatos, I. Methodology of scientific research programmes (hard core / protective belt / heuristics): "Falsification and the Methodology of Scientific Research Programmes" (1970), collected in The Methodology of Scientific Research Programmes (1978). (The essay mistakenly links Proofs and Refutations.)
- Stanford Encyclopedia of Philosophy, "Abduction." plato.stanford.edu
- Peirce (biography / triad). Internet Encyclopedia of Philosophy
- Platypus history (hoax, Shaw, scissors). Wikipedia · BBC Science Focus
- Nixon diamond. Wikipedia
- Beel et al. (2025). "An Evaluation of the AI Scientist" (Sakana critique). arXiv:2502.14297 — quotations verified against the PDF. arXiv
- Agrawal et al. (2026). "Can AI Conduct Autonomous Scientific Research? Case Studies on Two Real-World Tasks." bioRxiv, DOI 10.64898/2026.01.05.697809 — "sophisticated hallucinations" quote verified via Crossref metadata and indexed full text. bioRxiv
- Gottweis et al. (2026). "Accelerating scientific discovery with Co-Scientist." Nature 655:487–496. DOI 10.1038/s41586-026-10644-y. Nature
- Pauli / "not even wrong" (Peierls attribution). Wikipedia
- KISS Sorcar (agent framework used for the case-study mapping and to produce this review): Sen, K., "KISS Sorcar: A Stupidly-Simple General-Purpose and Software Engineering AI Assistant." GitHub · website
- The essay under review: "We Forgot to Teach AI Agents to Be Wrong on Purpose," Vislesy Ventures, 2026-06-18. vislesy.com