Jaqueline Martins Sovereign Chain Ltd, United
Kingdom jaqueline@hsp-protocol.com
Draft v1 — for submission to SafeGenAI @ NeurIPS 2027 workshop.
Multi-LLM consensus systems are increasingly deployed in regulated professional services (immigration advisory, financial compliance, medical, legal), where practitioners hope that agreement among independent models will reduce hallucination risk. We document the opposite failure: when sub-models share overlapping training corpora with stale snapshots of a regulated domain, they converge on the same fabrication, and semantic-agreement scoring correctly reports high confidence in the wrong answer. On 2026-06-28, six sub-models produced 5-out-of-5 fabricated claims about UK immigration adviser rules at 78% consensus confidence.
We introduce a two-layer safety guard: (1) a domain-aware regex heuristic flagging fabrication shapes (SOC codes, precise monetary figures, invented citations, regulator renames) at zero LLM cost, and (2) an LLM-as-judge second line that, when the regex layer flags, calls a cheap model to identify the shared prior the sub-models are converging on and returns a calibration multiplier. The layers are fused multiplicatively.
We construct UK-EU-RegHalluc-v1, a piloto benchmark
of 10 claims across UK immigration and EU AI Act domains, each with
ground-truth answers verified against primary authoritative sources
(gov.uk, artificialintelligenceact.eu) on
2026-07-11. Every claim’s ground_truth_source URL is
included in the release so third parties can re-verify.
On this piloto, a single-model baseline (DeepSeek Chat) fabricates in 80% of prompts. Our two-layer guard reduces neither the fabrication rate nor the answer content — it flags 75% of fabrications for downstream caution at a marginal cost of £0.00005/query. We report this as calibrated risk classification, not correction. The guard is particularly relevant to Article 14 (human oversight) obligations under the EU AI Act, which take effect 2026-08-02 for high-risk systems.
Code, benchmark, and reproducer are released under FSL-1.1-Apache-2.0.
The 2024–2026 wave of multi-LLM consensus systems (Council Mode [Anonymous et al., 2026], Semantic Quorum Assurance [Author, 2026], multi-agent debate frameworks) is grounded in a simple intuition: if k independent models agree, the answer is more likely correct than if k = 1. This intuition holds for many domains, but breaks in a specific and important way for regulated professional services, where:
When these conditions hold, semantic agreement between sub-models is not evidence of factual truth — it is evidence of shared prior overlap on a stale snapshot. We call this failure mode convergent hallucination.
On 2026-06-28, a real production consensus query asking how a UK immigration adviser should register produced a 78%-confidence answer in which five of five verifiable factual claims were wrong:
Six of eleven sub-models agreed with each other inside this fabricated world. Standard semantic-agreement scoring reported 78% confidence and would have been reported as-such to a downstream reviewer.
Council Mode [Anonymous et al., 2026] describes a regex triage + lightweight classifier + 3-model panel architecture nearly identical to ours in shape, and explicitly acknowledges — but does not resolve — that “if all experts share the same misconception, the Council will reproduce it.” We operationalise a solution.
Hallucination Cascade [Anonymous et al., 2026] quantifies shared-prior propagation across multi-agent LLM systems using a sigmoid weighting formula. Their focus is measurement; ours is prevention at inference time.
Semantic Quorum Assurance [Anonymous, 2026]
introduces a quorum-based validator panel with adaptive assurance
scoring. The overlap with our work is significant; we differ in (a)
domain-aware regex first layer, (b) explicit
shared_priors[] extraction in the judge’s structured
output, and (c) release of a domain-verified benchmark.
HALoGEN [Ravichander et al., ACL 2025] provides the canonical taxonomy of hallucination categories (memory-based, incorrect-knowledge, fabrication) that we tag each benchmark claim with.
Cleanlab TLM [Cleanlab, 2024] and Patronus Lynx [Patronus, 2024] are commercial single-model uncertainty scorers. Neither targets multi-model consensus scenarios; neither surfaces shared priors.
Multi-agent debate [Du et al., 2023; Liang et al., 2024] and self-consistency [Wang et al., 2023] both use inter-model or intra-model disagreement as a signal, but assume disagreement is protective. Convergent hallucination is the class of failure where this assumption breaks.
The first layer is a domain-aware regex heuristic that flags known fabrication shapes in the prompt and answer. It runs offline (no LLM call), is deterministic, and is safe in the hot path. It costs approximately 0.5ms per query.
We flag six categories:
| Category | Example pattern |
|---|---|
| UK regulatory term | OISC, IAA, FCA, GMC, Companies House |
| Legal citation | § followed by dotted number; Article n(m); Appendix … § |
| Precise monetary figure | £X,XXX with no rounding |
| Official code | SOC XXXX, NACE letter+digits, ICD-10 |
| Versioned document | vYYYY-MM-DD |
| Regulator rename | keyword pair (OISC + IAA), (Cadre + Cadre Ordinaire) |
When flags fire, the guard emits a risk_level in
{low, elevated, high} based on flag density and reported
consensus confidence. Levels elevated and high
trigger the second layer.
When the regex guard flags, we call a single LLM (Haiku,
GPT-4o-mini or the cheapest meta-capable model in the pool)
with a strict JSON schema, side-by-side rendering of the n sub-model responses, and the
original prompt. The judge returns:
{
"is_verifiable_claim": bool,
"shared_priors": [str, ...],
"falsifiable_via": [str, ...],
"epistemic_calibration": float in [0,1],
"reasoning": str
}
The judge’s system prompt directs it not to answer the
original question, but to identify what assumption the
sub-models appear to be drawing from.
shared_priors explicitly names the assumption (e.g. “all
six models cite the pre-2025 name OISC”).
We fuse the two signals multiplicatively:
Cfinal = Coriginal ⋅ (1 − pregex) ⋅ cjudge
where pregex ∈ [0, 1] is the regex penalty and cjudge ∈ [0, 1] is the judge calibration.
Rationale: multiplication rather than mean or min ensures that a false positive in one layer costs some confidence but does not destroy correctness. Both layers must vote to trust for the score to survive intact. When the judge errors (provider down, unparseable JSON), the regex penalty applies alone.
We construct 10 piloto claims split as 7 UK immigration + 3 EU AI Act. Every claim contains:
correct_answer verified against a primary
authoritative source.ground_truth_source URL.ground_truth_verified_date (2026-07-11 for all piloto
claims).fabricated_variants — real observed
fabrications tagged with the model that produced them and the
observation date.fabrication_pattern list from a controlled
catalogue.Sources accepted for ground truth: gov.uk,
legislation.gov.uk,
artificialintelligenceact.eu,
eur-lex.europa.eu, fca.org.uk,
gmc-uk.org. Wikipedia rejected. Every URL is included in
the release so third parties can re-verify.
Per system-under-test we report:
For the piloto we use a rule-based downstream judge that extracts high-signal tokens (SOC codes, £ figures, dates, article/annex references, regulator names) from the ground-truth extract and from the observed fabricated variants. The judge marks an answer as fabricated if it contains any variant’s key tokens, or if it contradicts the ground-truth extract on those tokens.
This is transparent and reproducible but weaker than an LLM judge. In the expanded n=300 benchmark we will use an LLM judge with an inter-annotator agreement study.
single_deepseek — DeepSeek Chat, one call, no
guard.single_grok — Grok 4, one call, no guard (crashed on
API credit during the piloto run; results excluded).quorum_full — quorum.core.consensus with
regex guard and epistemic judge, using DeepSeek + Grok as the 2-model
pool.| System | Fabrication rate | Flag rate | £ / query | Latency (avg) |
|---|---|---|---|---|
single_deepseek |
80% | 0% | £0.00034 | 4.8s |
single_grok |
— | — | — | (excluded) |
quorum_full |
80% | 75% | £0.00039 | 8.1s |
The single-model baseline fabricates in 80% of prompts. With only 2
providers in the consensus pool, quorum_full does not
reduce the fabrication rate. What it does is flag 75% of the
fabrications for downstream caution at a marginal cost of
£0.00005/query.
Per-domain fabrication is uniform (80% immigration, 67% EU AI Act) — the convergent-hallucination surface is not artefactual to a single domain.
Example — EU AI Act Article 31: DeepSeek reports
that Article 31 “governs the obligations of deployers of high-risk AI
systems.” Ground truth (verified at
artificialintelligenceact.eu/article/31/) is that Article
31 establishes requirements for Notified Bodies.
Deployer obligations are Article 26. quorum_full flagged
this at risk_level=high and reduced the reported confidence
from 1.00 to 0.11.
Example — UK IHS: DeepSeek reports the Immigration
Health Surcharge as £624/year, the pre-Feb-2024 figure. Ground truth
(gov.uk) is £1,035. quorum_full flagged and reduced
confidence to 0.64.
Regulated professional advice is a high-stakes domain. Even a
75%-flag-rate guard should not be interpreted as a substitute for
qualified human review. Our system’s disclaimer field
(_disclaimer in every /v1/consensus response)
states verbatim: “Advisory only — not a conformity assessment under
Regulation (EU) 2024/1689.”
The dataset does not contain personally identifiable information. The
fabricated_variants are synthesised in the piloto; the
expanded n=300 benchmark will include real observed fabrications with
model identifiers logged.
Code and benchmark are released under Functional Source
License 1.1 (Apache-2.0-future) at
github.com/jaquelinejaque/sovereignchain. Reproduction:
git clone <repo>
cd sovereignchain/quorum
uv sync && source .venv/bin/activate
export DEEPSEEK_API_KEY=... XAI_API_KEY=...
python benchmarks/uk_regulated_hallucination_v1/runners/run_benchmark.py
python benchmarks/uk_regulated_hallucination_v1/runners/report.py results/*.jsonFollowing NeurIPS 2024+ disclosure guidelines: this work used
Anthropic Claude as a coding assistant for implementation of the
epistemic_judge module
(src/quorum/core/epistemic_judge.py), the benchmark runner
(benchmarks/.../run_benchmark.py), and for drafting
assistance on this manuscript. All experimental design decisions,
hypothesis formulation, benchmark construction, ground-truth source
verification, and result interpretation are the author’s.
(To be completed in v2. Placeholder anchors: Council Mode arXiv 2604.02923; Hallucination Cascade arXiv 2606.07937; Semantic Quorum Assurance arXiv 2606.08021; HALoGEN arXiv 2501.08292; Cleanlab TLM; Patronus Lynx; Du et al. 2023 multi-agent debate; Wang et al. 2023 self-consistency.)