The only AI safety infrastructure with empirical evidence of catching convergent hallucination.
When six LLMs agree on the same wrong answer about a regulated topic, standard "consensus" scoring hides the failure. Quorum catches it. Patent-pending. EU AI Act-aligned. Deployed in production.
Multi-LLM "consensus" was supposed to solve this. It doesn't.
When Anthropic, OpenAI, Google, DeepSeek and Grok all agree on the same answer, it feels safe. In regulated domains, it isn't. Their training data overlaps. Their stale snapshots align. The consensus scoring reports high confidence — on a shared fabrication.
Regulator renames
OISC was renamed to Immigration Advice Authority (Jan 2025). Six sub-models we tested still use "OISC" — with full confidence.
Fee updates
The UK Immigration Health Surcharge changed from £624 to £1,035 in Feb 2024. DeepSeek still returns £624.
Article misattribution
DeepSeek reports EU AI Act Article 31 as "deployer obligations." It actually governs Notified Bodies. Deployer obligations are Article 26.
Invented citations
Models fabricate section numbers ("Appendix Skilled Worker §3.2.1 v2026-06-15") that do not exist. All six sub-models cite it together.
Who this puts at risk
Legal firms giving immigration advice. Financial services making FCA compliance calls. Medical institutions triaging drug interactions. Any regulated professional whose advice is now increasingly LLM-assisted — and whose insurance excludes AI-caused errors.
This is the first published empirical evidence.
Peer-published in July 2026 with DOI 10.5281/zenodo.21314595. Full reproducer released. Every claim in the benchmark verified against gov.uk, FCA, artificialintelligenceact.eu. No Wikipedia. No LLM ground truth. Only primary authoritative sources.
What we built
A two-layer safety guard: (1) a domain-aware regex heuristic that flags fabrication shapes (SOC codes, precise monetary figures, invented citations, regulator renames) at zero LLM cost. (2) An LLM-as-judge second line that identifies the shared prior the sub-models are converging on, and returns a calibration multiplier.
The layers fuse multiplicatively. On our benchmark, this catches 75% of fabrications that single-model calls silently approve — at a marginal cost of £0.00005 per query.
Who we work with
Law firms (immigration, corporate)
- Custom fabrication detection for your jurisdiction
- OISC / IAA / SRA-aware regex layer
- Audit trail for professional indemnity
Financial services (FCA-regulated)
- FCA rulebook citation verification
- Threshold value drift detection
- Article 14 human-oversight compliance
Healthcare (MHRA / NHS)
- Drug interaction claim validation
- NICE guideline citation checks
- Prescribing-rule fabrication guards
AI-first startups (EU AI Act)
- Article 14 human oversight readiness
- Annex VI internal conformity assessment support
- Advisory PDF evidence records per query
Enterprise pricing
Three ways to work with us. Self-serve tiers are on api.quorum-ai.dev. Enterprise is here.
| Engagement | Price | What you get |
|---|---|---|
| Integration project 4–8 weeks |
£15k – £50k fixed fee |
Custom regex layer for your regulated domain, integration into your existing LLM stack (Claude, OpenAI, Azure, Vertex), staff training, and 30-day post-launch support. Delivered under NDA. Typical scope: one domain, up to 4 provider integrations. |
| Retainer monthly |
£3k – £15k per month |
Ongoing benchmark expansion for your domain, priority bug fixes, quarterly regulatory-drift reviews, and access to updated regex catalogues as UK/EU regulators change. Includes 2 hours/month of expert consultation. |
| On-prem deploy annual licence |
from £120k per year |
Full Quorum stack deployed in your own VPC / on-prem infrastructure. Suitable for regulated tenants who cannot use hosted APIs. Includes source code licence, SLA, named support engineer, and 4 hours/month expert consultation. Custom terms. |
All engagements include the patent-pending HSP tamper-evident traceability log (PCT/US26/11908). All prices exclude VAT. Discounts available for multi-year commitments.
Ready to see how many fabrications your current LLM stack is silencing?
Book a 20-minute discovery call. I'll run 10 prompts from your domain against your existing stack and Quorum side-by-side, and send you the results within 48 hours. No slide deck. No sales pitch. Just numbers.
Book a call →Or email directly: jaqueline@hsp-protocol.com