Risk Dashboard
Behavioral safety index and model vulnerability profiling
Failure Breakdown by Product Risk
Aggregated incidents weighted by product-risk multipliers
Loading live failure metrics...
Observatory Engine
Real-time heuristics & LLM probe
Continuous deterministic scanning for citation hallucinations, excessive confidence, negation violations, persuasive manipulation, memory loss, and recursive circular looping.
Adversarial LLM Testing & Inspection
Probe models with adversarial traps or inspect custom text
Inspector Diagnostics Verdict
Failure mode determination & evidence
Configure a prompt and trigger the test to view live diagnosis.
Live Incident History
All logged tests and detected anomalies in persistent storage
| Time | Provider / Model | Detected Category | Subtype | Confidence | Status | Action |
|---|---|---|---|---|---|---|
| Loading incidents... | ||||||
Reproducible Evaluation Suite
Deterministic benchmark cases for 6 core failure categories
Click 'Run All Evaluations' to execute local reproducible failure test cases.
AI Failure Taxonomy Classification
Formalized classification of Generative AI failure modes mapped to severity multipliers and mitigation strategies.