Metadata-Version: 2.4
Name: killpass
Version: 0.2.1
Summary: AI that double-checks itself before it speaks: a skeptic harness that attacks LLM claims against source documents.
Author: Srinivas Merugu
License: Apache-2.0
Project-URL: Homepage, https://github.com/srinu16it/killpass
Project-URL: Evidence, https://github.com/srinu16it/arin-research-notes
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Provides-Extra: pdf
Requires-Dist: pypdf>=4; extra == "pdf"
Dynamic: license-file

# killpass

**AI that double-checks itself before it speaks.**

![How killpass works](assets/how-it-works.svg)

## The problem

AI assistants believe things too easily.

Ask an AI: *"Did this company raise its forecast?"* It reads a headline that
says "Company Raises Outlook!" and answers **yes**. But the actual document
says the forecast was *cut*. The AI repeated the headline instead of checking
the source.

Every team building AI products has this problem. It is the main reason
people don't trust AI answers.

## What killpass does

> It gives your AI a built-in **skeptic** — a second AI whose only job is to
> try to prove the first one wrong, before the answer reaches a human.

Like spell-check, but for AI claims.

## How it works — three steps

1. **Your AI makes a claim.** *"This company raised its forecast."*
2. **killpass sends the claim to a skeptic AI** with one instruction:
   *don't trust it — go to the original document and try to destroy this
   claim.* If it can't be proven from the source, it fails.
3. **Back comes a simple verdict:**
   - **CONFIRMED** — the claim survived the attack, with the exact quote
     that proves it
   - **REFUTED** — here is the sentence in the document that kills it

The default is skepticism: when the evidence is unclear, the claim fails.

## Why believe this works

This isn't a theory. killpass is extracted from a live research system that
runs it every day. One real week:

The system nominated four stocks because headlines said "raised guidance."
Each claim went through the skeptic. **All four died** — one press release
literally said "Updates," not "Raises"; another buried a forecast *cut*
inside the same document; a third funded its good news with a new loan the
headline never mentioned. The skeptic caught what the reader-AI missed,
four out of four, by reading the actual SEC filings.

The receipts are public: [arin-research-notes](https://github.com/srinu16it/arin-research-notes)
— every kill is documented with dates and sources.

## It works on any news, not just finance

The same three lines handle any domain. The clearest demo is one study, two
claims (run it yourself: `examples/any_news_test.py`, tested with a free
local model at temperature 0):

| Claim | Verdict | The deciding quote |
|---|---|---|
| "A new study **proves** coffee **prevents** diabetes" | **REFUTED** | *"As an observational study, no causal conclusions can be drawn."* |
| "An observational study found coffee drinkers had a 12% lower incidence" | **CONFIRMED** | *"...showed a 12% lower incidence of type 2 diabetes..."* |

Same study. Same source. Opposite verdicts. The skeptic isn't anti-claim —
it's anti-**overclaim**. That's the whole product in one table.

(Honest limit: killpass verifies claims against the sources *you provide* —
it checks "does the document say this," it does not discover truth. If the
source lies, the verdict faithfully cites the lie. Every verdict is
auditable because it must quote its document.)

## Who needs this

Anyone whose AI answers questions from documents: legal tech, finance
tools, medical summaries, customer support, research assistants. They all
ship wrong answers today, and they all know it.

## The honest limit (read this)

killpass checks that a verdict is **grounded in a real, verbatim source
span** — not that the claim is *true*. A quote can exist in a document yet
be negated (*"we deny guidance was raised"*) or a third-party rumor. A
substring engine cannot catch those, so we **measure** the residual instead
of faking a filter. On the shipped benchmark (local model, temperature 0),
negation and rumor false-confirmed at **0%**, true positives were not
over-rejected, and a real fake-raise was correctly refuted. Full model:
[SECURITY.md](SECURITY.md) · frozen contract: [SCHEMA.md](SCHEMA.md).

## For engineers (30 seconds)

```python
from killpass import Skeptic

skeptic = Skeptic(llm=my_llm)  # my_llm: prompt -> text, any model

verdict = skeptic.attack(
    claim="Acme Corp raised its FY26 guidance",
    sources=[acme_press_release, acme_10q],
)

print(verdict.result)     # CONFIRMED | REFUTED | INSUFFICIENT
print(verdict.evidence)   # the exact quote that decided it
print(verdict.rationale)  # one paragraph, human-readable
```

Design principles (full detail in [DESIGN.md](DESIGN.md)):
- **Refute-first prompting** — the skeptic is rewarded for killing claims,
  not confirming them
- **Evidence or it didn't happen** — every verdict must quote its source
- **Unclear = failed** — ambiguity never passes
- **Model-agnostic** — bring any LLM; killpass is the harness, not the brain

## Bring any source — PDF, Word, web page, text

The judge only ever sees text; small loaders turn real-world documents into
text first (so what was judged is always inspectable):

```python
from killpass import Skeptic, load

sources = [
    load("acme_press_release.pdf"),          # PDF (pip install killpass[pdf])
    load("board_minutes.docx"),               # Word — no extra install
    load("https://ir.acme.com/news/q2.html"), # web page — no extra install
    open("notes.txt").read(),                 # plain text always works
]
verdict = Skeptic(llm=my_llm).attack("Acme raised its guidance", sources)
```

Fetching stays separate from judging on purpose: `load()` runs before the
skeptic, never during — verdicts remain reproducible and nothing browses
the web mid-judgment.

## Status

**v0.1.0 — installable and tested.** `pip install -e .` from source today;
PyPI release imminent. Five tests cover the contract, including the
grounding rule: a verdict whose "evidence" does not actually appear in the
sources is automatically downgraded. The verification patterns come from a
production research system; the receipts are public in
[arin-research-notes](https://github.com/srinu16it/arin-research-notes).

## License

Apache 2.0.
