No single "is this phishing?" call. A deterministic pre-pass reads the headers and links, a local decision model answers eight atomic yes/no signals in one forward pass, and a logistic head combines them into a score with its reasons attached.
Both runs below replay real measured values from the shipped model.
Pure stdlib parsing: reply-to mismatch, IP-literal and punycode link hosts, homoglyph brand domains, SPF/DKIM failures. Obfuscated artifacts become readable statements the model can weigh.
Contributions are coefficient × value in logit space: what actually pushed the score. The same pipeline on the legitimate email scores 0.053.
Each is one narrow yes/no question over the structured email state, phrased positively. Negated and forced-choice questions are the documented failure mode this decomposition exists to avoid.
Does the email tell the recipient to confirm their password or account login to restore or keep access?
Does the email send the recipient to click a link hosted on a domain unrelated to the sender or the brand it claims?
Act immediately or within hours, for example threatening consequences for delay?
Borrows the name or branding of a known company while the sending domain belongs to someone else?
Asks for a payment, transfer, gift cards or cryptocurrency to the sender or an account the sender names?
Asks the recipient to send personal identity data such as an ID number, card number or date of birth?
Promises an unearned prize, lottery win, inheritance or windfall?
Asks the recipient to keep the request secret from colleagues or bypass normal procedures?
A security drill saying "we never ask for your password" scores 0.11 on asks_credentials. Wording that broke this hard negative was rejected in testing.
The first draft of asks_credentials scored 0.59 on a textbook phish. The shipped one-sentence wording scores 0.92 as the top reason. Same model, same email.
| temporal test split, n=183 | keyword | forced-choice | composite |
|---|---|---|---|
| AUC | 0.5947 | 0.9444 | 0.9531 |
| Precision / Recall @ 0.5 | 1.000 / 0.051 | 0.902 / 0.590 | 0.913 / 0.808 |
| FPR @ 95% TPR | 1.000 | 0.124 | 0.210 |
| $ per 1,000 emails | $0 | $0 | $0 |
Nazario phishing vs Enron legitimate, MinHash-deduplicated across classes (95 of 700 removed) and split per-class by time. The composite beats the wide forced-choice question by 0.9 pt AUC and 21.8 pt recall at the operating point; the wide question still ranks better at the 95%-recall tail, and that limit is reported with the win.