laya-phishield / how it decides

Eight narrow questions, one verdict, every reason named.

No single "is this phishing?" call. A deterministic pre-pass reads the headers and links, a local decision model answers eight atomic yes/no signals in one forward pass, and a logistic head combines them into a score with its reasons attached.

Both runs below replay real measured values from the shipped model.

01 the email arrives
01email
02deterministic pre-pass
display name claims a brand it does not own no other header or URL flag fires

Pure stdlib parsing: reply-to mismatch, IP-literal and punycode link hosts, homoglyph brand domains, SPF/DKIM failures. Obfuscated artifacts become readable statements the model can weigh.

03eight atomic signals · one forward pass
04logistic head → verdict
0.990 phishing
Why
asks_credentials+2.56
urgency_pressure+1.86
external_link_risk+1.56

Contributions are coefficient × value in logit space: what actually pushed the score. The same pipeline on the legitimate email scores 0.053.

The eight signals

Each is one narrow yes/no question over the structured email state, phrased positively. Negated and forced-choice questions are the documented failure mode this decomposition exists to avoid.

asks_credentials

Does the email tell the recipient to confirm their password or account login to restore or keep access?

"confirm your password within 24 hours to restore your account"

external_link_risk

Does the email send the recipient to click a link hosted on a domain unrelated to the sender or the brand it claims?

"PayPal" linking to http://203.0.113.41/verify

urgency_pressure

Act immediately or within hours, for example threatening consequences for delay?

"within 24 hours or your account will be closed"

brand_impersonation

Borrows the name or branding of a known company while the sending domain belongs to someone else?

"PayPal Security" from paypal-account-verify.example

payment_gift_request

Asks for a payment, transfer, gift cards or cryptocurrency to the sender or an account the sender names?

"buy 8 x $100 gift cards and reply with the codes"

requests_pii

Asks the recipient to send personal identity data such as an ID number, card number or date of birth?

"send your full name, date of birth and phone number"

too_good_to_be_true

Promises an unearned prize, lottery win, inheritance or windfall?

"you have won GBP 2,500,000.00 - selected at random"

suspicious_instructions

Asks the recipient to keep the request secret from colleagues or bypass normal procedures?

"do not discuss this with accounting yet"

negation trap, held

A security drill saying "we never ask for your password" scores 0.11 on asks_credentials. Wording that broke this hard negative was rejected in testing.

wording is measured

The first draft of asks_credentials scored 0.59 on a textbook phish. The shipped one-sentence wording scores 0.92 as the top reason. Same model, same email.

Measured, not promised

temporal test split, n=183keywordforced-choicecomposite
AUC0.59470.94440.9531
Precision / Recall @ 0.51.000 / 0.0510.902 / 0.5900.913 / 0.808
FPR @ 95% TPR1.0000.1240.210
$ per 1,000 emails$0$0$0

Nazario phishing vs Enron legitimate, MinHash-deduplicated across classes (95 of 700 removed) and split per-class by time. The composite beats the wide forced-choice question by 0.9 pt AUC and 21.8 pt recall at the operating point; the wide question still ranks better at the 95%-recall tail, and that limit is reported with the win.