Verascan Report

Data Contamination & Leakage Audit

Threshold: {{ report.threshold }} Methods: {{ report.methods_used | join(', ') }} {{ timestamp }}
{% if report.contamination_rate == 0 %}

Clean — No Contamination Detected

0 of {{ report.eval_size }} evaluation examples leaked into the training set.

{% elif report.contamination_rate < 0.03 %}

Low Risk Contamination ({{ contamination_pct }})

{{ report.total_matches }} matching pairs detected across datasets.

{% else %}

Significant Contamination Detected ({{ contamination_pct }})

{{ report.total_matches }} matching pairs detected — evaluation results may be inflated.

{% endif %}
{{ contamination_pct }}
Evaluation Set Size
{{ report.eval_size | e }}
Total eval samples audited
Training Set Size
{{ report.train_size | e }}
Reference training corpus
Total Matches Flagged
{{ report.total_matches | e }}
Pairs meeting threshold
Detection Breakdown
Exact: {{ report.exact_matches }} {% if report.ngram_matches > 0 or 'ngram' in report.methods_used %} N-gram: {{ report.ngram_matches }} {% endif %} Fuzzy: {{ report.fuzzy_matches }} {% if report.semantic_matches > 0 or 'semantic' in report.methods_used %} Semantic: {{ report.semantic_matches }} {% endif %}
Across {{ report.methods_used | length }} engine(s)
{% if matches %}
{% if report.exact_matches > 0 %} {% endif %} {% if report.ngram_matches > 0 %} {% endif %} {% if report.fuzzy_matches > 0 %} {% endif %} {% if report.semantic_matches > 0 %} {% endif %}
{% for m in matches %}
#{{ loop.index }} {{ m.method }} Score: {{ '%.3f' | format(m.score) }}
Eval [#{{ m.eval_index }}] ↔ Train [#{{ m.train_index }}]
Word-Level Diff Comparison Train → Eval Diff
{{ m.diff_html }}
Evaluation Text Index #{{ m.eval_index }}
{{ m.eval_text_escaped }}
Training Text Index #{{ m.train_index }}
{{ m.train_text_escaped }}
{% endfor %}
{% else %}
✅
No contamination detected

None of the evaluation samples exceeded the similarity threshold against the training set.

{% endif %}