| # | paper | title | value |
|---|
For each paper, we extract plain text from paper.html (dropping <script>, <style>, <pre>, <code>, the auto-injected enhancement panels, and the References section). Then:
log((N+1)/(df+1)) + 1 for IDF.df ≤ 3 (used by ≤ 3 papers total).Source: tools/corpus_analyze.py. Stdlib-only; no sklearn or external models. Run on demand; output in assets/corpus_analysis.json · corpus_keyword_index.json · corpus_similarity.json.