ragpreflight — Corpus Readiness Report

NEEDS REVIEW Average 74/100 · 52 critical issue(s) require attention before ingestion.
73.8/100 175 documents · 0 duplicate group(s)
📁 olmOCR-bench validation sample (175 PDFs · 7 categories) 📄 175 docs ⚠ 52 critical 📋 547 total issues

Corpus-level Issues

Top Actions Required

Document Scores (click a filename to expand issues)

ScoreFileFormatIssuesCritical
45
148c79ae1507bf2a6ee7c11c91559ac6aa2f0556_page_9_processed.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
10a_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
10c_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
14d_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
14e_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
17_pg17_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
17_pg2_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
17_pg33_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
17_pg4_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
02a3e23da54b82c414d272051b0b5d8f44d8_page_3_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
06125479b00df48e91fe2fb616292d14116d_page_4_pg1.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
22.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
32.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
23.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
36.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
33.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
45.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
38.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
51.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
58.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
62.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
59.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
67.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
72.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
77.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
74.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
81.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
80.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
83.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
90.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
92.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
93.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
98.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
97.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
94.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
99.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
1_pg113.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
1_pg10.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
1_pg125.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
1_pg131.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
1_pg19.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
1_pg61.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
1_pg72.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
3_pg101.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
3_pg148.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
3_pg199.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
3_pg204.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
3_pg251.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
3_pg270.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
3_pg634.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
5_pg558.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
45
5_pg281.pdf ● CRITICAL
CRITICAL [content] Only 0% of pages have extractable text. This document is likely a scanned PDF.
→ Apply OCR (Tesseract, AWS Textract, Google Document AI) to the full document before ingestion.
WARNING [content] Page yields no extractable text — may be a scanned image. (page 1)
→ Run OCR (e.g. Tesseract) on image-only pages before ingestion.
WARNING [content] Low content density (0%) — document may be mostly whitespace or boilerplate.
→ Verify the document contains meaningful content before ingestion.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF41
82
1596c5fcd2f89b0fd3a3aa06c13034dc09317d6f_page_12.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "Geomagnetism and Aeronomy International, v.1, N 1, p.53-58, 19" · "odel of the upper atmosphere with variable latitudinal integr" · "o O.V., Namgaladze A.N. Global model of the upper atmosphere"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'ru' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
82
70b11ae28cefcbcdb69d51840f5f6c127821a0aa_page_5.pdf
WARNING [content] Math-heavy document: 27 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ ""R2  ) V:L  ±   cH ( N. ca"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ " 0! 1   ! -./ 23 Nabis capsiformis (Het., Nabidae)" · " 0! 1   ! -./ 23 Nabis capsiformis (Het., Nabidae) 4/"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'cy' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF50
83
2503.08925_pg16.pdf
WARNING [content] Math-heavy document: 48 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "ce a (2, 2)-isogeny Φ = 휙∗×휙′∗:  " · "tomorphism group 퐶2 × 퐶5, all families in"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "푄) 푅− Õ 푅∈휙′−1 (∞) 푅, whose kernel is contained in (퐸푡,푠×퐸푠,푡)" · "2 = 푥6 + 푡푥4 + 푠푥2 + 1 over a finite field F푝with 푝odd. The quo" · "AND WESOLOWSKI 6.1. Automorphism groups admitting a subgroup"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
83
2503.09299_pg16.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "( ′)) graphon Trunc. SVD(A′) Figure 1: Upper Left: difference" · "gure 1: Upper Left: difference between target functions for o"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 4 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
83
065d792cf3923fc76a17183cfd13480cd51d_page_9_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "nheim . . . . . . 11. Internationaler CIGR-Kongreß für Agr" · "Strukturwandlungen der Landwirtschaft und einige Auswirkun" · "Fahrzeuge) Strukturwandlungen der Landwirtschaft und einige"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'de' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
83
081875a1035e34dee439b9a2a3a55e319405_page_4_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "2 Los Alamos National Laboratory Our discussion" · "Alamos National Laboratory Our discussion is confined to macr"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 3 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
84
2503.06570_pg6.pdf
WARNING [content] Math-heavy document: 59 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "d by the formal sum α ⋆τ β = α ∪β + X d∈E" · "rmal sum α ⋆τ β = α ∪β + X d∈Eff̸=0, 0≤ℓ≤"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Continuous and Discrete Asymptotic" · "Continuous and Discrete Asymptotic Behavi"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'ca' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF50
84
2503.07373_pg21.pdf
WARNING [content] Math-heavy document: 324 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "We see eQ0(δχω) = −1 3!Q0(¯χγ3dω" · "] + 1 2Q0(¯χγaχeµ a)∂µ = 1 2[ξ, ϕ] + Lω ξ"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces, zero_as_o (page 1)
↳ "[c, δχω]) + eδ2 χω, hence obtaining eQ2 0ω = eQ0(ιξFω −dωc +" · "We see eQ0(δχω) = −1 3!Q0(¯χγ3dωψ" · "Q2 0 right away, obtaining Q2 0c = QP C(ιξδχω) + δχ 1 2ιξιξFω"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'el' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF50
84
2503.09472_pg15.pdf
WARNING [ocr] OCR substitution artifacts detected: l_I_1, broken_ligatures, mid_word_spaces (page 1)
↳ "    dx1 dt = x2 + b11 b10x1x2 + b20 b2 10x2 2 + b12 b10x2" · "x2 4 + v03x3 4 + O(∥x∥4) wherein: v11 = −b11 b10 , v02 = −b20" · "so we can get      dx1 dt ="
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
84
20_pg42_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: l_I_1, rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "sicObiomablendstheHomerianwithIgbomythJASON KEITH 1.BecomingM" · "in, King Arthur’s knights,/ Bernard and Philip, Giles, the res" · "breath in here, the trampled floor/ pungent and deep. The lit"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 2 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
84
00f6c2eea6b51fb1bf637b5f8a850588d7e2_page_10_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "(IJACSA) International Journal of Advanced Co" · "(IJACSA) International Journal of Advanced Compu" · "(IJACSA) International Journal of Advanced Computer Science a"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 3 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
84
b5c5b8661b5a272e7a175cdb20d49e67ba0d_pg4.pdf
WARNING [content] Math-heavy document: 7 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "path coefficients (β) support that brand"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "oud Lajevardi et al, 2014 Journal of Applied Science and Agri" · "i et al, 2014 Journal of Applied Science and Agriculture, 9(" · "Masoud Lajevardi et al, 2014 Journal of Applie"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 2 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
85
2503.04448_pg10.pdf
WARNING [content] Math-heavy document: 20 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "is given by: E[L] = λE[K] 2(1 −ρ)  α + ρ"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "ma 3. The average number of waiting customers in the polling" · "Lemma 3. The average number of waiting cust"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
85
2503.03903_pg9.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "E-SEM SCHUBERT POLYNOMIALS 9 Figure 5. Above is the bottom pi" · "POLYNOMIALS 9 Figure 5. Above is the bottom pipe dream for 1"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
85
2503.03879_pg4.pdf
WARNING [content] Math-heavy document: 56 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "indicator variable (η) for the switch poi" · "= 1, . . . , N −1 0 ≤(∆h+ l + ∆h− l ) ⊥ηl"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Fig. 2. Location of switch point" · "Fig. 2. Location of switch point at the end of"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
2503.05342_pg5.pdf
WARNING [content] Math-heavy document: 13 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "n a framing vector ⃗λ = (λ1, . . . , λc)," · "oup of the boundary ∂V of the solid torus"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "of the boundary ∂V of the solid torus V is the free abelian" · "It is known that the fundamental"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'da' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
85
2503.05436_pg25.pdf
WARNING [content] Math-heavy document: 22 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "(2;1) 25 Therefore σ conjugates ϕ and Q." · "follows that Fix(Q) ⊂Rm is sub vector spa"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "be a planar polynomial vector field of degree n as our polynom" · "FIELDS OF TYPE (2;1) 25 Therefore σ conjugates ϕ and Q. In p" · "efore σ conjugates ϕ and Q. In particular it follows that Fix"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'da' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF50
85
2503.06967_pg8.pdf
WARNING [content] Math-heavy document: 78 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "T], x′ 0, x0 ∈Rd0, α′ 0, α0 ∈A0, x′, x ∈" · "such that for all t ∈[0, T], x′ 0, x0 ∈Rd"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: O_0, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "able, and the map (x0, µ) 7→∂x0g0 is continuous. The map µ 7→" · "y is meant to serve as a mean field descrip- tion of the follo" · "8 (A3) There exists a constant cL > 0 such tha"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'ca' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF50
85
2503.07145_pg13.pdf
WARNING [content] Math-heavy document: 65 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "{x ∈dom(h) : h(x) ≤α} is closed for all" · "d λi yields ∂lijΦ = log(lij) −log(mij) + λi" · "le score, so that C ∩(0, ∞)n×n ̸= ∅. Then"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "onvex on idom(h). (3) h is co-finite because limt→∞ h(tx) t =" · "0, . . . , n −1) be an irreducible score, so that C ∩(0, ∞)n×" · "6. Let x ⪯(0, . . . , n −1) be an irreducible score, so that"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'nl' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF50
85
2503.06723_pg3.pdf
WARNING [content] Math-heavy document: 62 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "where φ(z) is described by" · "n what follows d, m ∈N will be two fixed"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "where φ(z) is described by the capacitary formula" · "where φ(z) is described by the capacitary fo"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'nl' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
85
062cc165e270a165b0660ec26b2478ac7a3fdab6_page_3.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "is a gradual change to wilderness and a sense of being in th" · "numerous industrial type structures associated" · "numerous industrial type structures ass"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
85
b79ccdcb9871dce02d3873bdf9690d9f6dd58cb4_page_19.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "55 INFEASIBILITY CRITERIA: Infiltration Setbacks Feature Set" · "ction Easement ≥ 20 feet Top of slopes >20% and over 10 fe"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
20_pg32_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: l_I_1, rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "d up to become an actorbecauseIloveallofit,Iwant to play all" · "ehind her”, Gough says. She warns against trying to understand" · "designer frocks, red carpets, flying film-festival visits, and"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
20_pg33_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: l_I_1, rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "ioPeris- Mencheta asJavier andIsabel Ripples but no waves in" · "s. Midway through, the film turns briefly into a farce as both" · "andmusic reviews and likes to flirt with his hostesses in the"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
20_pg35_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "the year the boy next doorlearnttoroar.EdSheer- an,thatsupers" · "r.EdSheer- an,thatsuperstarinaflannel shirt, left The Rolling" · "roar.EdSheer- an,thatsuperstarinaflannel shirt, left The Rolli"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
20_pg49_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "intheearlyyears,indulgent governmentsissuedfreepermitsand the" · "o a three- monthlowinDecember,figuresshow. The Purchasing Mana" · "for a tonne of CO2 and establishing a market cost for carbon"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
09f801e3a2ec90ef456d34ad571a46f36fce_page_35_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "n. L’enfance (Protection Maternelle Infantile, Aide Sociale à" · "SVERSAL DE L’USAGER” Les travailleurs sociaux sont en premièr" · "TRANSVERSAL DE L’USAGER” Les travailleurs sociaux sont en p"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [content] Document language detected as 'fr' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
85
2_pg39.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "III. ALGEBRAIC FUNCTIONS 25 Using V, with u = x2 + 1, v = x2" · "Ex. 6. x2 -\- xy — y 2 = 1. We can consider y a function of x"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
85
2_pg349.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Supplementary Exercises 167 215. Moment of inertia" · "a of a right circular cylinder about a line tangent to its ba"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
85
2_pg238.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "56 Simple Areas and Volumes Chap. 4" · "56 Simple Areas and Volumes Chap. 4 4. Find th"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
85
2_pg65.pdf
WARNING [content] Math-heavy document: 5 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "proved, d cos u = d sin ( - — u) = cosf - — u"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "TRANSCENDENTAL FUNCTIONS 51 Using the formula just proved, d" · "NSCENDENTAL FUNCTIONS 51 Using the formula just proved, d cos"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
0cda549ca8e0f36b6fbb414b9f19d580daa7_pg4.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Table 2 shows that participants with a suicide attempt" · "Table 2 shows that participants with a suici"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
56d7730c3751b4cde9e63d7355e1f6961d75_pg3.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "gin outline October 23: International structure and levels o" · "3 • Reading: Kim • Assignment: Final" · "Kim • Assignment: Finalize paper topic, begin outline O"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
5f20b8b654a2a30ed308fb87d69a636fec2a_pg4_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "nning, implementation, and learning obtained from the followin" · "consisting of planning, implementat" · "consisting of planning, implementation, a"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 2 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
6d48ac024cd71f0899fed96626b2b175c909_pg9.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "9 Campus Map (coming soon) Prayer Rooms T" · "9 Campus Map (coming soon) Prayer Rooms Ther"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 2 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
85
904a1b4e9f7cdc2169509e570c03b5e6507a_pg141.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "99 2000 2001 2002 Taxa de actividade económica (15 – 64 anos" · "1998 1999 2000 2001 2002 Taxa de actividade económica (15 –"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 2 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [content] Document language detected as 'pt' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
85
ebfe1f9a6792e1ba1c90c9f34ecf842504e2_pg29_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "29 application sets respectively. In p" · "29 application sets respectively. In particul"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
2503.05521_pg22.pdf
WARNING [content] Math-heavy document: 58 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "speration function τB×f F((p, x), (q, y)" · "peration function τB×f F((p, x), (q, y))"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "x), (q, y)) = 0. τB×f F satisfies the reverse △-inequality τB" · "22 CHRISTIAN KETTERER speration function τB×f F((p, x), (q," · "2 CHRISTIAN KETTERER speration function τB×f F((p, x), (q, y)"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'da' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
86
2503.05558_pg18.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "i i “main” — 2025/3/10 — 1:02 — page 7" · "i i i i 718 Michael R. Douglas and Cristofero Fraser-Taliente"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'da' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
2503.06334_pg4.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "4 KENNETH STEPHENSON The packing P can also be viewed as a c" · "4 KENNETH STEPHENSON The packing P can also be viewed a"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'it' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
2503.06716_pg12.pdf
WARNING [content] Math-heavy document: 34 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "er than or equal to α + O (1) →α > 0 as k" · "find a sequence (uk)k≥1 in ˙Hs RN such t"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "orem is not true. Then we can find a sequence (uk)k≥1 in ˙Hs " · "to present the proof of the Bianchi-Egnell type stability fo" · "12 CHAKRABORTY & SARKAR We are now ready to present the p"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'nl' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF50
86
2503.04045_pg2.pdf
WARNING [content] Math-heavy document: 22 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "xpressed as N = p + η, (1.2) where p is p" · "that every large N ≥1 can be rep- resent"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Theorem 1.2 (Chen). Every sufficiently large even number N can" · "OMAS Theorem 1.2 (Chen). Every sufficiently large even number N"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
2503.08910_pg26.pdf
WARNING [content] Math-heavy document: 66 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "´ES F. URIBE-ZAPATA σ(i′) = 1 for some i′" · "i′) = 1 for some i′ ∈J, then Ξ1(bσ(i′) i′"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "ers, B := P(ω), let B0 be the field of sets over ω generated b" · "i∈J bσ(i) i  = 0. In conclusion, Ξ(a) = Ξ0(g(σ0)), where σ0" · "F. URIBE-ZAPATA σ(i′) = 1 for some i′ ∈J, then Ξ1(bσ(i′) i′"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
86
2503.07310_pg2.pdf
WARNING [content] Math-heavy document: 5 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "ven by (Prob). min x∈X f(x) s.t. g(x, u)"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "ust feasible solution. An alternative approach is based on an" · "uncertainty. There are two key drivers" · "uncertainty. There are two key drivers for the se"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
2503.09337_pg14.pdf
WARNING [content] Math-heavy document: 14 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "∥W ⊙(D ∗t X −Y)∥F + λ∥X∥1, (17) such that" · ") such that Tα : RM1×M2×···×MN →RM1×M2×··"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: l_I_1, broken_ligatures, mid_word_spaces (page 1)
↳ "ge operator defined by: Tα(X)i1i2...iN = (|Xi1i2...iN | −α)+s" · "Then we define the completion problem as" · "Then we define the completion probl"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
2503.09374_pg3.pdf
WARNING [content] Math-heavy document: 46 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "ng model y = F(x) + η, (2.1) where F : Rd" · "= π(y|x)π0(x) Z(y) ∝exp(−Φ(x))π0(x), where," · "unknown parameter x ∈Rd from data y ∈Rn,"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "The paper is organized as follows: In section 2," · "The paper is organized as follows:"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
2503.09591_pg20.pdf
WARNING [content] Math-heavy document: 44 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "nduced subgraphs of ΛU. The base cases ar" · "se cases are when n ≤7, which are easily"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "en n ≤7, which are easily verified. For the inductive step, su" · "Upper Bound of Theorem 2 We aim to prove Theorem 2 by induct" · "4.3 Proof of Upper Bound of Theorem 2 We"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
2503.09469_pg13.pdf
WARNING [content] Math-heavy document: 90 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "therefore, λ1(B−ω0C′) = ⟨V 1(B)," · "⟩= 2µ1−ω0ξ since C′ ∈C. So the strict ine"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "rreducibility defined for alternating projections (2) that is" · "therefore, λ1(B−ω0C′) = ⟨V 1(B), (B" · "(B))⟩= 2µ1−ω0ξ since C′ ∈C. So the strict inequality is asser"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
86
0ccc00a50d5e4ab6920b2b6490b27d0c0f30df9d_page_29_processed.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Optimal Debt and Equity Values 136" · "Optimal Debt and Equity Values 1369 Figure"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
08f3aeec1eda1a569d0c9ef7bb892b5cb6a81b5b_page_1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Legal disclaimer Medix Biochemica prod" · "Legal disclaimer Medix Biochemica pr"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
68adea4d76e152a7abebfb39dc0d851fe5695a84_page_4.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "International Journal of Reproductiv" · "International Journal of Reproductive B" · "Joolayi et al 716"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
6c4fbd39f4d9a240edd464579b1e04d1ac687780_page_8.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "lish Center of Discovery & Learning • 33 South St. • Chicopee," · "en Dobry!, Vol. XV, No. 4, April 2014 — 8 ===== June 21, 2014" · "OW Polish Genealogical Society of Massachusetts Polish Center"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
b2ca8e0034542f65bd80e927f16e65c7d466b18c_page_2.pdf
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "less DSNs is largely in their flexible deployment and communic" · "98 P. Tarvainen et al. Such systems, hence" · "98 P. Tarvainen et al. Such systems, hencefort"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
e71f5c5980249ece8863fa1b43c0418e4b3244fd_page_6.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "www.iosrjournals.org" · "Evaluate the role of Tranexamic Acid in Joint Replacement Su" · "Evaluate the role of Tranexamic Acid in"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 2 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
13_pg162_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: l_I_1, rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "r in religious changes. Bist1oGRA PHY.—Richter, Die evangel" · "fessions under one common government, and, resulting from it," · "142 two different confessions under one"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
13_pg165_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "‘march from Vignamont to Tournay in face of the enemy. On h" · "ze on September 18, 1691. Again in the next campaign he cov" · "LUXEMBURG of England at Leuze on September 18, 1691"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
13_pg531_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "aks of the im- portance and ornamentation of Maltese dwelling" · "MALTA Carthaginian times, continued in Malta" · "MALTA Carthaginian times, continued in Malta unde"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
20_pg58_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: l_I_1, rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "eepicksworthconsidering. DannyIngs(Southampton) Sincejoiningf" · "rue value for the Denmark international. The 26-year-old (righ" · "Football Joini’sleagueontheofficial FantasyPremierLeaguegamet"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
0005784d0d255f6652180433936fa2998188_page_4_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "4 Deployment and lessons learned We implemented the model d" · "information specified on the dial" · "information specified on the dialog act, b"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
02943986dff2d4a84b7c2dd81abfc51fceea_page_7_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "024 /NO.6 las mujeres que comienzan su carrera hoy. En Dina" · "V. 2023 – ENE. 2024 /NO.6 las mujeres que comienzan su carre"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'es' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
019a8841dddeb304774ba62a9f5d8dccb43d_page_1_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "es / La Revue de médecine interne 32S (2011) S99–S191 H. Jeuli" · "S174 Communications affichées / La Revue de médecine i" · "S174 Communications affichées / La Revue de"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'fr' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
01ed6dcc6a3d0237daa88b293d0c0667df1a_page_49_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "ation and Cell Cycle • Murine fibroblasts (NIH 3T3 and CRE BAG" · "Chapter 15 — Assays for Cell Viability, Proliferation and Fun" · "698 Chapter 15 — Assays for Cell Viability, Proliferat"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
04e88521bffcae275b79ed2c89e24c1f8855_page_2_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "Chennai - 600 010. Indian Journal of Otolaryngologv and Head" · "Papillary Thyroid Carcinoma in a T" · "Papillary Thyroid Carcinoma in a Thyroglossal Cyst 295 F"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
03b953c18a4f7d2bbe390940dec839395cdd_page_18_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "y aircraft checked, filled, turned, tightened, touched, replac" · "Is it time for an annual inspection? E" · "Is it time for an annual inspecti"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
055c2edec4677b71f7bb6a759ebacc701bb1_page_4_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "ivity takes place in the afternoon with a maximum between 15:" · "to lower CG activity and vice versa lower Z va" · "to lower CG activity and vice ver"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
03e922d514407b071d31bfa0121f23d9ecff_page_2_pg1.pdf
WARNING [content] Math-heavy document: 6 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "variables were used χ2-test and t-test." · "espondents was 11.31±1.483 years, the av"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "groups, heavy school bag, furniture design that is not suita" · "165 Mater Sociomed. 2016 Jun; 28(3): 164-167" · "• ORIGINAL PAPER Epidemiology of Musculoskeletal Disorders i"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
0621a0090414e4681e90a4e1ea543acca910_page_4_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "L’ornementation des bracelets de l’" · "L’ornementation des bracelets de l’âge du B" · "L’ornementation des bracelets de l’âge du Bron"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'fr' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
081875a1035e34dee439b9a2a3a55e319405_page_16_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "14 Los Alamos National Laboratory a comparativel" · "s Alamos National Laboratory a comparatively robust “cermet”("
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
081875a1035e34dee439b9a2a3a55e319405_page_20_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "18 Los Alamos National Laboratory Preliminary re" · "ational Laboratory Preliminary results of the tests show five"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
0845863d27c84b14011e17a388821525e783_page_4_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "s of the system, namely the Pernambuco, Patos, Portalegre, Pic" · "of topography generation, modification, and denudation. Contin" · "sand content and maturity of turbidite deposits. Such"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
09f801e3a2ec90ef456d34ad571a46f36fce_page_42_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "es Électeurs ont choisi l’alternance au Conseil départemental." · "Tribune de l’opposition Les Élect" · "Tribune de l’opposition Les Électeurs"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'fr' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
081875a1035e34dee439b9a2a3a55e319405_page_47_pg1.pdf
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "azingly bright” light that “turned yellow, then red, and then" · "A backward glance . . . to avoid the blinding flash” he expec" · "A backward glance . . . to avoid the blin"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
0b128decb905a38700baf486672b8634c308_page_4_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "19, 2017 http://circres.ahajournals.org/ Downloaded from" · "ach animal 96 hours after the first vitamin D injection in the" · "control animals (Table). Calcium levels m"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
0e5f0c3447c4f1335b330d58b5e14d6f15b1_page_4_pg1.pdf
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "For the composite outcome of death or ESRD (F" · "For the composite outcome of death"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
0d767c7cab3af7117e33d74252a524497957_page_4_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "then the pressure difference between the basement" · "then the pressure difference betwee"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
0be9ba925bc2e3164ce649b02e3e29a07239_page_5_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "Environment Conservation Journal Medicinal value of Lep" · "53 Environment Conservation Journal" · "rvation Journal Medicinal value of Lepidium latifolium:"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
2_pg189.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Art. 5 Curves with a Given Slope 1 Integration" · "Art. 5 Curves with a Given Slope 1 Integrati"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
4_pg433.pdf
WARNING [ocr] OCR substitution artifacts detected: O_0, broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "o |e d3+ mb, bz Cg Kee r0g Cx mb, 03 Cz But the la" · "ch element of a column is multiplied by the same number, and" · "DETERMINANTS 417 489. If each element of a column is mu"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
4_pg451.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "that 5 is a root of the equation wv’ —3a?—182%+40=0. Divid" · "EQUATIONS 435 Ex. 1. Prove that 5 is a root of the equati"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
4_pg186.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "170 ELEMENTARY ALGEBRA Quantities of one kind are said to b" · "ELEMENTARY ALGEBRA Quantities of one kind are said to be inv" · "Let the proportion be a:b=6:¢. Then Ore Con nS aLGLe) H"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
529eeb84602e4fba7df7a4b127af6c830f97_pg12.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "esigning building that turn corners well, so that both elevati" · "We recommend Creating streets that are principall" · "We recommend Creating streets tha"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
4705915137b1295c5750cdaf91e48e23c7b7_pg3_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "Am J Clin Pathol 2008;130:865-869" · "EWQ 867 © American Society for Clinical Pathology Microbi" · "1309/AJCP04IZAMPISEWQ 867 © American Society for Clinical"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
83f40aac04a692b7aacbd813cf4da85354c8_pg26_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces, zero_as_o (page 1)
↳ "library relies on several external libraries. We use the Armad" · "and 137.0s, respectively, to fit the lasso path for the three" · "Data set n m Ratio Cancer 3.9k 217 1.14 Amazon"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
981f5f245a1c3dcbafc67603d942144df82e_pg7.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Economics of Education Review 83 (202" · "Economics of Education Review 83 (2021)"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
86
e79bd5294680f89f222a69f37e5b2059cc40_pg166.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "148 Table 23 Food Security Index Value and Range f" · "23 Food Security Index Value and Range f Percent Cumula"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
86
bfbadac12e1f83721b97b9787c5613fb4665_pg1_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "tinue this week. In the International Competitions and Asses" · "ar parents, Great student achievement results from out of sc" · "Dear parents, Great student achiev"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
2503.09087_pg6.pdf
WARNING [content] Math-heavy document: 20 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "e24 such that Cab = α. 3.2. Counting dire" · ", o) = Inw(G, o)·Q v∈V (G)(deg+ o (v)−1)!"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "and we have to translate P to find a “non-saturated” vertex. T" · "v1 v2 v3 v4 (b) v1 v2 v3 v4 Figure 1. An example for K4. fro" · "4 (b) v1 v2 v3 v4 Figure 1. An example for K4. from v1. Initi"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
0805b42e6e18b9c2f2f3d31d8ecf73e1a422f182_page_2_processed.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Download and Read Holt Chemistry Study Guide Stoichiometry" · "Download and Read Holt Chemistry Study"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
29cf8f78dcf746f0d24796dea279a36234362a28_page_30.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures (page 1)
↳ "YubiKey 5 Series Technical Manual"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
4b6eb432bc6c5d780bd95418be74a5f98ad00feb_page_3.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "16, KAUNAS, LITHUANIA JVE Journal of Vibroengineering Aims" · "S, LITHUANIA JVE Journal of Vibroengineering Aims and Sco" · "KAUNAS, LITHUANIA JVE Journal of Vibroengineering Aims an"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
6a3e03935289ebe49abbf0809717974f80c0f7ca_page_9.pdf
WARNING [ocr] OCR substitution artifacts detected: l_I_1, rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "vent-related fMRI study. NeuroImage 37, 1445–1456. Landy, M.S" · "they perceived using a 7 alternative forced choice task. We p" · "In the second block the stimuli were presented again once"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
6b9ddb2bcd72226a740cf1f09a6ace52ea4d0fa7_page_4.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Subcourses in FYSA01, Physics 1: General Physics Applie" · "Subcourses in FYSA01, Physics 1: General"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
7f757df3765dcaae45710b624bb76553ded10b50_page_4.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "13816 Boletín Oficial de Canarias núm. 154, lu" · "13816 Boletín Oficial de Canarias núm. 154, lunes 11"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
ba6aebad479086f039aa97daa8208472439af6ec_page_1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "so increasingly using the internet for medical issues. This st" · "re center. The results of the first 2 pages of a Google search" · "ted on 21 Mar 2023 — The copyright holder is the author/funde"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
bab14288ca3310add63c8df5c705d7a6fb6116a9_page_6.pdf
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "t/ CB_O underwent a more significant decrease of 23 mV (from 0" · "respectively (Fig. 2a,b and Table 1)." · "respectively (Fig. 2a,b and Table 1). With CO-strippin"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
d76a8019c0c56d1c9e67b53e4abeb2d97fc4bf35_page_8.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "618 Ata Özdemirci / Procedia Social and Beha" · "ta Özdemirci / Procedia Social and Behavioral Sciences 24 (20"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
11_pg161_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "machine while the wheel is turning. North. PLUERE. Weeping. {" · "PLU 633 POC PLOWEFERE. Companion in play. (J.-S.) PLOWKKY. C" · "U 633 POC PLOWEFERE. Companion in play. (J.-S.) PLOWKKY. Cove"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
11_pg252_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ ", dates, Almaund rys, pomme-garnates, Kanel and setewale. Gy o" · "the hare's head to the goose-giblet, i. e., tit for tat. (10)" · "SET 724 SEW (7) To protect ; to accompany. Yorksh"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
11_pg36_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "ly much in fashion. The man turned the woman round several tim" · "AV 508 LAW Who lukes to the lefte syde, wheune his horse laun" · "LAV 508 LAW Who lukes to the lefte syde, wheun"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
cfbe79724feedf08f6e55cea253a979a9ccadd89_page_3.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Kimia itu Mudah | Binar Terang d" · "Kimia itu Mudah | Binar Terang di Su"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [content] Document language detected as 'id' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
11_pg418_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "pressed by the hand of the spinner. Forhy. (2) The skinnv pa" · "TRO 890 TRO and the yarn is pressed by the han"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
11_pg57_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "29 LOR (4) The privilege of turning out cattle on com- mons. N" · "LOP 529 LOR (4) The privilege of turning out cattle o" · "LOP 529 LOR (4) The privilege of turning out cattle on com-"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
12_pg174_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "lling in BAPS software for learning genetic structures of popu" · "anc, D., 2014. Accuracy and efficiency of algorithms for the d" · "Cohan, F.M., 1994. The effects of rare but promiscuous"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
20_pg18_pg1.pdf
WARNING [content] Possible PII: 3 email value(s) across 1 page(s). (pages 1)
↳ e.g. i@inews.co.uk · i@inews.co.uk · inq***@ipso.co.uk
→ Review whether email values should be redacted before ingestion into your RAG corpus.
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "m/theipaper A question of governance A desire to be governed b" · "TLEOVER, DERBYSHIRE Can I briefly add two more Americanisms th" · "18 @ TWEETS AND EMAILS Your View i@inews.co.uk Please includ"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
20_pg39_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "2.00 Steve Wright In The Afternoon 4.15 Steve Wright In The A" · "the taut little exploitation flick that made him a star, Mel" · "Y 4 JANUARY 2019 === What We Did On Our Holiday 10.50pm BBC2"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
07f4d706dd1736e247163f0128702c65bdd8_page_1_pg1.pdf
WARNING [content] Math-heavy document: 5 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "resolution (100–200 μm) or stent- related" · "s the separation of ≥1 stent strut from"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces, mojibake (page 1)
↳ "ttp://circinterventions.ahajournals.org DOI: 10.1161/CIRCINTE" · "1 C oronary stent malapposition is the separation of ≥1 s" · "1 C oronary stent malapposition is the sep"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF40
87
00e980a0c4645fc83f27c467d10bbdeb8661_pg64.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "T_4.DOC 4-33 TABLE 4-18 Milestones for Dechlorination Up" · "4-33 TABLE 4-18 Milestones for Dechlorination Upgrades at"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
0091c5b23fe54b3a7cbe5b0b8c26057b7d70_pg3_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ "NS For analysis of usage patterns, the 204 children were divid" · "Immunisation state and its documenta" · "Immunisation state and its documentation in"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
26f221a3e03a0084c486be718476250c9247_pg2_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Exhibitor List Booth # Company/Org" · "guardianfineart.com 104 Mutual of America Paul Wierzba mutual"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
371cfed48826fdffb18d9e1d5a9010ab0ccf_pg10_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "and Conditions (http://www.journals.uchicago.edu/t-and-c)." · "e of law mechanism (EU); Qualified majority voting (EU) Meso M" · "es Level Moral Reach Contestation Fundamental Type 1 Human ri"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
87
2d54e996d1247bd36695d872863eefebd8f8_pg4.pdf
WARNING [encoding] Unexpected control characters found — possible binary corruption. (page 1)
→ Strip or replace control characters before ingestion.
WARNING [ocr] OCR substitution artifacts detected: rn_m, fi_fl_ligature, broken_ligatures, mid_word_spaces (page 1)
↳ "ssive behavior (Bjorklund & Harnishfeger, 1995) or other behav" · "The dynamics of delay of gratification. In R.F. Baumeister & K" · "knowing the social rules, but are a"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
508eb2725fb94e86b737dd667ff4301ee252_pg13_pg1.pdf
WARNING [content] Math-heavy document: 9 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "um (two groups) 2 × 106 cells (n/a) n/"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "3 505 Table 5. Characteristics of studies included—prec" · "Table 5. Characteristics of studies included—preclinica"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
87
8ec7ac2f208b3942f460d76e3cb2149616b5_pg1_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "South London Freight Quality Partnership Steer" · "Mick Heduan Transport for London Edward Noble"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
88
0722235b45282fd9c4b9d92efc9dbf6e9704d798_page_2_processed.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Y OF SOCIAL SCIENCES of the University of the West Indies is" · "FACULTY OF SOCIAL SCIENCES of the University of the West In"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
88
0f727ae22d10919282c22063a6cade3e1fe0b3d2_page_1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures (page 1)
↳ "Physician Directory health.ucsd.edu"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
88
b4c3c4ac3d6f7b52a993cec7ca8b3ad43cecabad_page_3.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures (page 1)
↳ "John Bibb Tate Memoir, 1921-1983 Manu"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
88
4_pg380.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "RIABLES AND LIMITS 400. Functions (§ 292) are usually denote" · "TS 400. Functions (§ 292) are usually denoted by symbols of"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
88
7772a40a9b783abb88fe1a47252b09c40e0e_pg6.pdf
WARNING [ocr] OCR substitution artifacts detected: rn_m, broken_ligatures, mid_word_spaces (page 1)
↳ ") No No No No No Yes No California (2003) Yes No No No No No N" · "e 1 State templates and inclusion of bicycle information Stat" · "Table 1 State templates and inclusion of bic"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
88
de5a70a52f98b51f94ad02dda266d1a1e2b6_pg2.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "ALL CHAMPIONSHIP EURO2006 Preliminaries - Group D Match No:" · "23 R1 Foul Team A 32:36 R1 Not to recognisable Team A 33:33 R"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
88
e82a04c638821610957e5dae68f92e48b992_pg5.pdf
WARNING [content] Math-heavy document: 11 formula hits across 1 page(s). Equations often extract as garbled text. (pages 1)
↳ "1,627 (2564) 207.55 ±25 Gamma Hosp Includ"
→ Consider converting equations to plain-language descriptions, or use a math-aware parser (e.g. MathPix, GROBID) before ingestion.
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "per year Parameters Mean SE Units Distribution Source Surviva" · "Table 1 Base case input per year Parameters"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
88
ccc11e02dbe24bb40b9884a783927065db41_pg3_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures, mid_word_spaces (page 1)
↳ "Birmingham Public Schools, MI Ma" · "c Schools, MI Math Performance by Subgroup, Grades 3-8, 2019-"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30
89
39e0e582e2aeb751d615308c4486e5d9422a37db_page_2.pdf
INFO [content] Document language detected as 'pt' (expected 'en').
→ Ensure your embedding model supports this language. Mixed-language corpora may need separate embedding spaces.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF20
89
a6820dddf8e2940cd9bd8d9590253728e3a7_pg15_pg1.pdf
WARNING [ocr] OCR substitution artifacts detected: broken_ligatures (page 1)
↳ "UAL 2015 6.3 Item Quantity Part Number Description 1"
→ Re-scan the source document with higher DPI (≥300 DPI recommended) or post-process with an OCR correction tool.
WARNING [structure] 1 table(s) detected. Tables often chunk poorly as plain text.
→ Use a table-aware extractor (e.g. pdfplumber, Camelot, AWS Textract) and convert tables to Markdown or CSV before chunking.
INFO [metadata] PDF metadata is sparse or missing (title, author, date).
→ Add metadata to improve retrieval context. Use a PDF editor or exiftool to update document properties.
PDF30

Taxonomy Coverage — Garani 2026

Source: Garani, A. (2026). A Systematic Taxonomy of Failure Modes in Retrieval-Augmented Generation Systems. TrustNLP 2026. doi:10.18653/v1/2026.trustnlp-main.27

2direct
F3, F7
2proxy
F1, F11
2risk signal
F4, F23
21runtime required
6unsupported

2 direct + 4 proxy/risk = 6 of 33 modes assessed statically. A tool claiming 33/33 detection is lying.

What this audit cannot determine

Static pre-ingestion scanning cannot assess the following failure modes (Garani 2026 taxonomy) — they require live system traces, LLM outputs, or agent logs:

For runtime coverage: DeepEval · Ragas · RAGChecker · Phoenix / OpenInference