Metadata-Version: 2.5
Name: allread
Version: 0.1.0
Summary: Read any file as clean text in one line — local, private, encoding-safe.
Project-URL: Homepage, https://github.com/sepehrtaji-dev/allread
Project-URL: Issues, https://github.com/sepehrtaji-dev/allread/issues
Author: sepehrtaji-dev
License: MIT License
        
        Copyright (c) 2026 sepehrtaji-dev
        
        Permission is hereby granted, free of charge, to any person obtaining a copy
        of this software and associated documentation files (the "Software"), to deal
        in the Software without restriction, including without limitation the rights
        to use, copy, modify, merge, publish, distribute, sublicense, and/or sell
        copies of the Software, and to permit persons to whom the Software is
        furnished to do so, subject to the following conditions:
        
        The above copyright notice and this permission notice shall be included in all
        copies or substantial portions of the Software.
        
        THE SOFTWARE IS PROVIDED "AS IS", WITHOUT WARRANTY OF ANY KIND, EXPRESS OR
        IMPLIED, INCLUDING BUT NOT LIMITED TO THE WARRANTIES OF MERCHANTABILITY,
        FITNESS FOR A PARTICULAR PURPOSE AND NONINFRINGEMENT. IN NO EVENT SHALL THE
        AUTHORS OR COPYRIGHT HOLDERS BE LIABLE FOR ANY CLAIM, DAMAGES OR OTHER
        LIABILITY, WHETHER IN AN ACTION OF CONTRACT, TORT OR OTHERWISE, ARISING FROM,
        OUT OF OR IN CONNECTION WITH THE SOFTWARE OR THE USE OR OTHER DEALINGS IN THE
        SOFTWARE.
License-File: LICENSE
Keywords: csv,docx,encoding,files,markdown,pdf,read,text-extraction
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Topic :: Text Processing
Requires-Python: >=3.9
Provides-Extra: all
Requires-Dist: openpyxl>=3.1; extra == 'all'
Requires-Dist: pymupdf>=1.24; extra == 'all'
Requires-Dist: python-docx>=1.1; extra == 'all'
Provides-Extra: docx
Requires-Dist: python-docx>=1.1; extra == 'docx'
Provides-Extra: pdf
Requires-Dist: pymupdf>=1.24; extra == 'pdf'
Provides-Extra: xlsx
Requires-Dist: openpyxl>=3.1; extra == 'xlsx'
Description-Content-Type: text/markdown

# AllRead

**Read any file as clean text in one line — local, private, encoding-safe.**

```python
import allread

text = allread.read("report.csv")   # markdown table
text = allread.read("notes.txt")    # plain text
text = allread.read("page.html")    # extracted text + tables
```

No config. No encoding crashes. Works the same on Windows, macOS and Linux.

## Install

```bash
pip install allread            # core: txt, md, rst, log, csv, tsv, json, jsonl, xml, html
pip install "allread[pdf]"     # + PDF support (PyMuPDF)
pip install "allread[docx]"    # + Word documents
pip install "allread[xlsx]"    # + Excel workbooks
pip install "allread[all]"     # everything
```

Optional formats are lazy: `allread.read("report.pdf")` raises a clear
`ImportError` telling you exactly which extra to install.

## Why AllRead?

- **Encoding-safe by default** — UTF-8, UTF-16/32, UTF-8-BOM, cp1252 and
  legacy encodings are handled automatically; Persian/Arabic/any non-ASCII
  content just works.
- **Zero required dependencies** — the core is pure standard library.
  Heavy engines are optional extras, loaded lazily.
- **Structured files become Markdown** — CSV/TSV/Excel come back as
  Markdown tables, ready for LLM prompts, notebooks or docs.
- **Same API everywhere** — `read()` returns text, `load()` returns a
  `Document` with source and format metadata.

## API

| Function | Returns | Purpose |
|---|---|---|
| `allread.read(path)` | `str` | extracted text/markdown |
| `allread.load(path)` | `Document` | text + `source` + `format` |
| `allread.supported_formats()` | `dict` | formats and availability |

## Examples

```python
doc = allread.load("sales.xlsx")
print(doc.format)        # "xlsx"
print(doc.text)          # markdown tables, one section per sheet
```

## Roadmap

- [x] Core formats (txt, md, rst, log, csv, tsv, json, jsonl, xml, html)
- [x] Optional PDF / DOCX / XLSX engines with lazy imports
- [ ] OCR for scanned PDFs and images (`allread[ocr]`)
- [ ] Audio transcription (`allread[audio]`)
- [ ] CLI: `allread file.pdf > out.md`

## License

[MIT](LICENSE)
