Metadata-Version: 2.4
Name: webreader
Version: 2.3.8
Summary: Tree-based HTML reader for Python.
Author-email: morichan <morichan@gmail.com>
License-Expression: MIT
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: User Interfaces
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: pafer
Provides-Extra: dev
Requires-Dist: pytest; extra == "dev"
Requires-Dist: pytest-cov; extra == "dev"
Dynamic: license-file

![Python](https://img.shields.io/badge/python-3.8%2B-blue)
![Dependencies](https://img.shields.io/badge/dependencies-none-success)
![Tests](https://img.shields.io/badge/tests-pytest-informational)

## Features

- Tree-based parsing into an editable document tree
- CSS selectors: combinators, attributes, pseudo-classes, selector lists
- Extractors: metadata, links, tables, text, images, forms, headings,
  scripts, stylesheets
- Sanitization with a removal report
- Minify, pretty-print, Markdown conversion
- JSON / CSV / Markdown export
- CLI tool
- Standard library only at runtime

## Installation

```
pip install webreader
```

## Quickstart

```python
import webreader

html = '<h1 class="title">Hello</h1><a href="/x">link</a>'
doc = webreader.parse(html)

webreader.select(doc, 'h1.title')   # [{'tag': ..., 'attrs': ..., 'text': ..., 'html': ...}]
doc.find('a').get_attr('href')       # '/x'

webreader.minify(html)              # collapsed whitespace
webreader.pretty(html)              # re-indented, readable markup
webreader.to_markdown(html)         # convert to Markdown

clean, report = webreader.sanitize(html)
```

Bytes input is decoded using detected encoding: BOM sniffing first,
then `<meta charset>` in the first 2 KB, then the configured
`encoding` fallback.

In non-strict mode, malformed HTML does not raise; the warnings are
available on the document:

```python
doc = webreader.parse('<div><p>x</p>')   # unclosed <div>
doc.errors                                 # ['unclosed tag(s) at end of input: <div>']
```

## Extraction

```python
webreader.extract_metadata(html)    # title, OG/Twitter, canonical, charset, JSON-LD
webreader.extract_links(html)       # links and asset references
webreader.extract_tables(html)      # matrices with colspan/rowspan expanded
webreader.extract_text(html)        # readable text without boilerplate
webreader.extract_images(html)      # src, alt, dimensions, srcset candidates
webreader.extract_forms(html)       # form fields with types and defaults
webreader.extract_headings(html)    # h1-h6 outline + word count + reading time
webreader.extract_scripts(html)     # inline and external script inventory
webreader.extract_stylesheets(html) # link[rel~=stylesheet] and inline styles
```

## CSS selectors

Supported syntax: type selectors, `*`, `#id`, `.class`, attribute
selectors (`[attr]`, `[attr=value]` and the `^=`, `$=`, `*=`, `~=`,
`|=` operators), the pseudo-classes `:first-child`, `:last-child`,
`:nth-child(an+b)`, `:not(compound)` and `:contains(text)`, compounds
such as `div.content#main[href="x"]`, the combinators whitespace
(descendant), `>` (child), `+` (adjacent sibling) and `~` (general
sibling), and comma-separated selector lists.

```python
webreader.select(doc, 'a[href^="https://"]')
webreader.select(doc, 'ul > li:nth-child(2n+1):not(.skip)')
webreader.select(doc, 'h2 + p, blockquote p:contains("note")')
```

Invalid or unsupported selectors raise `SelectorSyntaxError`.

## Tree editing

Nodes support a mutation API, and the edited tree can be
re-serialized:

```python
doc = webreader.parse('<div><a href="/x">link</a></div>')
link = doc.find('a')
link.set_attr('rel', 'noopener').remove_attr('class')

new = webreader.Node('p')
new.children.append('added text')
doc.find('div').append_child(new)

link.unwrap()        # replace the <a> with its inner text
doc.to_html()        # '<div>link<p>added text</p></div>'
```

`Node` also provides `remove()` and `replace_with(*nodes)`.

## Markdown conversion

`to_markdown` renders headings, paragraphs, links, images,
emphasis/strong/strikethrough, inline and fenced code (with language
detection from `class="language-..."`), nested lists, blockquotes,
horizontal rules and tables.

```python
webreader.to_markdown('<h1>T</h1><p>Hi <b>bold</b></p>')
# '# T\n\nHi **bold**'
```

## Sanitization

`sanitize()` returns `(clean_html, report)`:

- tags outside the whitelist are unwrapped, inner content kept
- `script`, `style`, `noscript`, `template` are dropped entirely
- attributes filtered through a safe-list, `on*` handlers removed
- `javascript:`, `vbscript:` and non-image `data:` URLs blocked
- the report lists removed tags/attributes, blocked URLs and detected threats

## CLI

```
python -m webreader parse -f input.html
python -m webreader select -f input.html -s 'a[href]'
python -m webreader links -f input.html --base-url https://example.com
python -m webreader meta -f input.html
python -m webreader tables -f input.html --format csv
python -m webreader text -f input.html
python -m webreader images -f input.html
python -m webreader forms -f input.html
python -m webreader minify -f input.html --out min.html
python -m webreader pretty -f input.html --out pretty.html
python -m webreader md -f input.html --out page.md
python -m webreader sanitize -f input.html --out clean.html --allowed p,a,ul
```

Use `-f -` to read from stdin. Every output command accepts `--out`
to write to a file instead of stdout.

### Default settings

| Key | Default | Purpose |
|---|---|---|
| `strip_whitespace` | `True` | collapse whitespace in text nodes |
| `encoding` | `utf-8` | fallback when bytes input has no detectable encoding |
| `max_depth` | `100` | maximum element nesting |
| `preserve_comments` | `False` | keep comments in the tree |
| `strict_mode` | `False` | raise on malformed HTML |
| `remove_comments` | `True` | strip comments when minifying |
| `allowed_tags` | `p, br, b, i, a, ul, ol, li, div, span` | sanitize whitelist |
| `max_file_size_mb` | `10` | reject oversized input |
| `cache_size` | `50` | parse cache capacity |

## Exceptions

`HTMLPaferError` (base), `ConfigLoadError`, `SanitizationError`,
`SelectorSyntaxError`.
