Metadata-Version: 2.4
Name: soup5ever
Version: 0.1.0
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: Implementation :: CPython
Classifier: Programming Language :: Rust
Classifier: Topic :: Text Processing :: Markup :: HTML
Requires-Dist: beautifulsoup4>=4.13
Requires-Dist: maturin>=1.9,<2 ; extra == 'dev'
Requires-Dist: edwh ; extra == 'dev'
Requires-Dist: pytest>=8 ; extra == 'dev'
Requires-Dist: html5lib>=1.1 ; extra == 'dev'
Requires-Dist: hypothesis>=6 ; extra == 'dev'
Requires-Dist: tabulate ; extra == 'dev'
Provides-Extra: dev
License-File: LICENSE
Summary: A BeautifulSoup tree builder backed by Rust's html5ever HTML5 parser
Keywords: html,html5,parser,beautifulsoup,bs4,html5ever
License-Expression: MIT
Requires-Python: >=3.10
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Source, https://github.com/robinvandernoord/soup5ever

# soup5ever

A [BeautifulSoup 4](https://www.crummy.com/software/BeautifulSoup/) tree builder backed by
Rust's [html5ever](https://github.com/servo/html5ever). Same HTML5 trees as
`BeautifulSoup(html, "html5lib")`, 6–14x faster.

```python
from bs4 import BeautifulSoup
import soup5ever

soup = BeautifulSoup(html, "html5ever")
```

## Installation

```
pip install soup5ever
```

Wheels for Linux (x86_64, aarch64; glibc and musl) and macOS (arm64), one `abi3` wheel per
platform for CPython 3.10+. The only dependency is `beautifulsoup4>=4.13`; no Rust, libxml2
or html5lib needed.

## Usage

Importing `soup5ever` registers the `"html5ever"` parser (also `"soup5ever"`). The generic
features `"html5"` and `"html"` are registered at the lowest priority, so importing it doesn't
change what `BeautifulSoup(markup, "html5")` or `BeautifulSoup(markup)` select.

The usual options work (`from_encoding`, `element_classes`, `multi_valued_attributes`,
`store_line_numbers`, `attribute_dict_class`). Like html5lib, `parse_only` isn't supported.

## Compatibility

The target is `BeautifulSoup(html, "html5lib")`; where html5lib and the HTML standard
disagree, soup5ever follows the standard.

- BS4's own tests for its html5lib builder: 93/93 pass.
- html5lib-tests tree construction: 1,588/1,592 match the spec. Against html5lib the trees
  are identical in 1,417 cases; the other 175 are classified automatically (167 html5lib
  behind the spec, 4 BS4 html5lib-adapter bugs, 4 `<selectedcontent>` misses both share).
- 182 targeted differential cases and ~23k fuzz inputs (triaged with Chromium): every
  remaining difference is an html5lib bug or listed below.

Differences from html5lib:

- Current HTML standard where html5lib 1.1 (2020) is behind: the 2025 `<select>` rules,
  `<template>` in `<head>`, `</p>` in SVG/MathML, `<main>`/`<summary>`, `<search>`,
  `<dialog>`, ruby. See `tests/corpus.py` for each case.
- BS4's html5lib adapter never applies the Noah's Ark clause (`AttrList` has no `__eq__`);
  soup5ever does.
- Undeclared, valid UTF-8 bytes are decoded as UTF-8 (html5lib guesses windows-1252 without
  `chardet`). Decoding follows the WHATWG Encoding Standard, not Python's codecs.
- Lone surrogates in a `str` become U+FFFD.
- `sourceline`/`sourcepos` match html5lib for ~98.5% of elements.

Shared with html5lib: no `parse_only`, no `Script`/`Stylesheet` string subclasses, no
`<selectedcontent>` cloning.

## Benchmarks

`BeautifulSoup(markup, parser)` on an in-memory `str`, median of 7 rounds, 4-vCPU Linux VM,
CPython 3.11:

| document             | html5lib | html5ever | speedup |
| -------------------- | -------- | --------- | ------- |
| small page (3 KiB)   | 2.4 ms   | 250 µs    | 9.6x    |
| large page (1.2 MiB) | 888 ms   | 101 ms    | 8.8x    |
| malformed tag soup   | 483 ms   | 34 ms     | 14.2x   |
| deeply nested        | 243 ms   | 26 ms     | 9.3x    |
| table-heavy          | 507 ms   | 69 ms     | 7.3x    |
| SVG/MathML-heavy     | 323 ms   | 51 ms     | 6.4x    |

Parsing in Rust is 14–32% of soup5ever's time; the rest is BeautifulSoup's own object
constructors. To reproduce:

```
uv pip install -e .[dev]
python benchmarks/run.py
```

## Development

```
uv venv venv && source venv/bin/activate
uv pip install -e .[dev]                # builds the Rust extension
python scripts/fetch_upstream_tests.py  # BS4's builder tests + html5lib-tests
pytest
```

Re-run `uv pip install -e .[dev]` after changing Rust code.

```
edwh fmt && edwh lint                   # Python (ruff)
cargo fmt && cargo clippy               # Rust (lint levels in Cargo.toml)
```

`pytest` includes a 3000-input differential fuzz run; `pytest --fuzz 20000 --fuzz-seed 7`
runs a bigger one.

## License

MIT

