Metadata-Version: 2.4
Name: exhash
Version: 0.4.18
Classifier: Programming Language :: Rust
Classifier: Programming Language :: Python :: Implementation :: CPython
Requires-Dist: fastcore>=2.2.29
Requires-Dist: httpx>=0.28.1
Requires-Dist: fastship>=0.0.11 ; extra == 'dev'
Requires-Dist: maturin>=1.0,<2.0 ; extra == 'dev'
Requires-Dist: pytest ; extra == 'dev'
Provides-Extra: dev
License-File: LICENSE
Summary: Verified line-addressed file editor using lnhash addresses
Home-Page: https://github.com/AnswerDotAI/exhash
Author-email: Jeremy Howard <j@fast.ai>
License-Expression: Apache-2.0
Requires-Python: >=3.10
Description-Content-Type: text/markdown; charset=UTF-8; variant=GFM
Project-URL: Homepage, https://github.com/AnswerDotAI/exhash
Project-URL: Issues, https://github.com/AnswerDotAI/exhash/issues
Project-URL: Repository, https://github.com/AnswerDotAI/exhash

# exhash: Verified Line-Addressed File Editor

exhash combines Can Bölük's very clever [line number + hash editing system](https://blog.can.ac/2026/02/12/the-harness-problem/) with the powerful and expressive syntax of the classic [ex editor](https://en.wikipedia.org/wiki/Ex_(text_editor)).

Install via pip to get a convenient Python API, an IPython cell magic, and the `exhash`/`lnhashview` CLI commands:

```bash
pip install exhash
```

## lnhash format

We refer to an *lnhash* as a tag of the form `lineno|hash|`, where `hash` is the low 12 bits of CRC-32 (IEEE) over the line's UTF-8 content, encoded as two Base64url characters (`A–Z`, `a–z`, `0–9`, `-`, `_`), high six bits first.

Address forms:

- `lineno|hash|`: hash-verified address
- `$`: last line (no hash)
- `%`: whole file (`1,$`, no hashes)
- `lineno`: line number with no hash. Only `p` accepts it, including as the subcommand of `g`, `g!`, or `v`.

## CLI

The `exhash` and `lnhashview` commands are Python console scripts over the native Rust extension, installed into your PATH by pip.

### View

```bash
# Shows every line prefixed with its lnhash
lnhashview path/to/file.txt
# Optional line number range to show
lnhashview path/to/file.txt 10 20
```

If `end` is past EOF, `lnhashview` returns through the last available line instead of failing.

### Edit

```bash
# Substitute on one line
exhash file.txt '12|vN|s/foo/bar/g'

# Transliterate characters on one line
exhash file.txt '12|vN|y/abc/ABC/'

# Change one line with inline text (spaces after c are literal text)
exhash file.txt '12|vN|c    replacement line'

# Append multiline text from stdin (read through EOF; every line is literal)
exhash file.txt '12|vN|a' <<'EOF'
new line 1
new line 2
EOF

# Dry-run
exhash --dry-run file.txt '12|vN|d'

# Set shift width for < and >
exhash --sw 2 file.txt '12|vN|>1'

# Last line and whole file shorthands (no hash)
exhash file.txt '$d'
exhash file.txt '%j'

# Print lines 12 to 20: p accepts line numbers with no hashes
exhash file.txt '12,20p'

# Move a line to EOF using $ as the destination
exhash file.txt '12|vN|m$'

# Create a missing file by treating it as empty input
exhash new.txt '0|AA|a' <<'EOF'
first line
EOF
```

Substitute and global commands use [`fancy-regex`](https://docs.rs/fancy-regex/latest/fancy_regex/) syntax, including lookbehind (`(?<=...)`, `(?<!...)`), lookahead (`(?=...)`, `(?!...)`), and pattern backreferences (`\1`):

- Replacement syntax uses `$1`, `$0`, `${name}`, and `$$` for a literal dollar sign; `\1` in a replacement stays literal. See [`fancy_regex::Regex::try_replacen`](https://docs.rs/fancy-regex/latest/fancy_regex/struct.Regex.html#method.try_replacen).
- `\/` escapes the command delimiter in pattern/replacement
- Custom delimiters: `s`, `y`, `g`, `g!`, and `v` all accept any non-alphanumeric char as delimiter instead of `/`, e.g. `s@pat@rep@`, `g@pat@cmd`. Each command in a combo picks its own delimiter independently: `g@a/b@s/old/new/`
- For example, `s///` accepts newlines in pattern/replacement; replacement newlines split one line into multiple lines.
- Transliteration uses `y/src/dst/` and requires source/destination to have equal character counts
- A substitute whose pattern matches nothing in its addressed range fails (nothing is written), so a typo cannot silently no-op; substitutes running inside `g`/`g!`/`v` payloads stay lenient, since not every selected line need match
- Matching errors, including exceeding the engine's default backtracking limit, abort the command set before any files are written

When passing multiple commands, each command's lnhashes are verified immediately before it runs. A single-line address may match either the line's current hash or its call-start hash, so commands can stack on one line. Range addresses remain strict, and structural changes invalidate call-start records at and below their topmost affected line.

For CLI multiline `a/i/c` commands, omit inline text and provide the text block on stdin:

```bash
printf "new line 1\nnew line 2\n" | exhash file.txt "2|7v|a"
```

If the file does not exist and the command set is valid on empty input, exhash treats it as an empty file and writes the result. For example, `0|AA|a` can create a new file.

### Stdin filter mode

```bash
cat file.txt | exhash --stdin - '1|vN|s/foo/bar/'
```

In `--stdin` mode, multiline `a/i/c` text blocks are not available.

### Notebook cells

`lnhashview-cell` and `exhash-cell` apply the same workflow to notebook cells, addressed by exact or unique ID prefixes. Pass comma-separated IDs to view several cells together; the output adds a `# cell <id>` header to each group:

```bash
lnhashview-cell nbs/00_core.ipynb ab12cd34
lnhashview-cell nbs/00_core.ipynb ab12cd34,ef56ab78
exhash-cell nbs/00_core.ipynb ab12cd34 '3|7v|s/old/new/'
exhash-cell --dry-run nbs/00_core.ipynb ab12cd34 '3|7v|d'
printf 'replacement line\n' | exhash-cell nbs/00_core.ipynb ab12cd34 '3|7v|c'
```

### Document outlines

`exhash-open` opens Markdown, source code, notebooks, URLs, or stdin as a verified section tree. Its default output is the immediate outline; copy a displayed token back to read that section:

```bash
exhash-open README.md
exhash-open README.md '1.2.|21|Z1|,101|Js|'
exhash-open README.md --paths --depth 2
exhash-open README.md --search 'CLI|console'
exhash-open README.md --lnhashs
exhash-open https://example.com/llms.txt --links
```

## Python API

```py
from exhash import exhash, file_exhash, lnhash, lnhashview, lnhashview_file, line_hash
```

### Viewing

```py
text = "foo\nbar\n"
view = lnhashview(text)                        # ["1|Gy|foo", "2|PU|bar"]
view = lnhashview_file("f.py", start=1, end=260) # end past EOF is clamped
```

`lnhashview`/`lnhashview_file` return a `list` subclass whose repr shows the rows verbatim, one per line, so a bare call in IPython displays a ready-to-copy view.

### Editing

`exhash(text, cmds, sw=4)` takes the text and a required iterable of tuple command specs (use `[]` for no-op). Raw command strings are rejected by the Python API. `sw` controls how far `<` and `>` shift.

A command is usually `(addr, op)` or `(addr, op, payload)`. `addr` is an lnhash address string from `lnhash(...)`/`lnhashview(...)`; put ranges in that same string, e.g. `f"{a1},{a2}"`. Substitute uses `(addr, "s", pattern, replacement[, flags])`, so patterns and replacements can contain `/` without delimiter escaping.

Text fields can contain newlines. That covers multiline `a`/`i`/`c` payloads and substitute pattern/replacement. In `a`/`i`/`c` payloads a trailing newline ends the last line, as in CLI stdin blocks: `"x"` and `"x\n"` both insert one line, `"x\n\n"` inserts `x` and then a blank line, and an empty payload is no lines, so `c` with `""` deletes the addressed lines. Commands such as `d`, `m`, and `sort` do not take text.

```py
addr = lnhash(1, "foo")  # "1|Gy|"
res = exhash(text, [(addr, "s", "foo", "baz")])
print(res["lines"])    # ["baz", "bar"]
print(res["modified"]) # [1]

# Multiple commands
a1, a2 = lnhash(1, "foo"), lnhash(2, "bar")
res = exhash(text, [(a1, "s", "foo", "FOO"), (a2, "s", "bar", "BAR")])

# Hashes are checked just-in-time per command.
# If earlier commands change/shift a later target line, recompute lnhash first.

# Change one line; leading spaces are part of the replacement
res = exhash(text, [(addr, "c", "    replacement line")])

# Append multiline text in one tuple payload (no dot terminator)
res = exhash(text, [(addr, "a", "new line 1\nnew line 2")])

# Wrong for the Python API: the trailing "." would be inserted literally
# res = exhash(text, [(addr, "a", "new line 1\nnew line 2\n.")])

# Also wrong: do not split the inserted text into separate cmds entries
# res = exhash(text, [(addr, "a"), "new line 1", "new line 2"])

# Change shift width for < and >
res = exhash(text, [(addr, ">", "1")], sw=2)

# Literal / needs no delimiter escaping in tuple substitute fields
res = exhash("a/b\n", [(lnhash(1, "a/b"), "s", "a/b", "c/d")])

# Literal newlines in replacement split one line into multiple lines
res = exhash("foo\n", [(lnhash(1, "foo"), "s", "foo", "bar\nbaz")])
print(res["lines"])  # ["bar", "baz"]

# Literal newlines in pattern can match across lines
a1, a2 = lnhash(1, "foo"), lnhash(2, "bar")
res = exhash("foo\nbar\n", [(f"{a1},{a2}", "s", "foo\nbar", "replaced")])

# Global commands take a pattern plus a nested subcommand tuple (no address)
res = exhash("keep\nTODO x\n", [("%", "g", "TODO", ("d",))])

# Transliterate takes source/dest fields (equal character counts)
res = exhash("abc\n", [(lnhash(1, "abc"), "y", "abc", "ABC")])
```

### File helpers

`lnhashview_file` reads directly from one file path. All file paths, including file-qualified addresses, expand a leading `~` to your home directory. `file_exhash(path, *cmds, sw=4, inplace=True)` uses `path` as the default file context for unqualified addresses. Pass each command as its own tuple argument. Put file-qualified source and `m`/`t` destination addresses in the address/destination tuple fields:

```py
view = lnhashview_file("file.py")

# By default, writes changed files after every command succeeds
# and returns the combined diff string.
diff = file_exhash("file.py", (addr, "s", "foo", "bar"))

# With inplace=False, files are unchanged and a FileSetEditResult is returned.
res = file_exhash("file.py", (addr, "s", "foo", "bar"), inplace=False)
print(res.changed)          # ["file.py"]
print(res["file.py"].lines)
print(res.format_diff())    # includes --- file.py / +++ file.py headers

# Missing files are treated as empty only when the command is valid on empty input.
diff = file_exhash("new.py", ("0|AA|", "a", "print('hi')"))

# File-qualified addresses can edit or transfer lines across files.
diff = file_exhash("src/a.py",
    ("src/a.py:24|8S|,38|De|", "m", "src/b.py:$"),
    (r"src/a.py:5|Gq|", "s", r"from \.b import old", r"from \.b import helper"))
```

A file prefix is separated from the address with `:`. Escape literal colons in filenames as `\:` and literal backslashes as `\\`.

`file_exhash(..., inplace=False)` returns a `FileSetEditResult`:

- `res.files`: dict of path to `FileEditResult`
- `res.changed`: changed paths, in first-touch order
- `res.printed`: paths with lines addressed by `p` (`res[path].printed` gives the line numbers)
- `res.default_path`: the default path passed to `file_exhash`
- `res[path]`: shorthand for `res.files[path]`
- `res.format_diff(context=1)`: combined diff of the changed targets, with `--- path` / `+++ path` headers. It never includes printed lines.
- `str(res)`: each target's diff, followed by its printed lines (see [EditResult](#editresult))

### Notebook cells

`lnhashview_cell(path, cell_id, ...)` returns a normal lnhash view for one cell. `lnhashview_cells(path, *cell_ids, ...)` returns the requested cells in order, using `# cell <id>` headers before each cell's normal `lineno|hash|content` lines. `cell_exhash(path, cell_id, *cmds, sw=4, inplace=True)` edits one cell; pass each command as its own tuple argument. Like `file_exhash` it writes and returns a diff by default, and `inplace=False` previews the `EditResult` without touching the file.

### The `%%exhash` cell magic

Importing `exhash.skill` under IPython or Jupyter registers the `%%exhash` cell magic - the standard way to apply `a`/`i`/`c` payload commands interactively. The magic line is `%%exhash <path> [<cell_id>] <address> <a|i|c>`; the payload is everything below it, taken verbatim. Nothing in the payload is parsed as Python, so there is no quoting or escaping at all:

```
%%exhash notes.txt 2|7v|a
new line 1
new line 2
```

- `%%exhash new.py 0|AA| a` creates a missing file.
- `%%exhash f.py % c` replaces the whole file (`%` needs no hashes). With a cell id, `%%exhash nb.ipynb ab12 % c` replaces that notebook cell's source.
- `%%exhash f.py 12|Py|,15|HD| c` replaces just that range, both addresses from one `lnhashview_file` view.
- A trailing newline ends the last line, as in every `a`/`i`/`c` payload; to end the payload with a blank line, leave an empty line at the bottom of the cell.
- Each magic cell applies one command and returns the diff.

Tuple `a`/`i`/`c` payloads (as in the examples above) remain for scripts and tests, where magics don't exist. Interactively, prefer the magic: a Python string layer invites quoting mistakes.

### Pyskill

The package registers `exhash.skill` as a pyskill exposing the primary Python APIs with LLM-oriented workflow docs. Use `doc(exhash.skill)` after importing it through a pyskills host.

### EditResult

`exhash()` returns an `EditResult` with attributes (also accessible via `res["key"]`):

- `lines`: list of output lines
- `hashes`: lnhash for each output line
- `modified`: 1-based line numbers of modified/added lines
- `deleted`: 1-based line numbers of removed lines (in original)
- `origins`: for each output line, the 1-based original line number (None if inserted)
- `printed`: 1-based line numbers explicitly addressed by `p`

Lines addressed by `p` are not part of the diff. `res.format_printed()` returns them as a bare `lnhashview` with no headers. It never caps or truncates its rows. `str(res)` and the repr show the diff, then a `# printed` header, then the printed lines. A result that changed nothing shows the printed lines alone, with no header. `file_exhash` and `cell_exhash` return the same output. A `p`-only call writes nothing and returns the bare view. A call that changes and prints nothing returns `none: No changes.`, as `fastcore`'s editors do. When more than one target is reported, a target with only printed lines is headed by `# file <path>` or `# cell <id>`.

`res.format_diff(context=1)` returns a unified-diff-style summary showing only changed lines with context:

```py
res = exhash(text, [(addr, "s", "foo", "baz")])
print(res.format_diff())
# --- original
# +++ modified
# -1|Gy|foo
# +1|PU|baz
#  2|X2|bar
```

`format_diff(maxlen=n)` caps each diff row at `n` chars plus a closing `…`. Where a run of changed rows holds as many `-` rows as `+` rows, the nth `-` row pairs with the nth `+` row. A capped row of a pair starts 20 chars before the pair's first difference, with `…` after its address. Every other capped row keeps its start. The result reprs and the diffs that `file_exhash` and `cell_exhash` return use `maxlen=180`. `truncate_diff` then keeps their first 15 lines.

All diff strings returned by `format_diff`, `file_exhash`, and `cell_exhash` are fastcore `PrettyString`s, and the result objects' reprs show the diff too - so in IPython, ending a cell with the bare call displays the diff verbatim, no `print` needed.

## Document outlines

`open_doc` opens a file (`fname=`, or a `Path` as `src`; recorded for `refresh()` and edits), a URL (an `https?://` str, fetched), or any other str as text, and returns a `Section` tree: Markdown sections from headings, code sections (py, js, ts, tsx, rs, zig, swift) from tree-sitter definitions, and `.ipynb` sections from md-heading cells over cells. The bare repr is a fixed-width outline, one row per section:

```
1.6.|56|lq|,78|74| Release [725] Publishing is handled by GitHub Actions in `.github/workflows/ci.yml`…
```

The leading token is a verified address: the dotted addr (trailing dot; the root's is `.`) fused with the section's `start,end` lnhash boundary pair. `at(token)` navigates with the first hash verified, so a stale copy fails loudly; the boundary pair drops straight into a `file_exhash` range command, so a listing is also an edit address book. `find(title)`, `search(pat)`, `paths(depth)`, and numeric indexing (`d[1][6]`) traverse the live tree; `links(pat)` lists inline links numbered document-wide, and `open(n)` opens link `n` as a new tree (fetched or read relative to `base`). Previews join lines with `¶` and render links as `[text][n]`, so no URL is ever displayed. A markdown row's preview starts under its heading; a code row has no title, and its preview opens with the def line itself, signature included. `view()` returns a section's rendered text the same way (`.src` is the raw source); `view(*tokens)` returns the live sections at those verified addresses, displayed under `# token` headers when more than one; `nums=`/`lnhashs=` switch any view to stored lines with edit-ready addresses. Notebook section tokens carry the heading cell id, and `view(lnhashs=True)` emits `cellid:lineno|hash|` rows ready for `cell_exhash`.

`search(pat)` returns one hit per matching source line, not one row per section. Each 180-character row shows the containing section's verified token, the matching line's hash address, and a preview starting at that line and continuing across newlines as `¶`, up to the section's end. Section tokens repeat when several lines in the same section match: copy the section token to `view()` to read the whole section, or use the line address to edit the hit. Notebook line addresses include the cell ID. Each hit exposes `.section`, `.address`, and `.preview`; slicing the results preserves their display format.

## Tests

```bash
pytest -q
```

