Metadata-Version: 2.5
Name: pasteglint
Version: 0.1.0
Summary: Find invisible characters and explain copy-paste text differences.
License-Expression: MIT
License-File: LICENSE
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Requires-Python: >=3.10
Description-Content-Type: text/markdown

# pasteglint

Find the character that makes copied text behave differently from what you see.
Zero runtime dependencies. Python 3.10+. MIT licensed.

## Install

~~~console
python -m pip install pasteglint
~~~

For a local source checkout, run **python -m pip install .** in this directory.

## Command line

~~~console
pasteglint check snippet.py
pasteglint check snippet.py --json
pasteglint diff expected.txt actual.txt
pasteglint check -
~~~

An unexpected space might produce:

~~~text
4:8 U+00A0 NO-BREAK SPACE [space]; possible replacement: ' '
~~~

Input files and standard input are UTF-8. A byte-order mark is kept and reported.
Newline differences are preserved. A dash means standard input; only one input
can use it at a time. All commands also work as **python -m pasteglint**.

## Python

~~~python
from pasteglint import inspect_text, compare_text

(finding,) = inspect_text("x\u00a0= 1")
assert finding.codepoint == "U+00A0"
assert (finding.line, finding.column) == (1, 2)

changes = compare_text("API_KEY", "\uff21PI_KEY")
assert changes[0].left == "A"
assert changes[0].right == "\uff21"
~~~

**inspect_text(text)** returns immutable Finding objects with index, line,
column, codepoint, name, kind and an optional replacement. It reports selected
unusual spaces, smart quotes, dashes, fullwidth ASCII, Unicode format characters,
unusual controls, and Unicode line/paragraph separators.

**compare_text(left, right)** returns differing spans and their exact text.
Spans are zero-based and end-exclusive. Lines and columns are one-based Unicode
code-point positions, not byte offsets or rendered character widths.

Findings are observations: joiners, direction markers and curly punctuation can
be legitimate. Nothing is changed automatically. This is a debugging utility,
not a complete Unicode security scanner or programming-language parser. Exact
diff can be slow on very large or repetitive files; use it for text snippets.

## Exit codes

- 0: no configured findings, or identical texts.
- 1: findings or differences.
- 2: invalid input, unreadable files, or invalid UTF-8.

## Development

~~~console
python -m pip install -e .
python -m unittest discover -s tests -v
~~~
