word-extract
Copyright 2026 Ivan Ortega

This product includes software developed by Ivan Ortega: the `wordextract` package and the
`docextract-core` package in this repository. Portions were written with the assistance of AI
coding tools under the author's direction and review.

Licensed under the Apache License, Version 2.0 (see LICENSE).

Third-party material
--------------------

Porter stemming algorithm
  wordextract/stem.py is a transcription, written for this project, of the algorithm in
  M. F. Porter, "An algorithm for suffix stripping", Program 14(3):130-137, 1980. The
  algorithm is the author's published work; this repository contains no copy of any
  reference implementation of it.

Stemmer test vectors
  fixtures/stem/porter_reference_vectors.tsv lists words and the stems an independent
  implementation produced for them: NLTK's PorterStemmer(mode="ORIGINAL_ALGORITHM")
  (NLTK is licensed under the Apache License, Version 2.0; https://www.nltk.org/). Only the
  word/stem pairs are included. NLTK is used by a development-time tool
  (tools/make_stem_vectors.py) and is neither bundled nor a dependency.

Runtime dependencies (installed separately, not bundled)
  lxml (BSD-3-Clause), https://lxml.de/

Development dependencies (installed separately, not bundled)
  pytest (MIT), python-docx (MIT)

The test documents under fixtures/ are synthetic, generated by the scripts in tools/. They
contain no real customer, third-party or employer data. Real documents and term lists, when
used, live outside the repository (fixtures/real/ is git-ignored).
