Metadata-Version: 2.4
Name: agent-lexicon
Version: 0.8.2
Summary: A deterministic terminology layer for AI agents.
Project-URL: Homepage, https://github.com/SkeinRank/agent-lexicon
Project-URL: Repository, https://github.com/SkeinRank/agent-lexicon
Project-URL: Issues, https://github.com/SkeinRank/agent-lexicon/issues
Author: SkeinRank
License:                               Apache License
                                Version 2.0, January 2004
                             http://www.apache.org/licenses/
        
        TERMS AND CONDITIONS FOR USE, REPRODUCTION, AND DISTRIBUTION
        
        1. Definitions.
        
           "License" shall mean the terms and conditions for use, reproduction,
           and distribution as defined by Sections 1 through 9 of this document.
        
           "Licensor" shall mean the copyright owner or entity authorized by
           the copyright owner that is granting the License.
        
           "Legal Entity" shall mean the union of the acting entity and all
           other entities that control, are controlled by, or are under common
           control with that entity. For the purposes of this definition,
           "control" means (i) the power, direct or indirect, to cause the
           direction or management of such entity, whether by contract or
           otherwise, or (ii) ownership of fifty percent (50%) or more of the
           outstanding shares, or (iii) beneficial ownership of such entity.
        
           "You" (or "Your") shall mean an individual or Legal Entity
           exercising permissions granted by this License.
        
           "Source" form shall mean the preferred form for making modifications,
           including but not limited to software source code, documentation
           source, and configuration files.
        
           "Object" form shall mean any form resulting from mechanical
           transformation or translation of a Source form, including but
           not limited to compiled object code, generated documentation,
           and conversions to other media types.
        
           "Work" shall mean the work of authorship, whether in Source or
           Object form, made available under the License, as indicated by a
           copyright notice that is included in or attached to the work
           (an example is provided in the Appendix below).
        
           "Derivative Works" shall mean any work, whether in Source or Object
           form, that is based on (or derived from) the Work and for which the
           editorial revisions, annotations, elaborations, or other modifications
           represent, as a whole, an original work of authorship. For the purposes
           of this License, Derivative Works shall not include works that remain
           separable from, or merely link (or bind by name) to the interfaces of,
           the Work and Derivative Works thereof.
        
           "Contribution" shall mean any work of authorship, including
           the original version of the Work and any modifications or additions
           to that Work or Derivative Works thereof, that is intentionally
           submitted to Licensor for inclusion in the Work by the copyright owner
           or by an individual or Legal Entity authorized to submit on behalf of
           the copyright owner. For the purposes of this definition, "submitted"
           means any form of electronic, verbal, or written communication sent
           to the Licensor or its representatives, including but not limited to
           communication on electronic mailing lists, source code control systems,
           and issue tracking systems that are managed by, or on behalf of, the
           Licensor for the purpose of discussing and improving the Work, but
           excluding communication that is conspicuously marked or otherwise
           designated in writing by the copyright owner as "Not a Contribution."
        
           "Contributor" shall mean Licensor and any individual or Legal Entity
           on behalf of whom a Contribution has been received by Licensor and
           subsequently incorporated within the Work.
        
        2. Grant of Copyright License. Subject to the terms and conditions of
           this License, each Contributor hereby grants to You a perpetual,
           worldwide, non-exclusive, no-charge, royalty-free, irrevocable
           copyright license to reproduce, prepare Derivative Works of,
           publicly display, publicly perform, sublicense, and distribute the
           Work and such Derivative Works in Source or Object form.
        
        3. Grant of Patent License. Subject to the terms and conditions of
           this License, each Contributor hereby grants to You a perpetual,
           worldwide, non-exclusive, no-charge, royalty-free, irrevocable
           (except as stated in this section) patent license to make, have made,
           use, offer to sell, sell, import, and otherwise transfer the Work,
           where such license applies only to those patent claims licensable
           by such Contributor that are necessarily infringed by their
           Contribution(s) alone or by combination of their Contribution(s)
           with the Work to which such Contribution(s) was submitted. If You
           institute patent litigation against any entity (including a
           cross-claim or counterclaim in a lawsuit) alleging that the Work
           or a Contribution incorporated within the Work constitutes direct
           or contributory patent infringement, then any patent licenses
           granted to You under this License for that Work shall terminate
           as of the date such litigation is filed.
        
        4. Redistribution. You may reproduce and distribute copies of the
           Work or Derivative Works thereof in any medium, with or without
           modifications, and in Source or Object form, provided that You
           meet the following conditions:
        
           (a) You must give any other recipients of the Work or
               Derivative Works a copy of this License; and
        
           (b) You must cause any modified files to carry prominent notices
               stating that You changed the files; and
        
           (c) You must retain, in the Source form of any Derivative Works
               that You distribute, all copyright, patent, trademark, and
               attribution notices from the Source form of the Work,
               excluding those notices that do not pertain to any part of
               the Derivative Works; and
        
           (d) If the Work includes a "NOTICE" text file as part of its
               distribution, then any Derivative Works that You distribute must
               include a readable copy of the attribution notices contained
               within such NOTICE file, excluding those notices that do not
               pertain to any part of the Derivative Works, in at least one
               of the following places: within a NOTICE text file distributed
               as part of the Derivative Works; within the Source form or
               documentation, if provided along with the Derivative Works; or,
               within a display generated by the Derivative Works, if and
               wherever such third-party notices normally appear. The contents
               of the NOTICE file are for informational purposes only and
               do not modify the License. You may add Your own attribution
               notices within Derivative Works that You distribute, alongside
               or as an addendum to the NOTICE text from the Work, provided
               that such additional attribution notices cannot be construed
               as modifying the License.
        
           You may add Your own copyright statement to Your modifications and
           may provide additional or different license terms and conditions
           for use, reproduction, or distribution of Your modifications, or
           for any such Derivative Works as a whole, provided Your use,
           reproduction, and distribution of the Work otherwise complies with
           the conditions stated in this License.
        
        5. Submission of Contributions. Unless You explicitly state otherwise,
           any Contribution intentionally submitted for inclusion in the Work
           by You to the Licensor shall be under the terms and conditions of
           this License, without any additional terms or conditions.
           Notwithstanding the above, nothing herein shall supersede or modify
           the terms of any separate license agreement you may have executed
           with Licensor regarding such Contributions.
        
        6. Trademarks. This License does not grant permission to use the trade
           names, trademarks, service marks, or product names of the Licensor,
           except as required for reasonable and customary use in describing the
           origin of the Work and reproducing the content of the NOTICE file.
        
        7. Disclaimer of Warranty. Unless required by applicable law or
           agreed to in writing, Licensor provides the Work (and each
           Contributor provides its Contributions) on an "AS IS" BASIS,
           WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or
           implied, including, without limitation, any warranties or conditions
           of TITLE, NON-INFRINGEMENT, MERCHANTABILITY, or FITNESS FOR A
           PARTICULAR PURPOSE. You are solely responsible for determining the
           appropriateness of using or redistributing the Work and assume any
           risks associated with Your exercise of permissions under this License.
        
        8. Limitation of Liability. In no event and under no legal theory,
           whether in tort (including negligence), contract, or otherwise,
           unless required by applicable law (such as deliberate and grossly
           negligent acts) or agreed to in writing, shall any Contributor be
           liable to You for damages, including any direct, indirect, special,
           incidental, or consequential damages of any character arising as a
           result of this License or out of the use or inability to use the
           Work (including but not limited to damages for loss of goodwill,
           work stoppage, computer failure or malfunction, or any and all
           other commercial damages or losses), even if such Contributor
           has been advised of the possibility of such damages.
        
        9. Accepting Warranty or Additional Liability. While redistributing
           the Work or Derivative Works thereof, You may choose to offer,
           and charge a fee for, acceptance of support, warranty, indemnity,
           or other liability obligations and/or rights consistent with this
           License. However, in accepting such obligations, You may act only
           on Your own behalf and on Your sole responsibility, not on behalf
           of any other Contributor, and only if You agree to indemnify,
           defend, and hold each Contributor harmless for any liability
           incurred by, or claims asserted against, such Contributor by reason
           of your accepting any such warranty or additional liability.
        
        END OF TERMS AND CONDITIONS
        
        APPENDIX: How to apply the Apache License to your work.
        
           To apply the Apache License to your work, attach the following
           boilerplate notice, with the fields enclosed by brackets "[]"
           replaced with your own identifying information. (Don't include
           the brackets!)  The text should be enclosed in the appropriate
           comment syntax for the file format. We also recommend that a
           file or class name and description of purpose be included on the
           same "printed page" as the copyright notice for easier
           identification within third-party archives.
        
        Copyright [yyyy] [name of copyright owner]
        
        Licensed under the Apache License, Version 2.0 (the "License");
        you may not use this file except in compliance with the License.
        You may obtain a copy of the License at
        
            http://www.apache.org/licenses/LICENSE-2.0
        
        Unless required by applicable law or agreed to in writing, software
        distributed under the License is distributed on an "AS IS" BASIS,
        WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
        See the License for the specific language governing permissions and
        limitations under the License.
License-File: LICENSE
Keywords: ai-agents,governance,lexicon,rag,terminology,tool-use
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Typing :: Typed
Requires-Python: >=3.10
Provides-Extra: completion
Requires-Dist: argcomplete<4.0,>=3.0; extra == 'completion'
Provides-Extra: oov
Requires-Dist: huggingface-hub<1.0,>=0.20; extra == 'oov'
Requires-Dist: tokenizers<1.0,>=0.15; extra == 'oov'
Provides-Extra: semantic
Requires-Dist: sentence-transformers<4.0,>=2.6; extra == 'semantic'
Description-Content-Type: text/markdown

<h1 align="center">Agent Lexicon</h1>

<p align="center">
  <strong>A deterministic terminology layer for AI agents</strong><br>
  One canonical vocabulary across agents, branches, and tool calls.
</p>

<p align="center">
  <img alt="CI" src="https://img.shields.io/github/actions/workflow/status/SkeinRank/agent-lexicon/ci.yml?branch=main&label=CI">
  <img alt="PyPI" src="https://img.shields.io/pypi/v/agent-lexicon">
  <img alt="Python" src="https://img.shields.io/pypi/pyversions/agent-lexicon">
  <img alt="License" src="https://img.shields.io/github/license/SkeinRank/agent-lexicon">
</p>

<p align="center">
  <a href="#proof">Proof</a> ·
  <a href="docs/quickstart.md">Quickstart</a> ·
  <a href="docs/concepts.md">Concepts</a> ·
  <a href="#how-it-works">How it works</a>
</p>

When many agents work a long coding session, each one quietly invents its own names. One branch writes `accessToken`, another `authToken`, a third `bearer_token` — all the same concept. By merge time the service speaks five dialects of itself. Agent Lexicon gives every agent a single canonical vocabulary to read from, resolves the words they actually use back to that canon, and flags terminology that drifted before it lands in `main`.

It is dependency-free, runs locally, and is deterministic by design: the same input always produces the same output, and every decision carries a reason you can audit.

As a command-line tool, install it with [pipx](https://pipx.pypa.io) so `agent-lexicon` and the short `alex` alias are available in every project:

```bash
pipx install agent-lexicon
pipx install "agent-lexicon[completion]"   # with shell tab-completion
```

To use it as a library inside a project, install it with pip into that project's environment instead:

```bash
pip install agent-lexicon
```

Requires Python 3.10+. Apache 2.0. Zero runtime dependencies.

---

## Proof

### Benchmark: less terminology drift

Paired synthetic benchmark: the same coding tasks run with and without an agent-lexicon context brief.

**Setup:** Claude Sonnet 4.6, temperature 0, 10 tasks × 3 repeats × 2 conditions = 60 runs. Scored by an independent regex scorer over raw model output.

| Metric | No lexicon | With agent-lexicon |
|---|---:|---:|
| Exact canonical name (strict) | 0% | 60% |
| Canonical term inside compound (substring) | 10% | 100% |

This is an early synthetic benchmark, not a universal claim. Harness, tasks, scorer, and raw results: [SkeinRank/agent-lexicon-benchmark](https://github.com/SkeinRank/agent-lexicon-benchmark).

### Live output

Real output from the shipped example lexicon (`examples/customer_limits/lexicon.yaml`). Two terms share the surface word *limit* — `billing.credit_limit` and `api.rate_limit`. This is exactly where an agent drifts and calls the wrong tool.

**An ambiguous word stops the agent instead of guessing:**

```console
$ agent-lexicon resolve examples/customer_limits/lexicon.yaml "please raise the limit"
Status: ambiguous
Action: ask_clarification
Message: Found 2 possible canonical terms.
Lexicon snapshot: sha256:98b7c5324a20c58926ea8e3413f87851c6d8c354e93197a69c56f1e142ea962e
Candidates:
- api.rate_limit (rate limit) scopes=api matches='limit'
- billing.credit_limit (credit limit) scopes=billing matches='limit'
```

**The same word, scoped, resolves cleanly:**

```console
$ agent-lexicon resolve examples/customer_limits/lexicon.yaml "please raise the limit" --scope billing
Status: resolved
Action: use_terms
Message: Resolved to billing.credit_limit.
Lexicon snapshot: sha256:98b7c5324a20c58926ea8e3413f87851c6d8c354e93197a69c56f1e142ea962e
Candidates:
- billing.credit_limit (credit limit) scopes=billing matches='limit'
```

**A wrong tool call is blocked before it runs:**

```console
$ agent-lexicon guard examples/customer_limits/lexicon.yaml "raise the credit limit" --tool api.update_rate_limit
Status: blocked
Action: block
Allowed: no
Reason: Requested tool is not allowed for the resolved terminology.
Resolution: resolved
Lexicon snapshot: sha256:98b7c5324a20c58926ea8e3413f87851c6d8c354e93197a69c56f1e142ea962e
Matched terms:
- billing.credit_limit
Allowed tools:
- billing.update_credit_limit
```

No model was called. No embedding was computed. Run it again and you get the same answer, byte for byte.

---

## Why this exists

The longer and wider an agent session runs, the more the shared vocabulary drifts. This is not a hallucination problem — the agents are not inventing facts. They are naming the *same* concept inconsistently, in branches that never see each other until merge. The result is a codebase where one idea lives under several names, and nobody decided that on purpose.

Existing tools do not close this gap:

- **Knowledge graphs** model how concepts relate, but require pre-built structure and do not gate a tool call at runtime on raw text.
- **LLM or embedding similarity** can guess that two names mean the same thing, but the guess is non-deterministic and cannot be reproduced or audited a year later.
- **Linters** catch inconsistent identifiers in code, but have no notion of a canonical term and do not work on prose, comments, or tool-call text.

Agent Lexicon is a different layer: it takes raw text in, normalizes it, resolves it against a reviewed canonical vocabulary, and returns a structured, deterministic decision. Optional semantics sit on top — as a *suggestion to a human*, never as the thing that decides.

---

## How it works

Three pieces, each doing one job.

**Resolve** — Given a span of text, find the canonical terms and aliases inside it. Matching uses a dependency-free Aho-Corasick trie, so it is fast and works on prose, comments, and code-style identifiers (`accessToken`, `access_token`, `ACCESS_TOKEN` all resolve to the same term). Input is Unicode-normalized first, so invisible separators, full-width characters, and bidi-control tricks cannot slip a different term past the matcher.

**Guard** — Given resolved terminology and a tool the agent wants to call, decide whether that call is allowed. Ambiguous terminology returns `ask_clarification`. A tool that is not permitted for the resolved term returns `block`. Bidi-control characters in the triggering text are surfaced as a high-risk finding and block by default.

**Drift detection at merge** — Read the added lines between two git refs and classify every identifier: already known, a likely alias of an existing term, or a genuinely new term that nobody reviewed. The dangerous class — a coined name with no canonical neighbour — is what surfaces by default.

```console
$ agent-lexicon check-merge --root . --base main --head feature-branch --include 'src/**'
Git merge terminology check: 1 files, 6 added lines
Range: main...feature-branch
Lexicon: lexicon/lexicon.yaml
Lexicon snapshot: sha256:98b7c5324a20c58926ea8e3413f87851c6d8c354e93197a69c56f1e142ea962e
Summary: known=2, likely_alias=0, likely_new_term=3, unresolved_unknown=0, hidden_unresolved=1
Known terminology:
- auth.py:2 'authToken' -> auth.access_token (access token) scopes=auth
New terminology candidates:
- auth.py:3 'credentialBlob' unknown; possible new term
- auth.py:4 'sessionKey' unknown; possible new term
- auth.py:5 'quuxHandle' unknown; possible new term
Hidden unresolved identifiers: 1. Use --include-unresolved-unknowns to inspect low-signal identifiers.
```

Add `--fail-on-review` to make this a blocking CI check that returns a non-zero exit code when unreviewed drift appears.

---

## Three ways to use it

**Command line** — the full local loop, no code required. Every command is also available under the short alias `alex`, so `alex resolve …` works the same as `agent-lexicon resolve …`.

```bash
agent-lexicon init                      # create lexicon/, workspace, policy, and scan config
agent-lexicon scan                      # discover candidate terms from configured paths
agent-lexicon scan README.md docs src   # or override paths explicitly
agent-lexicon review                    # open the local web inbox to accept/reject
agent-lexicon publish --update-lexicon  # publish accepted decisions and update lexicon.yaml
agent-lexicon resolve <lexicon> "text"  # resolve terminology in any text
agent-lexicon guard   <lexicon> "text" --tool <name>   # gate a tool call
agent-lexicon context <lexicon>         # print the canonical vocabulary brief for an agent
agent-lexicon lint-diff --stdin         # lint a working diff for terminology drift
agent-lexicon check-merge --base main --head <branch>  # detect drift at merge
agent-lexicon check-merge --base main --head <branch> --semantic-check  # CI-style pass/fail
```

### In an agent workflow

Two of these commands are built for wrapping an AI coding agent:

**Before a task** — hand the agent the project's canonical vocabulary so it starts with the right language:

```bash
agent-lexicon context lexicon/lexicon.yaml
# Use these canonical terms:
# - ContextSpace
# - RuntimeSnapshot
#
# Avoid:
# - WorkspaceScope (use "ContextSpace" instead)
```

**During a task** — check a working diff for terminology drift before it is committed, with layered severity:

```bash
git diff | agent-lexicon lint-diff --stdin
# Terminology lint: 2 files, 18 added lines
#
# Deprecated terms (fail):
# - src/session.py:14 WorkspaceScope -> use "ContextSpace"
#
# Possible typos / near-misses (warn):
# - docs/api.md:7 ContextSapce -> did you mean "ContextSpace"?
#
# New project terms (info):
# - src/memory.py:22 TaskMemoryProfile
# (exit code 1: a deprecated term was used)
```

Level 1 (declared deprecated terms) fails the check. Level 2 (lexical near-misses) warns, or fails under `--strict`. Level 3 (unknown project terms) is reported for awareness only. An optional `--semantic` flag adds probabilistic suggestions but never changes the exit code — enforcement stays deterministic.

**At merge / PR** — a deterministic terminology gate alongside your other CI checks:

```bash
agent-lexicon check-merge --base main --head HEAD --semantic-check
# Terminology check: 3 files, 42 added lines
# Semantic conflicts detected (1):
# - customer cap vs credit limit (use "credit limit")
# (exit code 1, so CI fails)
```

Both are deterministic: they flag terms the lexicon *already declares* (deprecated aliases and near-misses to canonical terms), never guesses.

**Python library** — call the same logic inline.

```python
from agent_lexicon import load_lexicon, resolve_text, guard_tool_call

lexicon = load_lexicon("lexicon/lexicon.yaml")

decision = resolve_text(lexicon, "please raise the limit", scopes=["billing"])
print(decision.status)        # ResolutionStatus.RESOLVED
print(decision.action)        # ResolutionAction.USE_TERMS

guard = guard_tool_call(
    lexicon,
    "raise the credit limit",
    tool_name="api.update_rate_limit",
)
print(guard.status)           # ToolGuardStatus.BLOCKED
```

**MCP server** — expose the lexicon to any MCP-compatible agent over stdio.

```bash
agent-lexicon mcp serve --root . --lexicon lexicon/lexicon.yaml
```

The server exposes six tools: `resolve_term`, `check_language`, `guard_tool_call`, `find_evidence`, `submit_proposal`, and `get_snapshot`. List their full definitions with `agent-lexicon mcp tools`.

### Repository scan config

`agent-lexicon init` creates `.agent-lexicon/config.yaml` so common repository scans do not need long CLI commands. By default, `agent-lexicon scan` starts from documentation and common source roots, applies language-aware include globs for popular stacks, and respects the repository `.gitignore`.

```yaml
scan:
  paths:
    - README.md
    - docs
    - src
    - app
    - packages
    - lib
    - services
  include:
    - "docs/**/*.md"
    - "docs/**/*.txt"
    - "**/*.py"
    - "**/*.ts"
    - "**/*.tsx"
    - "**/*.go"
    - "**/*.rs"
    - "**/*.java"
    - "**/*.kt"
    - "**/*.cs"
    - "**/*.sql"
    - "**/*.yaml"
  exclude:
    - ".venv/**"
    - "node_modules/**"
    - "dist/**"
    - "**/generated/**"
  respect_gitignore: true
  max_file_bytes: 1000000
```

`.gitignore` is treated as the first line of repository-specific ignore behavior. Use `scan.exclude` for Agent Lexicon-specific rules such as generated fixtures that are still tracked in Git.

CLI flags still win when you need a one-off run:

```bash
agent-lexicon scan docs src --include "src/**/*.py" --exclude "src/generated/**"
agent-lexicon scan --no-gitignore
agent-lexicon check-merge --base main --head HEAD --exclude "docs/generated/**"
```

---

## GitHub Actions workflow

The repository includes a terminology review workflow for pull requests. It validates the tracked lexicon and runs merge-time drift detection against the PR diff:

```yaml
name: Agent Lexicon Terminology Review

on:
  pull_request:
    branches: [main]

jobs:
  terminology-review:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
        with:
          fetch-depth: 0
      - uses: actions/setup-python@v5
        with:
          python-version: "3.11"
      - run: python -m pip install --upgrade poetry==2.1.1
      - run: poetry install --with dev
      - run: poetry run agent-lexicon validate lexicon/lexicon.yaml --lint --strict-lint
      - run: poetry run agent-lexicon check-merge --root . --base origin/${{ github.base_ref }} --head HEAD
```

The checked-in workflow is review-first by default: it prints terminology drift without blocking every PR. Set `base_ref`, `head_ref`, and `fail_on_review=true` for an on-demand blocking run, or add `--fail-on-review` when your team is ready to make terminology review a required merge gate.

---

## The dictionary is code

The canonical vocabulary lives in a git-tracked YAML file. A term has a canonical form, aliases, the scopes it belongs to, and optionally the tools that are allowed to act on it.

```yaml
version: 1
scopes:
  - id: billing
    label: Billing
  - id: api
    label: API
terms:
  - id: billing.credit_limit
    canonical: credit limit
    scopes: [billing]
    tools: [billing.update_credit_limit]
    aliases:
      - surface: customer cap
      - surface: account limit
  - id: api.rate_limit
    canonical: rate limit
    scopes: [api]
    tools: [api.update_rate_limit]
    aliases:
      - surface: requests per minute
```

Because it is just a file in the repo, the vocabulary versions, diffs, and reviews the same way your code does. Runtime decisions also carry a content-addressed snapshot reference (`sha256:<digest>`), so the same text can be replayed later against the exact same vocabulary content. A built-in linter warns when a surface is broad enough to over-trigger or to affect a guard decision:

```console
$ agent-lexicon lint lexicon/lexicon.yaml
Lexicon lint: warnings (1 warning)
[warning] tool_broad_surface: tool-routed term uses a broad surface that can
  affect guard decisions (term=data.primary_key; surface='PK'). Hint: Use
  explicit tool-facing aliases and avoid bare words on terms with tools.
```

Review and publish decisions are kept as workspace provenance records and can be exported as JSONL for audit or handoff:

```console
$ agent-lexicon workspace export-decision-log --root . --action review_decision_saved
{"action":"review_decision_saved","actor":"local","rule_id":"human_review",...}
```

---

## Optional semantics, kept honest

When the deterministic heuristics are confident, they decide alone. When a new identifier lands in a gray zone — close to an existing term but not a clear match — an optional semantic reranker can suggest the most likely canonical neighbour, so a reviewer sees *"`authToken` might be your `access token`"* instead of an unsorted pile of unknowns.

```bash
pip install "agent-lexicon[oov]"        # tokenizer-backed out-of-vocabulary scoring
pip install "agent-lexicon[semantic]"   # semantic near-miss reranking
```

This is deliberately a **suggestion to a human, marked as non-deterministic**, never an autonomous decision. The semantic layer never commits a term on its own. The thing that decides stays deterministic and auditable; the thing that suggests is allowed to be smart. That boundary is the point — it is what keeps every committed decision reproducible.

---

## Design guarantees

These hold on the deterministic runtime and local review paths:

- **Deterministic.** The same text against the same immutable lexicon snapshot always produces the same decision. No model, no embedding, no randomness on the resolve and guard paths.
- **Reproducible.** Runtime and merge reports include a content-addressed `lexicon_snapshot_ref` (`sha256:<digest>`), so a decision can be replayed later against the exact same vocabulary content.
- **Auditable.** Every runtime decision reports its reason — which surface matched, at which span, in which scope, and why a tool was allowed or blocked. Local review and publish decisions are also written to an append-only provenance log with actor, action, rule, result, and lexicon snapshot metadata.
- **Dependency-free core.** The resolver and matcher have zero runtime dependencies and run entirely in memory. Optional extras are opt-in and never touch the hot path.
- **Safe by construction.** Local writes are atomic (a reader sees a complete file or none), and the workspace database is configured for concurrent access without torn reads.
- **Storage boundary.** The local workspace is SQLite-backed by default, but workflow code depends on a small `WorkspaceStore` boundary so future shared storage can be added without changing the deterministic runtime.

---

## Documentation

- [Quickstart](docs/quickstart.md) — local setup, scan, review, publish, and runtime usage.
- [Concepts](docs/concepts.md) — terms, aliases, scopes, resolution, guard decisions, and merge-time drift detection.
- [Python API reference](docs/api.md) — `load_lexicon`, `resolve_text`, `guard_tool_call`, and the decision and enum types they return.
- [MCP server reference](docs/mcp.md) — the six MCP tools, their arguments, and their return values.

Contributing, security, and community:

- [Contributing](CONTRIBUTING.md) — development setup and how to send a change.
- [Security policy](SECURITY.md) — how to report a vulnerability privately.
- [Code of conduct](CODE_OF_CONDUCT.md) — community standards.

---

## Status

Agent Lexicon is an early, actively developed project (0.7.x). The core — resolve, guard, near-miss, dictionary-as-code, and merge-time drift detection — is well tested (326 passing tests) and used through the CLI, the Python API, and the local MCP server. Scaling it across many processes or a networked deployment is on the roadmap, not yet proven in production.

If terminology consistency across long, multi-agent sessions is a real cost for you — especially in regulated domains where decisions must be reproducible and auditable — this is built for exactly that.

---

## License

Apache 2.0. Free for commercial use.
