Metadata-Version: 2.4
Name: wikidata-ner-classifier
Version: 0.4.0
Summary: Hierarchical Wikidata item classification and mention-focused LLM type inference
Author: Roberto Avogadro
License-Expression: MIT
Project-URL: Homepage, https://github.com/roby-avo/ner-wikidata
Project-URL: Issues, https://github.com/roby-avo/ner-wikidata/issues
Project-URL: Source, https://github.com/roby-avo/ner-wikidata
Keywords: wikidata,ner,entity-linking,knowledge-graph,classification
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# Wikidata NER Classifier 0.4.0

Classification into one shared retrieval-oriented hierarchy through two
complementary paths:

1. `WikidataNERClassifier` classifies Wikidata items deterministically from
   P31/P279 token clues and optional descriptions.
2. `OpenRouterNERClassifier` infers the type of one target mention from free
   text, a structured record, or tabular context using an LLM.

Both paths return classes from the packaged hierarchy:

```text
coarse_type -> fine_type -> subtype -> specific_type
```

The mention classifier cannot invent labels: structured outputs enumerate the
library's classes, and returned branches, subtypes, and facets are validated
again locally.

## What changed in 0.4.0

- Added mention-focused type inference for free text and structured/tabular data.
- Added a dependency-free OpenRouter client with strict JSON-schema output.
- Added branch-local LLM refinement using the same subtype and facet rules as
  Wikidata item classification.
- Added exact mention-span marking, target-focus validation, and safe abstention.
- Retained the deterministic Wikidata token/clue classifier unchanged.

The QID is retained as an identifier and is never used as a lookup key.

## Two classification paths

### Wikidata items: deterministic token clues

Use this path when P31/P279 labels are already available:

```python
from wikidata_ner import WikidataNERClassifier

classifier = WikidataNERClassifier()
prediction = classifier.predict(
    qid="Q3441181",
    types=[{"id": "Q11424", "name": "film"}],
    description="1964 sword-and-sandal film directed by Giuseppe Vari",
)
```

The primary branch is selected with the library's deterministic token/clue
rules. Descriptions can refine that branch but cannot replace its P31/P279
anchor.

### Input mentions: LLM inference through OpenRouter

Use this path when the input is a mention and its type must be inferred from
context:

```bash
export OPENROUTER_API_KEY='...'
```

```python
from wikidata_ner import OpenRouterNERClassifier

classifier = OpenRouterNERClassifier(
    model="openai/gpt-oss-120b",
)

prediction = classifier.predict_text(
    "Rome Against Rome is a 1964 sword-and-sandal film directed by "
    "Giuseppe Vari; the story is set partly in Rome.",
    mention="Rome Against Rome",
)

print(prediction.fine_type)       # FILM
print(prediction.specific_type)   # SWORD_AND_SANDAL_FILM
```

Only `Rome Against Rome` is classified. The later `Rome` is contextual evidence
about a different mention and cannot become the prediction target.

For a structured record:

```python
prediction = classifier.predict_record(
    {
        "label": "Rome Against Rome",
        "types": [{"name": "film"}],
        "description": "1964 sword-and-sandal film",
    }
)
```

For a table cell:

```python
prediction = classifier.predict_table_cell(
    "acetylsalicylic acid",
    column_header="active ingredient",
    row_context={
        "drug": "Aspirin",
        "molecular_formula": "C9H8O4",
    },
    same_column_values=["ibuprofen", "paracetamol", "naproxen"],
)
```

The cell is always the target. Headers, row attributes, and same-column samples
are evidence about the cell, never alternative targets.

For contextual free text, `mention=` is required. You can disambiguate repeated
surface forms with an exact character span:

```python
prediction = classifier.predict_text(
    text,
    mention="Rome",
    mention_span=(start, end),
)
```

Use `preview_prompts(...)` to inspect the normalized mention, exact prompts, and
JSON schemas without making an API call:

```python
preview = classifier.preview_prompts(
    text,
    mention="Rome Against Rome",
    assumed_fine_type="FILM",
)
```

The predictor makes two calls when the selected branch has refinement rules:

1. Select exactly one controlled coarse/fine branch or abstain.
2. Select only subtypes and facets allowed inside that fine branch.

OpenRouter is called at
`https://openrouter.ai/api/v1/chat/completions` with strict JSON-schema output,
`provider.require_parameters=true`, temperature zero, and optional response
healing. The package continues to have no runtime dependencies. You may inject a
custom `client=` for testing or infrastructure integration.

`MentionPrediction` supports the same retrieval conveniences as deterministic
predictions:

```python
payload = prediction.to_dict()
fields = prediction.to_retrieval_fields(prefix="ner")
query_filter = prediction.elasticsearch_filter()
```

The selected OpenRouter model must support structured outputs. Pin a model slug
in production and store the returned model, prompt version, taxonomy version,
usage, evidence, and confidence with each result.

## Deterministic evidence policy

The default evidence policy is now:

1. `types[].name` selects `coarse_type` and `fine_type`.
2. `ancestor_types[].name`, when supplied, provides lower-weight class ancestry.
3. Direct type labels and `description` refine only the selected branch.
4. `context_string` is ignored by the classifier by default because it often
   contains related people, organizations, countries, genres, and formats.
5. Description evidence cannot change a `FILM` branch into a location, company,
   person, or another unrelated branch.
6. Unsupported specificity is not invented.

The packaged configuration contains:

- 187 fine-type rules;
- 295 structural subtype rules;
- 33 controlled facet rules;
- 6 branch-local composite-type templates.

## Installation

```bash
python -m pip install wikidata-ner-classifier
```

Python 3.10 or newer is required. The library has no runtime dependencies.

## Deterministic basic use

```python
from wikidata_ner import WikidataNERClassifier

classifier = WikidataNERClassifier()

prediction = classifier.predict(
    qid="Q3441181",
    types=[{"id": "Q11424", "name": "film"}],
    description="1964 sword-and-sandal film directed by Giuseppe Vari",
)

print(prediction.to_dict())
```

Relevant output:

```json
{
  "coarse_type": "CREATIVE_WORK",
  "fine_type": "FILM",
  "subtype": null,
  "specific_type": "SWORD_AND_SANDAL_FILM",
  "specific_types": [
    "SWORD_AND_SANDAL_FILM"
  ],
  "facets": {
    "genre": [
      "SWORD_AND_SANDAL"
    ]
  },
  "refinement_sources": [
    "description"
  ]
}
```

The description adds specificity only inside the already established `FILM`
branch.

## Alpaca or Elasticsearch entities

Both source objects and complete Elasticsearch hits are accepted:

```python
prediction = classifier.predict_entity(hit_or_source)
```

```python
{
  "qid": "Q3441181",
  "types": [{"name": "film"}],
  "description": "1964 sword-and-sandal film directed by Giuseppe Vari"
}
```

```python
{
  "_id": "Q3441181",
  "_source": {
    "qid": "Q3441181",
    "types": [{"name": "film"}],
    "description": "1964 sword-and-sandal film directed by Giuseppe Vari"
  }
}
```

A complete Elasticsearch response can be processed with:

```python
predictions = classifier.predict_elasticsearch_response(response)
```

## Why both subtype and specific type exist

A subtype describes a structural kind. A facet describes an independent
characteristic. A specific type is a retrieval-oriented composition.

```python
prediction = classifier.predict(
    "Q1",
    [
        {"name": "film"},
        {"name": "feature film"},
        {"name": "comedy film"},
    ],
)
```

This can produce:

```json
{
  "fine_type": "FILM",
  "subtype": "FEATURE_FILM",
  "specific_type": "FEATURE_FILM",
  "specific_types": [
    "FEATURE_FILM",
    "COMEDY_FILM"
  ],
  "facets": {
    "genre": [
      "COMEDY"
    ]
  }
}
```

`FEATURE_FILM` and `COMEDY_FILM` are compatible. They may be combined by the
retriever rather than forced into a single mutually exclusive label.

## Example refinements from the Alpaca query

| Type labels | Description | Fine type | Subtype | Most specific retrieval type |
|---|---|---|---|---|
| `film` | `1951 film directed by Luigi Zampa` | `FILM` | none | `FILM` |
| `film` | `1964 sword-and-sandal film ...` | `FILM` | none | `SWORD_AND_SANDAL_FILM` |
| `album` | `album by Holger Czukay` | `MUSICAL_WORK_SONG_OR_ALBUM` | `MUSIC_ALBUM` | `MUSIC_ALBUM` |
| `literary work` | `Alternative history, military science fiction story` | `BOOK_OR_WRITTEN_WORK` | `FICTION_STORY` | `MILITARY_SCIENCE_FICTION_LITERARY_WORK` |
| `pencil drawing` | `1953 work of art ...` | `VISUAL_ARTWORK_PHOTOGRAPH_OR_COMIC` | `PENCIL_DRAWING` | `PENCIL_DRAWING` |

A generic description cannot justify an invented subtype. A generic film remains
`FILM` when neither its type labels nor description contain a safe refinement.

## Retrieval indexing helpers

Store the prediction alongside each entity using stable keyword fields:

```python
fields = prediction.to_retrieval_fields(prefix="ner")
```

Example fields:

```json
{
  "ner_coarse_type": "CREATIVE_WORK",
  "ner_fine_type": "FILM",
  "ner_subtype": null,
  "ner_specific_type": "SWORD_AND_SANDAL_FILM",
  "ner_specific_types": [
    "SWORD_AND_SANDAL_FILM"
  ],
  "ner_facets": {
    "genre": [
      "SWORD_AND_SANDAL"
    ]
  }
}
```

A deterministic Elasticsearch filter can be generated with:

```python
query_filter = prediction.elasticsearch_filter(
    field="ner_specific_types",
    require_all=True,
)
```

For several compatible specific types, `require_all=True` emits one term filter
per type. Use `require_all=False` to emit a `terms` disjunction.

Index-time and query-time predictions should use the same library and rule-file
version.

## Description and context controls

Description refinement is enabled by default:

```python
classifier = WikidataNERClassifier(
    use_description_for_refinement=True,
    use_context_string_for_refinement=False,
)
```

Disable it when only class labels should be considered:

```python
classifier = WikidataNERClassifier(
    use_description_for_refinement=False,
)
```

Noisy context refinement is available only as an explicit opt-in:

```python
classifier = WikidataNERClassifier(
    use_context_string_for_refinement=True,
)
```

The separate `use_description=True` option allows description text to add
low-weight support to the primary coarse/fine scorer. It is disabled by default.
Descriptions therefore do not rescue a missing or unknown type anchor unless the
caller explicitly changes that policy.

## Live Alpaca notebook

Open `examples/alpaca_live_test.ipynb`.

The notebook:

1. issues the supplied Alpaca Elasticsearch request;
2. extracts QID, type labels, and description;
3. ignores `context_string` during classification;
4. displays `predicted_subtype`, `predicted_specific_type`, all compatible
   `specific_types`, facets, and confidence values;
5. demonstrates a retrieval filter generated from the prediction.

Set the bearer token before starting Jupyter:

```bash
export ALPACA_TOKEN='your-token'
```

The notebook also supports a hidden token prompt when the environment variable is
not set.

## CLI

```bash
wikidata-ner response.json > predictions.json
cat response.json | wikidata-ner
```

Relevant flags:

```text
--no-description-refinement
--context-refinement
--description-for-primary
--entity-label
```

## Validation

The source package includes unit tests for:

- direct type-label classification;
- description-only branch refinement;
- context exclusion by default;
- explicit context opt-in;
- subtype and facet compatibility;
- generic fallback behavior;
- QID independence;
- Elasticsearch hit input;
- retrieval-field generation;
- generated Elasticsearch filters.


## Mention-focused OpenRouter notebook

`examples/openrouter_mention_focused_ner.ipynb` classifies one explicit target
mention at a time from free text, structured records, or table cells. Context is
used only as evidence for that target. Contextual free text requires `mention=`
or an exact span; table helpers make the selected cell the target. The installable
library now exposes the same workflow through `OpenRouterNERClassifier`; the
notebook remains useful as an expanded prompt inspection and evaluation example.
