Metadata-Version: 2.4
Name: wikidata-ner-classifier
Version: 0.8.0
Summary: Hierarchical Wikidata item classification and mention-focused LLM type inference
Author: Roberto Avogadro
License-Expression: MIT
Project-URL: Homepage, https://github.com/roby-avo/ner-wikidata
Project-URL: Issues, https://github.com/roby-avo/ner-wikidata/issues
Project-URL: Source, https://github.com/roby-avo/ner-wikidata
Keywords: wikidata,ner,entity-linking,knowledge-graph,classification
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# Wikidata NER Classifier 0.8.0

`wikidata-ner-classifier` predicts retrieval-oriented NER types in two ways:

1. **Wikidata items:** deterministic prediction from P31/P279 token clues and an
   optional description.
2. **Input data:** LLM prediction from a target mention and its context, supplied
   as free text or tabular data.

Both paths return types from the same hierarchy:

```text
coarse_type -> fine_type -> subtype -> specific_type
```

The prediction can be used to narrow the candidate-retrieval space before entity
linking. The library does not make the final identity decision. LLM predictions
may include unverified Wikipedia and DBpedia URLs as search hints; a downstream
linker must retrieve and verify them. Wikidata QIDs and URLs are deliberately
excluded because they would be unique-identity guesses.

## Installation

```bash
pip install wikidata-ner-classifier
```

## 1. Predict NER types for Wikidata items

Use `WikidataNERClassifier` when the input is already a Wikidata item and its
P31/P279 type labels or aliases are available. Prediction is deterministic and
does not require an LLM or network request.

```python
from wikidata_ner import WikidataNERClassifier

classifier = WikidataNERClassifier()

prediction = classifier.predict(
    qid="Q3441181",
    types=[
        {"id": "Q11424", "name": "film"},
    ],
    description="1964 sword-and-sandal film directed by Giuseppe Vari",
)

print(prediction.coarse_type)    # CREATIVE_WORK
print(prediction.fine_type)      # FILM
print(prediction.specific_type)  # SWORD_AND_SANDAL_FILM
print(prediction.retrieval_key)
# CREATIVE_WORK/FILM/SWORD_AND_SANDAL_FILM
```

The P31/P279 token clues select the semantic branch. The description may refine
the result, but it does not replace the type-token evidence.

A complete entity mapping can also be passed directly:

```python
prediction = classifier.predict_entity(
    {
        "qid": "Q3441181",
        "types": [{"id": "Q11424", "name": "film"}],
        "description": "1964 sword-and-sandal film",
    }
)
```

For multiple Wikidata items, use `predict_batch()`:

```python
predictions = classifier.predict_batch(
    [
        {
            "qid": "Q3441181",
            "types": [{"name": "film"}],
            "description": "1964 sword-and-sandal film",
        },
        {
            "qid": "Q7259",
            "types": [{"name": "human"}],
            "description": "English mathematician and writer",
        },
    ]
)
```

## 2. Predict NER types for input data with an LLM

Use `OpenRouterNERClassifier` when the input is a mention whose type must be
inferred from context. The context can be free text, a structured record, or a
table cell.

Set an OpenRouter API key:

```bash
export OPENROUTER_API_KEY="..."
```

Create the classifier:

```python
from wikidata_ner import OpenRouterNERClassifier

classifier = OpenRouterNERClassifier(
    model="openai/gpt-oss-120b",
    provider="cerebras",
    allow_fallbacks=False,
    reasoning_effort="low",
)
```

### Free text

```python
prediction = classifier.predict_text(
    "Rome Against Rome is a 1964 sword-and-sandal film.",
    mention="Rome Against Rome",
)

print(prediction.coarse_type)    # CREATIVE_WORK
print(prediction.fine_type)      # FILM
print(prediction.specific_type)  # SWORD_AND_SANDAL_FILM
```

Only the supplied target mention is classified. The surrounding sentence is
contextual evidence.

The same LLM call also returns backend-neutral candidate-retrieval metadata:

```python
prediction = classifier.predict_text(
    "Rmoe is the capital and largest city of Italy.",
    mention="Rmoe",
)

print(prediction.retrieval_metadata.corrected_mention)  # Rome
print(prediction.retrieval_metadata.context_keywords)   # e.g. ("capital city", "Italy")
print(prediction.retrieval_metadata.wikipedia_urls)
# e.g. ("https://en.wikipedia.org/wiki/Rome",)
print(prediction.retrieval_metadata.dbpedia_urls)
# e.g. ("https://dbpedia.org/resource/Rome",)
print(prediction.high_level_reason)
```

Reference URLs are model predictions, not verified links. Invalid URL shapes,
Wikidata URLs, and QID-bearing values are removed locally, and every serialized
metadata object is explicitly marked `unverified`.

### Structured input

```python
prediction = classifier.predict_record(
    {
        "label": "Chrysler Cirrus",
        "description": "mid-size four-door sedan model",
        "manufacturer": "Chrysler",
    }
)

print(prediction.coarse_type)    # PRODUCT
print(prediction.fine_type)      # VEHICLE_WEAPON_OR_EQUIPMENT_MODEL
print(prediction.subtype)        # CAR_MODEL, when supported by the evidence
```

Existing QIDs, URLs, popularity, priors, and previous NER fields are not used as
prediction evidence. Newly predicted reference URLs remain optional search
hints and cannot determine the semantic type.

### Tabular input

For one table cell, provide the column meaning and bounded row/column context:

```python
prediction = classifier.predict_table_cell(
    "Germany",
    column_header="country name",
    row_context={
        "manufacturer": "Daimler AG",
        "vehicle_model": "Chrysler Cirrus",
        "assembly_location": "Sterling Heights, Michigan",
    },
    same_column_values=[
        "Germany",
        "United States",
        "Canada",
    ],
    table_name="vehicle_production.csv",
)

print(prediction.coarse_type)  # LOCATION
print(prediction.fine_type)    # COUNTRY_OR_SOVEREIGN_STATE
```

For multiple cells, use `TableCellTask` and `predict_table_cells()`:

```python
from wikidata_ner import TableCellTask

tasks = [
    TableCellTask(
        cell="Germany",
        column_header="country name",
        row_context={"manufacturer": "Daimler AG"},
        same_column_values=["Germany", "United States", "Canada"],
    ),
    TableCellTask(
        cell="United States",
        column_header="country name",
        row_context={"manufacturer": "General Motors"},
        same_column_values=["Germany", "United States", "Canada"],
    ),
]

predictions = classifier.predict_table_cells(tasks)
```

The production batch limit is 8 targets per physical request. Larger iterables
are split into multiple requests automatically.

## Prediction output

All prediction paths expose retrieval-oriented type fields such as:

- `coarse_type`
- `fine_type`
- `subtype`
- `specific_type` and `specific_types`
- `retrieval_key`, `retrieval_path`, and `retrieval_tags`
- `confidence`
- `abstained` and `abstention_reason`

LLM-backed mention predictions additionally expose:

- `retrieval_metadata`, including corrected spelling, mention variants,
  disambiguating keywords, and optional Wikipedia/DBpedia reference URLs
- `high_level_reason`, one explanation covering the type and metadata choices

Type and metadata confidence values are computed locally as evidence-derived
posterior probabilities; numeric confidence self-ratings from the model are not
used at any stage. This includes fine alternatives, subtypes, occupation or other
facets, and retrieval metadata. Wikipedia and DBpedia URL hints are accepted only
when the model explicitly returns them and they pass local domain, canonical-path,
QID, and title checks; the library never constructs missing URLs. No other URL
family or unique entity identifier is predicted.

Controlled type paths, keys, and tags are still validated and constructed
locally. Only the bounded `retrieval_metadata` hints are model-predicted.

`to_candidate_retrieval_profile()` provides a generic query contract that can be
adapted to Elasticsearch, a vector database, a knowledge-graph lookup, or another
retrieval system. Keep its roles separate:

- use `mention_queries` for lexical candidate recall;
- use `context_keywords` as soft disambiguation signals;
- use hierarchy hints as confidence-aware filters or boosts;
- treat `reference_urls` as optional exact lookups that still require
  verification.

In particular, avoid concatenating every context keyword into the mention query:
that can reward labels containing the context words rather than the entity named
by the mention.

## Release history

See [CHANGELOG.md](CHANGELOG.md) for version details.

## License

MIT
