Metadata-Version: 2.4
Name: wikidata-ner-classifier
Version: 0.7.1
Summary: Hierarchical Wikidata item classification and mention-focused LLM type inference
Author: Roberto Avogadro
License-Expression: MIT
Project-URL: Homepage, https://github.com/roby-avo/ner-wikidata
Project-URL: Issues, https://github.com/roby-avo/ner-wikidata/issues
Project-URL: Source, https://github.com/roby-avo/ner-wikidata
Keywords: wikidata,ner,entity-linking,knowledge-graph,classification
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Typing :: Typed
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
Dynamic: license-file

# Wikidata NER Classifier 0.7.1

`wikidata-ner-classifier` predicts retrieval-oriented NER types in two ways:

1. **Wikidata items:** deterministic prediction from P31/P279 token clues and an
   optional description.
2. **Input data:** LLM prediction from a target mention and its context, supplied
   as free text or tabular data.

Both paths return types from the same hierarchy:

```text
coarse_type -> fine_type -> subtype -> specific_type
```

The prediction can be used to narrow the candidate-retrieval space before entity
linking. The library does not identify or retrieve a Wikidata QID.

## Installation

```bash
pip install wikidata-ner-classifier
```

## 1. Predict NER types for Wikidata items

Use `WikidataNERClassifier` when the input is already a Wikidata item and its
P31/P279 type labels or aliases are available. Prediction is deterministic and
does not require an LLM or network request.

```python
from wikidata_ner import WikidataNERClassifier

classifier = WikidataNERClassifier()

prediction = classifier.predict(
    qid="Q3441181",
    types=[
        {"id": "Q11424", "name": "film"},
    ],
    description="1964 sword-and-sandal film directed by Giuseppe Vari",
)

print(prediction.coarse_type)    # CREATIVE_WORK
print(prediction.fine_type)      # FILM
print(prediction.specific_type)  # SWORD_AND_SANDAL_FILM
print(prediction.retrieval_key)
# CREATIVE_WORK/FILM/SWORD_AND_SANDAL_FILM
```

The P31/P279 token clues select the semantic branch. The description may refine
the result, but it does not replace the type-token evidence.

A complete entity mapping can also be passed directly:

```python
prediction = classifier.predict_entity(
    {
        "qid": "Q3441181",
        "types": [{"id": "Q11424", "name": "film"}],
        "description": "1964 sword-and-sandal film",
    }
)
```

For multiple Wikidata items, use `predict_batch()`:

```python
predictions = classifier.predict_batch(
    [
        {
            "qid": "Q3441181",
            "types": [{"name": "film"}],
            "description": "1964 sword-and-sandal film",
        },
        {
            "qid": "Q7259",
            "types": [{"name": "human"}],
            "description": "English mathematician and writer",
        },
    ]
)
```

## 2. Predict NER types for input data with an LLM

Use `OpenRouterNERClassifier` when the input is a mention whose type must be
inferred from context. The context can be free text, a structured record, or a
table cell.

Set an OpenRouter API key:

```bash
export OPENROUTER_API_KEY="..."
```

Create the classifier:

```python
from wikidata_ner import OpenRouterNERClassifier

classifier = OpenRouterNERClassifier(
    model="openai/gpt-oss-120b",
    provider="cerebras",
    allow_fallbacks=False,
    reasoning_effort="low",
)
```

### Free text

```python
prediction = classifier.predict_text(
    "Rome Against Rome is a 1964 sword-and-sandal film.",
    mention="Rome Against Rome",
)

print(prediction.coarse_type)    # CREATIVE_WORK
print(prediction.fine_type)      # FILM
print(prediction.specific_type)  # SWORD_AND_SANDAL_FILM
```

Only the supplied target mention is classified. The surrounding sentence is
contextual evidence.

### Structured input

```python
prediction = classifier.predict_record(
    {
        "label": "Chrysler Cirrus",
        "description": "mid-size four-door sedan model",
        "manufacturer": "Chrysler",
    }
)

print(prediction.coarse_type)    # PRODUCT
print(prediction.fine_type)      # VEHICLE_WEAPON_OR_EQUIPMENT_MODEL
print(prediction.subtype)        # CAR_MODEL, when supported by the evidence
```

Existing QIDs, URLs, popularity, priors, and previous NER fields are not used as
prediction evidence.

### Tabular input

For one table cell, provide the column meaning and bounded row/column context:

```python
prediction = classifier.predict_table_cell(
    "Germany",
    column_header="country name",
    row_context={
        "manufacturer": "Daimler AG",
        "vehicle_model": "Chrysler Cirrus",
        "assembly_location": "Sterling Heights, Michigan",
    },
    same_column_values=[
        "Germany",
        "United States",
        "Canada",
    ],
    table_name="vehicle_production.csv",
)

print(prediction.coarse_type)  # LOCATION
print(prediction.fine_type)    # COUNTRY_OR_SOVEREIGN_STATE
```

For multiple cells, use `TableCellTask` and `predict_table_cells()`:

```python
from wikidata_ner import TableCellTask

tasks = [
    TableCellTask(
        cell="Germany",
        column_header="country name",
        row_context={"manufacturer": "Daimler AG"},
        same_column_values=["Germany", "United States", "Canada"],
    ),
    TableCellTask(
        cell="United States",
        column_header="country name",
        row_context={"manufacturer": "General Motors"},
        same_column_values=["Germany", "United States", "Canada"],
    ),
]

predictions = classifier.predict_table_cells(tasks)
```

The production batch limit is 8 targets per physical request. Larger iterables
are split into multiple requests automatically.

## Prediction output

Both classifiers expose retrieval-oriented fields such as:

- `coarse_type`
- `fine_type`
- `subtype`
- `specific_type` and `specific_types`
- `retrieval_key`, `retrieval_path`, and `retrieval_tags`
- `confidence`
- `abstained` and `abstention_reason`

The model predicts semantic types only. QIDs and retrieval metadata are never
generated by the LLM; the library validates the prediction and constructs the
retrieval fields locally.

## Release history

See [CHANGELOG.md](CHANGELOG.md) for version details.

## License

MIT
