Metadata-Version: 2.4
Name: tajiknlp
Version: 1.0.3
Summary: Production-ready NLP library for Tajik language (Cyrillic script)
Author-email: Arabov Mullosharaf Kurbonovich <cool.araby@gmail.com>
Maintainer-email: Arabov Mullosharaf Kurbonovich <cool.araby@gmail.com>
License-Expression: MIT
Project-URL: Homepage, https://github.com/TajikNLPWorld/tajiknlp
Project-URL: Repository, https://github.com/TajikNLPWorld/tajiknlp
Project-URL: Issues, https://github.com/TajikNLPWorld/tajiknlp/issues
Project-URL: Documentation, https://TajikNLPWorld.github.io/tajiknlp
Project-URL: Changelog, https://github.com/TajikNLPWorld/tajiknlp/blob/main/CHANGELOG.md
Keywords: tajik,nlp,natural-language-processing,tokenization,lemmatization,low-resource,cyrillic,morphology,text-processing,persian-family
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Intended Audience :: Education
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.8
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Topic :: Text Processing :: Linguistic
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Classifier: Natural Language :: Persian
Classifier: Operating System :: OS Independent
Requires-Python: >=3.8
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: regex>=2023.0
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: ruff>=0.6.0; extra == "dev"
Requires-Dist: build>=1.0; extra == "dev"
Requires-Dist: twine>=4.0; extra == "dev"
Provides-Extra: docs
Requires-Dist: mkdocs>=1.5; extra == "docs"
Requires-Dist: mkdocs-material>=9.0; extra == "docs"
Requires-Dist: mkdocstrings[python]>=0.24; extra == "docs"
Dynamic: license-file

# 🇹🇯 TajikNLP

**Production-ready NLP library for the Tajik language** (Cyrillic script).

[![PyPI version](https://img.shields.io/pypi/v/tajiknlp.svg)](https://pypi.org/project/tajiknlp/)
[![Python 3.8+](https://img.shields.io/badge/python-3.8%2B-blue)](https://www.python.org/downloads/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![Ruff](https://img.shields.io/badge/code%20style-ruff-261230.svg)](https://github.com/astral-sh/ruff)
[![CI](https://github.com/TajikNLPWorld/tajiknlp/actions/workflows/ci.yml/badge.svg)](https://github.com/TajikNLPWorld/tajiknlp/actions/workflows/ci.yml)
[![Downloads](https://static.pepy.tech/badge/tajiknlp)](https://pepy.tech/project/tajiknlp)

TajikNLP provides a complete suite of text processing tools specifically designed for the unique challenges of the Tajik language (Cyrillic alphabet). Whether you're building a search engine, a chatbot, or conducting linguistic research, TajikNLP gives you robust, customizable components that work out of the box.

---

## ✨ Features (v1.0.3)

### 🧹 Text Cleaning (100% Complete)
- Remove URLs, emails, HTML tags, mentions, and hashtags
- Strip invisible Unicode characters and normalize whitespace
- Smart bracket cleaning (empty brackets and punctuation‑only brackets are removed)
- Remove spaces before punctuation marks

### 🔤 Normalization (100% Complete)
- Tajik Cyrillic standardization (e.g., `љ` → `ҷ`, `ї` → `ӣ`)
- Persian/Arabic digit conversion (`۱۲۳` → `123`)
- Quote and dash normalization (`« »` → `"`, `—` → `-`)
- Optional lowercasing and custom character replacement maps
- **NEW:** `check_normalization()` to validate text normalization

### ✂️ Tokenization (100% Complete)
- Regex‑based word tokenizer aware of Tajik‑specific letters (`ғ`, `ӣ`, `қ`, `ӯ`, `ҳ`, `ҷ`)
- Handles hyphenated compounds (`китоб-ҳо`)
- Preserves punctuation as separate tokens (optional)
- Provides accurate character offsets for each token
- Morpheme-level tokenization for linguistic analysis

### 📝 Sentence Splitting (100% Complete)
- Rule-based sentence boundary detection
- Protects Tajik-specific abbreviations (`т.д.`, `и.т.п.`, `ва ғ.`)
- Handles initials, decimals, dates, URLs, and emails

### 🏷️ Part‑of‑Speech Tagging (100% Complete)
- Rule‑based tagger with dictionary lookup and morphological fallback
- Category priority system from suffix rules
- Works on raw text without prior annotation

### 🔍 Morpheme Segmentation (100% Complete)
- Splits words into prefixes, root, and suffixes
- Uses linguistic rules from external JSON files (fully customizable)
- Returns structured morpheme data for each token

### 📚 Lemmatization (100% Complete)
- Hybrid approach: direct dictionary lookup + iterative affix stripping
- POS‑aware suffix removal for higher accuracy
- Built‑in dictionary of irregular forms

### 🌿 Stemming (100% Complete)
- Conservative-to-aggressive rule-based stemmer
- Safe mode for grammatical reductions
- Deep mode for derivational stripping
- Corpus-aware scoring for better accuracy

### 🛑 Stop Words Filtering (100% Complete)
- Comprehensive stop word list (100+ items)
- Automatically removes function words while preserving content words

### 🎯 Named Entity Recognition (NER) Ready (100% Complete)
- Character-level span alignment
- BIO tag generation for entity spans
- Offset preservation for training data preparation

### 🛠️ Utilities (100% Complete)
- **Script detection** – identify if text is Cyrillic, Latin, Arabic, or mixed
- **Language detection** – distinguish Tajik from Russian, Persian, English, Arabic
- **Quality scoring** – evaluate how "Tajik‑like" a piece of text is
- **Validation helpers** – quickly check if a string is valid Tajik text
- **NEW: Evaluation Metrics** – Precision, Recall, F1, Edit Distance, Word Error Rate, BLEU score, Classification Report

---

## 📊 Project Status

| Module               | Status        | Tests   | Coverage |
|----------------------|---------------|---------|----------|
| Preprocessing        | ✅ Complete   | 33/33   | 90%+     |
| Tokenization         | ✅ Complete   | 22/22   | 88%      |
| Sentence Splitting   | ✅ Complete   | 7/7     | 93%      |
| POS Tagging          | ✅ Complete   | 9/9     | 71%      |
| Morpheme Segmentation| ✅ Complete   | 8/8     | 72%      |
| Lemmatization        | ✅ Complete   | 23/23   | 79%      |
| Stemming             | ✅ Complete   | 6/6     | 85%      |
| Stop Words Filtering | ✅ Complete   | 6/6     | 71%      |
| NER / Alignment      | ✅ Complete   | 5/5     | 60%+     |
| Script Detection     | ✅ Complete   | 47/47   | 76%      |
| Validation & Metrics | ✅ Complete   | 18/18   | 98%      |
| **Overall**          | **✅ 100%**   | **171** | **79%**   |

---

## 📦 Installation

TajikNLP requires Python 3.8 or later.

### Basic Installation

```bash
pip install tajiknlp
```

### Development Installation

```bash
pip install "tajiknlp[dev]"
```

### Full Installation (with ML dependencies)

```bash
pip install "tajiknlp[full]"
```

---

## 🧪 Try It Now in Google Colab

Click the badge below to run a complete demo notebook directly in your browser:

[![Open In Colab](https://github.com/TajikNLPWorld/tajiknlp/tree/main/demo/tajiknlp_demo.ipynb)](https://github.com/TajikNLPWorld/tajiknlp/tree/main/demo/tajiknlp_demo.ipynb)

---

## 🚀 Quick Start

### Load the default pipeline

```python
from tajiknlp import load_pipeline

pipe = load_pipeline("default")
text = "Китобҳоямонро хондам, аммо нафаҳмидам."
doc = pipe(text)

for token in doc.tokens:
    print(f"{token.text:<15} → lemma: {token.lemma:<10} pos: {token.pos}")
```

**Output:**
```
китобҳоямонро   → lemma: китоб      pos: NOUN
хондам          → lemma: хондан     pos: VERB
,               → lemma: ,          pos: PUNCT
аммо            → lemma: аммо       pos: CONJ
нафаҳмидам      → lemma: фаҳмидан   pos: VERB
.               → lemma: .          pos: PUNCT
```

### Available Pipeline Presets

| Preset     | Components                                                       |
|------------|------------------------------------------------------------------|
| `minimal`  | Cleaner + Tokenizer + Morpheme splitter                          |
| `default`  | Full preprocessing + POS + Morphemes + Stopwords + Lemmatizer    |
| `full`     | Alias for `default`                                              |
| `stemming` | Same as default but with stemmer instead of lemmatizer           |
| `ner`      | Optimized for NER (preserves case)                               |

### Use individual components

```python
from tajiknlp import Doc
from tajiknlp.components.tokenizers import RegexTokenizer

tokenizer = RegexTokenizer(lowercase=True, keep_punct=True)
doc = tokenizer(Doc(text="Китоб-ҳо хондам!"))
print([t.text for t in doc.tokens])
# Output: ['китоб-ҳо', 'хондам', '!']
```

### Validate and score text quality

```python
from tajiknlp import quality_score, detect_script, check_normalization

score = quality_score("Ман китоб хондам.")
print(f"Score: {score['score']}")        # e.g., 1.00
print(f"Valid: {score['is_valid']}")     # True
print(f"Language: {score['language']}")  # tajik

script = detect_script("Салом!")
print(script)  # Script.CYRILLIC

norm_check = check_normalization("Салом , дӯст!")
print(norm_check["issues"])  # ['Contains space before punctuation.']
```

### Evaluate predictions with metrics

```python
from tajiknlp import precision_recall_f1, bleu_score

pred = ["NOUN", "VERB", "PUNCT"]
gold = ["NOUN", "ADJ",  "PUNCT"]
metrics = precision_recall_f1(pred, gold)
print(metrics)  # {'precision': 0.67, 'recall': 0.67, 'f1': 0.67}

bleu = bleu_score("Ман дар Душанбе", ["Ман дар Душанбе"])
print(bleu)  # 1.0
```

### Custom pipeline

```python
from tajiknlp import TajikPipeline
from tajiknlp.components.cleaners import TextCleaner
from tajiknlp.components.tokenizers import RegexTokenizer
from tajiknlp.components.stemmers import DictStemmer

pipe = TajikPipeline()
pipe.add_component(TextCleaner())
pipe.add_component(RegexTokenizer())
pipe.add_component(DictStemmer(deep_stemming=True))

doc = pipe("Ман китобҳоро хондам.")
```

---

## 📖 Documentation

Full documentation is available at **[TajikNLPWorld.github.io/tajiknlp](https://TajikNLPWorld.github.io/tajiknlp)**.

It includes:
- Detailed API references
- Component configuration guides
- Advanced usage examples
- NER and alignment tutorials

---

## 📁 Project Structure

```
tajiknlp/
├── src/tajiknlp/           # Main package source
│   ├── alignment/          # Span alignment utilities
│   ├── components/         # NLP components
│   │   ├── cleaners/       # Text cleaning
│   │   ├── embeddings/     # Word embeddings
│   │   ├── filters/        # Token filtering
│   │   ├── lemmatizers/    # Lemmatization
│   │   ├── normalizers/    # Text normalization
│   │   ├── sentencizers/   # Sentence splitting
│   │   ├── stemmers/       # Stemming
│   │   ├── taggers/        # POS tagging
│   │   └── tokenizers/     # Tokenization
│   ├── core/               # Base classes
│   ├── data/               # Static resources
│   ├── pipeline/           # Pipeline system
│   ├── resources/          # Resource management
│   └── utils/              # Utility functions
├── tests/                  # Test suite
├── examples/               # Usage examples
├── docs/                   # Documentation
└── pyproject.toml          # Project configuration
```

---

## 🧪 Running Tests

If you cloned the repository, you can run the test suite with:

```bash
# Install development dependencies
pip install -e ".[dev]"

# Run all tests
pytest tests/ -v

# Run with coverage report
pytest tests/ --cov=tajiknlp --cov-report=html
```

---

## 🤝 Contributing

We welcome contributions! Please see [CONTRIBUTING.md](CONTRIBUTING.md) for guidelines on:

- Setting up the development environment
- Coding standards and style guides
- Testing requirements
- Pull request process

---

## 📄 License

This project is licensed under the MIT License – see the [LICENSE](LICENSE) file for details.

Copyright (c) 2026 TajikNLPWorld
Copyright (c) 2026 Arabov Mullosharaf Kurbonovich

---

## 🙏 Acknowledgements

TajikNLP was inspired by the need for robust, open‑source NLP tools for low‑resource languages. Special thanks to all contributors and the Tajik linguistic community.

---

## 📊 Citation

If you use TajikNLP in your research, please cite:

```bibtex
@software{tajiknlp2026,
  author = {Arabov, Mullosharaf Kurbonovich},
  title = {TajikNLP: Production-ready NLP library for Tajik language},
  year = {2026},
  publisher = {GitHub},
  url = {https://github.com/TajikNLPWorld/tajiknlp}
}
```

---

**Made with ❤️ for the Tajik language.**
