Metadata-Version: 2.4
Name: langid-chat
Version: 0.1.0
Summary: A portable language detector (15 languages incl. code-mixed Indian) with a chat web UI
Author: orewamash
License: MIT
Project-URL: Homepage, https://github.com/orewamash/langid
Keywords: language-detection,nlp,machine-learning,chatbot
Classifier: Programming Language :: Python :: 3
Classifier: Natural Language :: English
Classifier: Topic :: Text Processing :: Linguistic
Requires-Python: >=3.9
Description-Content-Type: text/markdown
Requires-Dist: scikit-learn>=1.2
Requires-Dist: numpy

# Language Detection AI

A classical machine-learning language detector: feed it a sentence, it tells
you which language it is written in. Built with Python + scikit-learn using
character n-grams + Naive Bayes / Linear SVC.

Includes detection of **Roman-script code-mixed Indian languages** (Tanglish,
Hinglish, Teluglish, Kanglish, Manglish).

## Install & run the chatbot UI

The UI ships as a pip-installable package (`langid-chat`) that bundles the
trained model, so anyone can run it with a single command — no repo clone
required.

```bash
pip install langid-chat

# Launch the chat UI (http://127.0.0.1:8000)
langid-web

# Or use the CLI
langid "Kya kar rahe ho bhai"
langid --interactive
```

`langid-web` supports `--host` and `--port` (e.g. `langid-web --port 9000`).

To install from a source checkout instead:

```bash
pip install -e .
langid-web
# or without installing at all:
python web/app.py
```

## Quick start

```bash
# 0. (Optional) Re-fetch the code-mixed corpora into data/raw (cached by default)
python src/fetch_codemix.py

# 1. Prepare data -- the deployed model is trained on the FULL dataset with
#    balanced class weights (keeps all code-mixed signal):
python src/preprocess.py --no-balance
python src/train.py --class-weight balanced

# 2. Test new sentences
python detect.py "Kya kar rahe ho bhai"
python detect.py --interactive
```

> On Windows set `$env:PYTHONIOENCODING='utf-8'` first so multilingual text
> prints correctly.

`detect.py` also reports a **confidence score** per guess. Below a calibrated
threshold it says `LOW CONFIDENCE` instead of guessing — short or ambiguous
input (e.g. `ok`, `vanakkam`) gets flagged rather than answered wrongly.
Tune with `--min-confidence 0.5`, or force an always-guess with `0`.

## Languages supported (15)

Tamil, Russian, Arabic (distinct scripts) · Spanish, Portuguese, French,
Italian (linguistically close) · English, German, Dutch · plus code-mixed:
Tanglish, Hinglish, Teluglish, Kanglish, Manglish.

## Results

Best model (LinearSVC, full data + balanced class weights) reaches **~98% test
accuracy** across the 15 classes; script-distinct languages (Arabic, Russian,
Tamil) hit ~100%. High-confidence calls (>= 45%) are ~99% correct on the test
split; the only reliable misdetections left are 1–2 word Roman-script
fragments, which are exactly the ones the confidence gate flags.

**Limitations:** it only knows the 15 trained languages and still guesses on
anything fed to it without confidences if you lower the threshold. Very short
text is inherently unreliable and is flagged, not silently misanswered.

See `results/results.md` for the full write-up and `project.md`/`tasks.md` for
scope and tasks.
