Metadata-Version: 2.4
Name: maltopic
Version: 1.5.2
Summary: A multi-agent LLM topic modeling library.
License: MIT
License-File: LICENSE
Keywords: topic modeling,LLM,multi-agent,survey analysis,data enrichment,topic deduplication,semantic analysis
Author: Yash Sharma
Author-email: yash91sharma@gmail.com
Requires-Python: >=3.12,<4.0
Classifier: License :: OSI Approved :: MIT License
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Programming Language :: Python :: 3.15
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Dist: openai (>=1.79.0,<2.0.0)
Requires-Dist: pandas (>=2.2.3,<3.0.0)
Requires-Dist: streamlit (>=1.28,<2.0)
Requires-Dist: tiktoken (>=0.9.0,<0.10.0)
Requires-Dist: tqdm (>=4.67.1,<5.0.0)
Project-URL: Homepage, https://github.com/yash91sharma/MALTopic-py
Project-URL: Repository, https://github.com/yash91sharma/MALTopic-py
Description-Content-Type: text/markdown

# MALTopic: Multi-Agent LLM Topic Modeling Library

[![PyPI Version](https://img.shields.io/pypi/v/maltopic.svg)](https://pypi.org/project/maltopic/)
[![License: MIT](https://img.shields.io/badge/License-MIT-green.svg)](https://opensource.org/licenses/MIT)
[![DOI: 10.1109/WorldAI-IoT65487.2025.11105319](https://img.shields.io/badge/DOI-10.1109%2FWorldAI--IoT65487.2025.11105319-blue.svg)](https://ieeexplore.ieee.org/document/11105319)

**MALTopic** is a Python library designed for topic modeling using a coordinated multi-agent LLM framework. First framework to enhance the analysis of survey responses and open-ended text by integrating structured and categorical metadata with unstructured free-text responses.

The foundational research [**MALTopic paper**](https://ieeexplore.ieee.org/document/11105319).

---

## Overview & Architecture

Traditional topic modeling approaches (such as LDA or BERTopic) analyze text in isolation and often miss demographic, behavioral, or categorical context. MALTopic decomposes topic extraction into agents working collaboratively:

```
┌────────────────────────────────────────────────────────────────────────┐
│                          Raw Survey Data (CSV)                         │
│             [ Free-Text Column + Structured Metadata Columns ]          │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                        1. DATA ENRICHMENT AGENT                        │
│   Injects demographic/structured features into free text responses     │
│       without hallucinations or sentiment distortion. Generates:       │
│                     `{free_text_column}_enriched`                      │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                         2. TOPIC MINING AGENT                          │
│   Extracts structured topics (name, description, relevance, keywords)   │
│   • Auto-batching via tiktoken token counting for large datasets       │
│   • Automatic fallback on context-window errors                        │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                     3. TOPIC DEDUPLICATION AGENT                       │
│    Semantically reconciles & merges overlapping topics (>80% overlap)   │
│             into concise, comprehensive, non-redundant themes           │
└───────────────────────────────────┬────────────────────────────────────┘
                                    │
                                    ▼
┌────────────────────────────────────────────────────────────────────────┐
│                             Final Outputs                              │
│   • Enriched DataFrame (CSV)       • Clean Topics JSON                 │
│   • Usage Statistics (Tokens, Latency, Calls, Cost Breakdown)          │
└────────────────────────────────────────────────────────────────────────┘
```

---

## Installation

Install MALTopic via `pip`:

```bash
pip install maltopic
```

Or using `poetry`:

```bash
poetry add maltopic
```

### Dependencies
MALTopic is built for Python **3.12+** and **3.13+** and depends on:
- `openai>=1.79.0`
- `pandas>=2.2.3`
- `tiktoken>=0.9.0`
- `tqdm>=4.67.1`
- `streamlit>=1.28`

---

## Quickstart Tutorial

Here is a complete, copy-pasteable example analyzing customer survey feedback:

```python
import os
import pandas as pd
from maltopic import MALTopic

# 1. Prepare sample survey data
df = pd.DataFrame({
    "customer_id": [101, 102, 103, 104],
    "plan_tier": ["Premium", "Basic", "Enterprise", "Basic"],
    "region": ["North America", "Europe", "North America", "Asia-Pacific"],
    "feedback": [
        "The checkout process is smooth, but I wish customer support answered faster on weekends.",
        "Mobile app crashes frequently whenever I attempt to upload receipt attachments.",
        "Enterprise API response times are fantastic, but billing documentation is confusing.",
        "Navigation is intuitive, but pages take too long to load on cellular data.",
    ]
})

# 2. Initialize MALTopic client
client = MALTopic(
    api_key=os.environ.get("OPENAI_API_KEY", "your-api-key"),
    default_model_name="gpt-4o-mini",
    llm_type="openai",
)

# 3. Step 1: Enrich free-text feedback with categorical metadata
enriched_df = client.enrich_free_text_with_structured_data(
    survey_context="Quarterly product feedback survey assessing customer satisfaction across subscription tiers.",
    free_text_column="feedback",
    structured_data_columns=["plan_tier", "region"],
    df=df,
    examples=[
        "Checkout is slow, Basic, Europe -> A Basic tier user in Europe experienced delays during checkout."
    ],
)

# 4. Step 2: Generate latent topics from enriched responses
topics = client.generate_topics(
    topic_mining_context="Identify core pain points, product feature requests, and service feedback.",
    df=enriched_df,
    enriched_column="feedback_enriched",
)

# 5. Step 3: Semantically deduplicate and consolidate similar topics
clean_topics = client.deduplicate_topics(
    topics=topics,
    survey_context="Customer feedback survey across multiple subscription tiers.",
)

# 6. Display results
for i, topic in enumerate(clean_topics, 1):
    print(f"\n[{i}] {topic['name']}")
    print(f"    Description: {topic['description']}")
    print(f"    Relevance:   {topic['relevance']}")
    print(f"    Keywords:    {', '.join(topic['representative_words'])}")

# 7. Print comprehensive token and latency telemetry
client.print_stats()
```

---

## GUI (No-Code Web Interface)

MALTopic comes bundled with a modern, interactive web application built with Streamlit.

### Launching the GUI

After installing `maltopic`, launch the GUI directly from your terminal:

```bash
maltopic-gui
```

Or via Python:

```bash
python -m maltopic.gui
```

### 6-Step Visual Pipeline Wizard

1. **Step 1: Configure API**  
   Enter your OpenAI API key, target model name (e.g. `gpt-4o`, `gpt-4o-mini`, `o1-mini`), and optional custom parameter overrides. Your key is stored only in transient session memory.
2. **Step 2: Upload Data**  
   Upload any CSV file (up to 200 MB), preview rows, and select the free-text column alongside one or more structured metadata columns.
3. **Step 3: Enrich Text**  
   Define the survey context and optional few-shot examples. Watch the real-time progress bar as responses are contextualized. Download the enriched CSV at any time.
4. **Step 4: Generate Topics**  
   Input topic mining guidance. Topics are extracted and displayed as responsive visual cards displaying topic titles, descriptions, target relevance groups, and representative keyword tags.
5. **Step 5: Deduplicate Topics (Optional)**  
   Optionally run semantic deduplication to merge overlapping themes, with before-and-after comparison metrics.
6. **Step 6: Results & Export**  
   Inspect the final topic cards, view the interactive usage statistics dashboard (total tokens, call counts, success rate, response times, model breakdown), and export:
   - Topics as JSON (`maltopic_topics.json`)
   - Enriched dataset as CSV (`maltopic_enriched.csv`)
   - Session telemetry as JSON (`maltopic_stats.json`)

---

## Core Multi-Agent Pipeline

### 1. Data Enrichment Agent

Survey respondents frequently submit terse or ambiguous free-text (e.g., *"Takes too long"* or *"Support was unhelpful"*). The **Data Enrichment Agent** combines free-text with respondent metadata (demographics, subscription levels, satisfaction scores) to build a unified context string.

- **Sentiment Preservation**: The agent maintains the original sentiment and does not inject unsolicited assumptions or extrapolations.
- **Output Column**: Creates a new column named `{free_text_column}_enriched` in the returned DataFrame.
- **Few-Shot Prompting**: Supports optional domain-specific examples via the `examples` argument.

### 2. Topic Mining Agent & Automatic Batching

The **Topic Mining Agent** synthesizes responses to generate structured topic profiles adhering to four strict criteria:
- **Uniqueness**: Distinct concepts without redundant definitions.
- **Exhaustiveness**: Complete coverage of issues raised across all responses.
- **Non-Overlapping**: Clear boundaries between extracted themes.
- **Respondent-Awareness**: Explicitly highlights patterns across respondent segments.

#### Automatic Token Limit Batching
When dealing with thousands of survey responses, prompts can exceed the LLM's maximum context length. MALTopic handles this automatically:
1. **Error Interception**: Automatically catches token limit exceptions (`maximum context length`, `context window`, `too many tokens`).
2. **Dynamic Splitting**: Uses `tiktoken` to partition responses into optimal batches under token limits (default: 100k tokens per batch, with a 2,000-token instruction buffer).
3. **Batch Processing**: Displays a visual `tqdm` progress bar as each batch is processed independently.
4. **Topic Consolidation**: Merges and consolidates extracted topics into a unified set.

### 3. Semantic Topic Deduplication Agent

Traditional topic deduplication relies on exact keyword matching or string similarity (e.g., Levenshtein distance). MALTopic's **Deduplication Agent** performs **semantic understanding**:
- Detects topics with >80% conceptual overlap even if names use completely different words (e.g. *"Slow App Performance"* and *"High Latency & Lag"*).
- Merges descriptions and relevance scopes to avoid losing subtle respondent nuances.
- Deduplicates and combines representative word lists.
- Preserves genuinely distinct topics unaltered.
- Fails gracefully: If deduplication encounters an API issue, it returns the original topics without interrupting your pipeline.

### 4. Telemetry & Statistics Tracking

Every call to `MALTopic` is monitored in real-time by an internal `MALTopicStats` instance.

#### Tracked Metrics
- **Token Counts**: Input/prompt tokens, output/completion tokens, and total tokens.
- **Call Volume**: Successful calls, failed calls, total calls made.
- **Success Rate**: Exact percentage of successful calls (`0.0%` to `100.0%`).
- **Latency & Throughput**: Average response time per call and total session uptime.
- **Per-Model Breakdowns**: Token consumption and call statistics categorized per model.
- **Recent Call History**: Log of the last API calls with timestamps, latency, and status.

#### Estimating API Costs
You can use `client.get_stats()` to calculate exact API costs:

```python
stats = client.get_stats()
overview = stats["overview"]

# Example: calculate cost for gpt-4o ($2.50 per 1M input tokens, $10.00 per 1M output tokens)
input_cost = (overview["total_input_tokens"] / 1_000_000) * 2.50
output_cost = (overview["total_output_tokens"] / 1_000_000) * 10.00
total_cost = input_cost + output_cost

print(f"Total API Cost: ${total_cost:.4f}")
```

---

## Model Compatibility & Parameter Control

### Standard Chat Models vs. Reasoning Models

OpenAI provides two distinct classes of chat completion models:

| Model Category | Examples | Parameter Restrictions |
|---|---|---|
| **Standard Models** | `gpt-4`, `gpt-4o`, `gpt-4o-mini`, `gpt-4.1-nano` | Supports `temperature`, `top_p`, `seed` |
| **Reasoning Models** | `o1`, `o1-mini`, `o1-preview`, `o3`, `o3-mini`, `gpt-5`, `gpt-5-mini` | Rejects `temperature`, `top_p`, `seed`, `presence_penalty` (HTTP 400 if passed) |

MALTopic includes **automatic model detection**:
- When calling standard models with `override_model_params=None` (the default), MALTopic automatically sets `temperature=0.2`, `top_p=0.9`, and `seed=12345` for reproducibility.
- When calling reasoning models (`o1`, `o3`, `gpt-5`), MALTopic automatically omits these parameters to prevent API errors.

### Custom Parameter Overrides

You can pass custom parameters via `override_model_params` during `MALTopic` initialization:

```python
# Custom sampling configuration (completely replaces defaults)
client = MALTopic(
    api_key="your_api_key",
    default_model_name="gpt-4o",
    llm_type="openai",
    override_model_params={
        "temperature": 0.7,
        "max_tokens": 1000,
        "frequency_penalty": 0.2,
    }
)
```

To send only the bare minimum parameters (`model`, `messages`, `store=False`), pass an empty dictionary:

```python
client = MALTopic(
    api_key="your_api_key",
    default_model_name="o1-mini",
    llm_type="openai",
    override_model_params={},
)
```

---

## Complete Python API Reference

### `MALTopic` Class

```python
from maltopic import MALTopic
```

#### `MALTopic(api_key, default_model_name, llm_type, override_model_params=None)`
Initializes the MALTopic pipeline client.

- **Parameters**:
  - `api_key` (*str*): API credentials for authentication.
  - `default_model_name` (*str*): Model identifier (e.g., `"gpt-4o"`, `"o1-mini"`).
  - `llm_type` (*str*): LLM provider (currently `"openai"`).
  - `override_model_params` (*dict[str, Any] | None*, optional): Optional raw parameters dictionary to send to the provider API. Default: `None` (automatic parameter selection).

#### `enrich_free_text_with_structured_data(survey_context, free_text_column, structured_data_columns, df, examples=[]) -> pd.DataFrame`
Enriches unstructured survey responses with structured column values.

- **Parameters**:
  - `survey_context` (*str*): Background description of the survey.
  - `free_text_column` (*str*): Column containing raw text responses.
  - `structured_data_columns` (*list[str]*): Column names containing structured metadata.
  - `df` (*pd.DataFrame*): Input Pandas DataFrame.
  - `examples` (*list[str]*, optional): Few-shot examples illustrating the enrichment format.
- **Returns**: `pd.DataFrame` containing the original data plus `{free_text_column}_enriched`.

#### `generate_topics(topic_mining_context, df, enriched_column) -> list[dict[str, Any]]`
Extracts unique and non-overlapping topics from enriched text.

- **Parameters**:
  - `topic_mining_context` (*str*): Instructions describing desired topic focus.
  - `df` (*pd.DataFrame*): DataFrame containing enriched text.
  - `enriched_column` (*str*): Column name containing enriched responses.
- **Returns**: `list[dict[str, Any]]` where each dictionary has:
  - `"name"` (*str*): Topic name.
  - `"description"` (*str*): Explanation of the topic.
  - `"relevance"` (*str*): Target respondent profiles and patterns.
  - `"representative_words"` (*list[str]*): Representative NLP keywords.

#### `deduplicate_topics(topics, survey_context) -> list[dict[str, Any]]`
Semantically merges overlapping topics into unified themes.

- **Parameters**:
  - `topics` (*list[dict[str, Any]]*): List of topic dictionaries to deduplicate.
  - `survey_context` (*str*): Background context to guide merging decisions.
- **Returns**: `list[dict[str, Any]]` matching the topic schema.

#### `get_stats() -> dict[str, Any]`
Returns a dictionary containing session telemetry (`overview`, `averages`, `model_breakdown`, `recent_calls`).

#### `print_stats() -> None`
Prints formatted statistics directly to stdout.

#### `reset_stats() -> None`
Resets all tracked metrics to zero.

---

### `MALTopicStats` Class

```python
from maltopic import MALTopicStats
```

| Property / Method | Return Type | Description |
|---|---|---|
| `total_tokens_used` | `int` | Total tokens consumed across successful calls |
| `total_input_tokens` | `int` | Total prompt/input tokens |
| `total_output_tokens` | `int` | Total completion/output tokens |
| `total_calls_made` | `int` | Total calls attempted (`successful + failed`) |
| `successful_calls` | `int` | Number of successful calls |
| `failed_calls` | `int` | Number of failed calls |
| `success_rate` | `float` | Percentage of successful calls (`0.0` - `100.0`) |
| `average_tokens_per_call` | `float` | Mean tokens per successful call |
| `average_response_time` | `float` | Mean execution latency in seconds |
| `uptime` | `float` | Total tracker uptime in seconds |
| `get_model_breakdown()` | `dict[str, dict[str, Any]]` | Statistics partitioned by model name |
| `get_recent_calls(limit=10)` | `list[LLMCallStats]` | Recent call records |
| `get_summary()` | `dict[str, Any]` | Full nested telemetry dictionary |
| `print_summary()` | `None` | Prints summary table to stdout |
| `reset()` | `None` | Resets all telemetry |

---

### `LLMCallStats` Dataclass

```python
from maltopic import LLMCallStats
```

Individual call record stored in `stats.get_recent_calls()`:
- `timestamp` (*float*): Unix epoch timestamp.
- `model_name` (*str*): Model invoked.
- `success` (*bool*): Whether the call succeeded.
- `input_tokens` (*int*): Input tokens used.
- `output_tokens` (*int*): Output tokens generated.
- `total_tokens` (*int*): Total tokens consumed.
- `response_time` (*float*): Latency in seconds.
- `error_message` (*str | None*): Error string if failed.
- `metadata` (*dict[str, Any]*): Response metadata (e.g. `finish_reason`, `system_fingerprint`).

---

### Utility Functions (`maltopic.utils`)

```python
from maltopic import utils
```

- `validate_dataframe(df: pd.DataFrame, required_columns: list[str]) -> None`: Verifies non-empty DataFrame and presence of required columns.
- `count_tokens(text: str, model_name: str = "gpt-4") -> int`: Calculates exact tokens using `tiktoken`.
- `split_text_into_batches(labeled_responses: list[str], max_tokens_per_batch: int = 100000, model_name: str = "gpt-4") -> list[list[str]]`: Splits text into token-bounded batches.
- `is_token_limit_error(error: Exception) -> bool`: Identifies whether an exception is due to context window limits.
- `validate_topic_structure(topics: list[dict[str, Any]]) -> None`: Validates topic dictionary schema.
- `parse_topics_response(raw_response: str) -> list[dict[str, Any]]`: Parses and cleans LLM JSON response.
- `consolidate_topics(all_topics: list[dict[str, Any]]) -> list[dict[str, Any]]`: Fast deduplication by topic name.

---

## Data Privacy & Local Execution

MALTopic is engineered with data privacy in mind:
- **No Telemetry**: MALTopic does not transmit telemetry, analytics, or usage data to third parties.
- **Local Computing**: All DataFrame manipulation, token counting, batching, and GUI rendering occur 100% locally on your machine.
- **Provider Only**: The only network requests made are directly between your machine and your chosen LLM provider (e.g., OpenAI) using your own API key.
- **Ephemeral Credentials**: API keys entered into the web GUI are kept strictly in session memory and are never persisted to disk or logs.

---

## Troubleshooting & Best Practices

### Token Limit & Context Window Errors
- If your dataset is very large, MALTopic automatically detects context window exceedances and switches to batching.
- If you encounter rate limits (`429 Too Many Requests`), try setting `max_tokens` or choosing a tier-appropriate model (e.g. `gpt-4o-mini`).

### Quality of Extracted Topics
- Provide detailed `survey_context` explaining **why** the survey was conducted and **who** the audience is.
- In `topic_mining_context`, specify the granularity desired (e.g. *"Focus on actionable UX issues and billing complaints; avoid generic praise"*).

### Reasoning Models (o1, o3, GPT-5)
- Do not pass `temperature`, `top_p`, or `seed` to reasoning models. Leave `override_model_params=None` and MALTopic will automatically filter them out for you.

---

## Local Development & Testing

We welcome contributions to MALTopic! Follow these steps to set up your local development environment:

### 1. Clone the repository
```bash
git clone https://github.com/yash91sharma/MALTopic-py.git
cd MALTopic-py
```

### 2. Set up virtual environment
```bash
python3.13 -m venv .venv
source .venv/bin/activate
pip install -e .
pip install pytest mypy ruff
```

### 3. Run the test suite
```bash
pytest
```
All 151 unit and integration tests run in under 2 seconds.

### 4. Run type checking
```bash
mypy src tests
```

### 5. Run linter and formatting
```bash
ruff check src tests
ruff format --check src tests
```

---

## Changelog

For detailed release history and migration notes, see [CHANGELOG.md](CHANGELOG.md).

---

## Contributing

Contributions are warmly welcomed! Please submit pull requests or file issues on [GitHub](https://github.com/yash91sharma/MALTopic-py/issues).

---

## License

This project is licensed under the **MIT License**. See the [LICENSE](LICENSE) file for complete details.

---

## Citation

If you use MALTopic in your academic research or applications, please cite our IEEE publication:

```bibtex
@inproceedings{sharma2025maltopic,
  author    = {Sharma, Yash},
  title     = {MALTopic: A Multi-Agent LLM Framework for Survey Analysis and Context-Aware Topic Modeling},
  booktitle = {2025 World AI-IoT Congress},
  year      = {2025},
  doi       = {10.1109/WorldAI-IoT65487.2025.11105319},
  url       = {https://ieeexplore.ieee.org/document/11105319}
}
```

