Metadata-Version: 2.4
Name: sonarwise
Version: 0.2.0
Summary: Pluggable audio perception engine. Hear. Search. Retrieve.
Author-email: Venkatkumar Rajan <venkatkumarr.vk99@gmail.com>
License: Apache 2.0
Project-URL: Homepage, https://github.com/VK-Ant/sonarwise
Project-URL: Documentation, https://github.com/VK-Ant/sonarwise#readme
Project-URL: Repository, https://github.com/VK-Ant/sonarwise
Project-URL: Issues, https://github.com/VK-Ant/sonarwise/issues
Keywords: audio,speech,retrieval,RAG,transcription,diarization,speaker-recognition,audio-events,whisper,CLAP,perception,search,embeddings,voice-activity-detection,sound-classification
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: Apache Software License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.9
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Topic :: Multimedia :: Sound/Audio :: Analysis
Classifier: Topic :: Multimedia :: Sound/Audio :: Speech
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.9
Description-Content-Type: text/markdown
License-File: LICENSE
Requires-Dist: numpy>=1.24.0
Requires-Dist: tqdm>=4.65.0
Requires-Dist: pyyaml>=6.0
Provides-Extra: whisper
Requires-Dist: openai-whisper>=20231117; extra == "whisper"
Provides-Extra: faster-whisper
Requires-Dist: faster-whisper>=1.0.0; extra == "faster-whisper"
Provides-Extra: clap
Requires-Dist: transformers>=4.30.0; extra == "clap"
Requires-Dist: torch>=2.0.0; extra == "clap"
Requires-Dist: torchaudio>=2.0.0; extra == "clap"
Provides-Extra: diarization
Requires-Dist: pyannote.audio>=3.1.0; extra == "diarization"
Requires-Dist: torch>=2.0.0; extra == "diarization"
Provides-Extra: speaker
Requires-Dist: speechbrain>=1.0.0; extra == "speaker"
Requires-Dist: torch>=2.0.0; extra == "speaker"
Provides-Extra: events
Requires-Dist: torch>=2.0.0; extra == "events"
Requires-Dist: torchaudio>=2.0.0; extra == "events"
Provides-Extra: preprocess
Requires-Dist: scipy>=1.10.0; extra == "preprocess"
Provides-Extra: live
Requires-Dist: sounddevice>=0.4.6; extra == "live"
Provides-Extra: vad
Requires-Dist: silero-vad>=5.0.0; extra == "vad"
Requires-Dist: torch>=2.0.0; extra == "vad"
Provides-Extra: all
Requires-Dist: openai-whisper>=20231117; extra == "all"
Requires-Dist: transformers>=4.30.0; extra == "all"
Requires-Dist: torch>=2.0.0; extra == "all"
Requires-Dist: torchaudio>=2.0.0; extra == "all"
Requires-Dist: pyannote.audio>=3.1.0; extra == "all"
Requires-Dist: speechbrain>=1.0.0; extra == "all"
Requires-Dist: sounddevice>=0.4.6; extra == "all"
Requires-Dist: silero-vad>=5.0.0; extra == "all"
Requires-Dist: scipy>=1.10.0; extra == "all"
Provides-Extra: dev
Requires-Dist: pytest>=7.0; extra == "dev"
Requires-Dist: pytest-cov>=4.0; extra == "dev"
Requires-Dist: ruff>=0.1.0; extra == "dev"
Dynamic: license-file

<p align="center">
  <img src="https://raw.githubusercontent.com/VK-Ant/sonarwise/main/assets/sonarwise_ai.png" alt="sonarwise" width="800"/>
</p>

<h1 align="center">sonarwise</h1>
<p align="center"><b>Pluggable audio perception engine. Hear. Search. Retrieve.</b></p>

<p align="center">
  <a href="https://pypi.org/project/sonarwise"><img src="https://img.shields.io/pypi/v/sonarwise?color=blue" alt="PyPI"/></a>
  <a href="https://github.com/VK-Ant/sonarwise/blob/main/LICENSE"><img src="https://img.shields.io/badge/license-Apache%202.0-blue" alt="License"></a>
  <a href="https://www.python.org/"><img src="https://img.shields.io/badge/python-3.9+-green" alt="Python"></a>
    <a href="https://github.com/VK-Ant/sonarwise/blob/main/demos/Sonarwise_demo.ipynb">
        <img src="https://colab.research.google.com/assets/colab-badge.svg" alt="Open In Colab">
</p>

---

**sonarwise** indexes any audio — meetings, calls, podcasts, factory floors, lectures — and makes it searchable by text, speaker, sound events, or audio similarity. Every component is pluggable: swap transcription, embedding, diarization, or storage without changing your code.

## What's New in v0.2.0

- **Audio Preprocessing Pipeline** — DC removal, bandpass filtering, spectral-gating noise reduction, normalization, dynamic range compression
- **Language Detection** — Whisper, faster-whisper, or script-based heuristic detection for 99 languages
- **CLAP Event Classifier** — Zero-shot audio event classification using CLAP text-audio similarity
- **Conversation Analytics** — Turn-taking stats, talk-time ratios, interruption detection, speaker overlap analysis
- **Audio Clip Extraction** — Extract segments to WAV files by time range, speaker, or search results
- **Webhook Notifications** — HTTP callbacks on indexing completion, keyword detection, event triggers
- **Health Dashboard** — Pipeline health monitoring, component status, memory/latency tracking
- **Batch Operations** — Batch segment insertion for faster indexing of large files
- **Improved Robustness** — Thread-safe SQLite store, batched search (prevents OOM on large datasets), input validation, error recovery

## Features

- **Transcription** — Whisper, Faster Whisper, or bring your own ASR
- **Audio Embeddings** — CLAP joint text-audio space for semantic search
- **Speaker Diarization** — know who said what (pyannote)
- **Speaker Registry** — track speakers across files by voiceprint
- **Audio Event Detection** — detect alarms, machinery, glass breaks, and 50+ sound types
- **Audio Preprocessing** — noise reduction, normalization, bandpass filtering, compression
- **Language Detection** — automatic spoken language identification (99 languages)
- **Live Streaming** — real-time transcription with keyword and event callbacks
- **Conversation Analytics** — turn-taking, talk-time, interruptions, speaker overlap
- **Clip Extraction** — extract audio segments to WAV by time, speaker, or search
- **Export** — SRT, VTT, JSON, CSV, TXT, meeting notes
- **Pluggable Architecture** — every component swappable via base classes
- **CLI** — full command-line interface

## Install

```bash
# Core (no ML dependencies)
pip install sonarwise

# With all features
pip install sonarwise[all]

# Pick what you need
pip install sonarwise[whisper]          # Whisper transcription
pip install sonarwise[faster-whisper]   # Faster Whisper (CTranslate2)
pip install sonarwise[clap]            # CLAP audio embeddings
pip install sonarwise[diarization]     # Speaker diarization
pip install sonarwise[speaker]         # Speaker identification
pip install sonarwise[events]          # Audio event detection (PANNs)
pip install sonarwise[preprocess]      # Audio preprocessing (scipy)
pip install sonarwise[live]            # Live microphone capture
pip install sonarwise[vad]             # Silero VAD chunking
```

**Requires:** ffmpeg (`sudo apt install ffmpeg` or `brew install ffmpeg`)

## Quick Start

```python
from sonarwise import SonarWise

sw = SonarWise(diarization=True, events=True)

# Index audio
sw.index("meeting.wav")
sw.index_folder("./recordings/")

# Search by text
results = sw.query("budget discussion", top_k=5)
for r in results:
    print(f"[{r.speaker_name}] {r.transcript} (score: {r.score})")

# Search by audio similarity
results = sw.query_audio("alarm_clip.wav", top_k=5)

# Search by speaker
results = sw.query("budget", speaker="Ant", top_k=5)

# Search by event
results = sw.query_events(event="machine_fault", top_k=5)
```

## Audio Preprocessing (New in v0.2.0)

Clean and enhance audio before processing:

```python
from sonarwise.core.preprocessor import AudioPreprocessor, NoiseReducer, Normalizer
from sonarwise import AudioData

# Full pipeline
preprocessor = AudioPreprocessor(
    remove_dc=True,          # Remove DC offset
    bandpass=True,           # Bandpass filter (80Hz-7500Hz)
    noise_reduce=True,       # Spectral gating noise reduction
    normalize=True,          # Peak normalization
    compress=True,           # Dynamic range compression
)

# Or use convenience wrappers
reducer = NoiseReducer(strength=1.5)
normalizer = Normalizer(target_peak=0.95)
```

## Language Detection (New in v0.2.0)

Identify spoken language before or after transcription:

```python
from sonarwise.core.language_detector import (
    SimpleLanguageDetector,
    WhisperLanguageDetector,
    FasterWhisperLanguageDetector,
)

# Script-based detection (zero dependencies)
detector = SimpleLanguageDetector()
result = detector.detect_from_text("Bonjour le monde")
print(f"{result.language_name}: {result.confidence}")  # French: 0.5

# Whisper-based detection (highly accurate)
detector = WhisperLanguageDetector(model_size="base")
result = detector.detect(audio_segment)
print(f"{result.language_name}: {result.confidence}")  # French: 0.97
```

## Conversation Analytics (New in v0.2.0)

```python
# After indexing a meeting
analytics = sw.conversation_analytics("meeting.wav")

print(f"Total speakers: {analytics['total_speakers']}")
print(f"Total turns: {analytics['total_turns']}")
for speaker, stats in analytics['speaker_stats'].items():
    print(f"  {speaker}: {stats['talk_time_pct']:.1f}% talk time, "
          f"{stats['turn_count']} turns")
```

## Clip Extraction (New in v0.2.0)

```python
# Extract a specific time range
sw.extract_clip("meeting.wav", start_ms=30000, end_ms=60000,
                output="clip_30s_60s.wav")

# Extract all segments from a speaker
sw.extract_clip("meeting.wav", speaker="Ant",
                output="ant_segments.wav")
```

## Speaker Intelligence

```python
# Register a speaker
sw.register_speaker("Ant", reference_audio="ant_voice.wav")

# Query by speaker
results = sw.query("budget", speaker="Ant")

# Speaker timeline
timeline = sw.speaker_timeline("meeting.wav")

# Speaker stats
stats = sw.speaker_stats("meeting.wav")

# Find speaker across files
presence = sw.find_speaker_across(speaker="Ant", folders=["./meetings/"])
```

## Live Mode

```python
sw = SonarWise(mode="live", diarization=True, events=True)

@sw.on("transcript")
def on_speech(segment):
    print(f"[{segment.speaker_name}] {segment.transcript}")

@sw.on("keyword", words=["budget", "deadline", "risk"])
def on_keyword(segment):
    send_alert(segment.transcript)

@sw.on("sound_event", events=["alarm", "glass_break"])
def on_danger(event):
    trigger_alert(event)

sw.listen(source="microphone")
```

## Plug Any Model

Every component is swappable:

```python
from sonarwise import SonarWise
from sonarwise.core.transcriber import BaseTranscriber

class MyTranscriber(BaseTranscriber):
    def transcribe(self, audio):
        return my_model.process(audio)

sw = SonarWise(transcriber=MyTranscriber())
```

**Pluggable slots:**

| Component | Base Class | Default |
|-----------|-----------|---------|
| Transcriber | `BaseTranscriber` | Whisper |
| Audio Embedder | `BaseAudioEmbedder` | CLAP |
| Vector Store | `BaseVectorStore` | SQLite |
| Chunker | `BaseChunker` | Silero VAD |
| Diarizer | `BaseDiarizer` | pyannote |
| Speaker Embedder | `BaseSpeakerEmbedder` | ECAPA-TDNN |
| Event Classifier | `BaseEventClassifier` | PANNs / CLAP |
| Preprocessor | `BasePreprocessor` | AudioPreprocessor |
| Language Detector | `BaseLanguageDetector` | Whisper |
| Stream Listener | `BaseStreamListener` | sounddevice |

## CLI

```bash
sonarwise index meeting.wav
sonarwise query "budget discussion"
sonarwise speakers meeting.wav
sonarwise timeline meeting.wav
sonarwise listen --source microphone --on-keyword "budget,risk"
sonarwise export meeting.wav --format srt
sonarwise stats
```

## Export

```python
sw.export("meeting.wav", format="srt", output="subtitles.srt")
sw.export("meeting.wav", format="json", output="segments.json")
sw.export("meeting.wav", format="csv", output="data.csv")
sw.export("meeting.wav", format="txt", output="transcript.txt")
sw.export("meeting.wav", format="notes", output="meeting_notes.md")
```

## Ecosystem

sonarwise is part of the **Ant Intelligence Ecosystem**:

| Library | Purpose | Tagline |
|---------|---------|---------|
| [SightRAG](https://github.com/VK-Ant/SightRAG) | Visual perception | See. Search. Retrieve. |
| [sonarwise](https://github.com/VK-Ant/sonarwise) | Audio perception | Hear. Search. Retrieve. |
| [adaptive-intelligence](https://github.com/VK-Ant/adaptive-intelligence) | Reasoning & memory | Learn. Remember. Adapt. |
| [llmevalkit](https://github.com/VK-Ant/llmevalkit) | Evaluation | Evaluate. Score. Improve. |
| [docqwise](https://github.com/VK-Ant/docqwise) | Document intelligence | Read. Query. Understand. |
| [wavqwise](https://github.com/VK-Ant/wavqwise) | Audio Q&A | Ask. Listen. Answer. |
| [AntGuard](https://github.com/VK-Ant/AntGuard) | AI safety & guardrails | Guard. Filter. Protect. |

## License

Apache 2.0

## Author

Built by **Venkatkumar Rajan (VK-Ant)**

- GitHub: https://github.com/VK-Ant
- Portfolio: https://vk-ant.github.io/Venkatkumar
