Metadata-Version: 2.4
Name: fermion-research
Version: 0.2.6
Summary: Phonon speech recognition engine and Neutrino language models from Fermion Research, on CPU, Apple silicon and NVIDIA GPUs
Author: Fermion Research
Maintainer: Fermion Research
License-Expression: Apache-2.0
Project-URL: Homepage, https://fermionresearch.com
Project-URL: Documentation, https://fermionresearch.com/docs/
Project-URL: Repository, https://github.com/fermionresearch/phonon
Project-URL: Models, https://huggingface.co/FermionResearch
Project-URL: Receipts, https://github.com/fermionresearch
Project-URL: Issue Tracker, https://github.com/fermionresearch/phonon/issues
Keywords: llm,quantization,ternary,sub-2-bit,bitnet,inference,transformers,openai-api,cpu-inference,edge-ai,speech-recognition,asr,speech-to-text,mlx,apple-silicon
Classifier: Development Status :: 4 - Beta
Classifier: Environment :: Console
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: MacOS :: MacOS X
Classifier: Operating System :: Microsoft :: Windows
Classifier: Operating System :: POSIX :: Linux
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Classifier: Topic :: Software Development :: Libraries :: Python Modules
Requires-Python: >=3.10
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: torch>=2.3
Requires-Dist: transformers>=5.0
Requires-Dist: numpy>=1.26
Requires-Dist: huggingface_hub>=0.23
Dynamic: license-file

# fermion-research

The fermion command runs the Phonon speech models and the Neutrino language models from one install. It transcribes audio files, transcribes the microphone live on Apple silicon, and serves an OpenAI-compatible endpoint. Phonon-2 is the current speech model. It runs on Apple silicon through MLX, on CPUs under Linux (x86-64 and Arm), Windows and macOS, and on NVIDIA GPUs through the CUDA container.

## Models

| Model | Download | Weights |
|---|--:|---|
| **Phonon-2** (`phonon-2`) | 164 MB | [FermionResearch/Phonon-2](https://huggingface.co/FermionResearch/Phonon-2) |
| Phonon-1 (`phonon-1`) | 415 MB | [FermionResearch/Phonon-1](https://huggingface.co/FermionResearch/Phonon-1) |
| Phonon-1 Micro (`phonon-1-micro`) | 285 MB | [FermionResearch/Phonon-1-Micro](https://huggingface.co/FermionResearch/Phonon-1-Micro) |
| Phonon-1 Big (`phonon-1-big`) | 581 MB | [FermionResearch/Phonon-1-Big](https://huggingface.co/FermionResearch/Phonon-1-Big) |

Accuracy and speed for each model are on its model page ([fermionresearch.com/models/phonon-2](https://fermionresearch.com/models/phonon-2/) and the model cards above).

## Install

The package installs from PyPI.

```bash
pip install fermion-research
```

On Apple silicon, one more line adds the MLX speech runtime.

```bash
pip install mlx mlx-audio mlx-lm soundfile scipy zstandard
```

On Linux and Windows CPUs, the package runs with torch. On Linux, install torch from its CPU wheel index first, which skips the GPU build.

```bash
pip install --no-deps torch --index-url https://download.pytorch.org/whl/cpu   # Linux only
pip install fermion-research torch safetensors soundfile scipy zstandard
```

Supported CPUs: x86-64 with SSE4.1 or newer (AVX2, AVX-512 VNNI and AMX processors run faster tiers of the same kernels) and 64-bit Arm with NEON (dotprod and i8mm processors run faster tiers), on Linux and Windows; Apple silicon Macs run the MLX engine. `fermion describe` shows the features found on your machine and the tier it runs.

The CPU container runs on amd64 and arm64.

```bash
docker run --rm -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cpu:2.0.3 transcribe phonon-2 /audio/recording.wav
```

The CUDA container runs Phonon-2 on NVIDIA GPUs.

```bash
docker run --rm --gpus all -v "$PWD":/audio -v phonon-cache:/home/phonon/.cache ghcr.io/fermionresearch/phonon-cuda:1.0.4 transcribe phonon-2 /audio/recording.wav
```

## Run

Transcribe a file, transcribe the microphone live, or start a local OpenAI-compatible server. `phonon` runs Phonon-2, `phonon-1` runs Phonon-1, and `fermion <command> <model>` runs any model by name. Name the model. Phonon never guesses.

```bash
phonon transcribe meeting.wav               # transcribe a file with Phonon-2
phonon listen                               # live microphone transcription (Apple silicon)
phonon serve                                # OpenAI-compatible HTTP server on 127.0.0.1:8000
fermion transcribe phonon-2 meeting.wav     # the same, naming the model
```

Any OpenAI-compatible client can then send audio to the server.

```bash
curl -s http://127.0.0.1:8000/v1/audio/transcriptions \
  -F "file=@meeting.wav" \
  -F "model=phonon-2"
```

### Dictation tools

A dictation tool keeps one Phonon server running and sends each recording to it, over a local port or an owner-only Unix socket.

```bash
fermion serve phonon-2 --port 8010 --threads 4         # or: phonon serve --port 8010
fermion serve phonon-2 --unix-socket ~/.cache/fermion/phonon.sock   # owner-only socket in place of an API key
```

```toml
engine = "whisper"
[whisper]
backend = "remote"
remote_endpoint = "http://127.0.0.1:8010"
remote_model = "phonon-2"
```

Keeping the server running at login is covered in [docs/server.md](https://github.com/fermionresearch/phonon/blob/main/docs/server.md).

`fermion models` lists every model with its aliases and marks the ones already on the machine.

## Neutrino

The same install runs the Neutrino language models. `fermion chat` opens a conversation, `fermion generate` completes a prompt, and `fermion serve` starts an OpenAI-compatible chat endpoint, each with the model named.

```bash
fermion chat neutrino-8b                     # interactive
fermion generate neutrino-8b 'Write a haiku'
fermion serve neutrino-8b                    # OpenAI-compatible chat completions
```

## Documentation

- [The command line](https://github.com/fermionresearch/phonon/blob/main/docs/cli.md)
- [The HTTP server and its API](https://github.com/fermionresearch/phonon/blob/main/docs/server.md)
- [Installation on every platform](https://github.com/fermionresearch/phonon/blob/main/docs/install.md)
- [Running on CPUs, no GPU required](https://github.com/fermionresearch/phonon/blob/main/docs/cpu.md)
- [The NVIDIA CUDA image](https://github.com/fermionresearch/phonon/blob/main/docs/cuda.md)
- [Fixes for common problems](https://github.com/fermionresearch/phonon/blob/main/docs/troubleshooting.md)

## Licence

The Phonon-2 weights are released under CC-BY-4.0. They are a derivative of NVIDIA's parakeet-tdt-0.6b-v3, with the changes listed in the [NOTICE file of the weights repository](https://huggingface.co/FermionResearch/Phonon-2/blob/main/NOTICE). The Phonon-1 family weights and the command line are released under Apache-2.0, and Phonon-1 is built on [Qwen/Qwen3-ASR-0.6B](https://huggingface.co/Qwen/Qwen3-ASR-0.6B), also under Apache-2.0.
