Metadata-Version: 2.5
Name: nanoscope-lab
Version: 0.3.0
Summary: See what your language model learns
Project-URL: Homepage, https://github.com/almajd3713/nanoscope
Project-URL: Source, https://github.com/almajd3713/nanoscope
Project-URL: Issues, https://github.com/almajd3713/nanoscope/issues
Author: almajd3713
License-Expression: MIT
License-File: LICENSE
Keywords: education,interpretability,language-models,pytorch,research,transformers
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Education
Classifier: Intended Audience :: Science/Research
Classifier: Operating System :: OS Independent
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: <3.13,>=3.10
Requires-Dist: datasets<5,>=3.0
Requires-Dist: huggingface-hub<2,>=0.27
Requires-Dist: jsonschema<5,>=4.20
Requires-Dist: matplotlib<4,>=3.8
Requires-Dist: numpy<3,>=1.26
Requires-Dist: scipy<2,>=1.11
Requires-Dist: tiktoken<1,>=0.8
Requires-Dist: tokenizers<1,>=0.20
Requires-Dist: tomli-w>=1.0
Requires-Dist: tomli>=2; python_full_version < '3.11'
Requires-Dist: torch<3,>=2.2
Provides-Extra: dev
Requires-Dist: httpx<1,>=0.27; extra == 'dev'
Requires-Dist: hypothesis<7,>=6.100; extra == 'dev'
Requires-Dist: jsonschema<5,>=4.20; extra == 'dev'
Requires-Dist: libcst<2,>=1.1; extra == 'dev'
Requires-Dist: nbformat<6,>=5.10; extra == 'dev'
Requires-Dist: pyright>=1.1.390; extra == 'dev'
Requires-Dist: pytest-cov<8,>=6; extra == 'dev'
Requires-Dist: pytest<9,>=8.3; extra == 'dev'
Requires-Dist: ruff<1,>=0.9; extra == 'dev'
Provides-Extra: graph
Requires-Dist: libcst<2,>=1.1; extra == 'graph'
Provides-Extra: server
Requires-Dist: fastapi<1,>=0.115; extra == 'server'
Requires-Dist: libcst<2,>=1.1; extra == 'server'
Requires-Dist: ruff<1,>=0.9; extra == 'server'
Requires-Dist: uvicorn[standard]<1,>=0.30; extra == 'server'
Requires-Dist: watchfiles<2,>=0.24; extra == 'server'
Provides-Extra: wandb
Requires-Dist: wandb<1,>=0.19; extra == 'wandb'
Description-Content-Type: text/markdown

# nanoscope

See what your language model learns. You write an `nn.Module`; nanoscope handles the
data, the training loop, evaluation, checkpoints and comparison against baselines.

It has two levels that share one core. Learners call `run()`. Researchers write a `Study`
with seeds, budgets, parameter matching and preregistration. The level changes what you
see, never which code runs.

## Learn

```bash
pip install git+https://github.com/almajd3713/nanoscope
```

```python
import torch.nn as nn
from nanoscope import compare, run

class Bigram(nn.Module):
    def __init__(self, vocab_size: int, d_model: int = 32):
        super().__init__()
        self.token_embedding = nn.Embedding(vocab_size, d_model)
        self.head = nn.Linear(d_model, vocab_size, bias=False)

    def forward(self, idx):
        return self.head(self.token_embedding(idx))

result = run(Bigram, preset="tinystories-5min")   # about a minute on a laptop CPU
result.plot()
compare(result, "gpt2")                          # against a shipped 3-seed baseline
```

Work through the notebooks in order:

1. [`01-first-model`](notebooks/01-first-model.ipynb): write a bigram model and train it.
2. [`02-gpt2`](notebooks/02-gpt2.ipynb): a real transformer.
3. [`03-modern-block`](notebooks/03-modern-block.ipynb): RoPE, RMSNorm, SwiGLU, GQA, QK-norm, z-loss.
4. [`04-ablations`](notebooks/04-ablations.ipynb): which part matters, with seeds and confidence intervals.

Token data downloads from the Hub (`RedhouaneLazib/nanoscope-tokens`) when a preset has
it, and is tokenized locally otherwise. Everything is cached under `~/.nanoscope/data`.

## Research

See the [research guide](docs/research.md). The research program itself is in
[`docs/project-nanoscope.md`](docs/project-nanoscope.md).

## Build models from blocks

`GPT2` and `Modern` are short compositions of the blocks in `nanoscope.blocks`, and so can your
own models. See [docs/blocks.md](docs/blocks.md).

## Learn

Guided paths build the models step by step, with checks that say why. See
[docs/learn.md](docs/learn.md): `nanoscope learn list`, `learn start`, `learn check`.

## Command line

```bash
nanoscope presets                                   # available presets
nanoscope run nanoscope/models/gpt2.py:GPT2 --seeds 3
nanoscope compare modern gpt2 --preset tinystories-5min
nanoscope study studies/m1_ablation.py --devices cuda:0
nanoscope report studies/m1_ablation.py             # writes experiments/<name>/
nanoscope bench modern --compile reduce-overhead    # speed, and whether CPU or GPU is the limit
nanoscope status runs                               # what is running and how far along
nanoscope describe nanoscope/models/modern.py:Modern  # shapes, params, FLOPs, memory per module
nanoscope graph my_model.py                         # a model file's architecture, without running it
nanoscope blocks                                    # the blocks models are composed from
nanoscope prepare-data tinystories-5min             # download or tokenize now
nanoscope publish-data tinystories-5min user/repo   # upload tokens to a Hub dataset
```

## Develop

```bash
uv sync --all-extras
make test       # offline tests
make lint
make typecheck
make check      # lint, typecheck, then test: what CI runs
```

## Data credits

Tokens are derived from [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories)
(CDLA-Sharing-1.0) and [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceFW/fineweb-edu)
(ODC-By 1.0).
