Metadata-Version: 2.4
Name: scanpath-studio
Version: 0.31.0
Summary: Interactive Streamlit workbench for visualizing eye-tracking-while-reading scanpaths, computing reading measures, and exporting figures and tabular data.
Author: Keren Gruteke Klein, Maya Grossman, Ella Lion, Deborah N. Jakobi, David R. Reich, Lena Jäger, Yevgeni Berzak
Author-email: Omer Shubi <lacclab.technion@gmail.com>
License-Expression: MIT
Project-URL: Repository, https://github.com/lacclab/scanpath-studio
Project-URL: Documentation, https://lacclab.github.io/scanpath-studio/
Project-URL: Issues, https://github.com/lacclab/scanpath-studio/issues
Keywords: eye-tracking,scanpath,visualization,streamlit,reading,psycholinguistics
Classifier: Development Status :: 4 - Beta
Classifier: Programming Language :: Python
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3 :: Only
Classifier: Programming Language :: Python :: 3.11
Classifier: Programming Language :: Python :: 3.12
Classifier: Programming Language :: Python :: 3.13
Classifier: Programming Language :: Python :: 3.14
Classifier: Intended Audience :: Science/Research
Classifier: Topic :: Scientific/Engineering :: Visualization
Requires-Python: >=3.11
Description-Content-Type: text/markdown
License-File: LICENSE
License-File: NOTICE
Requires-Dist: streamlit>=1.64.0
Requires-Dist: pandas>=3.0
Requires-Dist: plotly>=7.1
Requires-Dist: numpy>=2.4
Requires-Dist: scipy>=1.17
Requires-Dist: pyarrow>=25.0
Requires-Dist: kaleido>=1.4
Requires-Dist: watchdog>=6.0.0
Requires-Dist: streamlit-sortables>=0.3.1
Requires-Dist: pillow>=12.3
Requires-Dist: imageio[ffmpeg]>=2.37
Requires-Dist: openpyxl>=3.1.5
Requires-Dist: xlrd>=2.0.2
Provides-Extra: test
Requires-Dist: pytest>=8.0.0; extra == "test"
Requires-Dist: pytest-cov>=4.0.0; extra == "test"
Requires-Dist: pytest-xdist>=3.0.0; extra == "test"
Requires-Dist: pytest-timeout>=2.1.0; extra == "test"
Provides-Extra: lint
Requires-Dist: ruff==0.16.8; extra == "lint"
Provides-Extra: docs
Requires-Dist: mkdocs-material>=9.5; extra == "docs"
Requires-Dist: mkdocstrings[python]>=0.26; extra == "docs"
Dynamic: license-file

# Scanpath Studio


[![PyPI](https://img.shields.io/pypi/v/scanpath-studio.svg)](https://pypi.org/project/scanpath-studio/)
[![Python versions](https://img.shields.io/pypi/pyversions/scanpath-studio.svg)](https://pypi.org/project/scanpath-studio/)
[![Live demo](https://img.shields.io/badge/Live_demo-Streamlit-FF4B4B?logo=streamlit&logoColor=white)](https://scanpath-studio.streamlit.app)
[![Docs](https://img.shields.io/badge/docs-mkdocs-blue)](https://lacclab.github.io/scanpath-studio/)
[![CI](https://github.com/lacclab/scanpath-studio/actions/workflows/ci.yml/badge.svg)](https://github.com/lacclab/scanpath-studio/actions/workflows/ci.yml)
[![Coverage](https://img.shields.io/endpoint?url=https%3A%2F%2Flacclab.github.io%2Fscanpath-studio%2Fcoverage%2Fbadge.json)](https://lacclab.github.io/scanpath-studio/coverage/)
[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://github.com/lacclab/scanpath-studio/blob/main/LICENSE)

An interactive workbench for visualizing **eye-tracking-while-reading** data.
Drop in a trial and see the scanpath the way the reader saw it — words at their
true on-screen positions, with fixations, saccades, a density heatmap, and
animated replay layered on top, all exportable as publication-ready figures.

It is **dataset-agnostic** (auto-detects EyeLink / Gazepoint / Tobii / SMI /
Pupil Labs / snake-case columns) and ships with a small [OneStop][onestop-paper] demo, so you can try it
with zero setup.

> **Authors:** Omer Shubi, Keren Gruteke Klein, Maya Grossman, Ella Lion, Deborah N. Jakobi,
> David R. Reich, Lena Jäger, and Yevgeni Berzak — Data and Decision Sciences
> (Technion) and Department of Computational Linguistics (University of Zurich).

![A reading scanpath replayed fixation by fixation](https://raw.githubusercontent.com/lacclab/scanpath-studio/main/assets/scanpath_animation.gif)

*A scanpath replayed fixation by fixation over the text the reader saw.*

## Try it

**Live demo (zero install):** <https://scanpath-studio.streamlit.app>

```bash
pip install scanpath-studio
scanpath-studio      # launches the app in your browser
```

## What it does

The scanpath plot is built from layers you toggle independently:

- **Text** drawn at the exact pixel coordinates the participant saw.
- **Fixations** sized and **colored by any column** in your data (duration, GPT-2 surprisal, word frequency, …).
- **Saccades**, with backward jumps (regressions) standing out.
- **Areas of interest** (word boxes from your data) and a word-level **heatmap** (total fixation duration, count, …).

On top of that:

- **Animated replay** — watch the scanpath unfold at real or scaled speed; export as interactive HTML, GIF, or MP4.
- **Compare readings** — overlay two trials on one canvas or place them side by side (e.g. ordinary vs. information-seeking, first vs. repeated, L1 vs. L2), including two trials from *different* datasets.
- **Critical-span, out-of-text & by-line** highlights — mark an answer span, flag fixations outside every word box, or color fixations by text line.
- **Triage** — star, tag, and annotate trials; save and restore everything as a JSON sidecar.
- **Bulk export** — one zip of per-trial PNG + SVG figures, plot settings, and tabular data across every filtered trial.
- **Author a scanpath** — draw fixations straight onto the stimulus canvas for a teaching figure or a schematic.
- **A computation register** — every derived value's formula, units and precedence, published as a [methodology page](https://lacclab.github.io/scanpath-studio/computations/).

![Two readers of the same paragraph, overlaid on one canvas](https://raw.githubusercontent.com/lacclab/scanpath-studio/main/assets/demo_dual_scanpath.png)

*Two readers of the same bundled-demo paragraph, overlaid on one canvas — 305
fixations between them
([watch it animated](https://raw.githubusercontent.com/lacclab/scanpath-studio/main/docs/assets/demo_dual_scanpath.gif)).*

The app is organized into three views, chosen from the navigation in the header
(💾 **Session** and ❓ **Help** sit beside them and open over whatever you are
looking at, rather than taking you away from it):

| View | What's there |
|------|--------------|
| 🗺️ **Scanpath** | The layered scanpath: a control line above the plot — dataset, trial picker, ◀ ▶ step, ⇅ sort and a filter funnel holding every way to narrow the pool (text, participant, conditions, annotations) — and, beside the figure, a right-hand **control rail** with **Animate** and **Compare** toggles plus the per-layer visualization controls (style each scanpath independently). The trial's key info shows as configurable chips above the plot. Below, subtabs: **Annotations**, **Stimulus & Context**, **Comparisons** (trials matching the selected trial on a field you choose), **Export** (single-trial *and* bulk — HTML / GIF / MP4 and figures / settings / tabular data), and **Share**. |
| 📊 **Corpus Analysis** | Four subtabs — **Per text**, **Per sentence**, **Per reader**, and **Groups** (profile one cohort, or compare two) — the question-oriented analysis views: metric distributions and word profiles, per-text heatmaps pooled over readers, reader summaries, and group differences with effect sizes. |
| 🗂️ **Data** | Two screens: 📂 **Available datasets** (the datasets this session holds, plus what's in the open one — data tables and summary statistics) and ✏️ **Edit dataset** (source and location, column mapping, recording setup, trial identity, stimulus images, and the participant/trial metadata tables). |

![The Scanpath Studio app](https://raw.githubusercontent.com/lacclab/scanpath-studio/main/docs/assets/app_screenshot.png)

## Roadmap

Planned and in-flight work is tracked in
[GitHub Issues](https://github.com/lacclab/scanpath-studio/issues). The archive
of everything closed before 2026-08-20 is browsable offline: double-click
`tracker/start.command` (macOS) or `tracker/start.bat` (Windows) — or run
`python3 tracker/server.py` (`python tracker\server.py` on Windows).

## Your data

Upload **CSV, TSV, TXT, Parquet, Feather, or Excel (.xlsx / .xls)** tables — or a
**.zip** of them — for words/AoIs, fixations, and (optionally) raw gaze. Columns
are auto-detected from common EyeLink, Gazepoint, Tobii, SMI, Pupil Labs, and
snake-case conventions; the **Column mapping** on 🗂️ Data → ✏️ **Edit
dataset** overrides any guess. The loader bends to fit real corpora — many files
per table (concatenated with a `source_file` tag), a single report (words- or
fixations-only), stimulus-level word boxes broadcast across readers, AoI-sequence
fixations placed at word/character-box centers, and trials recorded over several
screens. A separate one-row-per-reader table of **participant metadata** (and one
of trial metadata) attaches alongside, and its columns then behave like fields in
the data — filters, chips, trial sorting, and the export bundle.

If your data carries only raw fixations, the app computes the canonical per-word
measures itself — **FFD**, **FPRT** (gaze duration), **RPD** (go-past), **TFD**
(dwell), plus skips and regressions, following Rayner (1998) and Inhoff & Radach
(1998). Pre-aggregated EyeLink columns, when present, take precedence.

Several public corpora need no upload at all: **OneStop**,
[**PoTeC**](https://github.com/DiLi-Lab/PoTeC) (Potsdam Textbook Corpus) and
**MultiplEYE** have ready-made loaders, and thirty-one
[harmonised benchmark corpora](https://lacclab.github.io/scanpath-studio/benchmark-corpora/)
— German, Chinese, Persian, Danish, Spanish, Dutch, Russian, English, and the
multilingual MECO waves — load from one locally prepared bundle in a
single common schema, which is what makes cross-corpus comparison practical. The
PoTeC loader exercises that flexible pipeline end to end:

```python
import scanpath_studio as sps

words, fixations = sps.load_potec("data/PoTeC", download=True)  # ~45 MB on first call
fig = sps.plot_scanpath(words, fixations, "0", "b0", canvas_size=(1680, 1050))
```

## Command line & Python API

Everything the app draws is also available headless — same pipeline, same figure.

```bash
scanpath-studio render --sample --list-trials              # what's available
scanpath-studio render --sample -o scanpath.html           # interactive HTML
scanpath-studio render --words ia.csv --fixations fix.csv -p p1 -t t3 -o figure.png
scanpath-studio render --sample --animate -o replay.html   # animated replay
```

```python
import scanpath_studio as sps

words, fixations = sps.load_scanpath_data(
    "ia.csv", "fixations.csv"
)  # paths, globs, or lists; either table optional
sps.list_trials(words, fixations)
fig = sps.plot_scanpath(words, fixations, "p1", "t3")  # every layer toggle is a kwarg
sps.save_figure(fig, "scanpath.png")  # .html / .png / .svg / .pdf
measures = sps.compute_word_metrics(words, fixations)  # FFD / FPRT / RPD / TFD …
```

HTML export is browser-free; PNG/SVG/PDF/GIF/MP4 go through Kaleido (run
`plotly_get_chrome -y` once). See `scanpath-studio render --help` for all flags.

## Run from source

```bash
git clone https://github.com/lacclab/scanpath-studio.git
cd scanpath-studio
pip install -e ".[test]"          # or: uv sync --extra test --extra lint
streamlit run streamlit_app.py --server.address 127.0.0.1
```

Tested on Python 3.11–3.14. Run the tests with `pytest` (or
`uv run --extra test pytest`); see [AGENTS.md](https://github.com/lacclab/scanpath-studio/blob/main/AGENTS.md) for an architectural overview.

Joining the project? [CONTRIBUTING.md](https://github.com/lacclab/scanpath-studio/blob/main/CONTRIBUTING.md) is the whole of it —
setup, where the work is tracked
([GitHub Issues](https://github.com/lacclab/scanpath-studio/issues)), the checks
that gate CI, and how two people stay out of each other's way.

## Documentation

Full docs — getting started, the Python API, the CLI reference, data format, and
export/troubleshooting — are at **<https://lacclab.github.io/scanpath-studio/>**
(built from `docs/` with MkDocs Material). Build them locally with:

```bash
pip install -e ".[docs]"
mkdocs serve
```

## Citation

A system-demo paper is in preparation — **citation TBD**. Until then, cite the
software via GitHub's **"Cite this repository"** button (generated from
[`CITATION.cff`](https://github.com/lacclab/scanpath-studio/blob/main/CITATION.cff)).

If you use the bundled demo data, please cite the OneStop corpus:

```bibtex
@article{berzak2025onestop,
  title     = {{OneStop}: A 360-Participant {E}nglish Eye Tracking Dataset
               with Different Reading Regimes},
  author    = {Berzak, Yevgeni and Malmaud, Jonathan and Shubi, Omer
               and Meiri, Yoav and Lion, Ella and Levy, Roger},
  journal   = {Scientific Data},
  year      = {2025},
  publisher = {Nature Publishing Group},
  doi       = {10.1038/s41597-025-06272-2},
  url       = {https://www.nature.com/articles/s41597-025-06272-2},
}
```

The bundled demo is a subset of [OneStop Eye Movements][onestop-corpus], used
under its original license ([docs][onestop-docs]).

[onestop-paper]: https://www.nature.com/articles/s41597-025-06272-2
[onestop-corpus]: https://github.com/lacclab/OneStop-Eye-Movements
[onestop-docs]: https://lacclab.github.io/OneStop-Eye-Movements/

## Built with AI assistance

Much of this code was written with AI assistance. That is not the same as
bug-free. **Cross-check anything you publish against your own pipeline.** If
something looks wrong — or if you have a feature request or suggestion —
[open an issue](https://github.com/lacclab/scanpath-studio/issues).

## License

MIT — see [LICENSE](https://github.com/lacclab/scanpath-studio/blob/main/LICENSE); the bundled demo data carries its own licenses (see [NOTICE](https://github.com/lacclab/scanpath-studio/blob/main/NOTICE)).
