Metadata-Version: 2.4
Name: radiens-drive-catalog
Version: 0.0.14
License-File: LICENSE
Requires-Python: >=3.12
Requires-Dist: google-api-python-client>=2.188.0
Requires-Dist: google-auth-oauthlib>=1.2.4
Requires-Dist: google-auth>=2.47.0
Requires-Dist: pandas>=2.3.3
Requires-Dist: tqdm>=4.67.3
Description-Content-Type: text/markdown

# radiens-drive-catalog

A Python package for programmatically managing large neural datasets stored on Google Drive. It handles Drive scanning, local cataloging, and selective dataset download. Analysis is done locally — this package is purely about data management.

Documentation: <https://neuronexus.github.io/radiens-drive-catalog/latest/>

## Overview

Neural data is stored as xdat filesets (NeuroNexus format) on a shared Google Drive. Each dataset consists of 3 files sharing a common `base_name`:

```
{base_name}_data.xdat
{base_name}.xdat.json
{base_name}_timestamp.xdat
```

`radiens-drive-catalog` scans the Drive hierarchy, builds a local catalog indexed by `base_name`, and lets you query and download datasets selectively. Non-xdat content found alongside recordings — logs directories, PowerPoints, writeups — is also discovered and tracked as **Drive items**.

## Usage

### Recordings

```python
from radiens_drive_catalog import Catalog, Config

config = Config.from_file("config.json")
catalog = Catalog(config)

# Scan Drive and build the catalog (discovers recordings and Drive items)
result = catalog.scan()      # returns ScanResult with new/existing/removed counts
result = catalog.scan(flat=False)  # recursive traversal instead of flat file scan

# Query recordings using pandas directly
catalog.recordings_df
catalog.list_recordings()                                                                    # everything
catalog.list_recordings(drive_path="2026-02-15_batch/reaching")                             # exact folder
catalog.list_recordings(drive_path="2026-02-15_batch", drive_path_mode="prefix")            # full date subtree
catalog.list_recordings(drive_path="reaching", drive_path_mode="contains")                   # any depth
catalog.list_recordings(base_name="rat01_session3")                                          # exact base_name
catalog.list_recordings(base_name="rat01", base_name_mode="prefix")                           # base_name prefix
catalog.list_recordings(base_name="session3", base_name_mode="contains")                     # base_name substring

# Filters are ANDed together, e.g. all "rat01" recordings under "2026-02-15_batch"
catalog.list_recordings(drive_path="2026-02-15_batch", drive_path_mode="prefix", base_name="rat01", base_name_mode="prefix")

# Look up a recording by base_name (raises AmbiguousRecordingError if not globally
# unique — pass drive_path to disambiguate)
entry = catalog.get_recording("rat01_session3")
entry = catalog.get_recording("rat01_session3", drive_path="2026-02-15_batch/reaching")

# Download a recording (3 xdat files)
catalog.download_recording("2026-02-15_batch/reaching", "rat01_session3")

# Get the local path, downloading automatically if needed
path = catalog.get_recording_path("2026-02-15_batch/reaching", "rat01_session3")
```

### Drive items (non-xdat content)

Non-xdat files and folders (e.g. `logs/`, PowerPoints, writeups) are automatically cataloged as Drive items during `scan()`.

```python
# Query items using pandas directly
catalog.items_df
catalog.items_df[catalog.items_df["drive_path"].str.startswith("2026-02-15_batch")]
catalog.items_df[catalog.items_df["is_folder"] == True]

# Download an item (drive_path is the slash-joined path to the item's parent folder)
catalog.download_item("2026-02-15_batch/reaching", "logs")

# Get the local path, downloading automatically if needed
path = catalog.get_item_path("2026-02-15_batch/reaching", "logs")

# Render an annotated tree of the full Drive hierarchy
print(catalog.file_tree())

# Print a headline summary report: counts, download progress, date
# range, and Drive/local storage totals
print(catalog.summary())

# verbose=True adds an item-type breakdown, largest entries, itemized
# incomplete recordings, a per-folder breakdown, and a full per-entry listing
print(catalog.summary(verbose=True))
```

Items land under `local_data_dir/{drive_path}/{name}`, using the same Drive-mirroring convention as recordings.

## Configuration

Create a `config.json` (outside your repo — do not commit it):

```json
{
    "credentials_path": "/path/to/service_account.json",
    "root_folder_id": "your-drive-folder-id",
    "local_data_dir": "/path/to/local/data",
    "catalog_path": "/path/to/local/data/catalog.json"
}
```

`Config.from_file()` locates the config file using this resolution order:

1. Explicit `path` argument.
2. `RADIENS_DRIVE_CATALOG_CONFIG` environment variable.
3. `.secrets/config.json`, then `config.json`, searched starting in the current working directory and then each parent directory up to the filesystem root (closest directory wins) — the same convention git uses to locate `.git`, so this works the same whether you run from the repo root or a nested subdirectory (e.g. a notebook in `notebooks/`).
4. `~/.config/radiens-drive/config.json`.
5. `/etc/radiens-drive/config.json`.

```python
# Automatic discovery (env var or well-known paths)
config = Config.from_file()

# Explicit path
config = Config.from_file("/path/to/config.json")
```

The `root_folder_id` is the alphanumeric string in the Drive URL when you're inside the root data folder.

## Authentication

This package uses a Google service account for shared access among collaborators. To set it up:

1. Create a project in [Google Cloud Console](https://console.cloud.google.com)
2. Enable the Google Drive API
3. Create a service account and download its JSON credentials file
4. Share your root Drive data folder with the service account's email address (Viewer access is sufficient)
5. Point `credentials_path` in your config at the downloaded JSON file

Distribute the credentials file to collaborators securely — treat it like a password.

## Installation

This project uses `uv` for dependency management. If you don't have it:

**macOS / Linux:**

```bash
curl -LsSf https://astral.sh/uv/install.sh | sh
```

**Windows:**

```powershell
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
```

Then install the project:

```bash
uv sync
```

## Development

```bash
uv run pytest          # run tests
uv run mypy            # type checking
uv run ruff check .    # linting
uv run ruff format .   # formatting
```
