Metadata-Version: 2.4
Name: omnivision-research
Version: 0.1.0
Summary: MIT-licensed computer vision research framework
Author: Milo Evans
License-Expression: MIT
License-File: LICENSE
Keywords: computer-vision,model-configs,pytorch
Classifier: Development Status :: 3 - Alpha
Classifier: Intended Audience :: Developers
Classifier: Intended Audience :: Science/Research
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.12
Classifier: Topic :: Scientific/Engineering :: Artificial Intelligence
Requires-Python: >=3.12
Requires-Dist: numpy>=1.24.0
Requires-Dist: pillow>=10.0.0
Requires-Dist: pyyaml>=6.0
Requires-Dist: torch>=2.0.0
Provides-Extra: dev
Requires-Dist: build>=1.2.2; extra == 'dev'
Requires-Dist: pytest>=8.0.0; extra == 'dev'
Requires-Dist: twine>=5.0.0; extra == 'dev'
Description-Content-Type: text/markdown

# Omnivision

Omnivision is an MIT-licensed computer vision research framework for learning,
architecture experiments, and small local studies. It is not a deployment
runtime and does not vendor code, configs, checkpoints, or datasets from
license-restricted CV frameworks.

The publishable distribution is `omnivision-research`; the import package and
module CLI stay `omnivision`:

```bash
pip install omnivision-research
python -m omnivision --help
```

The public API keeps familiar one-object CV workflow ergonomics while staying
Omnivision-native and clean-room:

```python
from omnivision import Omni

model = Omni("examples/tiny_classifier.yaml")
run = model.train(data="synthetic:classification", epochs=2, out="scratch/run")
print(run["artifacts"]["summary"])
print(model.val(data="synthetic:classification"))
print(model.predict())
print(model.info())
model.save("scratch/tiny.pt")
same_arch = Omni("examples/tiny_classifier.yaml").load("scratch/tiny.pt")
restored = Omni.load("scratch/tiny.pt")
```

`Omni.load(path)` restores config, weights, and history. `Omni(config).load(path)`
loads compatible weights into the current config and keeps the current history.
In the CLI, `--checkpoint` restores a saved experiment for inspection,
reporting, validation, and prediction, while `--weights` initializes the current
training config before fitting.

Installed packages also include the same tiny configs:

```python
from omnivision import Omni, example_config_path

model = Omni(example_config_path("tiny_classifier"))
```

Use `Experiment.from_config(...)` when you want the explicit fit/evaluate/sample
object, or `Model(...)` when you only need the raw PyTorch module.

Architecture experiments can stay in Python:

```python
base = Omni("examples/mobile_classifier.yaml")
wider = base.variant(["stem.out_channels=16", "channel-mix.out_channels=24"], name="mobile-wide")
saved_variant = base.variant("stem.out_channels=16", name="mobile-wide", out="scratch/mobile-wide.yaml")
grid = base.sweep(["stem.out_channels=12,16", "channel-mix.out_channels=16,24"])
saved_grid = base.sweep(["stem.out_channels=12,16"], out="scratch/sweep")
print(base.compare(wider, out="scratch/compare"))
print(base.benchmark(wider, runs=3, warmup=1, out="scratch/bench"))
print(base.study(*grid, steps=1))
print(base.study(*grid, steps=1, out="scratch/study"))
print(saved_variant["path"])
print(saved_grid["artifacts"]["manifest"])
print(base.html_report("scratch/mobile-report.html"))
print(wider.report())
```

CLI commands accept normal flags and compact `key=value` shortcuts:

```bash
omni inspect --config examples/tiny_classifier.yaml
omni examples
omni new --template mobile_classifier --out scratch/mobile.yaml
omni data formats
omni data import --format coco-detect --source annotations/instances.json --images-root images --out scratch/detect.jsonl
omni data import --format voc-detect --source annotations --images-root images --out scratch/detect.jsonl
omni data import --format csv-classify --task classify --source labels.csv --images-root images --out scratch/classify.jsonl
omni data summarize --source scratch/detect.jsonl --task detect --check-files
omni ops --format markdown
omni diagram --config scratch/mobile.yaml
omni trace --config scratch/mobile.yaml
omni trace --config scratch/mobile.yaml --source path/to/image.png
omni report --config examples/branch_classifier.yaml
omni html-report --config examples/branch_classifier.yaml --out scratch/branch-report.html --source path/to/image.png
omni tasks
omni guide
omni variant --config scratch/mobile.yaml --set stem.out_channels=16 --out scratch/mobile-wide.yaml
omni sweep --config scratch/mobile.yaml --axis stem.out_channels=12,16 --axis channel-mix.out_channels=16,24 --out scratch/sweep
omni rank --manifest scratch/sweep --by parameters --top 3
omni compare --manifest scratch/sweep --out scratch/compare
omni benchmark --manifest scratch/sweep --runs 3 --warmup 1 --out scratch/bench
omni study --manifest scratch/sweep --steps 1 --out scratch/study
omni train --config examples/tiny_classifier.yaml --data synthetic:classification --epochs 2
omni mode=train model=examples/tiny_classifier.yaml data=synthetic:classification steps=1 batch=2
omni train --config examples/tiny_classifier.yaml --data synthetic:classification --steps 1 --seed 7 --out scratch/run
omni train --config examples/tiny_classifier.yaml --weights scratch/tiny.pt --data synthetic:classification --steps 1
omni mode=fit model=examples/tiny_classifier.yaml weights=scratch/tiny.pt data=synthetic:classification steps=1 batch=2
omni train --config examples/tiny_depth.yaml --data synthetic:depth --epochs 2
omni train --config examples/tiny_keypoints.yaml --data synthetic:keypoint --epochs 2
omni val --config examples/tiny_classifier.yaml --data synthetic:classification
omni predict --config examples/tiny_classifier.yaml
omni predict --config examples/tiny_classifier.yaml --source path/to/image.png
omni train --config examples/tiny_classifier.yaml --save scratch/tiny.pt
omni report --checkpoint scratch/tiny.pt
omni trace --checkpoint scratch/tiny.pt
omni predict --checkpoint scratch/tiny.pt --source path/to/image.png
```

`fit`, `evaluate`, and `sample` remain available as explicit aliases for
`train`, `val`, and `predict`.

## Graph YAML

Model configs are readable graph specs. Each row declares an operation by name:

```yaml
name: tiny-classifier
task: classify
input:
  channels: 3
  size: 32
classes: 3
graph:
  - id: stem
    op: conv
    out_channels: 12
    kernel: 3
  - op: activation
    kind: relu
  - op: global_avg_pool
  - op: flatten
  - op: linear
    out_features: 3
```

Built-in ops currently include `conv`, `conv_block`, `batchnorm`, `activation`,
`pool`, `global_avg_pool`, `flatten`, `dropout`, `identity`, `depthwise_conv`,
`separable_conv`, `residual_block`, `upsample`, `add`, `concat`, and `linear`.
Convolution-style ops accept `dilation` for receptive-field experiments.
Single-input ops also accept `repeat` to apply the same row several times in
sequence:

```yaml
graph:
  - id: encoder
    op: conv_block
    out_channels: 32
    kernel: 3
    repeat: 3
```

Configs can define small compound scales for architecture families:

```yaml
scale: small
scales:
  small: [0.50, 0.50, 64]  # depth, width, max channels
  medium: [1.00, 1.00, 128]
graph:
  - id: encoder
    op: conv_block
    out_channels: 64
    repeat: 3
  - id: dense-head
    op: conv
    out_channels: 2
    scale_width: false
```

Scaling resolves before model building. It adjusts `repeat` and `out_channels`,
caps channels at the optional third value, leaves semantic/depth output rows
unchanged when they already match the task contract, and leaves any row with
`scale_width: false` unchanged.
Run `omni new --template scaled_classifier --out scratch/scaled.yaml` for a
copyable scaled example.

Nodes run in order. If `inputs` is omitted, the node consumes the previous
node. Use `inputs` for branches and residual paths:

```yaml
graph:
  - id: stem
    op: conv
    out_channels: 12
  - id: smooth-path
    inputs: stem
    op: conv
    out_channels: 12
  - id: sharp-path
    inputs: stem
    op: conv
    out_channels: 12
    kernel: 1
    padding: 0
  - id: residual-mix
    inputs: [smooth-path, sharp-path]
    op: add
```

Add research blocks with:

```python
from torch import nn
from omnivision import Shape, register_op, unregister_op

def my_block(row, shape: Shape):
    module = nn.Conv2d(shape.channels, row["out_channels"], 1)
    return module, Shape(row["out_channels"], shape.height, shape.width)

register_op("my_block", my_block)
unregister_op("my_block")
```

An op factory receives the YAML row and current tensor shape, then returns a
PyTorch module plus the new shape. This keeps the extension point small and easy
to inspect.

## Learning Tools

`omni report --config ...` prints a Markdown table with each graph node, its
repeat count, inputs, output shape, parameter count, simple MAC estimate,
activation element count, output stride, and receptive-field estimate.
`omni compare --config ...`
compares multiple architecture candidates by task, input shape, class count,
detection slots, node count, parameters, MACs, activations, output stride,
receptive field, and deltas from the first config.
Rows include warnings when task, input shape, class count, or detection slots
differ from the baseline.
Add `--out scratch/compare` to write `results.json` plus `summary.md`.
`omni benchmark --config ... --out ...` adds a local zero-input forward-pass
timing check and writes `results.json` plus `summary.md` for rough research
comparisons on your current machine.
`omni new --template ... --out ...` copies a packaged starter config into a
local file you can edit.
`omni diagram --config ...` prints a Mermaid graph, and `omni study --config ...`
runs a deliberately short architecture comparison loop for research notes.
`omni tasks` prints the task contracts: expected model output, target format,
loss, metrics, and supported data inputs.
`omni validate-config --config ...` validates schema/build constraints and
reports non-fatal graph warnings such as unused branches.
`omni trace --config ...` runs one zero-input pass and reports tensor shapes and
simple activation stats for every node. Add `--source path/to/image.png` to trace
a real local image through the same graph.
`omni html-report --config ... --out report.html` writes a self-contained graph
and trace report that can be opened directly in a browser.
`report`, `html-report`, `diagram`, and `trace` also accept `--checkpoint` to
inspect the saved experiment config.
`omni study --out scratch/study ...` also writes `results.json` and `summary.md`
with per-candidate settings such as data, steps, batch size, learning rate,
device, and seed so short architecture studies can be reviewed later.
The JSON artifact preserves the full row data; the Markdown summary keeps the
main architecture metrics, seed, steps, and task metrics compact.
`omni train --out scratch/run ...` and `Omni.train(..., out="scratch/run")`
write `config.yaml`, `history.json`, `metrics.json`, and `summary.md` for a
single training run.
Add `--weights scratch/base.pt` to initialize a compatible config before the
fit loop, matching `Omni(config).load_weights(path).train(...)`.

Create architecture variants without editing YAML by hand:

```bash
omni variant \
  --config examples/mobile_classifier.yaml \
  --set stem.out_channels=16 \
  --set channel-mix.out_channels=24 \
  --name mobile-wide \
  --out scratch/mobile-wide.yaml
```

Variant edits use `node_id.field=value` for graph nodes, plus top-level paths
like `name=...`, `classes=...`, or `input.size=...`. They also support small
graph structure edits:

```bash
omni variant \
  --config examples/mobile_classifier.yaml \
  --set graph.delete.node_1_activation=true \
  --set 'graph.insert_after.stem={id: norm, op: batchnorm}' \
  --set 'graph.append={id: regularizer, op: dropout, p: 0.1}' \
  --out scratch/mobile-structural.yaml
```

Use `graph.append={...}` to append a graph row,
`graph.insert_after.node_id={...}` or `graph.insert_before.node_id={...}` to
insert near a node, and `graph.delete.node_id=true` to remove a row. Structural
edit mappings must include `op`. For rows without explicit ids, use the
generated id shown by `omni inspect`; prefer explicit ids for rows you plan to
edit again.

Create small architecture grids with:

```bash
omni sweep \
  --config examples/mobile_classifier.yaml \
  --set 'graph.append={id: regularizer, op: dropout, p: 0.1}' \
  --axis scale=small,medium \
  --axis stem.out_channels=12,16 \
  --axis channel-mix.out_channels=16,24 \
  --axis regularizer.p=0.0,0.2 \
  --max-params 2000 \
  --out scratch/mobile-sweep
```

Sweeps write one YAML per valid candidate plus `manifest.json`. Every candidate
is preflighted before files are written. Add `--set` to apply fixed edits before
the grid, and add `--max-params`, `--max-macs`, or `--max-activations` to skip
candidates outside a simple research budget. Use `omni rank --manifest
scratch/mobile-sweep --by parameters --top 3` to sort saved candidates by
parameters, MACs, or activations before selecting configs to study. The same
manifest can be passed to `omni compare`, `omni benchmark`, and `omni study`
with `--manifest scratch/mobile-sweep`.

## Data

Synthetic data is built in for smoke experiments. For local datasets, pass a
JSONL manifest to `--data`. Paths are resolved relative to the manifest.
Use `omni data import` to convert small external annotation files into this
native JSONL shape. The included detection importers are `coco-detect`, which
converts COCO annotations from absolute xywh boxes, and `voc-detect`, which
converts Pascal VOC XML annotations from absolute xyxy boxes. The included
classification importer is `csv-classify`, which converts `image,label` CSV
files into contiguous labels. Detection importers write normalized xyxy boxes
and contiguous labels. Imported image paths are written
relative to the output JSONL manifest when possible, because loaders resolve
relative paths from that manifest.

```python
from omnivision import import_dataset, list_dataset_formats
from omnivision import summarize_dataset

print(list_dataset_formats())
import_dataset(
    "annotations/instances.json",
    "scratch/detect.jsonl",
    format="coco-detect",
    images_root="images",
)
print(summarize_dataset("scratch/detect.jsonl", task="detect", check_files=True))
```

Imported manifests and external annotations are research data artifacts, not
publishable package files; keep them under ignored local output directories such
as `scratch/`.

Classification:

```json
{"image": "images/cat.png", "label": 0}
```

Semantic segmentation:

```json
{"image": "images/street.png", "mask": "masks/street.png"}
```

Monocular depth:

```json
{"image": "images/room.png", "depth": "depth/room.png"}
```

Detection:

```json
{"image": "images/ball.png", "boxes": [[0.1, 0.2, 0.8, 0.9]], "labels": [0]}
```

For multi-object detection experiments, add `detections: N` to the config and
return up to `N` normalized xyxy boxes per JSONL row:

```yaml
task: detect
classes: 3
detections: 2
```

```json
{"image": "images/scene.png", "boxes": [[0.1, 0.2, 0.4, 0.5], [0.55, 0.3, 0.9, 0.8]], "labels": [0, 2]}
```

Keypoint regression:

```json
{"image": "images/hand.png", "keypoints": [[0.2, 0.3], [0.7, 0.8]]}
```

For `task: keypoint`, `classes` is the number of keypoints and the model output
is a flat `classes * 2` normalized xy coordinate vector.

## Scope

This clean-room milestone supports classification, semantic segmentation,
monocular depth, fixed-slot detection, and normalized keypoint regression
experiments with synthetic or JSONL data, plus ImageFolder-style classification
directories. Rich detector heads and larger dataset formats should be added as
native Omnivision designs, not by copying another project structure.

## License Boundary

Run this before publishing:

```bash
python -m omnivision license-check
python -m omnivision release-check --dist dist
```

The check fails if inherited provenance markers, model weights, datasets, run
outputs, old experiment folders, or incompatible declared runtime dependency
licenses are present in the publishable tree.
`PROVENANCE.md` records the intended source boundary for reviewers and future
contributors.
