Metadata-Version: 2.4
Name: pax-detector
Version: 0.1.4
Summary: A vision pipeline for object and car make classification with CLI
Author-email: Gabriel Ayres <gabn2012@gmail.com>
License: MIT
Project-URL: Homepage, https://github.com/gabrielfnayres/pax-case
Project-URL: Repository, https://github.com/gabrielfnayres/pax-case
Keywords: computer-vision,object-detection,cli,yolo,vit
Classifier: Development Status :: 4 - Beta
Classifier: Intended Audience :: Developers
Classifier: License :: OSI Approved :: MIT License
Classifier: Programming Language :: Python :: 3
Classifier: Programming Language :: Python :: 3.10
Classifier: Topic :: Scientific/Engineering :: Image Recognition
Requires-Python: >=3.10
Description-Content-Type: text/markdown
Requires-Dist: datasets>=4.0.0
Requires-Dist: fastapi[standard]>=0.116.1
Requires-Dist: gradio>=5.44.1
Requires-Dist: huggingface-hub>=0.34.4
Requires-Dist: opencv-python>=4.11.0.86
Requires-Dist: pathlib>=1.0.1
Requires-Dist: pillow>=11.3.0
Requires-Dist: pytorch-lightning>=2.5.3
Requires-Dist: scikit-learn>=1.7.1
Requires-Dist: tensorboardx>=2.6.4
Requires-Dist: torch>=2.8.0
Requires-Dist: torchaudio>=2.8.0
Requires-Dist: torchmetrics>=1.8.1
Requires-Dist: transformers>=4.55.4
Requires-Dist: ultralytics>=8.3.188
Requires-Dist: uvicorn>=0.35.0

# Pax Case: Object and Car Make Classification

This project provides a vision pipeline for object and car make classification. It includes a command-line interface (CLI) for processing images and a Gradio-based web interface for interactive demonstrations.

## Problem Statement

The goal is to build a tool that can:
1.  **Object Classification**: Identify whether an image contains a car, truck, bicycle, or person.
2.  **Car Make Classification**: If the image contains a car, identify its make (e.g., Volkswagen, Chevrolet).

The solution must be scalable to handle up to 10 million images daily and be flexible enough to accommodate future classification tasks.

## Features

- **Two-Stage Pipeline**:
    - **Object Detection**: Uses a YOLOv8 model to detect objects (person, bicycle, car, truck).
    - **Car Make Classification**: Uses a Google SigLIP model fine-tuned on the Stanford Cars dataset to classify the make of detected cars.
- **CLI Tool** (`detector`):
    - Process single images or directories of images from the command line.
    - Supports configuration via a YAML file.
    - Adjustable confidence threshold for object detection.
- **Scalability and Extensibility**:
    - The architecture is designed to be scalable and extensible. See `SCALING_STRATEGY.md` for a detailed plan on scaling to 10M images/day and adding new classification types.

## Package usage
1. **Install PyPi package**
  ```bash
  pip install pax-detector
  ```

## Local Setup and Installation

1.  **Clone the repository**:
    ```bash
    git clone <repository-url>
    cd pax-case
    ```

2.  **Install dependencies**:
    It is recommended to use a virtual environment. This project uses `uv` https://docs.astral.sh/uv/getting-started/installation/ for package management.
    ```bash
    uv venv 
    source .venv/bin/activate

    uv pip install -e . # To build package locally
    uv sync # If want to install local dependencies 
    ```

## Usage

### Command-Line Interface (CLI)

The CLI tool (`detector`) is used for processing individual images.

**Basic Usage**:

To process a single image:
```bash
detector --image_path /path/to/your/image.jpg
```

To process all images in a directory:
```bash
detector --images_dir /path/to/your/images/
```

**Using a Configuration File**:

You can also run the CLI using a `config.yml` file to specify parameters.

*Example `config.yml`:*
```yaml
image_path: '/path/to/your/image.jpg'
confidence: 0.6
```

*Run with config*:
```bash
detector --config config.yml
```

**Arguments**:
- `--image_path`: Path to a single input image.
- `--images_dir`: Path to a directory of images.
- `--config`: Path to a YAML configuration file.
- `--confidence`: Confidence threshold for object detection (default: 0.5).

**Example**:
```bash
detector --image_path datasets/car-camera/images/00001.jpg --confidence 0.4
```

This will process the image and print the JSON output to the console.

## Extensibility

For details on how to add more classification types (e.g., helmet color, t-shirt color), please refer to the `SCALING_STRATEGY.md` document, which outlines the proposed architecture for extending the pipeline.
