my_project

SDOML Project

Automobile Market Analytics

Team #5

Team members

  • Miguel Zubitur
  • Magdalena Hristova
  • Diego Suarez


Project Information

What does this project do?

This project was developed for the Software Development Oriented to Machine Learning (SDOML) course.

The project analyzes the factors that influence the price of used cars, including vehicle characteristics such as age, mileage, condition, and technical features.

A machine learning model based on a neural network is also developed to predict the selling price of an automobile from its characteristics.

The project includes:

  • Data downloading
  • Data preprocessing
  • Exploratory data analysis
  • Machine learning model training
  • Model performance analysis

Project documentation

The project documentation is published using GitHub Pages:

Project Documentation

The documentation is generated using pdoc and is published from the docs/ directory.


Dataset

Data source

The dataset used in this project was obtained from Kaggle:

Automobile Market Analytics Dataset

The dataset is downloaded automatically using kagglehub.

The dataset is not stored in the repository. It can be downloaded again by following the workflow described below.


Installation

Requirements

The project requires:

  • Python 3.10 or newer
  • uv
  • Jupyter Notebook

The Python dependencies are managed through pyproject.toml and uv.lock.

Install the project

Clone the repository:

git clone https://github.com/Miguel153891/SDOML-project.git
cd SDOML-project

Install the project dependencies:

uv sync

This creates the required environment and installs the dependencies specified by the project.


How to Run the Project

The complete workflow should be executed in the following order.

1. Download the dataset

Run:

uv run python -m my_project.download_data

This downloads the dataset from Kaggle and creates:

data/
└── raw/
    └── automobile_dataset.csv

The data/raw/ directory is generated automatically if it does not exist.

2. Process and explore the data

Open:

notebooks/data_exploration.ipynb

and execute all cells.

The notebook performs the data preprocessing and exploratory data analysis.

The processed datasets are generated in:

data/
└── processed/
    ├── automobile_dataset.parquet
    └── automobile_test.parquet

The exploratory analysis also generates figures in:

reports/
└── figures/
    └── exploration/

3. Train the model

After the processed datasets have been generated, run:

uv run python -m my_project.train

The training script creates and trains the neural network.

The trained model is saved to:

models/
└── automobile_model.pth

The models/ directory is created automatically if it does not exist.

4. Analyze model performance

Open:

notebooks/performance_analysis.ipynb

and execute all cells.

The notebook loads the trained model and analyzes its predictions and performance.


Complete Workflow

For a clean setup, execute:

uv sync

Then download the dataset:

uv run python -m my_project.download_data

Then execute all cells in:

notebooks/data_exploration.ipynb

Then train the model:

uv run python -m my_project.train

Finally, execute:

notebooks/performance_analysis.ipynb

The complete workflow is:

Kaggle
   │
   ▼
download_data.py
   │
   ▼
data/raw/
   │
   ▼
data_exploration.ipynb
   │
   ├──► reports/figures/exploration/
   │
   ▼
data/processed/
   │
   ▼
train.py
   │
   ▼
models/automobile_model.pth
   │
   ▼
performance_analysis.ipynb

Project Structure

Repository structure

SDOML-project/
│
├── my_project/
│   ├── __init__.py
│   ├── dataset.py
│   ├── download_data.py
│   ├── fileManager.py
│   └── train.py
│
├── notebooks/
│   ├── data_exploration.ipynb
│   └── performance_analysis.ipynb
│
├── reports/
│   ├── export_notebooks/
│   └── figures/
│
├── docs/
│   └── Generated project documentation
│
├── LICENSE
├── README.md
├── Requirements.md
├── pyproject.toml
├── uv.lock
├── practice1.pdf
├── practice2.pdf
└── practice3.pdf

The following directories are generated locally and are not committed to Git:

data/
├── raw/
└── processed/

models/


Main Components

Python modules

my_project/download_data.py

Downloads the automobile dataset from Kaggle using kagglehub.

my_project/dataset.py

Defines the PyTorch AutomobileDataset, which loads the processed Parquet data and converts it into PyTorch tensors.

my_project/train.py

Contains the neural network and training functionality:

  • create_model() — creates the neural network.
  • train_model() — trains the model.
  • main() — loads the datasets, trains the model, and saves the trained model.

Jupyter notebooks

data_exploration.ipynb

Performs:

  • Data loading
  • Data preprocessing
  • Exploratory data analysis
  • Dataset preparation
  • Generation of exploratory figures

performance_analysis.ipynb

Loads the trained model and evaluates its predictions and performance.


Documentation Generation

Generate documentation locally

The API documentation is generated using pdoc.

Run the command

uv run pdoc -o docs my_project --docformat numpy

The generated documentation is stored in docs/ and published automatically through GitHub Pages.


Contributing

Contribution guidelines

Contributions should follow the project's Git workflow.

  1. Create a feature branch from main:
git checkout -b feature/my-change
  1. Make the required changes.
  2. Test the changes locally.
  3. Use descriptive commit messages:
git add .
git commit -m "Add descriptive change"
  1. Push the feature branch:
git push origin feature/my-change
  1. Open a pull request to main.

Changes should be focused and documented when appropriate.

Do not commit:

  • Downloaded datasets
  • Processed datasets
  • Trained model files
  • Python cache files
  • Other generated files excluded by .gitignore


License

This project is licensed under the MIT License.

See LICENSE for the complete license text.

 1import pathlib
 2
 3readme_path = pathlib.Path(__file__).parent.parent / "README.md"
 4
 5if readme_path.exists():
 6    __doc__ = readme_path.read_text(encoding="utf-8")
 7else:
 8    __doc__ = "Package documentation (README.md not found)."
 9    
10__all__ = ['dataset','train','fileManager']