my_project
SDOML Project
Automobile Market Analytics
Team #5
Team members
- Miguel Zubitur
- Magdalena Hristova
- Diego Suarez
Project Information
What does this project do?
This project was developed for the Software Development Oriented to Machine Learning (SDOML) course.
The project analyzes the factors that influence the price of used cars, including vehicle characteristics such as age, mileage, condition, and technical features.
A machine learning model based on a neural network is also developed to predict the selling price of an automobile from its characteristics.
The project includes:
- Data downloading
- Data preprocessing
- Exploratory data analysis
- Machine learning model training
- Model performance analysis
Project documentation
The project documentation is published using GitHub Pages:
The documentation is generated using pdoc and is published from the docs/ directory.
Dataset
Data source
The dataset used in this project was obtained from Kaggle:
Automobile Market Analytics Dataset
The dataset is downloaded automatically using kagglehub.
The dataset is not stored in the repository. It can be downloaded again by following the workflow described below.
Installation
Requirements
The project requires:
- Python 3.10 or newer
- uv
- Jupyter Notebook
The Python dependencies are managed through pyproject.toml and uv.lock.
Install the project
Clone the repository:
git clone https://github.com/Miguel153891/SDOML-project.git
cd SDOML-project
Install the project dependencies:
uv sync
This creates the required environment and installs the dependencies specified by the project.
How to Run the Project
The complete workflow should be executed in the following order.
1. Download the dataset
Run:
uv run python -m my_project.download_data
This downloads the dataset from Kaggle and creates:
data/
└── raw/
└── automobile_dataset.csv
The data/raw/ directory is generated automatically if it does not exist.
2. Process and explore the data
Open:
notebooks/data_exploration.ipynb
and execute all cells.
The notebook performs the data preprocessing and exploratory data analysis.
The processed datasets are generated in:
data/
└── processed/
├── automobile_dataset.parquet
└── automobile_test.parquet
The exploratory analysis also generates figures in:
reports/
└── figures/
└── exploration/
3. Train the model
After the processed datasets have been generated, run:
uv run python -m my_project.train
The training script creates and trains the neural network.
The trained model is saved to:
models/
└── automobile_model.pth
The models/ directory is created automatically if it does not exist.
4. Analyze model performance
Open:
notebooks/performance_analysis.ipynb
and execute all cells.
The notebook loads the trained model and analyzes its predictions and performance.
Complete Workflow
For a clean setup, execute:
uv sync
Then download the dataset:
uv run python -m my_project.download_data
Then execute all cells in:
notebooks/data_exploration.ipynb
Then train the model:
uv run python -m my_project.train
Finally, execute:
notebooks/performance_analysis.ipynb
The complete workflow is:
Kaggle
│
▼
download_data.py
│
▼
data/raw/
│
▼
data_exploration.ipynb
│
├──► reports/figures/exploration/
│
▼
data/processed/
│
▼
train.py
│
▼
models/automobile_model.pth
│
▼
performance_analysis.ipynb
Project Structure
Repository structure
SDOML-project/
│
├── my_project/
│ ├── __init__.py
│ ├── dataset.py
│ ├── download_data.py
│ ├── fileManager.py
│ └── train.py
│
├── notebooks/
│ ├── data_exploration.ipynb
│ └── performance_analysis.ipynb
│
├── reports/
│ ├── export_notebooks/
│ └── figures/
│
├── docs/
│ └── Generated project documentation
│
├── LICENSE
├── README.md
├── Requirements.md
├── pyproject.toml
├── uv.lock
├── practice1.pdf
├── practice2.pdf
└── practice3.pdf
The following directories are generated locally and are not committed to Git:
data/
├── raw/
└── processed/
models/
Main Components
Python modules
my_project/download_data.py
Downloads the automobile dataset from Kaggle using kagglehub.
my_project/dataset.py
Defines the PyTorch AutomobileDataset, which loads the processed Parquet data and converts it into PyTorch tensors.
my_project/train.py
Contains the neural network and training functionality:
create_model()— creates the neural network.train_model()— trains the model.main()— loads the datasets, trains the model, and saves the trained model.
Jupyter notebooks
data_exploration.ipynb
Performs:
- Data loading
- Data preprocessing
- Exploratory data analysis
- Dataset preparation
- Generation of exploratory figures
performance_analysis.ipynb
Loads the trained model and evaluates its predictions and performance.
Documentation Generation
Generate documentation locally
The API documentation is generated using pdoc.
Run the command
uv run pdoc -o docs my_project --docformat numpy
The generated documentation is stored in docs/ and published automatically through GitHub Pages.
Contributing
Contribution guidelines
Contributions should follow the project's Git workflow.
- Create a feature branch from
main:
git checkout -b feature/my-change
- Make the required changes.
- Test the changes locally.
- Use descriptive commit messages:
git add .
git commit -m "Add descriptive change"
- Push the feature branch:
git push origin feature/my-change
- Open a pull request to
main.
Changes should be focused and documented when appropriate.
Do not commit:
- Downloaded datasets
- Processed datasets
- Trained model files
- Python cache files
- Other generated files excluded by
.gitignore
License
This project is licensed under the MIT License.
See LICENSE for the complete license text.