<h1 id="Data-Exploration">Data Exploration¶</h1>
<h2 id="Setup">Setup¶</h2>
<h3 id="Imports">Imports¶</h3>
In [1]:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import math
In [2]:
import sys
sys.path.append("..")
from my_project.fileManager import FileManager
<h3 id="Auxiliar-functions">Auxiliar functions¶</h3>
In [3]:
def inspect(df):
print(f"Shape: {df.shape}")
print(f"Columns: {list(df.columns)}")
print(f"Data types:\n{df.dtypes}")
print()
<h3 id="Constants">Constants¶</h3>
In [4]:
RAW_FILE = "../data/raw/automobile_dataset"
PROCESSED_TRAIN_FILE = "../data/processed/automobile_dataset"
PROCESSED_TEST_FILE = "../data/processed/automobile_test"
EXPLORATION_IMAGES_FOLDER = "../reports/figures/exploration/"
<h2 id="Data-loading">Data loading¶</h2>
<p>We created a FileManager class to make it easier to save and load our dataset using different file formats. The class supports CSV, Parquet, Feather, and Pickle files. The class automatically adds the correct file extension, so we only need to provide the file name. This makes the code more flexible and avoids having to write different functions for each file format.</p>
In [5]:
fM = FileManager()
In [6]:
fM.set_format("csv")
df = fM.read(RAW_FILE)
display(df)
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
<thead>
<th></th>
<th>Make</th>
<th>Model</th>
<th>Year</th>
<th>Fuel_Type</th>
<th>Transmission</th>
<th>Engine_Size</th>
<th>Mileage</th>
<th>Horsepower</th>
<th>Torque</th>
<th>Owners</th>
<th>Accident_History</th>
<th>Service_History</th>
<th>Color</th>
<th>Body_Type</th>
<th>Drivetrain</th>
<th>Fuel_Efficiency</th>
<th>Location</th>
<th>Selling_Price</th>
</thead>
<tbody>
<th>0</th>
<th>1</th>
<th>2</th>
<th>3</th>
<th>4</th>
<th>...</th>
<th>5495</th>
<th>5496</th>
<th>5497</th>
<th>5498</th>
<th>5499</th>
</tbody>
<tbody>
<p>5500 rows × 18 columns</p>
| Mercedes-Benz | GLE | 2024 | Petrol | Automatic | 2.3 | 100 | 186.0 | 196.0 | 1 | 0.0 | NaN | Silver | SUV | AWD | 30.0 | IL | 64140 |
| Hyundai | Tucson | 2008 | Petrol | Automatic | 2.3 | 387035 | 189.0 | 190.0 | 4 | 0.0 | NaN | Blue | SUV | FWD | 35.0 | FL | 500 |
| Volkswagen | Golf | 2021 | Hybrid | NaN | 1.9 | 46054 | 158.0 | 153.0 | 1 | 0.0 | Partial Service | Gray | Hatchback | AWD | 43.0 | NY | 16429 |
| Chevrolet | Tahoe | 2005 | Petrol | NaN | 2.2 | 141302 | 169.0 | 165.0 | 5 | 1.0 | No Service | Blue | SUV | AWD | NaN | FL | 2199 |
| Toyota | Camry | 2022 | Petrol | Automatic | 1.9 | 32813 | 149.0 | 141.0 | 1 | 0.0 | Full Service | Brown | Sedan | FWD | 38.0 | CA | 21792 |
| ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |
| Volkswagen | Tiguan | 2022 | Petrol | Automatic | NaN | 29506 | 161.0 | NaN | 1 | 0.0 | Partial Service | Black | SUV | AWD | 33.0 | GA | 20333 |
| Nissan | Pathfinder | 2006 | Petrol | Automatic | 3.1 | 287332 | 243.0 | 227.0 | 5 | 1.0 | No Service | Gray | SUV | AWD | NaN | NY | 500 |
| BMW | X3 | 2013 | Petrol | Automatic | 2.2 | 95961 | 171.0 | 152.0 | 2 | 1.0 | No Service | Blue | SUV | AWD | 32.0 | MI | 9882 |
| Volkswagen | Golf | 2020 | Hybrid | Automatic | 1.5 | 56643 | 135.0 | 125.0 | 2 | 1.0 | No Service | White | Hatchback | FWD | 57.0 | NY | 10474 |
| Chevrolet | Tahoe | 2011 | Petrol | Automatic | 2.1 | 169929 | 168.0 | 158.0 | 3 | 1.0 | NaN | Gray | SUV | AWD | NaN | MI | 6193 |
<p>Our dataset focuses on analyzing the selling price of used cars. It contains 5500 records and 18 variables that describe different aspects of each vehicle: its specifications, ownership history, condition, and other contextual information. It contains a mixture of numerical and categorical values. The most important variable is the Selling_Price, which is the target value which our machine learning models will try to predict as new instances come through.</p>
In [7]:
inspect(df)
Shape: (5500, 18) Columns: ['Make', 'Model', 'Year', 'Fuel_Type', 'Transmission', 'Engine_Size', 'Mileage', 'Horsepower', 'Torque', 'Owners', 'Accident_History', 'Service_History', 'Color', 'Body_Type', 'Drivetrain', 'Fuel_Efficiency', 'Location', 'Selling_Price'] Data types: Make object Model object Year int64 Fuel_Type object Transmission object Engine_Size float64 Mileage int64 Horsepower float64 Torque float64 Owners int64 Accident_History float64 Service_History object Color object Body_Type object Drivetrain object Fuel_Efficiency float64 Location object Selling_Price int64 dtype: object
<h2 id="Memory-optimization">Memory optimization¶</h2>
<p>We check how much memory our dataset is using and then try to reduce it by optimizing the data types. First, we calculate the initial memory usage of the DataFrame. Then, we go through all the columns and change their data types to more efficient ones when possible. This helps make the dataset more memory-efficient, which can be useful when working with larger datasets.</p>
In [8]:
# Initial memory usage
initial_size = df.memory_usage(deep=True).sum() / (1024 ** 2)
print(f"Initial memory usage: {initial_size:.2f} MB")
# Optimize data types
for col in list(df.columns):
if df.dtypes[col] == "int64":
df[col] = pd.to_numeric(df[col], downcast='integer')
size = df.memory_usage(deep=True).sum() / (1024 ** 2)
#print(f"Memory usage after downcasting to integer: {size:.2f} MB")
if df.dtypes[col] == "float64":
df[col] = pd.to_numeric(df[col], downcast='float')
size = df.memory_usage(deep=True).sum() / (1024 ** 2)
#print(f"Memory usage after downcasting to float: {size:.2f} MB")
if df.dtypes[col] == "str" or df.dtypes[col] == "object":
df[col] = df[col].astype('category')
size = df.memory_usage(deep=True).sum() / (1024 ** 2)
#print(f"Memory usage after changing to categorical column: {size:.2f} MB")
# Optimized memory usage
final_size = df.memory_usage(deep=True).sum() / (1024 ** 2)
print(f"Final memory usage: {final_size:.2f} MB")
print(f"Reduction of {(1 - final_size / initial_size) * 100:.2f}%")
Initial memory usage: 2.90 MB Final memory usage: 0.22 MB Reduction of 92.51%
In [9]:
inspect(df)
Shape: (5500, 18) Columns: ['Make', 'Model', 'Year', 'Fuel_Type', 'Transmission', 'Engine_Size', 'Mileage', 'Horsepower', 'Torque', 'Owners', 'Accident_History', 'Service_History', 'Color', 'Body_Type', 'Drivetrain', 'Fuel_Efficiency', 'Location', 'Selling_Price'] Data types: Make category Model category Year int16 Fuel_Type category Transmission category Engine_Size float32 Mileage int32 Horsepower float32 Torque float32 Owners int8 Accident_History float32 Service_History category Color category Body_Type category Drivetrain category Fuel_Efficiency float32 Location category Selling_Price int32 dtype: object
<h2 id="Missing-values">Missing values¶</h2>
<p>Before doing anything, we checked the dataset for missing values using insull().sum(). This allows us to identify in which variables we had these kind of observations. The preprocessing step is important because most machine learning algorithms cannot work directly with missing data.</p>
- The numerical missing values have been replaced with the mean value of the corresponding column. This is because in the dataset there weren't extreme outliers.
- The categorical variables missing values have been dropped in order to avoid making assumptions about the missing values.
In [10]:
print("Missing values:")
print(df.isnull().sum() > 0)
print()
print("Fill missing numerical values:")
dictionary = {}
# Numerical fixes
dfk = df.isnull().sum() > 0
for col in list(df.columns):
if dfk[col] and df.dtypes[col] != "category":
dictionary[col] = dfk[col].mean()
df.fillna(dictionary)
# Categorical fixes
print("Drop rows with any missing categorical values:")
print("Categorical Fixes:")
df = df.dropna()
display(df)
# Missing values:
print(df.isnull().sum() > 0)
print()
Missing values: Make False Model False Year False Fuel_Type False Transmission True Engine_Size True Mileage False Horsepower True Torque True Owners False Accident_History True Service_History True Color True Body_Type False Drivetrain False Fuel_Efficiency True Location True Selling_Price False dtype: bool Fill missing numerical values: Drop rows with any missing categorical values: Categorical Fixes:
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
<thead>
<th></th>
<th>Make</th>
<th>Model</th>
<th>Year</th>
<th>Fuel_Type</th>
<th>Transmission</th>
<th>Engine_Size</th>
<th>Mileage</th>
<th>Horsepower</th>
<th>Torque</th>
<th>Owners</th>
<th>Accident_History</th>
<th>Service_History</th>
<th>Color</th>
<th>Body_Type</th>
<th>Drivetrain</th>
<th>Fuel_Efficiency</th>
<th>Location</th>
<th>Selling_Price</th>
</thead>
<tbody>
<th>4</th>
<th>7</th>
<th>15</th>
<th>17</th>
<th>21</th>
<th>...</th>
<th>5478</th>
<th>5479</th>
<th>5491</th>
<th>5497</th>
<th>5498</th>
</tbody>
<tbody>
<p>1329 rows × 18 columns</p>
| Toyota | Camry | 2022 | Petrol | Automatic | 1.9 | 32813 | 149.0 | 141.0 | 1 | 0.0 | Full Service | Brown | Sedan | FWD | 38.0 | CA | 21792 |
| Mercedes-Benz | C-Class | 2022 | Petrol | Automatic | 2.3 | 32133 | 180.0 | 162.0 | 1 | 0.0 | Partial Service | Green | Sedan | FWD | 35.0 | NY | 31416 |
| Mercedes-Benz | C-Class | 2010 | Petrol | Manual | 2.0 | 17138 | 169.0 | 160.0 | 3 | 0.0 | Partial Service | Blue | Sedan | FWD | 27.0 | NC | 11728 |
| Chevrolet | Tahoe | 2011 | Petrol | Automatic | 3.3 | 264441 | 258.0 | 252.0 | 4 | 0.0 | Partial Service | Gray | SUV | AWD | 22.0 | TX | 3395 |
| Audi | Q7 | 2016 | Petrol | Automatic | 3.1 | 111537 | 243.0 | 219.0 | 2 | 0.0 | Partial Service | Blue | SUV | AWD | 21.0 | FL | 23314 |
| ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |
| Audi | A6 | 2010 | Petrol | Manual | 2.1 | 155818 | 169.0 | 176.0 | 3 | 0.0 | Full Service | Red | Sedan | FWD | 41.0 | FL | 9714 |
| Honda | Accord | 2023 | Petrol | Automatic | 2.4 | 100 | 199.0 | 192.0 | 1 | 0.0 | Partial Service | Green | Sedan | FWD | 33.0 | GA | 23937 |
| Nissan | Pathfinder | 2020 | Petrol | Automatic | 2.7 | 79498 | 212.0 | 204.0 | 1 | 0.0 | Partial Service | Gray | SUV | AWD | 24.0 | OH | 18345 |
| BMW | X3 | 2013 | Petrol | Automatic | 2.2 | 95961 | 171.0 | 152.0 | 2 | 1.0 | No Service | Blue | SUV | AWD | 32.0 | MI | 9882 |
| Volkswagen | Golf | 2020 | Hybrid | Automatic | 1.5 | 56643 | 135.0 | 125.0 | 2 | 1.0 | No Service | White | Hatchback | FWD | 57.0 | NY | 10474 |
Make False Model False Year False Fuel_Type False Transmission False Engine_Size False Mileage False Horsepower False Torque False Owners False Accident_History False Service_History False Color False Body_Type False Drivetrain False Fuel_Efficiency False Location False Selling_Price False dtype: bool
<h2 id="Exploration">Exploration¶</h2>
<h4 id="1.-Car-Model-Count">1. Car Model Count¶</h4>
In [11]:
plt.figure(figsize=(10, 5))
car_counts = df["Make"].value_counts()
plt.bar(car_counts.index, car_counts.values , color='royalblue')
plt.xlabel("Car Model")
plt.ylabel("Count")
plt.xticks(rotation=90)
plt.title("Automobile Dataset: Car Model Count")
plt.tight_layout()
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}car_model_count.png", dpi=300)
plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
<p>This graph shows the number of cars for each brand in the dataset. We can see that the number of cars is quite similar between the different brands, so there is not a big difference in the representation of each brand. This means that the dataset is relatively balanced in terms of car brands, and it is not heavily dominated by one specific manufacturer</p>
<h4 id="2.-Heatmap">2. Heatmap¶</h4>
In [12]:
corr = df.select_dtypes(include=np.number).corr()
plt.figure(figsize=(7, 5))
sns.heatmap( corr, annot=True, cmap='coolwarm' , fmt=".2f")
plt.tight_layout()
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}heatmap.png", dpi=300)
# plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
<p>We can see some high correlations between several variables:</p>
- Horsepower and Torque have a very strong positive correlation (0.98). This means that cars with higher horsepower usually also have higher torque.
- Selling_Price and Year have a positive correlation (0.81). In general, newer cars tend to have a higher selling price.
- Selling_Price has a negative correlation with Mileage (-0.69) and Owners (-0.68). This means that cars with higher mileage or more previous owners tend to have a lower selling price.
- Engine_Size and Fuel_Efficiency have a negative correlation (-0.74). As the engine size increases, fuel efficiency tends to decrease.
- Engine_Size also has a positive correlation with Horsepower (0.67) and Torque (0.56). This suggests that cars with larger engines tend to have more horsepower and torque.
<h4 id="3.-Pruebas-de-Scatterplot">3. Pruebas de Scatterplot¶</h4>
In [13]:
makes = df["Make"].unique()
n = len(makes)
cols = 3
rows = math.ceil(n / cols)
fig, axes = plt.subplots(rows, cols, figsize=(18, 5 * rows))
axes = axes.flatten()
for ax, make in zip(axes, makes):
df_make = df[df["Make"] == make]
sns.scatterplot(
data=df_make,
x="Selling_Price",
y="Mileage",
color="skyblue",
ax=ax
)
ax.set_title(make)
# Eliminar ejes vacíos
for ax in axes[len(makes):]:
fig.delaxes(ax)
plt.tight_layout()
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}scatterplots.png", dpi=300)
# plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
<h4 id="%C2%A04.-Boxplot-by-Category"> 4. Boxplot by Category¶</h4>
In [14]:
sns.set_theme(style="whitegrid")
plt.figure(figsize=(7, 5))
sns.boxplot(data=df, x="Make", y="Selling_Price", palette="Set2")
plt.xticks(rotation=90)
plt.title("Selling Price vs Car ")
plt.tight_layout()
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}selling_vs_car.png", dpi=300)
# plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
/var/folders/mx/d1fc66x17bs8hc5yn2qp1c4h0000gn/T/ipykernel_2361/1602545604.py:3: FutureWarning: Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `x` variable to `hue` and set `legend=False` for the same effect. sns.boxplot(data=df, x="Make", y="Selling_Price", palette="Set2")
<p>Audi, BMW, and Mercedes-Benz have the highest selling prices in the dataset. These brands also include some of the most expensive vehicles, reflecting their position in the premium car market.</p>
<h4 id="5.-Seaborn-Pairplot-Experiments">5. Seaborn Pairplot Experiments¶</h4>
In [15]:
sns.set_theme(style="white")
sns.pairplot(df, hue="Selling_Price", height=2.5)
plt.suptitle("Iris dataset: Pairplot", y=1.02)
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}pairplots.png", dpi=300)
# plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
<h2 id="Categorical-to-numerical">Categorical to numerical¶</h2>
In [16]:
print(df["Make"].nunique())
print(df["Color"].nunique())
print(df["Model"].nunique())
print(df["Service_History"].nunique())
print(df["Body_Type"].nunique())
print(df["Drivetrain"].nunique())
print(df["Location"].nunique())
10 8 40 3 5 4 10
<p>Machine learning models cannot work directly with categorical data, so these variables must be converted into a numerical format before training. That's why we did some encoding.</p>
- For Service_History, we used ordinal encoding because the categories have a natural order. A vehicle with no service history provides less information than one with a partial service history.
- For the remaining categorical variables, we applied One-Hot Encoding. This method creates a separate binary column for each category and avoids introducing relationships between categories that do not actually exist.
In [17]:
# Ordinal for service (from no service to full service, hierarchy is preserved)
service_mapping = {
"No Service": 0,
"Partial Service": 1,
"Full Service": 2
}
df["Service_History"] = df["Service_History"].map(service_mapping)
# OHE for the other categorical columns
categorical_columns = [
"Make",
"Model",
"Fuel_Type",
"Transmission",
"Color",
"Body_Type",
"Drivetrain",
"Location"
]
df = pd.get_dummies(
df,
columns=categorical_columns,
dtype=int
)
In [18]:
print(df.head())
print(df.dtypes)
print(df.shape)
print(df.columns)
Year Engine_Size Mileage Horsepower Torque Owners Accident_History \
4 2022 1.9 32813 149.0 141.0 1 0.0
7 2022 2.3 32133 180.0 162.0 1 0.0
15 2010 2.0 17138 169.0 160.0 3 0.0
17 2011 3.3 264441 258.0 252.0 4 0.0
21 2016 3.1 111537 243.0 219.0 2 0.0
Service_History Fuel_Efficiency Selling_Price ... Location_CA \
4 2 38.0 21792 ... 1
7 1 35.0 31416 ... 0
15 1 27.0 11728 ... 0
17 1 22.0 3395 ... 0
21 1 21.0 23314 ... 0
Location_FL Location_GA Location_IL Location_MI Location_NC \
4 0 0 0 0 0
7 0 0 0 0 0
15 0 0 0 0 1
17 0 0 0 0 0
21 1 0 0 0 0
Location_NY Location_OH Location_PA Location_TX
4 0 0 0 0
7 1 0 0 0
15 0 0 0 0
17 0 0 0 1
21 0 0 0 0
[5 rows x 93 columns]
Year int16
Engine_Size float32
Mileage int32
Horsepower float32
Torque float32
...
Location_NC int64
Location_NY int64
Location_OH int64
Location_PA int64
Location_TX int64
Length: 93, dtype: object
(1329, 93)
Index(['Year', 'Engine_Size', 'Mileage', 'Horsepower', 'Torque', 'Owners',
'Accident_History', 'Service_History', 'Fuel_Efficiency',
'Selling_Price', 'Make_Audi', 'Make_BMW', 'Make_Chevrolet', 'Make_Ford',
'Make_Honda', 'Make_Hyundai', 'Make_Mercedes-Benz', 'Make_Nissan',
'Make_Toyota', 'Make_Volkswagen', 'Model_3 Series', 'Model_5 Series',
'Model_A4', 'Model_A6', 'Model_Accord', 'Model_Altima', 'Model_Atlas',
'Model_C-Class', 'Model_CR-V', 'Model_Camry', 'Model_Civic',
'Model_Corolla', 'Model_E-Class', 'Model_Elantra', 'Model_Equinox',
'Model_Escape', 'Model_Explorer', 'Model_F-150', 'Model_GLC',
'Model_GLE', 'Model_Golf', 'Model_Highlander', 'Model_Malibu',
'Model_Mustang', 'Model_Passat', 'Model_Pathfinder', 'Model_Pilot',
'Model_Q5', 'Model_Q7', 'Model_RAV4', 'Model_Rogue', 'Model_Santa Fe',
'Model_Sentra', 'Model_Silverado', 'Model_Sonata', 'Model_Tahoe',
'Model_Tiguan', 'Model_Tucson', 'Model_X3', 'Model_X5',
'Fuel_Type_Diesel', 'Fuel_Type_Electric', 'Fuel_Type_Hybrid',
'Fuel_Type_Petrol', 'Transmission_Automatic', 'Transmission_Manual',
'Color_Black', 'Color_Blue', 'Color_Brown', 'Color_Gray', 'Color_Green',
'Color_Red', 'Color_Silver', 'Color_White', 'Body_Type_Coupe',
'Body_Type_Hatchback', 'Body_Type_SUV', 'Body_Type_Sedan',
'Body_Type_Truck', 'Drivetrain_4WD', 'Drivetrain_AWD', 'Drivetrain_FWD',
'Drivetrain_RWD', 'Location_CA', 'Location_FL', 'Location_GA',
'Location_IL', 'Location_MI', 'Location_NC', 'Location_NY',
'Location_OH', 'Location_PA', 'Location_TX'],
dtype='object')
<h2 id="Save-processed-data">Save processed data¶</h2>
In [19]:
from sklearn.model_selection import train_test_split
# 80% training data
# 20% testing data
train_df, test_df = train_test_split(
df,
test_size=0.2,
random_state=22
)
In [20]:
fM.set_format("parquet")
fM.write(train_df, PROCESSED_TRAIN_FILE)
fM.write(test_df, PROCESSED_TEST_FILE)
<h2 id="Check-it-was-correctly-saved">Check it was correctly saved¶</h2>
In [21]:
fM.set_format("parquet")
fdf = fM.read(PROCESSED_TRAIN_FILE)
display(fdf)
inspect(fdf)
<style scoped>
.dataframe tbody tr th:only-of-type {
vertical-align: middle;
}
.dataframe tbody tr th {
vertical-align: top;
}
.dataframe thead th {
text-align: right;
}
</style>
<thead>
<th></th>
<th>Year</th>
<th>Engine_Size</th>
<th>Mileage</th>
<th>Horsepower</th>
<th>Torque</th>
<th>Owners</th>
<th>Accident_History</th>
<th>Service_History</th>
<th>Fuel_Efficiency</th>
<th>Selling_Price</th>
<th>...</th>
<th>Location_CA</th>
<th>Location_FL</th>
<th>Location_GA</th>
<th>Location_IL</th>
<th>Location_MI</th>
<th>Location_NC</th>
<th>Location_NY</th>
<th>Location_OH</th>
<th>Location_PA</th>
<th>Location_TX</th>
</thead>
<tbody>
<th>633</th>
<th>1028</th>
<th>4550</th>
<th>1056</th>
<th>1063</th>
<th>...</th>
<th>1461</th>
<th>4038</th>
<th>3372</th>
<th>497</th>
<th>3692</th>
</tbody>
<tbody>
<p>1063 rows × 93 columns</p>
| 2008 | 2.0 | 20347 | 167.0 | 151.0 | 5 | 1.0 | 1 | 35.0 | 3428 | ... | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2016 | 2.2 | 64385 | 162.0 | 167.0 | 4 | 1.0 | 1 | 29.0 | 5850 | ... | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 |
| 2013 | 2.5 | 158156 | 187.0 | 172.0 | 3 | 0.0 | 1 | 22.0 | 1281 | ... | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2013 | 2.5 | 131422 | 209.0 | 189.0 | 3 | 1.0 | 1 | 27.0 | 7586 | ... | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2008 | 3.5 | 270715 | 263.0 | 239.0 | 4 | 0.0 | 0 | 19.0 | 500 | ... | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... | ... |
| 2005 | 3.3 | 377675 | 250.0 | 231.0 | 3 | 0.0 | 1 | 24.0 | 500 | ... | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 |
| 2017 | 1.6 | 87365 | 122.0 | 103.0 | 1 | 0.0 | 0 | 43.0 | 15518 | ... | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| 2023 | 3.2 | 3623 | 243.0 | 246.0 | 1 | 0.0 | 1 | 23.0 | 24300 | ... | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| 2011 | 3.3 | 229650 | 238.0 | 223.0 | 3 | 1.0 | 1 | 24.0 | 7157 | ... | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 |
| 2005 | 2.2 | 410673 | 175.0 | 169.0 | 5 | 0.0 | 1 | 33.0 | 500 | ... | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
Shape: (1063, 93)
Columns: ['Year', 'Engine_Size', 'Mileage', 'Horsepower', 'Torque', 'Owners', 'Accident_History', 'Service_History', 'Fuel_Efficiency', 'Selling_Price', 'Make_Audi', 'Make_BMW', 'Make_Chevrolet', 'Make_Ford', 'Make_Honda', 'Make_Hyundai', 'Make_Mercedes-Benz', 'Make_Nissan', 'Make_Toyota', 'Make_Volkswagen', 'Model_3 Series', 'Model_5 Series', 'Model_A4', 'Model_A6', 'Model_Accord', 'Model_Altima', 'Model_Atlas', 'Model_C-Class', 'Model_CR-V', 'Model_Camry', 'Model_Civic', 'Model_Corolla', 'Model_E-Class', 'Model_Elantra', 'Model_Equinox', 'Model_Escape', 'Model_Explorer', 'Model_F-150', 'Model_GLC', 'Model_GLE', 'Model_Golf', 'Model_Highlander', 'Model_Malibu', 'Model_Mustang', 'Model_Passat', 'Model_Pathfinder', 'Model_Pilot', 'Model_Q5', 'Model_Q7', 'Model_RAV4', 'Model_Rogue', 'Model_Santa Fe', 'Model_Sentra', 'Model_Silverado', 'Model_Sonata', 'Model_Tahoe', 'Model_Tiguan', 'Model_Tucson', 'Model_X3', 'Model_X5', 'Fuel_Type_Diesel', 'Fuel_Type_Electric', 'Fuel_Type_Hybrid', 'Fuel_Type_Petrol', 'Transmission_Automatic', 'Transmission_Manual', 'Color_Black', 'Color_Blue', 'Color_Brown', 'Color_Gray', 'Color_Green', 'Color_Red', 'Color_Silver', 'Color_White', 'Body_Type_Coupe', 'Body_Type_Hatchback', 'Body_Type_SUV', 'Body_Type_Sedan', 'Body_Type_Truck', 'Drivetrain_4WD', 'Drivetrain_AWD', 'Drivetrain_FWD', 'Drivetrain_RWD', 'Location_CA', 'Location_FL', 'Location_GA', 'Location_IL', 'Location_MI', 'Location_NC', 'Location_NY', 'Location_OH', 'Location_PA', 'Location_TX']
Data types:
Year int16
Engine_Size float32
Mileage int32
Horsepower float32
Torque float32
...
Location_NC int64
Location_NY int64
Location_OH int64
Location_PA int64
Location_TX int64
Length: 93, dtype: object