<h1 id="Data-Exploration">Data Exploration¶</h1>
<h2 id="Setup">Setup¶</h2>
<h3 id="Imports">Imports¶</h3>
In [1]:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
import math
In [2]:
import sys
sys.path.append("..")
from my_project.fileManager import FileManager
<h3 id="Auxiliar-functions">Auxiliar functions¶</h3>
In [3]:
def inspect(df):
    print(f"Shape: {df.shape}")
    print(f"Columns: {list(df.columns)}")
    print(f"Data types:\n{df.dtypes}")
    print()
<h3 id="Constants">Constants¶</h3>
In [4]:
RAW_FILE = "../data/raw/automobile_dataset"
PROCESSED_TRAIN_FILE = "../data/processed/automobile_dataset"
PROCESSED_TEST_FILE = "../data/processed/automobile_test"
EXPLORATION_IMAGES_FOLDER = "../reports/figures/exploration/"
<h2 id="Data-loading">Data loading¶</h2>
<p>We created a FileManager class to make it easier to save and load our dataset using different file formats. The class supports CSV, Parquet, Feather, and Pickle files. The class automatically adds the correct file extension, so we only need to provide the file name. This makes the code more flexible and avoids having to write different functions for each file format.</p>
In [5]:
fM = FileManager()
In [6]:
fM.set_format("csv")
df = fM.read(RAW_FILE)

display(df)
<style scoped> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style> <thead> <th></th> <th>Make</th> <th>Model</th> <th>Year</th> <th>Fuel_Type</th> <th>Transmission</th> <th>Engine_Size</th> <th>Mileage</th> <th>Horsepower</th> <th>Torque</th> <th>Owners</th> <th>Accident_History</th> <th>Service_History</th> <th>Color</th> <th>Body_Type</th> <th>Drivetrain</th> <th>Fuel_Efficiency</th> <th>Location</th> <th>Selling_Price</th> </thead> <tbody> <th>0</th> <th>1</th> <th>2</th> <th>3</th> <th>4</th> <th>...</th> <th>5495</th> <th>5496</th> <th>5497</th> <th>5498</th> <th>5499</th> </tbody> <tbody></tbody>
Mercedes-Benz GLE 2024 Petrol Automatic 2.3 100 186.0 196.0 1 0.0 NaN Silver SUV AWD 30.0 IL 64140
Hyundai Tucson 2008 Petrol Automatic 2.3 387035 189.0 190.0 4 0.0 NaN Blue SUV FWD 35.0 FL 500
Volkswagen Golf 2021 Hybrid NaN 1.9 46054 158.0 153.0 1 0.0 Partial Service Gray Hatchback AWD 43.0 NY 16429
Chevrolet Tahoe 2005 Petrol NaN 2.2 141302 169.0 165.0 5 1.0 No Service Blue SUV AWD NaN FL 2199
Toyota Camry 2022 Petrol Automatic 1.9 32813 149.0 141.0 1 0.0 Full Service Brown Sedan FWD 38.0 CA 21792
... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ...
Volkswagen Tiguan 2022 Petrol Automatic NaN 29506 161.0 NaN 1 0.0 Partial Service Black SUV AWD 33.0 GA 20333
Nissan Pathfinder 2006 Petrol Automatic 3.1 287332 243.0 227.0 5 1.0 No Service Gray SUV AWD NaN NY 500
BMW X3 2013 Petrol Automatic 2.2 95961 171.0 152.0 2 1.0 No Service Blue SUV AWD 32.0 MI 9882
Volkswagen Golf 2020 Hybrid Automatic 1.5 56643 135.0 125.0 2 1.0 No Service White Hatchback FWD 57.0 NY 10474
Chevrolet Tahoe 2011 Petrol Automatic 2.1 169929 168.0 158.0 3 1.0 NaN Gray SUV AWD NaN MI 6193
<p>5500 rows × 18 columns</p>
<p>Our dataset focuses on analyzing the selling price of used cars. It contains 5500 records and 18 variables that describe different aspects of each vehicle: its specifications, ownership history, condition, and other contextual information. It contains a mixture of numerical and categorical values. The most important variable is the Selling_Price, which is the target value which our machine learning models will try to predict as new instances come through.</p>
In [7]:
inspect(df)
Shape: (5500, 18)
Columns: ['Make', 'Model', 'Year', 'Fuel_Type', 'Transmission', 'Engine_Size', 'Mileage', 'Horsepower', 'Torque', 'Owners', 'Accident_History', 'Service_History', 'Color', 'Body_Type', 'Drivetrain', 'Fuel_Efficiency', 'Location', 'Selling_Price']
Data types:
Make                 object
Model                object
Year                  int64
Fuel_Type            object
Transmission         object
Engine_Size         float64
Mileage               int64
Horsepower          float64
Torque              float64
Owners                int64
Accident_History    float64
Service_History      object
Color                object
Body_Type            object
Drivetrain           object
Fuel_Efficiency     float64
Location             object
Selling_Price         int64
dtype: object

<h2 id="Memory-optimization">Memory optimization¶</h2>
<p>We check how much memory our dataset is using and then try to reduce it by optimizing the data types. First, we calculate the initial memory usage of the DataFrame. Then, we go through all the columns and change their data types to more efficient ones when possible. This helps make the dataset more memory-efficient, which can be useful when working with larger datasets.</p>
In [8]:
# Initial memory usage
initial_size = df.memory_usage(deep=True).sum() / (1024 ** 2)
print(f"Initial memory usage: {initial_size:.2f} MB")

# Optimize data types
for col in list(df.columns):
    if df.dtypes[col] == "int64":
        df[col] = pd.to_numeric(df[col], downcast='integer')
        size = df.memory_usage(deep=True).sum() / (1024 ** 2)
        #print(f"Memory usage after downcasting to integer: {size:.2f} MB")
    if df.dtypes[col] == "float64":
        df[col] = pd.to_numeric(df[col], downcast='float')
        size = df.memory_usage(deep=True).sum() / (1024 ** 2)
        #print(f"Memory usage after downcasting to float: {size:.2f} MB")
    if df.dtypes[col] == "str" or df.dtypes[col] == "object":
        df[col] = df[col].astype('category')
        size = df.memory_usage(deep=True).sum() / (1024 ** 2)
        #print(f"Memory usage after changing to categorical column: {size:.2f} MB")

# Optimized memory usage
final_size = df.memory_usage(deep=True).sum() / (1024 ** 2)
print(f"Final memory usage: {final_size:.2f} MB")
print(f"Reduction of {(1 - final_size / initial_size) * 100:.2f}%")
Initial memory usage: 2.90 MB
Final memory usage: 0.22 MB
Reduction of 92.51%
In [9]:
inspect(df)
Shape: (5500, 18)
Columns: ['Make', 'Model', 'Year', 'Fuel_Type', 'Transmission', 'Engine_Size', 'Mileage', 'Horsepower', 'Torque', 'Owners', 'Accident_History', 'Service_History', 'Color', 'Body_Type', 'Drivetrain', 'Fuel_Efficiency', 'Location', 'Selling_Price']
Data types:
Make                category
Model               category
Year                   int16
Fuel_Type           category
Transmission        category
Engine_Size          float32
Mileage                int32
Horsepower           float32
Torque               float32
Owners                  int8
Accident_History     float32
Service_History     category
Color               category
Body_Type           category
Drivetrain          category
Fuel_Efficiency      float32
Location            category
Selling_Price          int32
dtype: object

<h2 id="Missing-values">Missing values¶</h2>
<p>Before doing anything, we checked the dataset for missing values using insull().sum(). This allows us to identify in which variables we had these kind of observations. The preprocessing step is important because most machine learning algorithms cannot work directly with missing data.</p>
  • The numerical missing values have been replaced with the mean value of the corresponding column. This is because in the dataset there weren't extreme outliers.
  • The categorical variables missing values have been dropped in order to avoid making assumptions about the missing values.
In [10]:
print("Missing values:")
print(df.isnull().sum() > 0)
print()

print("Fill missing numerical values:")
dictionary = {}

# Numerical fixes
dfk = df.isnull().sum() > 0
for col in list(df.columns):
    if dfk[col] and df.dtypes[col] != "category":
        dictionary[col] = dfk[col].mean()
df.fillna(dictionary)

# Categorical fixes
print("Drop rows with any missing categorical values:")
print("Categorical Fixes:")
df = df.dropna()
display(df)

# Missing values:
print(df.isnull().sum() > 0)
print()
Missing values:
Make                False
Model               False
Year                False
Fuel_Type           False
Transmission         True
Engine_Size          True
Mileage             False
Horsepower           True
Torque               True
Owners              False
Accident_History     True
Service_History      True
Color                True
Body_Type           False
Drivetrain          False
Fuel_Efficiency      True
Location             True
Selling_Price       False
dtype: bool

Fill missing numerical values:
Drop rows with any missing categorical values:
Categorical Fixes:
<style scoped> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style> <thead> <th></th> <th>Make</th> <th>Model</th> <th>Year</th> <th>Fuel_Type</th> <th>Transmission</th> <th>Engine_Size</th> <th>Mileage</th> <th>Horsepower</th> <th>Torque</th> <th>Owners</th> <th>Accident_History</th> <th>Service_History</th> <th>Color</th> <th>Body_Type</th> <th>Drivetrain</th> <th>Fuel_Efficiency</th> <th>Location</th> <th>Selling_Price</th> </thead> <tbody> <th>4</th> <th>7</th> <th>15</th> <th>17</th> <th>21</th> <th>...</th> <th>5478</th> <th>5479</th> <th>5491</th> <th>5497</th> <th>5498</th> </tbody> <tbody></tbody>
Toyota Camry 2022 Petrol Automatic 1.9 32813 149.0 141.0 1 0.0 Full Service Brown Sedan FWD 38.0 CA 21792
Mercedes-Benz C-Class 2022 Petrol Automatic 2.3 32133 180.0 162.0 1 0.0 Partial Service Green Sedan FWD 35.0 NY 31416
Mercedes-Benz C-Class 2010 Petrol Manual 2.0 17138 169.0 160.0 3 0.0 Partial Service Blue Sedan FWD 27.0 NC 11728
Chevrolet Tahoe 2011 Petrol Automatic 3.3 264441 258.0 252.0 4 0.0 Partial Service Gray SUV AWD 22.0 TX 3395
Audi Q7 2016 Petrol Automatic 3.1 111537 243.0 219.0 2 0.0 Partial Service Blue SUV AWD 21.0 FL 23314
... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ...
Audi A6 2010 Petrol Manual 2.1 155818 169.0 176.0 3 0.0 Full Service Red Sedan FWD 41.0 FL 9714
Honda Accord 2023 Petrol Automatic 2.4 100 199.0 192.0 1 0.0 Partial Service Green Sedan FWD 33.0 GA 23937
Nissan Pathfinder 2020 Petrol Automatic 2.7 79498 212.0 204.0 1 0.0 Partial Service Gray SUV AWD 24.0 OH 18345
BMW X3 2013 Petrol Automatic 2.2 95961 171.0 152.0 2 1.0 No Service Blue SUV AWD 32.0 MI 9882
Volkswagen Golf 2020 Hybrid Automatic 1.5 56643 135.0 125.0 2 1.0 No Service White Hatchback FWD 57.0 NY 10474
<p>1329 rows × 18 columns</p>
Make                False
Model               False
Year                False
Fuel_Type           False
Transmission        False
Engine_Size         False
Mileage             False
Horsepower          False
Torque              False
Owners              False
Accident_History    False
Service_History     False
Color               False
Body_Type           False
Drivetrain          False
Fuel_Efficiency     False
Location            False
Selling_Price       False
dtype: bool

<h2 id="Exploration">Exploration¶</h2>
<h4 id="1.-Car-Model-Count">1. Car Model Count¶</h4>
In [11]:
plt.figure(figsize=(10, 5))
car_counts = df["Make"].value_counts()
plt.bar(car_counts.index, car_counts.values , color='royalblue')
plt.xlabel("Car Model")
plt.ylabel("Count")
plt.xticks(rotation=90) 
plt.title("Automobile Dataset: Car Model Count")
plt.tight_layout()
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}car_model_count.png", dpi=300)

plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
<p>This graph shows the number of cars for each brand in the dataset. We can see that the number of cars is quite similar between the different brands, so there is not a big difference in the representation of each brand. This means that the dataset is relatively balanced in terms of car brands, and it is not heavily dominated by one specific manufacturer</p>
<h4 id="2.-Heatmap">2. Heatmap¶</h4>
In [12]:
corr = df.select_dtypes(include=np.number).corr()

plt.figure(figsize=(7, 5))

sns.heatmap( corr, annot=True, cmap='coolwarm' , fmt=".2f")

plt.tight_layout()
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}heatmap.png", dpi=300)

# plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
No description has been provided for this image
<p>We can see some high correlations between several variables:</p>
  • Horsepower and Torque have a very strong positive correlation (0.98). This means that cars with higher horsepower usually also have higher torque.
  • Selling_Price and Year have a positive correlation (0.81). In general, newer cars tend to have a higher selling price.
  • Selling_Price has a negative correlation with Mileage (-0.69) and Owners (-0.68). This means that cars with higher mileage or more previous owners tend to have a lower selling price.
  • Engine_Size and Fuel_Efficiency have a negative correlation (-0.74). As the engine size increases, fuel efficiency tends to decrease.
  • Engine_Size also has a positive correlation with Horsepower (0.67) and Torque (0.56). This suggests that cars with larger engines tend to have more horsepower and torque.
<h4 id="3.-Pruebas-de-Scatterplot">3. Pruebas de Scatterplot¶</h4>
In [13]:
makes = df["Make"].unique()
n = len(makes)

cols = 3
rows = math.ceil(n / cols)

fig, axes = plt.subplots(rows, cols, figsize=(18, 5 * rows))
axes = axes.flatten()

for ax, make in zip(axes, makes):
    df_make = df[df["Make"] == make]

    sns.scatterplot(
        data=df_make,
        x="Selling_Price",
        y="Mileage",
        color="skyblue",
        ax=ax
    )

    ax.set_title(make)

# Eliminar ejes vacíos
for ax in axes[len(makes):]:
    fig.delaxes(ax)

plt.tight_layout()
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}scatterplots.png", dpi=300)

# plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
No description has been provided for this image
<h4 id="%C2%A04.-Boxplot-by-Category"> 4. Boxplot by Category¶</h4>
In [14]:
sns.set_theme(style="whitegrid")
plt.figure(figsize=(7, 5))
sns.boxplot(data=df, x="Make", y="Selling_Price", palette="Set2")
plt.xticks(rotation=90) 
plt.title("Selling Price vs Car ")
plt.tight_layout()

plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}selling_vs_car.png", dpi=300)

# plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
/var/folders/mx/d1fc66x17bs8hc5yn2qp1c4h0000gn/T/ipykernel_2361/1602545604.py:3: FutureWarning: 

Passing `palette` without assigning `hue` is deprecated and will be removed in v0.14.0. Assign the `x` variable to `hue` and set `legend=False` for the same effect.

  sns.boxplot(data=df, x="Make", y="Selling_Price", palette="Set2")
No description has been provided for this image
<p>Audi, BMW, and Mercedes-Benz have the highest selling prices in the dataset. These brands also include some of the most expensive vehicles, reflecting their position in the premium car market.</p>
<h4 id="5.-Seaborn-Pairplot-Experiments">5. Seaborn Pairplot Experiments¶</h4>
In [15]:
sns.set_theme(style="white")
sns.pairplot(df, hue="Selling_Price", height=2.5)
plt.suptitle("Iris dataset: Pairplot", y=1.02)
plt.savefig(f"{EXPLORATION_IMAGES_FOLDER}pairplots.png", dpi=300)

# plt.close() # IF WE DONT WANT TO SHOW THE IMAGE.
plt.show() # IF WE WANT TO SHOW THE IMAGE
No description has been provided for this image
<h2 id="Categorical-to-numerical">Categorical to numerical¶</h2>
In [16]:
print(df["Make"].nunique())
print(df["Color"].nunique())
print(df["Model"].nunique())
print(df["Service_History"].nunique())
print(df["Body_Type"].nunique())
print(df["Drivetrain"].nunique())
print(df["Location"].nunique())
10
8
40
3
5
4
10
<p>Machine learning models cannot work directly with categorical data, so these variables must be converted into a numerical format before training. That's why we did some encoding.</p>
  • For Service_History, we used ordinal encoding because the categories have a natural order. A vehicle with no service history provides less information than one with a partial service history.
  • For the remaining categorical variables, we applied One-Hot Encoding. This method creates a separate binary column for each category and avoids introducing relationships between categories that do not actually exist.
In [17]:
# Ordinal for service (from no service to full service, hierarchy is preserved)
service_mapping = {
    "No Service": 0,
    "Partial Service": 1,
    "Full Service": 2
}

df["Service_History"] = df["Service_History"].map(service_mapping)


# OHE for the other categorical columns
categorical_columns = [
    "Make",
    "Model",
    "Fuel_Type",
    "Transmission",
    "Color",
    "Body_Type",
    "Drivetrain",
    "Location"
]

df = pd.get_dummies(
    df,
    columns=categorical_columns,
    dtype=int
)
In [18]:
print(df.head())
print(df.dtypes)
print(df.shape)
print(df.columns)
    Year  Engine_Size  Mileage  Horsepower  Torque  Owners  Accident_History  \
4   2022          1.9    32813       149.0   141.0       1               0.0   
7   2022          2.3    32133       180.0   162.0       1               0.0   
15  2010          2.0    17138       169.0   160.0       3               0.0   
17  2011          3.3   264441       258.0   252.0       4               0.0   
21  2016          3.1   111537       243.0   219.0       2               0.0   

   Service_History  Fuel_Efficiency  Selling_Price  ...  Location_CA  \
4                2             38.0          21792  ...            1   
7                1             35.0          31416  ...            0   
15               1             27.0          11728  ...            0   
17               1             22.0           3395  ...            0   
21               1             21.0          23314  ...            0   

    Location_FL  Location_GA  Location_IL  Location_MI  Location_NC  \
4             0            0            0            0            0   
7             0            0            0            0            0   
15            0            0            0            0            1   
17            0            0            0            0            0   
21            1            0            0            0            0   

    Location_NY  Location_OH  Location_PA  Location_TX  
4             0            0            0            0  
7             1            0            0            0  
15            0            0            0            0  
17            0            0            0            1  
21            0            0            0            0  

[5 rows x 93 columns]
Year             int16
Engine_Size    float32
Mileage          int32
Horsepower     float32
Torque         float32
                ...   
Location_NC      int64
Location_NY      int64
Location_OH      int64
Location_PA      int64
Location_TX      int64
Length: 93, dtype: object
(1329, 93)
Index(['Year', 'Engine_Size', 'Mileage', 'Horsepower', 'Torque', 'Owners',
       'Accident_History', 'Service_History', 'Fuel_Efficiency',
       'Selling_Price', 'Make_Audi', 'Make_BMW', 'Make_Chevrolet', 'Make_Ford',
       'Make_Honda', 'Make_Hyundai', 'Make_Mercedes-Benz', 'Make_Nissan',
       'Make_Toyota', 'Make_Volkswagen', 'Model_3 Series', 'Model_5 Series',
       'Model_A4', 'Model_A6', 'Model_Accord', 'Model_Altima', 'Model_Atlas',
       'Model_C-Class', 'Model_CR-V', 'Model_Camry', 'Model_Civic',
       'Model_Corolla', 'Model_E-Class', 'Model_Elantra', 'Model_Equinox',
       'Model_Escape', 'Model_Explorer', 'Model_F-150', 'Model_GLC',
       'Model_GLE', 'Model_Golf', 'Model_Highlander', 'Model_Malibu',
       'Model_Mustang', 'Model_Passat', 'Model_Pathfinder', 'Model_Pilot',
       'Model_Q5', 'Model_Q7', 'Model_RAV4', 'Model_Rogue', 'Model_Santa Fe',
       'Model_Sentra', 'Model_Silverado', 'Model_Sonata', 'Model_Tahoe',
       'Model_Tiguan', 'Model_Tucson', 'Model_X3', 'Model_X5',
       'Fuel_Type_Diesel', 'Fuel_Type_Electric', 'Fuel_Type_Hybrid',
       'Fuel_Type_Petrol', 'Transmission_Automatic', 'Transmission_Manual',
       'Color_Black', 'Color_Blue', 'Color_Brown', 'Color_Gray', 'Color_Green',
       'Color_Red', 'Color_Silver', 'Color_White', 'Body_Type_Coupe',
       'Body_Type_Hatchback', 'Body_Type_SUV', 'Body_Type_Sedan',
       'Body_Type_Truck', 'Drivetrain_4WD', 'Drivetrain_AWD', 'Drivetrain_FWD',
       'Drivetrain_RWD', 'Location_CA', 'Location_FL', 'Location_GA',
       'Location_IL', 'Location_MI', 'Location_NC', 'Location_NY',
       'Location_OH', 'Location_PA', 'Location_TX'],
      dtype='object')
<h2 id="Save-processed-data">Save processed data¶</h2>
In [19]:
from sklearn.model_selection import train_test_split

# 80% training data
# 20% testing data
train_df, test_df = train_test_split(
    df,
    test_size=0.2,
    random_state=22
)
In [20]:
fM.set_format("parquet")
fM.write(train_df, PROCESSED_TRAIN_FILE)
fM.write(test_df, PROCESSED_TEST_FILE)
<h2 id="Check-it-was-correctly-saved">Check it was correctly saved¶</h2>
In [21]:
fM.set_format("parquet")
fdf = fM.read(PROCESSED_TRAIN_FILE)

display(fdf)
inspect(fdf)
<style scoped> .dataframe tbody tr th:only-of-type { vertical-align: middle; } .dataframe tbody tr th { vertical-align: top; } .dataframe thead th { text-align: right; } </style> <thead> <th></th> <th>Year</th> <th>Engine_Size</th> <th>Mileage</th> <th>Horsepower</th> <th>Torque</th> <th>Owners</th> <th>Accident_History</th> <th>Service_History</th> <th>Fuel_Efficiency</th> <th>Selling_Price</th> <th>...</th> <th>Location_CA</th> <th>Location_FL</th> <th>Location_GA</th> <th>Location_IL</th> <th>Location_MI</th> <th>Location_NC</th> <th>Location_NY</th> <th>Location_OH</th> <th>Location_PA</th> <th>Location_TX</th> </thead> <tbody> <th>633</th> <th>1028</th> <th>4550</th> <th>1056</th> <th>1063</th> <th>...</th> <th>1461</th> <th>4038</th> <th>3372</th> <th>497</th> <th>3692</th> </tbody> <tbody></tbody>
2008 2.0 20347 167.0 151.0 5 1.0 1 35.0 3428 ... 0 1 0 0 0 0 0 0 0 0
2016 2.2 64385 162.0 167.0 4 1.0 1 29.0 5850 ... 0 0 0 0 1 0 0 0 0 0
2013 2.5 158156 187.0 172.0 3 0.0 1 22.0 1281 ... 0 1 0 0 0 0 0 0 0 0
2013 2.5 131422 209.0 189.0 3 1.0 1 27.0 7586 ... 0 0 0 1 0 0 0 0 0 0
2008 3.5 270715 263.0 239.0 4 0.0 0 19.0 500 ... 0 0 0 1 0 0 0 0 0 0
... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ... ...
2005 3.3 377675 250.0 231.0 3 0.0 1 24.0 500 ... 0 0 0 0 0 0 0 0 1 0
2017 1.6 87365 122.0 103.0 1 0.0 0 43.0 15518 ... 0 0 0 1 0 0 0 0 0 0
2023 3.2 3623 243.0 246.0 1 0.0 1 23.0 24300 ... 0 0 0 0 0 0 0 0 0 1
2011 3.3 229650 238.0 223.0 3 1.0 1 24.0 7157 ... 0 0 0 0 0 0 0 1 0 0
2005 2.2 410673 175.0 169.0 5 0.0 1 33.0 500 ... 0 0 0 0 0 0 0 0 0 1
<p>1063 rows × 93 columns</p>
Shape: (1063, 93)
Columns: ['Year', 'Engine_Size', 'Mileage', 'Horsepower', 'Torque', 'Owners', 'Accident_History', 'Service_History', 'Fuel_Efficiency', 'Selling_Price', 'Make_Audi', 'Make_BMW', 'Make_Chevrolet', 'Make_Ford', 'Make_Honda', 'Make_Hyundai', 'Make_Mercedes-Benz', 'Make_Nissan', 'Make_Toyota', 'Make_Volkswagen', 'Model_3 Series', 'Model_5 Series', 'Model_A4', 'Model_A6', 'Model_Accord', 'Model_Altima', 'Model_Atlas', 'Model_C-Class', 'Model_CR-V', 'Model_Camry', 'Model_Civic', 'Model_Corolla', 'Model_E-Class', 'Model_Elantra', 'Model_Equinox', 'Model_Escape', 'Model_Explorer', 'Model_F-150', 'Model_GLC', 'Model_GLE', 'Model_Golf', 'Model_Highlander', 'Model_Malibu', 'Model_Mustang', 'Model_Passat', 'Model_Pathfinder', 'Model_Pilot', 'Model_Q5', 'Model_Q7', 'Model_RAV4', 'Model_Rogue', 'Model_Santa Fe', 'Model_Sentra', 'Model_Silverado', 'Model_Sonata', 'Model_Tahoe', 'Model_Tiguan', 'Model_Tucson', 'Model_X3', 'Model_X5', 'Fuel_Type_Diesel', 'Fuel_Type_Electric', 'Fuel_Type_Hybrid', 'Fuel_Type_Petrol', 'Transmission_Automatic', 'Transmission_Manual', 'Color_Black', 'Color_Blue', 'Color_Brown', 'Color_Gray', 'Color_Green', 'Color_Red', 'Color_Silver', 'Color_White', 'Body_Type_Coupe', 'Body_Type_Hatchback', 'Body_Type_SUV', 'Body_Type_Sedan', 'Body_Type_Truck', 'Drivetrain_4WD', 'Drivetrain_AWD', 'Drivetrain_FWD', 'Drivetrain_RWD', 'Location_CA', 'Location_FL', 'Location_GA', 'Location_IL', 'Location_MI', 'Location_NC', 'Location_NY', 'Location_OH', 'Location_PA', 'Location_TX']
Data types:
Year             int16
Engine_Size    float32
Mileage          int32
Horsepower     float32
Torque         float32
                ...   
Location_NC      int64
Location_NY      int64
Location_OH      int64
Location_PA      int64
Location_TX      int64
Length: 93, dtype: object