๐ข Titanic Dataset – Data Preprocessing
Preparing raw Titanic data for Machine Learning
๐ What is Data Preprocessing?
Data preprocessing is the process of cleaning and transforming raw data into a suitable format for Machine Learning algorithms. Real-world datasets may contain missing values, categorical variables, unnecessary columns, and features with different scales.
```In the Titanic dataset, preprocessing is required before applying Machine Learning models. The main steps include data inspection, handling missing values, encoding categorical variables, removing unnecessary features, and preparing the final feature set.
```1️⃣ Load the Dataset
```First, the Titanic dataset is loaded using Pandas. Pandas provides functions for reading, inspecting and manipulating tabular data.
import pandas as pd
```
import numpy as np
df = pd.read_csv("train.csv")
print(df.head())
```
``` PassengerId Survived Pclass Name Sex Age SibSp Parch Fare 0 1 0 3 ... male 22.0 1 0 7.2500 1 2 1 1 ... female 38.0 1 0 71.2833 2 3 1 3 ... female 26.0 0 0 7.9250 3 4 1 1 ... female 35.0 1 0 53.1000 4 5 0 3 ... male 35.0 0 0 8.0500
2️⃣ Understand the Dataset
```Before preprocessing, we inspect the structure, data types and number of observations in the dataset.
print(df.shape)
```
print(df.info())
print(df.describe())
```
``` Dataset shape: (891, 12) Columns: PassengerId Survived Pclass Name Sex Age SibSp Parch Ticket Fare Cabin Embarked
Interpretation: The dataset contains passenger information and the target variable Survived, where 0 represents non-survival and 1 represents survival.
```3️⃣ Detect Missing Values
```Missing values can prevent Machine Learning algorithms from processing the data correctly. Therefore, we identify the number of missing values in each column.
print(df.isnull().sum())
``` Age 177 Cabin 687 Embarked 2 Name 0 Sex 0 Fare 0 Pclass 0
The output shows that Age, Cabin and Embarked contain missing observations.
```4️⃣ Handling Missing Values
```Missing numerical values can be replaced using statistical measures such as the mean or median. Categorical missing values can commonly be replaced with the most frequent category (mode).
# Fill missing Age values with median
```
df["Age"] = df["Age"].fillna(df["Age"].median())
# Fill missing Embarked values with mode
df["Embarked"] = df["Embarked"].fillna(
df["Embarked"].mode()[0]
)
```
``` Age missing values after preprocessing: 0 Embarked missing values after preprocessing: 0
5️⃣ Remove Features That Are Not Required
```Some variables may contain identifiers, free text or a large amount of missing information. Such variables may not be directly useful for the selected Machine Learning model.
df = df.drop(
["PassengerId", "Name", "Ticket", "Cabin"],
axis=1
```
)
print(df.head())
```
``` Survived Pclass Sex Age SibSp Parch Fare Embarked 0 0 3 male 22.0 1 0 7.2500 S 1 1 1 female 38.0 1 0 71.2833 C 2 1 3 female 26.0 0 0 7.9250 S
6️⃣ Convert Categorical Data into Numerical Data
```Machine Learning algorithms generally require numerical input. Therefore, categorical variables such as Sex and Embarked need to be converted into numerical representations.
df = pd.get_dummies(
df,
columns=["Sex", "Embarked"],
drop_first=True
```
)
print(df.head())
```
``` Survived Pclass Age SibSp Parch Fare Sex_male Embarked_Q Embarked_S 0 0 3 22.0 1 0 7.2500 1 0 1 1 1 1 38.0 1 0 71.2833 0 0 0 2 1 3 26.0 0 0 7.9250 0 0 1
Interpretation: The categorical values have now been represented using numerical features that can be supplied to a Machine Learning model.
```7️⃣ Separate Features and Target
```The target variable is the value that we want the Machine Learning model to predict. In the Titanic problem, the target is Survived.
X = df.drop("Survived", axis=1)
```
y = df["Survived"]
print("Feature shape:", X.shape)
print("Target shape:", y.shape)
```
``` Feature shape: (891, 8) Target shape: (891,)
8️⃣ Feature Selection Concept
```Feature selection means selecting the most useful input variables while removing irrelevant, redundant or unnecessary features.
The referenced Feature Selection for Machine Learning repository demonstrates several approaches, including constant-feature removal, quasi-constant-feature removal, duplicate-feature removal and correlation-based feature selection.
# Example: checking the correlation between features
```
correlation_matrix = X.corr()
print(correlation_matrix)
```
Pclass Age SibSp Parch Fare
```
Pclass 1.000 ...
Age ... 1.000 ...
SibSp ... ... 1.000 ...
Parch ... ... ... 1.000
Fare ... ... ... ... 1.000 ๐ฏ Preprocessing Pipeline
```After preprocessing, the dataset is transformed into a cleaner numerical representation that can be used for Machine Learning model development.
```✅ Final Preprocessed Data
```print(df.head())
```
print("Missing values:")
print(df.isnull().sum())
```
``` Dataset successfully preprocessed. ✔ Missing values handled ✔ Unnecessary features removed ✔ Categorical variables encoded ✔ Target variable separated ✔ Dataset ready for Machine Learning
No comments:
Post a Comment