Total Pageviews

Monday, September 28, 2026

๐Ÿšข Titanic Dataset – Data Preprocessing

๐Ÿšข Titanic Dataset – Data Preprocessing

Preparing raw Titanic data for Machine Learning

๐Ÿ“š What is Data Preprocessing?

Data preprocessing is the process of cleaning and transforming raw data into a suitable format for Machine Learning algorithms. Real-world datasets may contain missing values, categorical variables, unnecessary columns, and features with different scales.

```

In the Titanic dataset, preprocessing is required before applying Machine Learning models. The main steps include data inspection, handling missing values, encoding categorical variables, removing unnecessary features, and preparing the final feature set.

```

1️⃣ Load the Dataset

```

First, the Titanic dataset is loaded using Pandas. Pandas provides functions for reading, inspecting and manipulating tabular data.

import pandas as pd
```

import numpy as np

df = pd.read_csv("train.csv")

print(df.head())
```
Output:
```

PassengerId  Survived  Pclass     Name     Sex   Age  SibSp  Parch     Fare
0            1         0       3  ...       male  22.0      1      0   7.2500
1            2         1       1  ...     female  38.0      1      0  71.2833
2            3         1       3  ...     female  26.0      0      0   7.9250
3            4         1       1  ...     female  35.0      1      0  53.1000
4            5         0       3  ...       male  35.0      0      0   8.0500 

2️⃣ Understand the Dataset

```

Before preprocessing, we inspect the structure, data types and number of observations in the dataset.

print(df.shape)
```

print(df.info())
print(df.describe())
```
Output:
```

Dataset shape:
(891, 12)

Columns:
PassengerId
Survived
Pclass
Name
Sex
Age
SibSp
Parch
Ticket
Fare
Cabin
Embarked 
```

Interpretation: The dataset contains passenger information and the target variable Survived, where 0 represents non-survival and 1 represents survival.

```

3️⃣ Detect Missing Values

```

Missing values can prevent Machine Learning algorithms from processing the data correctly. Therefore, we identify the number of missing values in each column.

print(df.isnull().sum())
Example Output:
```

Age         177
Cabin       687
Embarked      2
Name          0
Sex           0
Fare          0
Pclass        0 
```

The output shows that Age, Cabin and Embarked contain missing observations.

```

4️⃣ Handling Missing Values

```

Missing numerical values can be replaced using statistical measures such as the mean or median. Categorical missing values can commonly be replaced with the most frequent category (mode).

# Fill missing Age values with median
```

df["Age"] = df["Age"].fillna(df["Age"].median())

# Fill missing Embarked values with mode

df["Embarked"] = df["Embarked"].fillna(
df["Embarked"].mode()[0]
)
```
Output:
```

Age missing values after preprocessing: 0
Embarked missing values after preprocessing: 0 

5️⃣ Remove Features That Are Not Required

```

Some variables may contain identifiers, free text or a large amount of missing information. Such variables may not be directly useful for the selected Machine Learning model.

df = df.drop(
["PassengerId", "Name", "Ticket", "Cabin"],
axis=1
```

)

print(df.head())
```
Output:
```

Survived  Pclass     Sex   Age  SibSp  Parch     Fare Embarked
0         0       3    male  22.0      1      0   7.2500        S
1         1       1  female  38.0      1      0  71.2833        C
2         1       3  female  26.0      0      0   7.9250        S 

6️⃣ Convert Categorical Data into Numerical Data

```

Machine Learning algorithms generally require numerical input. Therefore, categorical variables such as Sex and Embarked need to be converted into numerical representations.

df = pd.get_dummies(
df,
columns=["Sex", "Embarked"],
drop_first=True
```

)

print(df.head())
```
Output:
```

Survived  Pclass   Age  SibSp  Parch     Fare  Sex_male  Embarked_Q  Embarked_S
0         0       3  22.0      1      0   7.2500         1           0           1
1         1       1  38.0      1      0  71.2833         0           0           0
2         1       3  26.0      0      0   7.9250         0           0           1 
```

Interpretation: The categorical values have now been represented using numerical features that can be supplied to a Machine Learning model.

```

7️⃣ Separate Features and Target

```

The target variable is the value that we want the Machine Learning model to predict. In the Titanic problem, the target is Survived.

X = df.drop("Survived", axis=1)
```

y = df["Survived"]

print("Feature shape:", X.shape)
print("Target shape:", y.shape)
```
Output:
```

Feature shape: (891, 8)
Target shape: (891,) 

8️⃣ Feature Selection Concept

```

Feature selection means selecting the most useful input variables while removing irrelevant, redundant or unnecessary features.

The referenced Feature Selection for Machine Learning repository demonstrates several approaches, including constant-feature removal, quasi-constant-feature removal, duplicate-feature removal and correlation-based feature selection.

# Example: checking the correlation between features
```

correlation_matrix = X.corr()

print(correlation_matrix)
```
Output:
          Pclass       Age     SibSp     Parch      Fare
```

Pclass       1.000      ...
Age           ...       1.000      ...
SibSp         ...        ...      1.000      ...
Parch         ...        ...       ...      1.000
Fare          ...        ...       ...       ...      1.000 

๐ŸŽฏ Preprocessing Pipeline

```
Raw Data → Inspect → Handle Missing Values → Remove Unnecessary Features → Encode Categories → Feature Selection → ML-Ready Data

After preprocessing, the dataset is transformed into a cleaner numerical representation that can be used for Machine Learning model development.

```

✅ Final Preprocessed Data

```
print(df.head())
```

print("Missing values:")
print(df.isnull().sum())
```
Expected Result:
```

Dataset successfully preprocessed.

✔ Missing values handled
✔ Unnecessary features removed
✔ Categorical variables encoded
✔ Target variable separated
✔ Dataset ready for Machine Learning 

No comments:

Post a Comment