Total Pageviews

Monday, September 28, 2026

Logistic Regression

 For Logistic Regression, the important change is that Boston Housing's original MEDV is a continuous house-price target, so we need to convert it into classes. 

  • 0 = Low Price
  • 1 = High Price
  • Median MEDV as the classification threshold


# ============================================================
# BOSTON HOUSING DATASET
# LOGISTIC REGRESSION
# COMPLETE DATA PREPROCESSING
# GOOGLE COLAB VERSION
# ============================================================


# ============================================================
# 1. IMPORT LIBRARIES
# ============================================================

import numpy as np
import pandas as pd

import matplotlib.pyplot as plt
import seaborn as sns

from sklearn.datasets import fetch_openml

from sklearn.model_selection import (
    train_test_split,
    cross_val_score
)

from sklearn.pipeline import Pipeline

from sklearn.compose import ColumnTransformer

from sklearn.impute import SimpleImputer

from sklearn.preprocessing import (
    StandardScaler,
    MinMaxScaler,
    RobustScaler,
    OneHotEncoder
)

from sklearn.feature_selection import (
    SelectKBest,
    f_classif
)

from sklearn.linear_model import LogisticRegression

from sklearn.metrics import (
    accuracy_score,
    precision_score,
    recall_score,
    f1_score,
    confusion_matrix,
    classification_report,
    ConfusionMatrixDisplay,
    roc_curve,
    roc_auc_score
)

import warnings

warnings.filterwarnings("ignore")


# ============================================================
# 2. LOAD BOSTON HOUSING DATASET
# ============================================================

boston = fetch_openml(
    name="boston",
    version=1,
    as_frame=True
)

df = boston.frame.copy()

print("Boston Housing Dataset Loaded Successfully!")


# ============================================================
# 3. DISPLAY FIRST 5 ROWS
# ============================================================

print("\nFirst 5 Rows:")

display(
    df.head()
)


# ============================================================
# 4. DATASET SHAPE
# ============================================================

print("\nDataset Shape:")

print(
    df.shape
)


# ============================================================
# 5. COLUMN NAMES
# ============================================================

print("\nColumn Names:")

print(
    df.columns.tolist()
)


# ============================================================
# 6. DATA INFORMATION
# ============================================================

print("\nDataset Information:")

df.info()


# ============================================================
# 7. DATA TYPES
# ============================================================

print("\nData Types:")

print(
    df.dtypes
)


# ============================================================
# 8. CHECK MISSING VALUES
# ============================================================

print("\nMissing Values:")

print(
    df.isnull().sum()
)


# ============================================================
# 9. MISSING VALUE PERCENTAGE
# ============================================================

missing_percentage = (

    df.isnull()
      .mean()
      .mul(100)
      .sort_values(
          ascending=False
      )

)

print("\nMissing Value Percentage:")

print(
    missing_percentage
)


# ============================================================
# 10. CHECK DUPLICATE RECORDS
# ============================================================

duplicate_count = df.duplicated().sum()

print(
    "\nNumber of Duplicate Rows:",
    duplicate_count
)


# ============================================================
# 11. REMOVE DUPLICATES
# ============================================================

df = df.drop_duplicates()

print(
    "\nShape After Duplicate Removal:"
)

print(
    df.shape
)


# ============================================================
# 12. STATISTICAL SUMMARY
# ============================================================

print(
    "\nStatistical Summary:"
)

display(
    df.describe()
)


# ============================================================
# 13. UNIQUE VALUES
# ============================================================

print(
    "\nNumber of Unique Values:"
)

print(
    df.nunique()
)


# ============================================================
# 14. TARGET VARIABLE
# ============================================================

target = "MEDV"

print(
    "\nOriginal Target Variable:",
    target
)


# ============================================================
# 15. TARGET DISTRIBUTION
# ============================================================

plt.figure(
    figsize=(9,5)
)

sns.histplot(
    df[target],
    kde=True
)

plt.title(
    "Distribution of Boston House Prices"
)

plt.xlabel(
    "MEDV"
)

plt.ylabel(
    "Frequency"
)

plt.show()


# ============================================================
# 16. OUTLIER DETECTION
# ============================================================

numeric_columns = (

    df.select_dtypes(
        include=np.number
    ).columns

)


plt.figure(
    figsize=(14,6)
)

sns.boxplot(
    data=df[
        numeric_columns
    ]
)

plt.title(
    "Boxplot of Boston Housing Features"
)

plt.xticks(
    rotation=45
)

plt.show()


# ============================================================
# 17. IQR OUTLIER DETECTION
# ============================================================

print(
    "\nIQR OUTLIER ANALYSIS"
)

for column in numeric_columns:

    Q1 = df[column].quantile(
        0.25
    )

    Q3 = df[column].quantile(
        0.75
    )

    IQR = Q3 - Q1

    lower_bound = (
        Q1 - 1.5 * IQR
    )

    upper_bound = (
        Q3 + 1.5 * IQR
    )

    outliers = df[
        (df[column] < lower_bound)
        |
        (df[column] > upper_bound)
    ]

    print(
        column,
        "→",
        len(outliers),
        "outliers"
    )


# ============================================================
# 18. OUTLIER CAPPING
# ============================================================
#
# We do not automatically delete unusual observations.
# Instead, this demonstrates IQR-based capping.
# ============================================================

df_capped = df.copy()

for column in numeric_columns:

    Q1 = df_capped[column].quantile(
        0.25
    )

    Q3 = df_capped[column].quantile(
        0.75
    )

    IQR = Q3 - Q1

    lower_bound = (
        Q1 - 1.5 * IQR
    )

    upper_bound = (
        Q3 + 1.5 * IQR
    )

    df_capped[column] = (

        df_capped[column]
        .clip(
            lower_bound,
            upper_bound
        )

    )


print(
    "\nIQR Outlier Capping Completed."
)


# ============================================================
# 19. CORRELATION MATRIX
# ============================================================

correlation = df.corr(
    numeric_only=True
)


print(
    "\nCorrelation with MEDV:"
)

print(
    correlation[target]
    .sort_values(
        ascending=False
    )
)


# ============================================================
# 20. CORRELATION HEATMAP
# ============================================================

plt.figure(
    figsize=(12,9)
)

sns.heatmap(
    correlation,
    annot=True,
    fmt=".2f",
    cmap="coolwarm"
)

plt.title(
    "Boston Housing Correlation Matrix"
)

plt.show()


# ============================================================
# 21. CONVERT REGRESSION TARGET INTO CLASSIFICATION TARGET
# ============================================================
#
# Boston Housing MEDV is continuous.
#
# Logistic Regression requires categorical classes.
#
# Median price is used as the threshold.
#
# MEDV < median → 0 → Low Price
#
# MEDV >= median → 1 → High Price
# ============================================================

median_price = df[target].median()

print(
    "\nMedian House Price (MEDV):",
    median_price
)


df["Price_Class"] = (

    df[target] >= median_price

).astype(int)


# ============================================================
# 22. DISPLAY CLASS DISTRIBUTION
# ============================================================

print(
    "\nPrice Class Distribution:"
)

print(
    df["Price_Class"]
    .value_counts()
)


# ============================================================
# 23. VISUALIZE CLASS DISTRIBUTION
# ============================================================

plt.figure(
    figsize=(7,5)
)

sns.countplot(
    x="Price_Class",
    data=df
)

plt.title(
    "Low Price vs High Price Houses"
)

plt.xlabel(
    "Price Class"
)

plt.ylabel(
    "Number of Houses"
)

plt.xticks(
    [0,1],
    [
        "Low Price",
        "High Price"
    ]
)

plt.show()


# ============================================================
# 24. SEPARATE FEATURES AND TARGET
# ============================================================

X = df.drop(
    columns=[
        "MEDV",
        "Price_Class"
    ]
)

y = df["Price_Class"]


print(
    "\nFeature Shape:"
)

print(
    X.shape
)


print(
    "\nTarget Shape:"
)

print(
    y.shape
)


# ============================================================
# 25. IDENTIFY NUMERICAL FEATURES
# ============================================================

numeric_features = (

    X.select_dtypes(
        include=np.number
    )
    .columns
    .tolist()

)


categorical_features = (

    X.select_dtypes(
        exclude=np.number
    )
    .columns
    .tolist()

)


print(
    "\nNumerical Features:"
)

print(
    numeric_features
)


print(
    "\nCategorical Features:"
)

print(
    categorical_features
)


# ============================================================
# 26. TRAIN-TEST SPLIT
# ============================================================

X_train, X_test, y_train, y_test = (

    train_test_split(

        X,
        y,

        test_size=0.20,

        random_state=42,

        stratify=y

    )

)


print(
    "\nTraining Data:",
    X_train.shape
)

print(
    "Testing Data:",
    X_test.shape
)


# ============================================================
# 27. NUMERICAL PREPROCESSING
# ============================================================

numeric_pipeline = Pipeline(

    steps=[

        (
            "imputer",

            SimpleImputer(
                strategy="median"
            )

        ),

        (
            "scaler",

            StandardScaler()
        )

    ]

)


# ============================================================
# 28. CATEGORICAL PREPROCESSING
# ============================================================

categorical_pipeline = Pipeline(

    steps=[

        (
            "imputer",

            SimpleImputer(
                strategy="most_frequent"
            )

        ),

        (
            "encoder",

            OneHotEncoder(
                handle_unknown="ignore"
            )

        )

    ]

)


# ============================================================
# 29. COMBINE PREPROCESSING
# ============================================================

preprocessor = ColumnTransformer(

    transformers=[

        (
            "numeric",

            numeric_pipeline,

            numeric_features

        ),

        (
            "categorical",

            categorical_pipeline,

            categorical_features

        )

    ]

)


# ============================================================
# 30. LOGISTIC REGRESSION PIPELINE
# ============================================================

logistic_model = Pipeline(

    steps=[

        (
            "preprocessing",

            preprocessor

        ),

        (
            "model",

            LogisticRegression(
                max_iter=2000
            )

        )

    ]

)


# ============================================================
# 31. TRAIN MODEL
# ============================================================

logistic_model.fit(

    X_train,

    y_train

)


print(
    "\nLogistic Regression Model Trained Successfully!"
)


# ============================================================
# 32. MAKE PREDICTIONS
# ============================================================

y_pred = logistic_model.predict(
    X_test
)


# ============================================================
# 33. PREDICT PROBABILITIES
# ============================================================

y_probability = (

    logistic_model
    .predict_proba(
        X_test
    )[:, 1]

)


# ============================================================
# 34. ACCURACY
# ============================================================

accuracy = accuracy_score(

    y_test,

    y_pred

)


# ============================================================
# 35. PRECISION
# ============================================================

precision = precision_score(

    y_test,

    y_pred,

    zero_division=0

)


# ============================================================
# 36. RECALL
# ============================================================

recall = recall_score(

    y_test,

    y_pred,

    zero_division=0

)


# ============================================================
# 37. F1 SCORE
# ============================================================

f1 = f1_score(

    y_test,

    y_pred,

    zero_division=0

)


# ============================================================
# 38. ROC-AUC
# ============================================================

roc_auc = roc_auc_score(

    y_test,

    y_probability

)


# ============================================================
# 39. DISPLAY RESULTS
# ============================================================

print("\n")
print("============================================")
print("       LOGISTIC REGRESSION RESULTS")
print("============================================")

print(
    "Accuracy  :",
    round(accuracy, 4)
)

print(
    "Precision :",
    round(precision, 4)
)

print(
    "Recall    :",
    round(recall, 4)
)

print(
    "F1 Score  :",
    round(f1, 4)
)

print(
    "ROC-AUC   :",
    round(roc_auc, 4)
)

print("============================================")


# ============================================================
# 40. CONFUSION MATRIX
# ============================================================

cm = confusion_matrix(

    y_test,

    y_pred

)


print(
    "\nConfusion Matrix:"
)

print(
    cm
)


# ============================================================
# 41. CONFUSION MATRIX VALUES
# ============================================================

TN, FP, FN, TP = cm.ravel()


print(
    "\nTrue Negative (TN):",
    TN
)

print(
    "False Positive (FP):",
    FP
)

print(
    "False Negative (FN):",
    FN
)

print(
    "True Positive (TP):",
    TP
)


# ============================================================
# 42. CONFUSION MATRIX VISUALIZATION
# ============================================================

plt.figure(
    figsize=(7,6)
)

disp = ConfusionMatrixDisplay(

    confusion_matrix=cm,

    display_labels=[
        "Low Price",
        "High Price"
    ]

)

disp.plot()

plt.title(
    "Boston Housing - Logistic Regression"
)

plt.show()


# ============================================================
# 43. CONFUSION MATRIX HEATMAP
# ============================================================

plt.figure(
    figsize=(7,6)
)

sns.heatmap(

    cm,

    annot=True,

    fmt="d",

    cmap="Blues",

    xticklabels=[
        "Predicted Low",
        "Predicted High"
    ],

    yticklabels=[
        "Actual Low",
        "Actual High"
    ]

)

plt.title(
    "Confusion Matrix Heatmap"
)

plt.xlabel(
    "Predicted Class"
)

plt.ylabel(
    "Actual Class"
)

plt.show()


# ============================================================
# 44. CLASSIFICATION REPORT
# ============================================================

print(
    "\n============================================"
)

print(
    "          CLASSIFICATION REPORT"
)

print(
    "============================================"
)

print(

    classification_report(

        y_test,

        y_pred,

        target_names=[
            "Low Price",
            "High Price"
        ],

        zero_division=0

    )

)


# ============================================================
# 45. PERFORMANCE TABLE
# ============================================================

metrics_table = pd.DataFrame({

    "Metric": [

        "Accuracy",

        "Precision",

        "Recall",

        "F1 Score",

        "ROC-AUC"

    ],

    "Score": [

        accuracy,

        precision,

        recall,

        f1,

        roc_auc

    ]

})


metrics_table["Score"] = (

    metrics_table["Score"]
    .round(4)

)


print(
    "\nPerformance Metrics:"
)

display(
    metrics_table
)


# ============================================================
# 46. CROSS VALIDATION
# ============================================================

cv_scores = cross_val_score(

    logistic_model,

    X,

    y,

    cv=5,

    scoring="accuracy"

)


print(
    "\n5-Fold Cross Validation Accuracy:"
)

print(
    cv_scores
)


print(
    "\nMean Cross Validation Accuracy:"
)

print(
    round(
        cv_scores.mean(),
        4
    )
)


# ============================================================
# 47. ROC CURVE
# ============================================================

fpr, tpr, thresholds = roc_curve(

    y_test,

    y_probability

)


plt.figure(
    figsize=(8,6)
)

plt.plot(

    fpr,

    tpr,

    label=f"ROC-AUC = {roc_auc:.4f}"

)

plt.plot(

    [0,1],

    [0,1],

    linestyle="--"

)

plt.xlabel(
    "False Positive Rate"
)

plt.ylabel(
    "True Positive Rate"
)

plt.title(
    "ROC Curve - Logistic Regression"
)

plt.legend()

plt.show()


# ============================================================
# 48. PRECISION / RECALL / F1 GRAPH
# ============================================================

metric_names = [

    "Accuracy",

    "Precision",

    "Recall",

    "F1 Score"

]


metric_values = [

    accuracy,

    precision,

    recall,

    f1

]


plt.figure(
    figsize=(9,6)
)

plt.bar(

    metric_names,

    metric_values

)

plt.ylim(
    0,
    1
)

plt.ylabel(
    "Score"
)

plt.title(
    "Logistic Regression Performance"
)

plt.show()


# ============================================================
# 49. ACTUAL VS PREDICTED TABLE
# ============================================================

prediction_table = pd.DataFrame({

    "Actual Class": y_test.values,

    "Predicted Class": y_pred,

    "High Price Probability":
        y_probability

})


prediction_table["Actual Class"] = (

    prediction_table["Actual Class"]
    .map({
        0: "Low Price",
        1: "High Price"
    })

)


prediction_table["Predicted Class"] = (

    prediction_table["Predicted Class"]
    .map({
        0: "Low Price",
        1: "High Price"
    })

)


prediction_table[
    "High Price Probability"
] = (

    prediction_table[
        "High Price Probability"
    ].round(4)

)


print(
    "\nActual vs Predicted:"
)

display(
    prediction_table.head(20)
)


# ============================================================
# 50. TEST A NEW HOUSE
# ============================================================

# ============================================================
# 50. TEST A NEW HOUSE - CORRECTED VERSION
# ============================================================

print(
    "\n============================================"
)

print(
    "        NEW HOUSE CLASSIFICATION"
)

print(
    "============================================"
)


# ------------------------------------------------------------
# Create a new house using the MEDIAN/MODE of each feature
# ------------------------------------------------------------

new_house = pd.DataFrame(
    index=[0],
    columns=X.columns
)


# Fill numerical columns with their median
numeric_cols = X.select_dtypes(
    include=np.number
).columns


for column in numeric_cols:

    new_house.loc[0, column] = X[
        column
    ].median()


# Fill categorical columns with their most frequent value
categorical_cols = X.select_dtypes(
    exclude=np.number
).columns


for column in categorical_cols:

    new_house.loc[0, column] = X[
        column
    ].mode()[0]


# Make sure categorical columns retain their proper type
for column in categorical_cols:

    if str(X[column].dtype) == "category":

        new_house[column] = pd.Categorical(
            new_house[column],
            categories=X[column].cat.categories
        )


# ------------------------------------------------------------
# Display the new house data
# ------------------------------------------------------------

print("\nNew House Input:")

display(
    new_house
)


# ------------------------------------------------------------
# Predict class
# ------------------------------------------------------------

new_prediction = logistic_model.predict(
    new_house
)[0]


# ------------------------------------------------------------
# Predict probability
# ------------------------------------------------------------

new_probability = logistic_model.predict_proba(
    new_house
)[0][1]


# ------------------------------------------------------------
# Convert prediction to readable result
# ------------------------------------------------------------

if new_prediction == 1:

    result = "High Price"

else:

    result = "Low Price"


# ------------------------------------------------------------
# Display result
# ------------------------------------------------------------

print(
    "\nPredicted Class:",
    result
)

print(
    "Probability of High Price:",
    round(
        new_probability,
        4
    )
)

print(
    "Probability of Low Price:",
    round(
        1 - new_probability,
        4
    )
)


print(
    "\n============================================"
)
# ============================================================
# 51. FINAL SUMMARY
# ============================================================

print("""
============================================================

BOSTON HOUSING LOGISTIC REGRESSION
COMPLETE WORKFLOW

1. Load Boston Housing Dataset
2. Inspect Dataset
3. Check Data Types
4. Check Missing Values
5. Check Duplicate Records
6. Remove Duplicate Records
7. Statistical Summary
8. Distribution Analysis
9. Outlier Detection
10. IQR Analysis
11. Outlier Capping
12. Correlation Analysis
13. Convert MEDV into Classes
14. Low Price = 0
15. High Price = 1
16. Train-Test Split
17. Missing Value Imputation
18. Standardization
19. Categorical Encoding
20. Logistic Regression
21. Prediction
22. Accuracy
23. Precision
24. Recall
25. F1 Score
26. ROC-AUC
27. Confusion Matrix
28. Classification Report
29. Cross Validation
30. ROC Curve

============================================================
"""
)

LINEAR REGRESSION IN BOSTON HOUSING DATASET

 

# ============================================================
# BOSTON HOUSING DATASET
# COMPLETE DATA PREPROCESSING + LINEAR REGRESSION
# Google Colab Version
# ============================================================


# ============================================================
# 1. IMPORT LIBRARIES
# ============================================================

import numpy as np
import pandas as pd

import matplotlib.pyplot as plt
import seaborn as sns

from sklearn.datasets import fetch_openml

from sklearn.model_selection import train_test_split, cross_val_score

from sklearn.compose import ColumnTransformer

from sklearn.pipeline import Pipeline

from sklearn.impute import SimpleImputer

from sklearn.preprocessing import (
    StandardScaler,
    MinMaxScaler,
    RobustScaler,
    OneHotEncoder,
    PolynomialFeatures
)

from sklearn.feature_selection import SelectKBest, f_regression

from sklearn.linear_model import LinearRegression

from sklearn.metrics import (
    mean_absolute_error,
    mean_squared_error,
    r2_score
)

import warnings
warnings.filterwarnings("ignore")


# ============================================================
# 2. LOAD BOSTON HOUSING DATASET
# ============================================================

boston = fetch_openml(
    name="boston",
    version=1,
    as_frame=True
)

df = boston.frame.copy()

print("Dataset Loaded Successfully")


# ============================================================
# 3. DISPLAY DATASET
# ============================================================

print("\nFirst 5 Rows:")
display(df.head())


# ============================================================
# 4. DATASET SHAPE
# ============================================================

print("\nDataset Shape:")
print(df.shape)


# ============================================================
# 5. COLUMN NAMES
# ============================================================

print("\nColumn Names:")
print(df.columns.tolist())


# ============================================================
# 6. DATA INFORMATION
# ============================================================

print("\nDataset Information:")
df.info()


# ============================================================
# 7. DATA TYPES
# ============================================================

print("\nData Types:")
print(df.dtypes)


# ============================================================
# 8. CHECK MISSING VALUES
# ============================================================

print("\nMissing Values:")
print(df.isnull().sum())


# ============================================================
# 9. MISSING VALUE PERCENTAGE
# ============================================================

missing_percentage = (
    df.isnull().mean() * 100
).sort_values(ascending=False)

print("\nMissing Value Percentage:")
print(missing_percentage)


# ============================================================
# 10. DUPLICATE RECORDS
# ============================================================

print("\nNumber of Duplicate Rows:")
print(df.duplicated().sum())


# Remove duplicates if present

df = df.drop_duplicates()

print("\nShape After Duplicate Removal:")
print(df.shape)


# ============================================================
# 11. STATISTICAL SUMMARY
# ============================================================

print("\nStatistical Summary:")
display(df.describe())


# ============================================================
# 12. CHECK UNIQUE VALUES
# ============================================================

print("\nNumber of Unique Values:")
print(df.nunique())


# ============================================================
# 13. TARGET VARIABLE
# ============================================================

# Boston Housing target is MEDV

target = "MEDV"

print("\nTarget Variable:")
print(target)

print("\nTarget Statistics:")
print(df[target].describe())


# ============================================================
# 14. HISTOGRAM OF TARGET
# ============================================================

plt.figure(figsize=(9,5))

sns.histplot(
    df[target],
    kde=True
)

plt.title("Distribution of House Prices")
plt.xlabel("MEDV")
plt.ylabel("Frequency")

plt.show()


# ============================================================
# 15. CHECK OUTLIERS USING BOXPLOT
# ============================================================

plt.figure(figsize=(12,6))

sns.boxplot(
    data=df
)

plt.title("Boxplot of Boston Housing Features")
plt.xticks(rotation=45)

plt.show()


# ============================================================
# 16. OUTLIER DETECTION USING IQR
# ============================================================

numeric_columns = df.select_dtypes(
    include=np.number
).columns


def find_iqr_outliers(data, column):

    Q1 = data[column].quantile(0.25)

    Q3 = data[column].quantile(0.75)

    IQR = Q3 - Q1

    lower_bound = Q1 - 1.5 * IQR

    upper_bound = Q3 + 1.5 * IQR

    outliers = data[
        (data[column] < lower_bound) |
        (data[column] > upper_bound)
    ]

    return outliers


print("\nIQR Outlier Count:")

for column in numeric_columns:

    count = len(
        find_iqr_outliers(df, column)
    )

    print(
        column,
        ":",
        count
    )


# ============================================================
# 17. OUTLIER TREATMENT USING CAPPING
# ============================================================

# We preserve observations instead of deleting rows.

df_capped = df.copy()

for column in numeric_columns:

    Q1 = df_capped[column].quantile(0.25)

    Q3 = df_capped[column].quantile(0.75)

    IQR = Q3 - Q1

    lower_bound = Q1 - 1.5 * IQR

    upper_bound = Q3 + 1.5 * IQR

    df_capped[column] = df_capped[
        column
    ].clip(
        lower_bound,
        upper_bound
    )


print("\nOutlier capping completed.")


# ============================================================
# 18. CORRELATION MATRIX
# ============================================================

correlation = df.corr(
    numeric_only=True
)

print("\nCorrelation with MEDV:")

print(
    correlation[target]
    .sort_values(
        ascending=False
    )
)


# ============================================================
# 19. CORRELATION HEATMAP
# ============================================================

plt.figure(
    figsize=(12,9)
)

sns.heatmap(
    correlation,
    annot=True,
    fmt=".2f",
    cmap="coolwarm"
)

plt.title(
    "Boston Housing Correlation Heatmap"
)

plt.show()


# ============================================================
# 20. FEATURE / TARGET SEPARATION
# ============================================================

X = df.drop(
    columns=[target]
)

y = df[target]


print("\nFeature Shape:")
print(X.shape)

print("\nTarget Shape:")
print(y.shape)


# ============================================================
# 21. IDENTIFY NUMERICAL AND CATEGORICAL FEATURES
# ============================================================

numeric_features = X.select_dtypes(
    include=np.number
).columns.tolist()

categorical_features = X.select_dtypes(
    exclude=np.number
).columns.tolist()


print("\nNumerical Features:")
print(numeric_features)

print("\nCategorical Features:")
print(categorical_features)


# ============================================================
# 22. TRAIN-TEST SPLIT
# ============================================================

X_train, X_test, y_train, y_test = train_test_split(

    X,
    y,

    test_size=0.20,

    random_state=42
)


print("\nTraining Data:")
print(X_train.shape)

print("\nTesting Data:")
print(X_test.shape)


# ============================================================
# 23. NUMERICAL PREPROCESSING
# ============================================================

numeric_pipeline = Pipeline(

    steps=[

        (
            "imputer",
            SimpleImputer(
                strategy="median"
            )
        ),

        (
            "scaler",
            StandardScaler()
        )

    ]
)


# ============================================================
# 24. CATEGORICAL PREPROCESSING
# ============================================================

categorical_pipeline = Pipeline(

    steps=[

        (
            "imputer",
            SimpleImputer(
                strategy="most_frequent"
            )
        ),

        (
            "encoder",
            OneHotEncoder(
                handle_unknown="ignore"
            )
        )

    ]
)


# ============================================================
# 25. COMBINE PREPROCESSING
# ============================================================

preprocessor = ColumnTransformer(

    transformers=[

        (
            "numeric",
            numeric_pipeline,
            numeric_features
        ),

        (
            "categorical",
            categorical_pipeline,
            categorical_features
        )

    ]

)


# ============================================================
# 26. CREATE LINEAR REGRESSION PIPELINE
# ============================================================

linear_regression_model = Pipeline(

    steps=[

        (
            "preprocessing",
            preprocessor
        ),

        (
            "model",
            LinearRegression()
        )

    ]

)


# ============================================================
# 27. TRAIN LINEAR REGRESSION MODEL
# ============================================================

linear_regression_model.fit(
    X_train,
    y_train
)

print(
    "\nLinear Regression Model Trained Successfully!"
)


# ============================================================
# 28. PREDICTION
# ============================================================

y_pred = linear_regression_model.predict(
    X_test
)


# ============================================================
# 29. MODEL EVALUATION
# ============================================================

mae = mean_absolute_error(
    y_test,
    y_pred
)

mse = mean_squared_error(
    y_test,
    y_pred
)

rmse = np.sqrt(mse)

r2 = r2_score(
    y_test,
    y_pred
)


print("\n==============================")
print("LINEAR REGRESSION RESULTS")
print("==============================")

print(
    "Mean Absolute Error (MAE):",
    mae
)

print(
    "Mean Squared Error (MSE):",
    mse
)

print(
    "Root Mean Squared Error (RMSE):",
    rmse
)

print(
    "R² Score:",
    r2
)


# ============================================================
# 30. ACTUAL VS PREDICTED VALUES
# ============================================================

result = pd.DataFrame({

    "Actual Price": y_test.values,

    "Predicted Price": y_pred

})

print("\nActual vs Predicted:")
display(
    result.head(20)
)


# ============================================================
# 31. ACTUAL VS PREDICTED GRAPH
# ============================================================

plt.figure(
    figsize=(8,6)
)

plt.scatter(
    y_test,
    y_pred,
    alpha=0.7
)

plt.xlabel(
    "Actual House Price"
)

plt.ylabel(
    "Predicted House Price"
)

plt.title(
    "Actual vs Predicted House Prices"
)

# Perfect prediction reference line

minimum = min(
    y_test.min(),
    y_pred.min()
)

maximum = max(
    y_test.max(),
    y_pred.max()
)

plt.plot(
    [minimum, maximum],
    [minimum, maximum],
    linestyle="--"
)

plt.show()


# ============================================================
# 32. RESIDUAL ANALYSIS
# ============================================================

residuals = y_test - y_pred

plt.figure(
    figsize=(9,5)
)

sns.scatterplot(
    x=y_pred,
    y=residuals
)

plt.axhline(
    0,
    linestyle="--"
)

plt.xlabel(
    "Predicted Values"
)

plt.ylabel(
    "Residuals"
)

plt.title(
    "Residual Plot"
)

plt.show()


# ============================================================
# 33. RESIDUAL DISTRIBUTION
# ============================================================

plt.figure(
    figsize=(9,5)
)

sns.histplot(
    residuals,
    kde=True
)

plt.title(
    "Distribution of Residuals"
)

plt.xlabel(
    "Residual"
)

plt.show()


# ============================================================
# 34. CROSS VALIDATION
# ============================================================

cv_scores = cross_val_score(

    linear_regression_model,

    X,

    y,

    cv=5,

    scoring="r2"

)

print("\n5-Fold Cross Validation R² Scores:")

print(cv_scores)

print(
    "\nMean Cross Validation R²:",
    cv_scores.mean()
)


# ============================================================
# 35. TRY MIN-MAX SCALING
# ============================================================

minmax_pipeline = Pipeline(

    steps=[

        (
            "imputer",
            SimpleImputer(
                strategy="median"
            )
        ),

        (
            "scaler",
            MinMaxScaler()
        )

    ]
)


minmax_preprocessor = ColumnTransformer(

    transformers=[

        (
            "numeric",
            minmax_pipeline,
            numeric_features
        )

    ]

)


minmax_model = Pipeline(

    steps=[

        (
            "preprocessing",
            minmax_preprocessor
        ),

        (
            "model",
            LinearRegression()
        )

    ]

)


minmax_model.fit(
    X_train,
    y_train
)


minmax_prediction = minmax_model.predict(
    X_test
)


minmax_r2 = r2_score(
    y_test,
    minmax_prediction
)


print(
    "\nR² using Min-Max Scaling:",
    minmax_r2
)


# ============================================================
# 36. TRY ROBUST SCALING
# ============================================================

robust_pipeline = Pipeline(

    steps=[

        (
            "imputer",
            SimpleImputer(
                strategy="median"
            )
        ),

        (
            "scaler",
            RobustScaler()
        )

    ]

)


robust_preprocessor = ColumnTransformer(

    transformers=[

        (
            "numeric",
            robust_pipeline,
            numeric_features
        )

    ]

)


robust_model = Pipeline(

    steps=[

        (
            "preprocessing",
            robust_preprocessor
        ),

        (
            "model",
            LinearRegression()
        )

    ]

)


robust_model.fit(
    X_train,
    y_train
)


robust_prediction = robust_model.predict(
    X_test
)


robust_r2 = r2_score(
    y_test,
    robust_prediction
)


print(
    "\nR² using Robust Scaling:",
    robust_r2
)


# ============================================================
# 37. COMPARE SCALING METHODS
# ============================================================

comparison = pd.DataFrame({

    "Preprocessing": [
        "StandardScaler",
        "MinMaxScaler",
        "RobustScaler"
    ],

    "R2 Score": [
        r2,
        minmax_r2,
        robust_r2
    ]

})


print("\nScaling Comparison:")

display(
    comparison
)


# ============================================================
# 38. FEATURE SELECTION
# ============================================================

# Select top 8 numerical features

feature_selector = Pipeline(

    steps=[

        (
            "imputer",
            SimpleImputer(
                strategy="median"
            )
        ),

        (
            "scaler",
            StandardScaler()
        ),

        (
            "selection",
            SelectKBest(
                score_func=f_regression,
                k=8
            )
        )

    ]

)


feature_selection_model = Pipeline(

    steps=[

        (
            "preprocessing",
            feature_selector
        ),

        (
            "model",
            LinearRegression()
        )

    ]

)


feature_selection_model.fit(
    X_train[numeric_features],
    y_train
)


feature_prediction = (
    feature_selection_model.predict(
        X_test[numeric_features]
    )
)


feature_r2 = r2_score(
    y_test,
    feature_prediction
)


print(
    "\nR² After Feature Selection:",
    feature_r2
)


# ============================================================
# 39. FINAL MODEL SUMMARY
# ============================================================

print("\n")
print("======================================")
print("       FINAL MODEL SUMMARY")
print("======================================")

print(
    "Number of Training Samples:",
    len(X_train)
)

print(
    "Number of Testing Samples:",
    len(X_test)
)

print(
    "MAE:",
    round(mae, 4)
)

print(
    "MSE:",
    round(mse, 4)
)

print(
    "RMSE:",
    round(rmse, 4)
)

print(
    "R²:",
    round(r2, 4)
)

print(
    "Cross Validation Mean R²:",
    round(cv_scores.mean(), 4)
)

print("======================================")


# ============================================================
# 40. COMPLETE PREPROCESSING WORKFLOW
# ============================================================

print("""
============================================================

COMPLETE DATA PREPROCESSING WORKFLOW

1. Load Dataset
2. Inspect Dataset
3. Check Data Types
4. Check Missing Values
5. Handle Missing Values
6. Check Duplicate Records
7. Remove Duplicates
8. Statistical Summary
9. Distribution Analysis
10. Outlier Detection
11. IQR Outlier Analysis
12. Outlier Capping
13. Correlation Analysis
14. Feature / Target Separation
15. Train-Test Split
16. Numerical Imputation
17. Categorical Imputation
18. Categorical Encoding
19. Standardization
20. Min-Max Scaling
21. Robust Scaling
22. Feature Selection
23. Linear Regression
24. Prediction
25. MAE
26. MSE
27. RMSE
28. R² Score
29. Residual Analysis
30. Cross Validation

============================================================
""")


# ============================================================
# 41. CLASSIFICATION VERSION OF BOSTON HOUSING
# ============================================================
#
# IMPORTANT:
# Linear Regression predicts continuous house prices.
# Therefore, Precision, Recall and Confusion Matrix are
# not directly applicable to the original regression target.
#
# For educational purposes, we convert MEDV into two classes:
#
# 0 = Low Price
# 1 = High Price
#
# Median house price is used as the threshold.
# ============================================================


from sklearn.metrics import (
    accuracy_score,
    precision_score,
    recall_score,
    f1_score,
    confusion_matrix,
    classification_report,
    ConfusionMatrixDisplay
)


# ============================================================
# 42. CREATE CLASSIFICATION TARGET
# ============================================================

classification_df = df.copy()

median_price = classification_df["MEDV"].median()

print("Median House Price:", median_price)


classification_df["Price_Class"] = (
    classification_df["MEDV"] >= median_price
).astype(int)


print("\nClassification Target:")
print(classification_df["Price_Class"].value_counts())


# ============================================================
# 43. CREATE FEATURES AND TARGET
# ============================================================

X_class = classification_df.drop(
    columns=["MEDV", "Price_Class"]
)

y_class = classification_df["Price_Class"]


# ============================================================
# 44. TRAIN-TEST SPLIT
# ============================================================

X_train_class, X_test_class, y_train_class, y_test_class = train_test_split(

    X_class,
    y_class,

    test_size=0.20,

    random_state=42,

    stratify=y_class

)


print("\nClassification Training Shape:")
print(X_train_class.shape)

print("\nClassification Testing Shape:")
print(X_test_class.shape)


# ============================================================
# 45. PREPROCESSING PIPELINE
# ============================================================

numeric_features_class = X_class.select_dtypes(
    include=np.number
).columns.tolist()


categorical_features_class = X_class.select_dtypes(
    exclude=np.number
).columns.tolist()


numeric_pipeline_class = Pipeline(

    steps=[

        (
            "imputer",
            SimpleImputer(
                strategy="median"
            )
        ),

        (
            "scaler",
            StandardScaler()
        )

    ]

)


categorical_pipeline_class = Pipeline(

    steps=[

        (
            "imputer",
            SimpleImputer(
                strategy="most_frequent"
            )
        ),

        (
            "encoder",
            OneHotEncoder(
                handle_unknown="ignore"
            )
        )

    ]

)


classification_preprocessor = ColumnTransformer(

    transformers=[

        (
            "numeric",
            numeric_pipeline_class,
            numeric_features_class
        ),

        (
            "categorical",
            categorical_pipeline_class,
            categorical_features_class
        )

    ]

)


# ============================================================
# 46. CLASSIFICATION MODEL
# ============================================================

# Logistic Regression is used because precision,
# recall and confusion matrix are classification metrics.

from sklearn.linear_model import LogisticRegression


classification_model = Pipeline(

    steps=[

        (
            "preprocessing",
            classification_preprocessor
        ),

        (
            "model",
            LogisticRegression(
                max_iter=2000
            )
        )

    ]

)


# ============================================================
# 47. TRAIN CLASSIFICATION MODEL
# ============================================================

classification_model.fit(

    X_train_class,

    y_train_class

)


print(
    "\nClassification Model Trained Successfully!"
)


# ============================================================
# 48. CLASSIFICATION PREDICTIONS
# ============================================================

y_class_pred = classification_model.predict(
    X_test_class
)


# ============================================================
# 49. ACCURACY
# ============================================================

accuracy = accuracy_score(

    y_test_class,

    y_class_pred

)


print("\n==============================")
print("CLASSIFICATION RESULTS")
print("==============================")

print(
    "Accuracy:",
    round(accuracy, 4)
)


# ============================================================
# 50. PRECISION
# ============================================================

precision = precision_score(

    y_test_class,

    y_class_pred,

    zero_division=0

)


print(
    "Precision:",
    round(precision, 4)
)


# ============================================================
# 51. RECALL
# ============================================================

recall = recall_score(

    y_test_class,

    y_class_pred,

    zero_division=0

)


print(
    "Recall:",
    round(recall, 4)
)


# ============================================================
# 52. F1 SCORE
# ============================================================

f1 = f1_score(

    y_test_class,

    y_class_pred,

    zero_division=0

)


print(
    "F1 Score:",
    round(f1, 4)
)


# ============================================================
# 53. CONFUSION MATRIX
# ============================================================

cm = confusion_matrix(

    y_test_class,

    y_class_pred

)


print("\nConfusion Matrix:")
print(cm)


# ============================================================
# 54. DISPLAY CONFUSION MATRIX
# ============================================================

plt.figure(
    figsize=(7,6)
)

disp = ConfusionMatrixDisplay(

    confusion_matrix=cm,

    display_labels=[
        "Low Price",
        "High Price"
    ]

)

disp.plot()

plt.title(
    "Boston Housing Confusion Matrix"
)

plt.show()


# ============================================================
# 55. EXTRACT TP, TN, FP, FN
# ============================================================

TN, FP, FN, TP = cm.ravel()


print("\n==============================")
print("CONFUSION MATRIX VALUES")
print("==============================")

print(
    "True Negative (TN):",
    TN
)

print(
    "False Positive (FP):",
    FP
)

print(
    "False Negative (FN):",
    FN
)

print(
    "True Positive (TP):",
    TP
)


# ============================================================
# 56. MANUAL PRECISION CALCULATION
# ============================================================

manual_precision = TP / (
    TP + FP
) if (TP + FP) != 0 else 0


print(
    "\nManual Precision:",
    round(manual_precision, 4)
)


# ============================================================
# 57. MANUAL RECALL CALCULATION
# ============================================================

manual_recall = TP / (
    TP + FN
) if (TP + FN) != 0 else 0


print(
    "Manual Recall:",
    round(manual_recall, 4)
)


# ============================================================
# 58. MANUAL F1 SCORE
# ============================================================

manual_f1 = (

    2 * manual_precision * manual_recall

    / (manual_precision + manual_recall)

    if (manual_precision + manual_recall) != 0

    else 0

)


print(
    "Manual F1 Score:",
    round(manual_f1, 4)
)


# ============================================================
# 59. CLASSIFICATION REPORT
# ============================================================

print("\n==============================")
print("CLASSIFICATION REPORT")
print("==============================")

print(

    classification_report(

        y_test_class,

        y_class_pred,

        target_names=[
            "Low Price",
            "High Price"
        ],

        zero_division=0

    )

)


# ============================================================
# 60. METRICS SUMMARY TABLE
# ============================================================

metrics_table = pd.DataFrame({

    "Metric": [

        "Accuracy",

        "Precision",

        "Recall",

        "F1 Score"

    ],

    "Score": [

        accuracy,

        precision,

        recall,

        f1

    ]

})


metrics_table["Score"] = metrics_table[
    "Score"
].round(4)


print("\n==============================")
print("MODEL PERFORMANCE TABLE")
print("==============================")

display(
    metrics_table
)


# ============================================================
# 61. CONFUSION MATRIX TABLE
# ============================================================

confusion_table = pd.DataFrame(

    cm,

    index=[
        "Actual Low Price",
        "Actual High Price"
    ],

    columns=[
        "Predicted Low Price",
        "Predicted High Price"
    ]

)


print("\n==============================")
print("CONFUSION MATRIX TABLE")
print("==============================")

display(
    confusion_table
)


# ============================================================
# 62. COMPLETE METRICS SUMMARY
# ============================================================

final_metrics = pd.DataFrame({

    "Metric": [

        "MAE",
        "MSE",
        "RMSE",
        "R² Score",
        "Classification Accuracy",
        "Classification Precision",
        "Classification Recall",
        "Classification F1 Score"

    ],

    "Value": [

        mae,
        mse,
        rmse,
        r2,
        accuracy,
        precision,
        recall,
        f1

    ]

})


final_metrics["Value"] = final_metrics[
    "Value"
].round(4)


print("\n======================================")
print("COMPLETE BOSTON HOUSING MODEL METRICS")
print("======================================")

display(
    final_metrics
)


# ============================================================
# 63. VISUALIZE PRECISION, RECALL AND F1
# ============================================================

classification_metrics = {

    "Precision": precision,

    "Recall": recall,

    "F1 Score": f1,

    "Accuracy": accuracy

}


plt.figure(
    figsize=(9,6)
)

plt.bar(

    classification_metrics.keys(),

    classification_metrics.values()

)

plt.ylim(
    0,
    1
)

plt.ylabel(
    "Score"
)

plt.xlabel(
    "Metric"
)

plt.title(
    "Classification Performance Metrics"
)

plt.show()


# ============================================================
# 64. CONFUSION MATRIX HEATMAP
# ============================================================

plt.figure(
    figsize=(7,6)
)

sns.heatmap(

    cm,

    annot=True,

    fmt="d",

    cmap="Blues",

    xticklabels=[
        "Predicted Low",
        "Predicted High"
    ],

    yticklabels=[
        "Actual Low",
        "Actual High"
    ]

)

plt.title(
    "Confusion Matrix Heatmap"
)

plt.xlabel(
    "Predicted Class"
)

plt.ylabel(
    "Actual Class"
)

plt.show()


# ============================================================
# 65. INTERPRETATION
# ============================================================

print("""
============================================================
INTERPRETATION OF CLASSIFICATION METRICS
============================================================

Accuracy:
Measures the proportion of all predictions that are correct.

Precision:
Among observations predicted as High Price, precision tells us
how many were actually High Price.

Recall:
Among observations that were actually High Price, recall tells
us how many were correctly identified as High Price.

F1 Score:
Combines Precision and Recall into a single measure.

True Positive (TP):
Actual High Price and predicted High Price.

True Negative (TN):
Actual Low Price and predicted Low Price.

False Positive (FP):
Actual Low Price but predicted High Price.

False Negative (FN):
Actual High Price but predicted Low Price.

============================================================
IMPORTANT:
============================================================

The original Boston Housing problem is a REGRESSION problem.

Therefore:

MAE, MSE, RMSE and R²
        ↓
are the appropriate regression metrics.

Precision, Recall, F1 Score and Confusion Matrix
        ↓
require a CLASSIFICATION problem.

For teaching purposes, the continuous MEDV target was converted
into:

0 = Low Price
1 = High Price

using the median house price as the threshold.

============================================================
""")

BOSTON DATA HOUSE

๐Ÿค– MACHINE LEARNING PROGRAMS

Step-by-Step Implementation Using Python & Machine Learning

01
๐Ÿ“ˆ Linear Regression
▶ CLICK HERE
02
๐Ÿงฎ Logistic Regression
▶ CLICK HERE
03
๐ŸŽฏ K-Nearest Neighbors (KNN)
▶ CLICK HERE
04
๐Ÿ”ต K-Medoid
▶ CLICK HERE
05
๐ŸŸ  K-Means Clustering
▶ CLICK HERE
06
๐ŸŒณ Hierarchical Clustering
▶ CLICK HERE
07
๐Ÿ” DBSCAN
▶ CLICK HERE
08
๐Ÿงฉ Feature Selection
▶ CLICK HERE
09
๐Ÿ”ฌ Feature Extraction
▶ CLICK HERE
10
๐Ÿ”„ Cross Validation
▶ CLICK HERE
09
๐Ÿ”ฌ DECISION TREE 
▶ CLICK HERE
10
๐Ÿ”„ --
▶ CLICK HERE

Data Preprocessing in Machine Learning step by step theory


๐Ÿ“Š

Data Preprocessing in Machine Learning

Complete Theory, Techniques, Importance, Best Practices, Advantages, Disadvantages and Frequently Asked Questions

๐Ÿงน
Data Cleaning Removes errors and inconsistencies
๐Ÿ”
Data Analysis Understands patterns and quality
๐Ÿ› ️
Transformation Converts data into useful forms
๐Ÿท️
Encoding Converts categories into numerical form
๐Ÿ“
Scaling Makes numerical features comparable
๐ŸŽฏ
Feature Selection Keeps useful variables

1. What is Data Preprocessing?

๐Ÿ“– Definition

Data preprocessing is the process of cleaning, organizing, transforming and preparing raw data before it is supplied to a machine learning model.

Real-world datasets are rarely perfect. They may contain missing values, duplicate records, incorrect entries, inconsistent formats, noisy observations, irrelevant variables and extreme values.

๐Ÿ’ก Simple Answer:

Data preprocessing converts raw and possibly problematic data into a cleaner and more suitable form so that a machine learning algorithm can learn meaningful patterns.

2. Why is Data Preprocessing Important?

✅

Improves Data Quality

Helps correct incomplete, inaccurate and inconsistent records.

⚡

Faster Training

Organized data can reduce unnecessary complexity and help algorithms train more efficiently.

๐ŸŽฏ

Better Predictions

Cleaner data allows models to learn more meaningful relationships.

๐Ÿ“

Feature Consistency

Scaling and encoding put different feature types into forms algorithms can process.

๐Ÿ›ก️

Generalization

Properly prepared datasets can help reduce the risk of learning noise instead of useful patterns.

๐Ÿ”Ž

Better Interpretation

Clean and structured data makes patterns, relationships and feature importance easier to examine.

3. Common Data Quality Problems

❓

Missing Values

A field has no recorded value because information may have been unavailable, omitted, incorrectly entered, or lost during data collection.

๐Ÿ“‹

Duplicate Records

The same observation appears more than once and can give excessive importance to particular observations.

๐Ÿ”Š

Noisy Data

Data containing random errors, incorrect measurements, typing mistakes or other unwanted variations.

⚠️

Outliers

Observations that are substantially different from the majority of observations.

๐Ÿ”„

Inconsistent Formats

Different representations of dates, units, text, capitalization or other values.

๐Ÿšซ

Irrelevant Features

Variables that provide little or no useful information for the machine learning task.

4. Complete Data Preprocessing Pipeline

1. Data Collection
→
2. Data Understanding
→
3. Data Cleaning
→
4. Missing Values
→
5. Duplicate Removal
→
6. Transformation
→
7. Encoding
→
8. Feature Selection
→
9. Scaling

5. Data Collection and Understanding

๐Ÿ“– Definition

This stage involves obtaining relevant and reliable data and understanding its structure, variables, distributions, feature types and quality problems.

Exploratory Data Analysis (EDA)

Exploratory Data Analysis is used to investigate the dataset before modeling. It helps reveal relationships, distributions, correlations, outliers, anomalies and possible class imbalance.

๐Ÿ’ก Answer:

Before cleaning or transforming data, first understand what the data contains and what problem the model needs to solve.

6. Data Cleaning

๐Ÿงน Definition

Data cleaning is the process of identifying and correcting errors, inconsistencies, invalid values, duplicate records and unsuitable formats in a dataset.

Important Data Cleaning Activities

  • Correct invalid values.
  • Correct spelling mistakes.
  • Standardize date formats.
  • Standardize measurement units.
  • Remove duplicate records.
  • Identify impossible values.
  • Correct errors introduced during data collection.

7. Handling Missing Values

❓ Definition

Missing values occur when a particular observation does not contain a value for one or more variables.

Methods of Handling Missing Values

๐Ÿ—‘️ Deletion

Rows or columns containing excessive missing information may be removed when doing so is appropriate.

๐Ÿ“Š Mean Imputation

A missing numerical value can be replaced with the mean of the relevant feature.

๐Ÿ“ˆ Median Imputation

Missing numerical values can be replaced using the median, which is often less influenced by extreme values.

๐Ÿท️ Mode Imputation

Missing categorical values can be replaced with the most frequently occurring category.

๐Ÿค– Predictive Imputation

Statistical or machine learning methods can estimate missing values from other available information.

⚠️ Important:

The choice of missing-value treatment depends on the amount of missing information, the feature's nature, possible bias and the importance of the variable.

8. Removing Duplicates and Irrelevant Records

Duplicate Records

Duplicate observations may cause certain samples to receive excessive influence during training. They can also distort feature distributions and evaluation results.

Irrelevant Features

Irrelevant variables add unnecessary complexity without contributing useful information to the prediction task. Removing them can simplify the model and improve interpretability.

๐Ÿ’ก Answer:

Remove duplicate or irrelevant information when there is sufficient evidence that it does not represent useful information for the specific machine learning task.

9. Data Transformation

๐Ÿ”„ Definition

Data transformation changes raw feature values into a representation that is more appropriate for machine learning algorithms.

๐Ÿ“ Scaling

Changes numerical features to comparable scales so that large-valued variables do not dominate algorithms that are sensitive to magnitude.

๐Ÿ“ Normalization

Rescales numerical values to a fixed range, commonly 0 to 1.

๐Ÿ“Š Standardization

Centers a numerical feature around zero and scales it according to its standard deviation.

๐Ÿ“‰ Log Transformation

Compresses large numerical values and can reduce positive skewness and stabilize variance.

๐Ÿชฃ Binning

Converts continuous numerical values into meaningful groups or intervals.

๐Ÿงฎ Mathematical Transformation

Mathematical functions can transform variables into representations that are more useful for analysis.

10. Normalization — Min-Max Scaling

๐Ÿ“– Definition

Normalization rescales numerical values to a fixed interval, commonly from 0 to 1.

๐ŸŽฏ Common Applications:

Normalization is often useful for algorithms where distances or feature magnitudes matter, including KNN, K-Means, SVM and some neural-network applications.

⚠️ Limitation:

Min-Max scaling is sensitive to extreme values because the minimum and maximum determine the resulting range.

11. Standardization — Z-Score Scaling

๐Ÿ“– Definition

Standardization transforms numerical features so that their center is around zero and their scale is measured in standard-deviation units.

Common Applications

  • Linear Regression
  • Logistic Regression
  • Support Vector Machines
  • Principal Component Analysis
  • Other algorithms benefiting from comparable feature scales
⚠️ Important:

Standardization is generally less sensitive to outliers than Min-Max scaling, but extreme observations can still influence the mean and standard deviation.

12. Normalization vs Standardization

Feature Normalization Standardization
Purpose Places values in a fixed range Centers and scales values
Typical Range 0 to 1 No fixed range
Based On Minimum and maximum Mean and standard deviation
Outlier Sensitivity High Can still be affected
Typical Uses Distance-based algorithms and neural networks Regression, PCA and SVM

13. Log Transformation

๐Ÿ“‰ Definition

Log transformation applies a logarithmic transformation to numerical data to compress large values and reduce positive skewness.

Benefits

  • Can reduce skewness.
  • Can stabilize variance.
  • Compresses very large values.
  • Can reduce the influence of extreme magnitudes.
⚠️ Important:

Ordinary logarithms cannot directly process zero or negative values. Appropriate adjustments are required before applying a logarithmic transformation.

14. Encoding Categorical Variables

๐Ÿท️ Definition

Categorical encoding converts non-numerical categories into numerical representations that machine learning algorithms can process.

๐Ÿ”ข

Label Encoding

Assigns an integer value to each category.

It is particularly appropriate when categories have a meaningful natural order.

Example:

Poor → Fair → Good → Excellent

๐Ÿ”ฒ

One-Hot Encoding

Creates a separate binary column for each category. Each observation receives 0 or 1 in those columns.

It is suitable for nominal categories that do not have a natural ranking.

Example:

Department: IT, HR, Finance

15. Choosing the Wrong Encoding Method

⚠️ Why is this a problem?

Applying numerical labels to categories that have no natural order can accidentally make an algorithm interpret the assigned numbers as meaningful rankings.

On the other hand, one-hot encoding a feature with a very large number of unique categories can create a huge number of columns and increase memory and computational requirements.

๐Ÿ’ก Selection should consider:
  • Whether the variable is nominal or ordinal.
  • The number of unique categories.
  • The size of the dataset.
  • The machine learning algorithm.

16. Feature Selection

๐ŸŽฏ Definition

Feature selection identifies useful variables and removes redundant or less relevant variables from a dataset.

Benefits

  • Reduces unnecessary complexity.
  • Can improve model interpretability.
  • Can reduce training time.
  • Can reduce the risk of overfitting.
  • Allows the model to concentrate on useful information.

17. Dimensionality Reduction

๐Ÿ“‰ Definition

Dimensionality reduction reduces the number of variables or dimensions used to represent a dataset while attempting to preserve important information.

๐Ÿ’ก Main Purpose:

It can simplify complex datasets, reduce computational requirements and accelerate machine learning workflows.

18. Outliers — Keep or Remove?

⚠️ Definition

An outlier is an observation that is substantially different from the general pattern of the remaining observations.

When Should an Outlier Be Kept?

An unusual observation may represent a genuine and important event. Examples can include unusual customer behavior or legitimate rare events.

When Can an Outlier Be Removed?

If investigation shows that an extreme value was caused by a measurement problem, equipment failure, data-entry mistake or another known error, removal or correction may be appropriate.

Important Rule:

Never remove an outlier simply because it looks unusual. First investigate its source and the domain context.

19. Feature Engineering

๐Ÿง  Definition

Feature engineering is the process of creating or modifying variables so that they contain useful information for a machine learning model.

Feature engineering can include creating interaction variables, mathematical transformations, discretization, binning and other domain-specific representations.

20. Automated Preprocessing Workflows

๐Ÿ”— Pipeline

A pipeline combines multiple preprocessing and modeling operations into a reusable sequence.

This helps ensure that the same transformations are consistently applied to training, validation and deployment data.

๐Ÿงฉ ColumnTransformer

ColumnTransformer allows different preprocessing operations to be applied to different groups of columns.

For example, numerical variables can be scaled while categorical variables are encoded in the same workflow.

๐Ÿ’ก Main Benefits:
  • Consistency
  • Reproducibility
  • Fewer preprocessing errors
  • Less repeated code
  • Easier deployment
  • Easier maintenance

21. Tools That Can Automate Data Preprocessing

☁️ Google AutoML

Can automate parts of the machine learning workflow, including preprocessing and model development.

☁️ Azure AutoML

Can evaluate preprocessing strategies and models while supporting experiment tracking and deployment workflows.

๐Ÿค– H2O.ai

Provides automated machine learning capabilities including preprocessing, feature engineering and model selection.

๐Ÿงฐ Scikit-Learn

Provides preprocessing utilities, imputers, scalers, encoders, feature-selection tools and pipelines.

22. Data Preprocessing Best Practices

๐Ÿ” 1. Explore First

Examine statistics and visualizations before making major preprocessing decisions.

๐Ÿง  2. Use Domain Knowledge

Understand the real-world meaning of variables, missing values and unusual observations.

๐Ÿ“ 3. Document Changes

Record imputation, scaling, encoding, feature selection and transformation decisions.

๐ŸŽฏ 4. Match the ML Task

Preprocessing should reflect whether the problem is classification, regression or another machine learning task.

23. Advantages of Data Preprocessing

  • Improves the quality of data supplied to machine learning models.
  • Helps models learn useful patterns rather than data errors.
  • Can improve generalization by reducing irrelevant noise.
  • Handles missing values, duplicates and inconsistencies.
  • Reduces unnecessary noise.
  • Can reduce training time by removing irrelevant information.
  • Makes datasets easier to analyze and visualize.
  • Can improve model interpretability.

24. Disadvantages of Data Preprocessing

  • Preprocessing can require substantial time and effort.
  • An inappropriate preprocessing decision can reduce model quality.
  • Large datasets can require considerable computational resources.
  • Excessive cleaning may accidentally remove useful information.
  • Automated tools can hide assumptions made during preprocessing.
  • Complex workflows may require additional maintenance and debugging.
  • Good preprocessing often requires domain knowledge.

25. Common Data Preprocessing Mistakes

๐Ÿšจ Data Leakage

Fitting an imputer, scaler or encoder using the entire dataset before separating training and test data can allow information from the test set to influence training.

๐Ÿšจ Blind Outlier Removal

Removing every unusual observation without investigating its cause can eliminate valuable information.

๐Ÿšจ Wrong Encoding

Using ordinal-style numerical labels for nominal categories can create artificial relationships.

๐Ÿšจ High-Cardinality One-Hot Encoding

One-hot encoding variables with hundreds or thousands of categories can produce very large sparse feature spaces.

๐Ÿšจ No Test Validation

Preprocessing should be checked on unseen data to make sure transformations behave consistently.

26. Preprocessing According to ML Task

Classification Regression
May require attention to class imbalance. Often requires attention to skewness and influential observations.
Categorical target labels may need appropriate encoding. The target is usually continuous.
Feature transformations should support class prediction. Feature and target distributions should be examined carefully.

27. Automation vs Manual Preprocessing

๐Ÿค– Automation

Useful for repetitive operations, large datasets and standardized workflows.

๐Ÿ‘จ‍๐Ÿซ Manual Decisions

Important when unusual data patterns, domain rules, business requirements or critical decisions require human judgment.

๐Ÿ’ก Key Point:

Automated tools can reduce repetitive work, but preprocessing decisions should still be reviewed using statistical reasoning and domain knowledge.

28. Frequently Asked Questions

Q1. What is data preprocessing in simple words?
Answer: Data preprocessing means cleaning, organizing and transforming raw data so that it can be effectively used by a machine learning model.
Q2. What are the main steps of data preprocessing?
Answer: The major steps include data understanding, data cleaning, handling missing values, removing duplicates and irrelevant records, transformation, categorical encoding, feature selection and dimensionality reduction.
Q3. What is the difference between normalization and standardization?
Answer: Normalization maps values to a fixed range such as 0–1. Standardization centers values around zero and scales them using standard deviation.
Q4. Why is preprocessing important?
Answer: It helps machine learning algorithms work with cleaner, more consistent and appropriately represented data, improving the reliability of the learning process.
Q5. What tools are used for data preprocessing?
Answer: Common tools include Scikit-Learn preprocessing utilities, Pipeline, ColumnTransformer and automated machine learning platforms such as Google AutoML, Azure AutoML and H2O.ai.
Q6. Should every outlier be removed?
Answer: No. An outlier may be a genuine and meaningful observation. It should be investigated before deciding whether it should be retained, corrected or removed.
Q7. Why should preprocessing be fitted only on training data?
Answer: Fitting preprocessing transformations using test data can cause data leakage and make evaluation results appear better than they would be on truly unseen data.
Q8. What is one-hot encoding?
Answer: One-hot encoding converts each category of a nominal variable into a separate binary feature containing 0 or 1.
Q9. What is label encoding?
Answer: Label encoding assigns a unique integer to each category. It is most appropriate when the categories have a meaningful order.
Q10. How much of an ML project can preprocessing require?
Answer: The referenced Online Manipal article states that preprocessing can account for roughly 60–80% of the total project time in real-world scenarios, although the actual proportion varies considerably by project.

29. Quick Revision — One Page Summary

Concept Simple Meaning
Data Preprocessing Preparing raw data for machine learning.
Data Cleaning Correcting errors and inconsistencies.
Missing Values Handling unavailable observations.
Duplicates Removing repeated records when appropriate.
Outliers Investigating unusually different observations.
Normalization Scaling values to a fixed range.
Standardization Centering and scaling using mean and standard deviation.
Log Transformation Compressing large values and reducing skewness.
One-Hot Encoding Binary columns for nominal categories.
Label Encoding Integer labels for ordered categories.
Feature Selection Keeping relevant variables.
Dimensionality Reduction Representing data with fewer dimensions.
Pipeline Combining preprocessing operations into a reusable workflow.

๐Ÿ“š Learning Reference

This educational section is based on the concepts discussed in the Online Manipal article on data preprocessing, including preprocessing steps, numerical and categorical techniques, best practices, advantages, disadvantages, common mistakes and FAQs.

```