Total Pageviews

Thursday, October 1, 2026

Hierarchical Clustering

Hierarchical Clustering

Hierarchical Clustering is an unsupervised machine learning technique used to group similar data points into clusters. Instead of directly producing only one partition, it builds a hierarchy of clusters.

There are two main approaches:

  • Agglomerative Clustering – Bottom-Up approach
  • Divisive Clustering – Top-Down approach

1.

a )  Agglomerative Hierarchical Clustering

Agglomerative clustering starts by treating every data point as an individual cluster. The closest clusters are then repeatedly merged until all observations belong to one large cluster.

This is therefore called a Bottom-Up approach.


b )  Divisive Clustering

Divisive Clustering is known as top-down approach.

In this approach we take on huge cluster and starts breaking it up into smaller clusters until it reaches individual data points (or single point clusters).


2. Basic Working

  1. Initially, every data point is considered as a separate cluster.
  2. Calculate the distance between clusters.
  3. Find the two closest clusters.
  4. Merge those two clusters.
  5. Recalculate the distances between the new cluster and the remaining clusters.
  6. Continue the process until one large cluster remains.
  7. Represent the merging process using a Dendrogram.
  8. Choose the desired number of clusters by cutting the dendrogram at an appropriate height.

3. Distance Calculation

Euclidean Distance:

d(P₁,P₂) = √[(X₁ − X₂)² + (Y₁ − Y₂)²]

Example:
If P₁ = (170,56) and P₂ = (168,60)

d = √[(170−168)² + (56−60)²]
d = √(4 + 16)
d = √20 ≈ 4.47

4. Linkage Methods

Single Linkage

Distance between two clusters is based on the closest pair of points.

Complete Linkage

Distance is based on the farthest pair of points between two clusters.

Average Linkage

Distance is calculated using the average distance between points of two clusters.

Ward Linkage

Merges clusters while minimizing the increase in within-cluster variance.

5. Example Dataset

We will use a small Height and Weight dataset similar to the example described in the reference article.

Point Height Weight
P0 165 55
P1 170 56
P2 168 60
P3 178 70
P4 166 55
P5 180 72

6. Python Implementation

Google Colab / Python Code

# ============================================
# HIERARCHICAL CLUSTERING
# AGGLOMERATIVE APPROACH
# ============================================

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

from sklearn.cluster import AgglomerativeClustering
from scipy.cluster.hierarchy import dendrogram, linkage

# --------------------------------------------
# 1. Create Dataset
# --------------------------------------------

data = {
    "Height": [165, 170, 168, 178, 166, 180],
    "Weight": [55, 56, 60, 70, 55, 72]
}

df = pd.DataFrame(data)

print("Dataset:")
print(df)


# --------------------------------------------
# 2. Select Features
# --------------------------------------------

X = df[["Height", "Weight"]]


# --------------------------------------------
# 3. Create Dendrogram
# --------------------------------------------

linked = linkage(
    X,
    method="single",
    metric="euclidean"
)

plt.figure(figsize=(10,6))

dendrogram(
    linked,
    labels=["P0","P1","P2","P3","P4","P5"]
)

plt.title("Hierarchical Clustering Dendrogram")
plt.xlabel("Data Points")
plt.ylabel("Euclidean Distance")

plt.show()


# --------------------------------------------
# 4. Agglomerative Clustering
# --------------------------------------------

model = AgglomerativeClustering(
    n_clusters=2,
    metric="euclidean",
    linkage="single"
)

labels = model.fit_predict(X)


# --------------------------------------------
# 5. Add Cluster Labels
# --------------------------------------------

df["Cluster"] = labels

print("\nClustered Dataset:")
print(df)


# --------------------------------------------
# 6. Visualize Clusters
# --------------------------------------------

plt.figure(figsize=(8,6))

plt.scatter(
    df["Height"],
    df["Weight"],
    c=df["Cluster"],
    s=120
)

plt.xlabel("Height")
plt.ylabel("Weight")
plt.title("Agglomerative Hierarchical Clustering")

plt.grid(True)

plt.show()

7. Dendrogram

A dendrogram is a tree-like diagram that shows the sequence in which clusters are merged.

The vertical axis represents the distance at which clusters are merged. A horizontal cut through the dendrogram can be used to determine the number of clusters.

8. How to Select Number of Clusters?

  1. Generate the dendrogram.
  2. Look for the largest vertical gap that does not intersect another horizontal cluster line.
  3. Draw a horizontal line through that gap.
  4. Count the number of vertical branches crossed by the line.
  5. The number of branches represents the selected number of clusters.

9. Using Different Linkage Methods

Compare Single, Complete, Average and Ward Linkage

import matplotlib.pyplot as plt
from scipy.cluster.hierarchy import linkage, dendrogram

methods = [
    "single",
    "complete",
    "average",
    "ward"
]

plt.figure(figsize=(16,10))

for i, method in enumerate(methods):

    plt.subplot(2,2,i+1)

    linked = linkage(
        X,
        method=method,
        metric="euclidean"
    )

    dendrogram(
        linked,
        labels=["P0","P1","P2","P3","P4","P5"]
    )

    plt.title(method.capitalize() + " Linkage")
    plt.xlabel("Data Points")
    plt.ylabel("Distance")

plt.tight_layout()
plt.show()

10. Complete Agglomerative Clustering Code

Single Google Colab Code

# ============================================
# COMPLETE HIERARCHICAL CLUSTERING
# ============================================

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

from sklearn.cluster import AgglomerativeClustering
from scipy.cluster.hierarchy import dendrogram, linkage

# Dataset
data = pd.DataFrame({
    "Height": [165,170,168,178,166,180],
    "Weight": [55,56,60,70,55,72]
})

print("Original Dataset")
print(data)


# Features
X = data[["Height","Weight"]]


# --------------------------------------------
# Dendrogram
# --------------------------------------------

Z = linkage(
    X,
    method="single",
    metric="euclidean"
)

plt.figure(figsize=(10,6))

dendrogram(
    Z,
    labels=["P0","P1","P2","P3","P4","P5"]
)

plt.title("Hierarchical Clustering Dendrogram")
plt.xlabel("Data Points")
plt.ylabel("Distance")

plt.show()


# --------------------------------------------
# Agglomerative Model
# --------------------------------------------

hc = AgglomerativeClustering(
    n_clusters=2,
    metric="euclidean",
    linkage="single"
)

data["Cluster"] = hc.fit_predict(X)


# --------------------------------------------
# Display Result
# --------------------------------------------

print("\nFinal Cluster Result")
print(data)


# --------------------------------------------
# Visualization
# --------------------------------------------

plt.figure(figsize=(8,6))

for cluster in sorted(data["Cluster"].unique()):

    subset = data[data["Cluster"] == cluster]

    plt.scatter(
        subset["Height"],
        subset["Weight"],
        s=150,
        label="Cluster " + str(cluster)
    )

    for index, row in subset.iterrows():

        plt.annotate(
            "P" + str(index),
            (row["Height"], row["Weight"]),
            xytext=(5,5),
            textcoords="offset points"
        )

plt.xlabel("Height")
plt.ylabel("Weight")

plt.title(
    "Agglomerative Hierarchical Clustering"
)

plt.legend()
plt.grid(True)

plt.show()

11. Advantages

No Need to Start With K

The hierarchy can be examined before deciding how many clusters to retain.

Dendrogram

The dendrogram provides a visual representation of the merging process.

Different Linkages

Single, complete, average and Ward linkage can be used depending on the data and objective.

Hierarchical Structure

The algorithm provides information about clusters at multiple levels.

12. Limitations

  • Computational cost can become high for very large datasets.
  • The result can depend strongly on the selected linkage method.
  • Once clusters are merged, the basic agglomerative process does not normally undo that merge.
  • Different distance measures can produce different clustering structures.

13. Summary

Hierarchical Clustering builds a hierarchy of clusters. In the Agglomerative approach, every observation starts as its own cluster and the closest clusters are repeatedly merged. A dendrogram records this merging process and can be used to select a desired number of clusters.

The reference article demonstrates this process using Euclidean distance and single linkage, including the formation of the distance matrix, successive merging of clusters, dendrogram construction and selection of two clusters.

📊 Clustering Algorithms Notebook

Four Major Clustering Techniques

K-Means • K-Medoids • DBSCAN • Hierarchical Clustering

Common Dataset Used by All Algorithms

Dataset: The same six two-dimensional points are used for all four clustering algorithms so that their results can be compared fairly.
import numpy as np

X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

print(X)
Dataset:
[[ 1  2]
 [ 2  2]
 [ 2  3]
 [ 8  7]
 [ 8  8]
 [25 80]]

1. K-Means Clustering

Concept: K-Means divides the dataset into a predefined number of clusters. It calculates the distance between data points and cluster centroids and repeatedly updates the centroids until the clusters become stable.
Important Parameter: n_clusters=2
This tells K-Means to create two clusters.

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create K-Means model
kmeans = KMeans(
    n_clusters=2,
    random_state=42,
    n_init=10
)

# 3. Fit model
kmeans.fit(X)

# 4. Get labels
labels = kmeans.labels_

print("K-Means Labels:", labels)

# 5. Get centroids
print("Centroids:")
print(kmeans.cluster_centers_)

# 6. Plot
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.scatter(
    kmeans.cluster_centers_[:,0],
    kmeans.cluster_centers_[:,1],
    marker='X',
    s=250,
    c='red',
    label='Centroids'
)

plt.title("K-Means Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.legend()
plt.show()

# 7. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)





2. K-Medoids Clustering

Concept: K-Medoids is similar to K-Means, but instead of calculating the average position of a cluster, it selects an actual data point as the representative center called a medoid. Unlike K-Means, the center of a cluster must be one of the existing data points.
Important Parameters:
n_clusters=2 → Number of clusters
random_state=42 → Makes the result reproducible

Python / Google Colab Code

!pip install scikit-learn-extra -q

import numpy as np
import matplotlib.pyplot as plt

from sklearn_extra.cluster import KMedoids
from sklearn.metrics import silhouette_score

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create K-Medoids model
kmedoids = KMedoids(
    n_clusters=2,
    random_state=42
)

# 3. Fit model
kmedoids.fit(X)

# 4. Get labels
labels = kmedoids.labels_

print("K-Medoids Labels:", labels)

# 5. Get medoids
print("Medoids:")
print(kmedoids.cluster_centers_)

# 6. Plot
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.scatter(
    kmedoids.cluster_centers_[:,0],
    kmedoids.cluster_centers_[:,1],
    marker='X',
    s=250,
    c='red',
    label='Medoids'
)

plt.title("K-Medoids Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.legend()
plt.show()

# 7. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)
Important Difference: K-Means uses calculated centroids, while K-Medoids uses actual data points as cluster centers. Therefore, K-Medoids is generally less sensitive to extreme values than K-Means.








3. DBSCAN Clustering

Concept: DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise. Instead of specifying the number of clusters, DBSCAN groups points according to their density. It can also identify isolated points as noise or anomalies. In this dataset, the point [25, 80] is far away from the other observations and can be identified as noise.
Important Parameters:
eps=3 → Maximum neighborhood distance
min_samples=2 → Minimum number of points required to form a dense region
label=-1 → Noise / anomaly

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import DBSCAN
from sklearn.metrics import silhouette_score

# 1. Our dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Setup and fit DBSCAN
dbscan = DBSCAN(
    eps=3,
    min_samples=2
)

dbscan.fit(X)

# 3. Get cluster labels
# -1 means noise/anomaly
labels = dbscan.labels_

print("Assigned Labels:", labels)

# 4. Plot results
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.colorbar(
    scatter,
    label='Cluster ID (-1 = Noise)'
)

plt.title("DBSCAN Clustering Results")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)

plt.show()

# 5. Silhouette Score
# Remove noise before calculating score
mask = labels != -1

if len(set(labels[mask])) >= 2:
    score = silhouette_score(
        X[mask],
        labels[mask]
    )
    print("Silhouette Score:", score)
else:
    print("Silhouette Score cannot be calculated.")
Important: DBSCAN is different from K-Means and K-Medoids because we do not need to specify the number of clusters beforehand. DBSCAN can also identify noise points using the label -1.



4. Hierarchical Clustering

Concept: Hierarchical clustering creates a hierarchy of clusters. In agglomerative hierarchical clustering, every data point initially forms its own cluster. The closest clusters are then repeatedly merged until the required number of clusters is obtained. A dendrogram can be used to visualize this merging process.
Important Parameters:
n_clusters=2 → Number of final clusters
linkage='ward' → Method used to calculate cluster distance

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import AgglomerativeClustering
from sklearn.metrics import silhouette_score
from scipy.cluster.hierarchy import dendrogram, linkage

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create hierarchical model
hierarchical = AgglomerativeClustering(
    n_clusters=2,
    linkage='ward'
)

# 3. Fit and predict
labels = hierarchical.fit_predict(X)

print("Hierarchical Labels:", labels)

# 4. Plot clusters
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.title("Hierarchical Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)

plt.show()

# 5. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)

# 6. Create dendrogram
linked = linkage(X, method='ward')

plt.figure(figsize=(8,5))

dendrogram(linked)

plt.title("Hierarchical Clustering Dendrogram")
plt.xlabel("Data Points")
plt.ylabel("Distance")

plt.show()







Comparison of Four Clustering Techniques

Technique Number of Clusters Required? Cluster Center Noise Detection Main Idea
K-Means Yes Centroid No Distance from centroid
K-Medoids Yes Actual data point No Distance from medoid
DBSCAN No Density Yes Density-based grouping
Hierarchical Yes No fixed center No Repeated cluster merging

Dataset Observation

The dataset contains two groups of nearby observations and one far-away point:

[1,2] [2,2] [2,3]

[8,7] [8,8]

[25,80]  ← isolated point

This makes the dataset useful for demonstrating an important difference between the algorithms, particularly DBSCAN's ability to identify an isolated observation as noise.


DBSCAN Clustering Algorithm

DBSCAN Clustering Algorithm

DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise. It is an unsupervised machine learning clustering algorithm that groups closely packed data points and identifies isolated points as noise.

1. What is DBSCAN?

DBSCAN creates clusters by looking at the density of data points. Instead of asking how many clusters should be created, DBSCAN searches for regions where points are sufficiently close together.

```

One important advantage is that DBSCAN can discover non-spherical and arbitrary-shaped clusters. It can also identify observations that do not belong to any dense region as noise or outliers.

Unlike K-Means, DBSCAN does not require the number of clusters to be specified beforehand.

```

2. Important DBSCAN Concepts

Core Point

A point is a core point when it has at least the required number of neighboring points within the specified epsilon distance.

Border Point

A border point is close enough to a core point to belong to its cluster, but it does not itself contain enough neighboring points to satisfy the minimum-density requirement.

Noise Point

A noise point is neither a core point nor a border point. DBSCAN can therefore leave this observation outside the discovered clusters.

3. Main DBSCAN Parameters

```
eps Maximum distance used to determine whether two points are neighbors.
min_samples Minimum number of samples required in a neighborhood for a point to be considered a core point.
metric Distance function used by DBSCAN. Euclidean distance is the default, but Manhattan, Cosine, Haversine and other metrics can be used.
algorithm Method used for nearest-neighbor search, such as auto, ball_tree, kd_tree or brute.
leaf_size Controls the leaf size used by tree-based nearest-neighbor searches.
p Power parameter for the Minkowski distance when applicable.
```

4. How DBSCAN Works

Step 1 — Select Parameters

Choose values for eps and min_samples.

Step 2 — Select an Unvisited Point

DBSCAN begins by selecting an unvisited observation.

Step 3 — Examine Its Neighborhood

All points within the epsilon distance are identified.

Step 4 — Identify a Core Point

If enough points exist within the neighborhood, the point becomes a core point and a new cluster can be started.

Step 5 — Expand the Cluster

Neighboring core points are recursively examined and their neighbors are added to the same cluster.

Step 6 — Identify Border Points

Points close to a core point but without enough neighbors themselves can become border points.

Step 7 — Identify Noise

Points that cannot be associated with a density-connected cluster remain classified as noise.

5. Density Reachability

Density reachability describes how one point can be reached from another through a chain of sufficiently dense points.

```

If point A is a core point and point B is within its epsilon neighborhood, B can be directly density-reachable from A. A chain of such relationships can connect points across an entire cluster.

```

6. Density Connectivity

Two points are density-connected when they can both be reached through density-based connections from a suitable core point.

```

Density connectivity is therefore one of the foundations used by DBSCAN to determine which observations belong to the same cluster.

```

7. Choosing the eps Parameter

Selecting an appropriate eps value is important. One systematic approach is the k-distance graph.

```
  1. Choose a value of k related to min_samples.
  2. Calculate each observation's distance to its k-th nearest neighbor.
  3. Sort these distances.
  4. Plot the sorted distances.
  5. Look for an elbow in the graph.
  6. Use the elbow as a possible starting point for eps.
```

8. Choosing min_samples

A commonly suggested starting point is:

```
min_samples = 2 × number_of_features

This is only a starting guideline. The appropriate value depends on the dataset, its dimensionality, noise level and clustering objective.

```

9. Distance Metrics

Metric Typical Use
Euclidean General numerical data and geometric distance.
Manhattan Useful for grid-like or coordinate-based distances.
Cosine Useful for high-dimensional vectors such as text representations.
Haversine Useful when working with latitude and longitude coordinates.

10. Why Feature Scaling Matters

Important: DBSCAN depends directly on distance. If one feature has a much larger numerical range than another, it can dominate the distance calculation.

Two commonly used scaling techniques are:

```
  • StandardScaler — standardizes features around zero.
  • MinMaxScaler — scales features into a specified range, commonly 0 to 1.
```

11. Python Implementation of DBSCAN

The following example follows the workflow demonstrated in the DataCamp tutorial using a two-moon dataset. The example uses eps=0.15 and min_samples=5.

import numpy as np
import matplotlib.pyplot as plt

from sklearn.datasets import make_moons
from sklearn.cluster import DBSCAN
from sklearn.neighbors import NearestNeighbors

# -----------------------------------
# 1. Create the dataset
# -----------------------------------

X, y = make_moons(
    n_samples=200,
    noise=0.05,
    random_state=42
)

# -----------------------------------
# 2. Visualize original data
# -----------------------------------

plt.figure(figsize=(8, 5))

plt.scatter(
    X[:, 0],
    X[:, 1]
)

plt.title("Original Moon Dataset")
plt.xlabel("Feature 1")
plt.ylabel("Feature 2")
plt.show()

# -----------------------------------
# 3. K-distance graph
# -----------------------------------

k = 5

neigh = NearestNeighbors(n_neighbors=k)

neigh.fit(X)

distances, indices = neigh.kneighbors(X)

distances = np.sort(distances[:, k-1])

plt.figure(figsize=(8, 5))

plt.plot(distances)

plt.xlabel("Points")
plt.ylabel("5th Nearest Neighbor Distance")
plt.title("K-Distance Graph")

plt.show()

# -----------------------------------
# 4. DBSCAN
# -----------------------------------

epsilon = 0.15
min_samples = 5

dbscan = DBSCAN(
    eps=epsilon,
    min_samples=min_samples
)

clusters = dbscan.fit_predict(X)

# -----------------------------------
# 5. Visualize clusters
# -----------------------------------

plt.figure(figsize=(8, 5))

plt.scatter(
    X[:, 0],
    X[:, 1],
    c=clusters
)

plt.title("DBSCAN Clustering")
plt.xlabel("Feature 1")
plt.ylabel("Feature 2")

plt.show()

# -----------------------------------
# 6. Count clusters
# -----------------------------------

n_clusters = len(
    set(clusters)
) - (1 if -1 in clusters else 0)

n_noise = list(clusters).count(-1)

print("Number of clusters:", n_clusters)
print("Number of noise points:", n_noise)
In scikit-learn, DBSCAN represents noise points using the label -1. The DataCamp example reports two clusters and five noise points for its demonstrated configuration.

12. DBSCAN vs Other Clustering Techniques

Feature K-Means K-Medoids Hierarchical DBSCAN
Type Centroid-based Medoid-based Hierarchy-based Density-based
Number of clusters Specify K Specify K Can choose using dendrogram/cut Not required beforehand
Cluster shape Generally convex/spherical Generally compact clusters Depends on linkage and distance Can discover arbitrary shapes
Noise detection No explicit noise class No explicit noise class No explicit DBSCAN-style noise class Yes
Outlier handling Every point assigned Every point assigned Every point placed in hierarchy Can label points as noise
Non-linear clusters Usually difficult Usually difficult Can handle many structures Strong capability
Main parameters K K Linkage/distance eps, min_samples

The DataCamp comparison particularly highlights the differences between DBSCAN and K-Means in cluster shape, predefined cluster count, noise handling, scalability and parameter sensitivity.

13. When DBSCAN is Useful

  • When the number of clusters is unknown.
  • When clusters are not spherical.
  • When the dataset contains possible outliers.
  • When dense regions are more meaningful than distance from a centroid.
  • When arbitrary-shaped clusters need to be discovered.

14. Limitations of DBSCAN

  • Results can be sensitive to eps and min_samples.
  • Very different cluster densities can be difficult for standard DBSCAN.
  • Distance becomes less intuitive in very high-dimensional data.
  • Large datasets can require substantial computational resources.
  • Feature scaling is important when features have different ranges.
```

Alternatives such as OPTICS and HDBSCAN can be considered for datasets with challenging or varying density structures.

```

15. Single Python Program — K-Means vs K-Medoids vs Hierarchical vs DBSCAN

The following program runs four clustering techniques on the same two-moon dataset so that their behavior can be compared using the same input data.

# ==========================================================
# COMPARISON OF CLUSTERING TECHNIQUES
# K-MEANS
# K-MEDOIDS
# HIERARCHICAL CLUSTERING
# DBSCAN
# ==========================================================

import numpy as np
import matplotlib.pyplot as plt

from sklearn.datasets import make_moons
from sklearn.preprocessing import StandardScaler

from sklearn.cluster import (
    KMeans,
    AgglomerativeClustering,
    DBSCAN
)

from sklearn_extra.cluster import KMedoids


# ----------------------------------------------------------
# 1. Generate Dataset
# ----------------------------------------------------------

X, y = make_moons(
    n_samples=200,
    noise=0.05,
    random_state=42
)


# ----------------------------------------------------------
# 2. Feature Scaling
# ----------------------------------------------------------

scaler = StandardScaler()

X_scaled = scaler.fit_transform(X)


# ----------------------------------------------------------
# 3. K-MEANS
# ----------------------------------------------------------

kmeans = KMeans(
    n_clusters=2,
    random_state=42,
    n_init=10
)

kmeans_labels = kmeans.fit_predict(X_scaled)


# ----------------------------------------------------------
# 4. K-MEDOIDS
# ----------------------------------------------------------

kmedoids = KMedoids(
    n_clusters=2,
    random_state=42
)

kmedoids_labels = kmedoids.fit_predict(X_scaled)


# ----------------------------------------------------------
# 5. HIERARCHICAL CLUSTERING
# ----------------------------------------------------------

hierarchical = AgglomerativeClustering(
    n_clusters=2,
    linkage="ward"
)

hierarchical_labels = hierarchical.fit_predict(X_scaled)


# ----------------------------------------------------------
# 6. DBSCAN
# ----------------------------------------------------------

dbscan = DBSCAN(
    eps=0.3,
    min_samples=5
)

dbscan_labels = dbscan.fit_predict(X_scaled)


# ----------------------------------------------------------
# 7. Count DBSCAN Clusters and Noise
# ----------------------------------------------------------

dbscan_clusters = len(
    set(dbscan_labels)
) - (1 if -1 in dbscan_labels else 0)

dbscan_noise = list(
    dbscan_labels
).count(-1)


# ----------------------------------------------------------
# 8. Display Results
# ----------------------------------------------------------

print("====================================")
print("CLUSTERING COMPARISON")
print("====================================")

print("K-Means clusters       : 2")
print("K-Medoids clusters     : 2")
print("Hierarchical clusters  : 2")

print("DBSCAN clusters        :", dbscan_clusters)
print("DBSCAN noise points    :", dbscan_noise)


# ----------------------------------------------------------
# 9. Visualization
# ----------------------------------------------------------

fig, axes = plt.subplots(
    2,
    2,
    figsize=(14, 10)
)


# K-Means
axes[0, 0].scatter(
    X_scaled[:, 0],
    X_scaled[:, 1],
    c=kmeans_labels
)

axes[0, 0].set_title("K-Means")


# K-Medoids
axes[0, 1].scatter(
    X_scaled[:, 0],
    X_scaled[:, 1],
    c=kmedoids_labels
)

axes[0, 1].set_title("K-Medoids")


# Hierarchical
axes[1, 0].scatter(
    X_scaled[:, 0],
    X_scaled[:, 1],
    c=hierarchical_labels
)

axes[1, 0].set_title("Hierarchical Clustering")


# DBSCAN
axes[1, 1].scatter(
    X_scaled[:, 0],
    X_scaled[:, 1],
    c=dbscan_labels
)

axes[1, 1].set_title("DBSCAN")


for ax in axes.flat:
    ax.set_xlabel("Feature 1")
    ax.set_ylabel("Feature 2")

plt.tight_layout()

plt.show()

16. Important Installation for K-Medoids

pip install scikit-learn-extra
Note: K-Medoids is commonly imported from sklearn_extra.cluster, while K-Means, DBSCAN and Agglomerative Clustering are available through scikit-learn.

17. Overall Workflow

Dataset → Preprocessing → Scaling → Choose Algorithm → Clustering → Visualization → Interpretation

For DBSCAN specifically:

```
Dataset
```

↓
Feature Selection
↓
Feature Scaling
↓
Choose min_samples
↓
Create K-Distance Graph
↓
Select eps
↓
Apply DBSCAN
↓
Identify Core / Border / Noise
↓
Visualize Clusters
↓
Interpret Results

18. Quick Summary

Algorithm Main Idea Needs K? Noise Detection
K-Means Groups observations around centroids. Yes No
K-Medoids Groups observations around representative medoid points. Yes No explicit noise class
Hierarchical Builds a hierarchy of nested clusters. Can be selected by cutting hierarchy No explicit DBSCAN-style noise class
DBSCAN Groups points according to density. No Yes

kmedoid

K-Medoids Clustering

Theory, Python Implementation and Comparison with K-Means

1. What is K-Medoids?

K-Medoids is an unsupervised machine learning clustering algorithm. It divides a dataset into a predefined number of groups called clusters.

The main idea of K-Medoids is to select an actual observation from the dataset as the representative point of each cluster. This representative point is called a medoid.

Important: A medoid must always be an actual data point from the dataset.

How K-Medoids Works

  1. Choose the number of clusters, K.
  2. Select K initial medoids.
  3. Calculate the distance between every data point and every medoid.
  4. Assign each data point to its nearest medoid.
  5. Calculate the total clustering cost.
  6. Try replacing a medoid with another data point.
  7. If the replacement reduces the cost, keep the replacement.
  8. Repeat until no useful replacement improves the clustering.

Manhattan Distance

In the example below, Manhattan distance is used.

Distance = |X₂ − X₁| + |Y₂ − Y₁|

Example

Suppose the medoid is (9,3) and the point is (8,4).

Distance = |8 − 9| + |4 − 3| = 1 + 1 = 2

Therefore, the distance between the two points is 2.

2. Dataset Used in the Example

The following 10-point dataset is used for both K-Means and K-Medoids comparison.

Point X Y
P193
P284
P346
P485
P525
P638
P758
P844
P9104
P1096
Number of clusters K = 2
Initial Medoid 1 = (4,6)
Initial Medoid 2 = (9,3)

3. Python Implementation of K-Medoids

The following code implements K-Medoids using the same dataset. It uses Manhattan distance and creates two clusters.

# ==========================================
# K-MEDOIDS CLUSTERING
# ==========================================

# Install library in Google Colab
!pip install scikit-learn-extra -q

# Import libraries
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

from sklearn_extra.cluster import KMedoids


# Dataset
data = np.array([
    [9, 3],
    [8, 4],
    [4, 6],
    [8, 5],
    [2, 5],
    [3, 8],
    [5, 8],
    [4, 4],
    [10, 4],
    [9, 6]
])


# Create DataFrame
df = pd.DataFrame(data, columns=["X", "Y"])

print("Dataset:")
print(df)


# Number of clusters
k = 2


# Create K-Medoids model
model = KMedoids(
    n_clusters=k,
    metric="manhattan",
    random_state=42
)


# Train model
model.fit(data)


# Cluster labels
labels = model.labels_

print("\nCluster Labels:")
print(labels)


# Get medoids
medoid_indices = model.medoid_indices_
medoids = data[medoid_indices]

print("\nMedoids:")

for i, medoid in enumerate(medoids):
    print(
        "Medoid", i + 1,
        "=", tuple(medoid)
    )


# Display clusters

for cluster in range(k):

    print("\nCluster", cluster + 1)

    cluster_points = data[labels == cluster]

    print(cluster_points)


# Total clustering cost
print("\nTotal Cost:")
print(model.inertia_)


# Plot clusters

plt.figure(figsize=(9, 7))

for cluster in range(k):

    points = data[labels == cluster]

    plt.scatter(
        points[:, 0],
        points[:, 1],
        s=100,
        label=f"Cluster {cluster + 1}"
    )


# Plot medoids
plt.scatter(
    medoids[:, 0],
    medoids[:, 1],
    s=250,
    marker="X",
    label="Medoids"
)


# Point labels
for i, point in enumerate(data):

    plt.annotate(
        f"P{i + 1}",
        (point[0], point[1]),
        xytext=(5, 5),
        textcoords="offset points"
    )


plt.xlabel("X")
plt.ylabel("Y")
plt.title("K-Medoids Clustering")

plt.legend()
plt.grid(True)

plt.show()

4. K-Means vs K-Medoids

Both algorithms are clustering techniques, but they calculate their cluster representatives differently.

Feature K-Means K-Medoids
Type Unsupervised Unsupervised
Representative Centroid Medoid
Representative is actual data point? No Yes
Center calculation Mean Actual observation
Distance in this example Euclidean Manhattan
Effect of outliers More sensitive Generally more robust
Speed Usually faster Usually more computationally expensive
Center can be a new point? Yes No
K-Means:
Calculates the average of points to create a centroid.

K-Medoids:
Selects an existing observation as the medoid.

5. Result Comparison for This Dataset

Cluster K-Means K-Medoids
Cluster 1 P3, P5, P6, P7, P8 P3, P5, P6, P7, P8
Cluster 2 P1, P2, P4, P9, P10 P1, P2, P4, P9, P10
Cluster 1 Representative (3.6, 6.2) (4, 6)
Cluster 2 Representative (8.8, 4.4) (9, 3)
Important observation: For this dataset and initialization, both methods produce the same two groups, but their representatives are different.

K-Means produces the new centroids:

C1 = (3.6, 6.2)     C2 = (8.8, 4.4)

K-Medoids uses actual observations:

M1 = (4, 6)     M2 = (9, 3)

6. Single Python Program: K-Means + K-Medoids

The following program performs both algorithms on the same dataset and displays their results side-by-side.

# ==========================================================
# K-MEANS AND K-MEDOIDS ON THE SAME DATASET
# ==========================================================

# Install K-Medoids library in Google Colab
!pip install scikit-learn-extra -q


# ----------------------------------------------------------
# IMPORT LIBRARIES
# ----------------------------------------------------------

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

from sklearn.cluster import KMeans
from sklearn_extra.cluster import KMedoids


# ----------------------------------------------------------
# DATASET
# ----------------------------------------------------------

data = np.array([
    [9, 3],
    [8, 4],
    [4, 6],
    [8, 5],
    [2, 5],
    [3, 8],
    [5, 8],
    [4, 4],
    [10, 4],
    [9, 6]
])


# ----------------------------------------------------------
# DISPLAY DATASET
# ----------------------------------------------------------

df = pd.DataFrame(
    data,
    columns=["X", "Y"]
)

print("ORIGINAL DATASET")
print(df)


# ----------------------------------------------------------
# NUMBER OF CLUSTERS
# ----------------------------------------------------------

K = 2


# ==========================================================
# K-MEANS
# ==========================================================

kmeans = KMeans(
    n_clusters=K,
    init=np.array([
        [4, 6],
        [9, 3]
    ]),
    n_init=1,
    random_state=42
)

kmeans.fit(data)


# K-Means labels
kmeans_labels = kmeans.labels_


# K-Means centroids
kmeans_centroids = kmeans.cluster_centers_


# K-Means cost
kmeans_cost = kmeans.inertia_


print("\n")
print("=" * 50)
print("K-MEANS RESULTS")
print("=" * 50)

print("\nCluster Labels:")
print(kmeans_labels)

print("\nCentroids:")

for i, centroid in enumerate(kmeans_centroids):

    print(
        "Centroid", i + 1,
        "=", tuple(np.round(centroid, 2))
    )

print("\nK-Means SSE:")
print(round(kmeans_cost, 2))


# Display K-Means clusters

for cluster in range(K):

    print("\nK-Means Cluster", cluster + 1)

    points = data[kmeans_labels == cluster]

    print(points)


# ==========================================================
# K-MEDOIDS
# ==========================================================

kmedoids = KMedoids(
    n_clusters=K,
    metric="manhattan",
    init=np.array([
        [4, 6],
        [9, 3]
    ]),
    random_state=42
)

kmedoids.fit(data)


# K-Medoids labels
kmedoids_labels = kmedoids.labels_


# K-Medoids medoid indices
medoid_indices = kmedoids.medoid_indices_


# K-Medoids medoids
kmedoids_medoids = data[medoid_indices]


# K-Medoids cost
kmedoids_cost = kmedoids.inertia_


print("\n")
print("=" * 50)
print("K-MEDOIDS RESULTS")
print("=" * 50)

print("\nCluster Labels:")
print(kmedoids_labels)

print("\nMedoids:")

for i, medoid in enumerate(kmedoids_medoids):

    print(
        "Medoid", i + 1,
        "=", tuple(medoid)
    )

print("\nK-Medoids Cost:")
print(kmedoids_cost)


# Display K-Medoids clusters

for cluster in range(K):

    print("\nK-Medoids Cluster", cluster + 1)

    points = data[kmedoids_labels == cluster]

    print(points)


# ==========================================================
# SIDE-BY-SIDE COMPARISON
# ==========================================================

print("\n")
print("=" * 70)
print("K-MEANS vs K-MEDOIDS")
print("=" * 70)

print("\nK-MEANS")

for cluster in range(K):

    points = data[kmeans_labels == cluster]

    print(
        "Cluster",
        cluster + 1,
        ":",
        [tuple(p) for p in points]
    )


print("\nK-MEDOIDS")

for cluster in range(K):

    points = data[kmedoids_labels == cluster]

    print(
        "Cluster",
        cluster + 1,
        ":",
        [tuple(p) for p in points]
    )


print("\nK-Means Centroids:")

for centroid in kmeans_centroids:

    print(tuple(np.round(centroid, 2)))


print("\nK-Medoids Medoids:")

for medoid in kmedoids_medoids:

    print(tuple(medoid))


# ==========================================================
# PLOT K-MEANS
# ==========================================================

plt.figure(figsize=(8, 6))

for cluster in range(K):

    points = data[kmeans_labels == cluster]

    plt.scatter(
        points[:, 0],
        points[:, 1],
        s=100,
        label=f"Cluster {cluster + 1}"
    )


plt.scatter(
    kmeans_centroids[:, 0],
    kmeans_centroids[:, 1],
    s=250,
    marker="X",
    label="Centroids"
)


for i, point in enumerate(data):

    plt.annotate(
        f"P{i + 1}",
        (point[0], point[1]),
        xytext=(5, 5),
        textcoords="offset points"
    )


plt.xlabel("X")
plt.ylabel("Y")
plt.title("K-Means Clustering")
plt.legend()
plt.grid(True)

plt.show()


# ==========================================================
# PLOT K-MEDOIDS
# ==========================================================

plt.figure(figsize=(8, 6))

for cluster in range(K):

    points = data[kmedoids_labels == cluster]

    plt.scatter(
        points[:, 0],
        points[:, 1],
        s=100,
        label=f"Cluster {cluster + 1}"
    )


plt.scatter(
    kmedoids_medoids[:, 0],
    kmedoids_medoids[:, 1],
    s=250,
    marker="X",
    label="Medoids"
)


for i, point in enumerate(data):

    plt.annotate(
        f"P{i + 1}",
        (point[0], point[1]),
        xytext=(5, 5),
        textcoords="offset points"
    )


plt.xlabel("X")
plt.ylabel("Y")
plt.title("K-Medoids Clustering")
plt.legend()
plt.grid(True)

plt.show()

7. Conclusion

K-Means and K-Medoids are both unsupervised clustering algorithms used to divide observations into groups.

  • K-Means calculates a centroid using the mean of the observations in a cluster.
  • K-Medoids selects an actual observation as the representative of a cluster.
  • In this example, both methods produce the same two groups.
  • However, K-Means produces new center points, whereas K-Medoids uses existing observations.
  • K-Medoids can be useful when the use of an actual observation as the cluster representative is important.
Easy way to remember:

K-Means → Mean = Centroid
K-Medoids → Existing Data Point = Medoid

Tuesday, September 29, 2026

ttt

Online MCQ Test

10 Question MCQ Test

1. Which organelle is called the powerhouse of the cell?
2. What is the SI unit of electric current?
3. Which gas is most abundant in Earth's atmosphere?
4. What is 15 × 6?
5. Who wrote Julius Caesar?
6. Which blood cells help fight infections?
7. What is the chemical symbol for sodium?
8. Which planet is known as the Red Planet?
9. What process do green plants use to make food?
10. Which of these is a prime number?

Monday, September 28, 2026

KNN

 # ============================================================

# BOSTON HOUSING DATASET

# K-NEAREST NEIGHBORS (KNN)

# COMPLETE DATA PREPROCESSING + EVALUATION

# GOOGLE COLAB VERSION

# ============================================================



# ============================================================

# 1. IMPORT LIBRARIES

# ============================================================


import numpy as np

import pandas as pd


import matplotlib.pyplot as plt

import seaborn as sns


from sklearn.datasets import fetch_openml


from sklearn.model_selection import (

    train_test_split,

    cross_val_score

)


from sklearn.compose import ColumnTransformer


from sklearn.pipeline import Pipeline


from sklearn.impute import SimpleImputer


from sklearn.preprocessing import (

    StandardScaler,

    OneHotEncoder

)


from sklearn.neighbors import KNeighborsClassifier


from sklearn.metrics import (

    accuracy_score,

    precision_score,

    recall_score,

    f1_score,

    confusion_matrix,

    classification_report,

    ConfusionMatrixDisplay,

    roc_curve,

    roc_auc_score

)


import warnings


warnings.filterwarnings("ignore")



# ============================================================

# 2. LOAD BOSTON HOUSING DATASET

# ============================================================


boston = fetch_openml(

    name="boston",

    version=1,

    as_frame=True

)


df = boston.frame.copy()


print("Boston Housing Dataset Loaded Successfully!")



# ============================================================

# 3. DISPLAY FIRST 5 ROWS

# ============================================================


print("\nFirst 5 Rows:")


display(

    df.head()

)



# ============================================================

# 4. DATASET SHAPE

# ============================================================


print("\nDataset Shape:")


print(

    df.shape

)



# ============================================================

# 5. COLUMN NAMES

# ============================================================


print("\nColumn Names:")


print(

    df.columns.tolist()

)



# ============================================================

# 6. DATA INFORMATION

# ============================================================


print("\nDataset Information:")


df.info()



# ============================================================

# 7. DATA TYPES

# ============================================================


print("\nData Types:")


print(

    df.dtypes

)



# ============================================================

# 8. MISSING VALUES

# ============================================================


print("\nMissing Values:")


print(

    df.isnull().sum()

)



# ============================================================

# 9. MISSING VALUE PERCENTAGE

# ============================================================


missing_percentage = (


    df.isnull()

      .mean()

      .mul(100)

      .sort_values(

          ascending=False

      )


)


print("\nMissing Value Percentage:")


print(

    missing_percentage

)



# ============================================================

# 10. DUPLICATE VALUES

# ============================================================


print(

    "\nNumber of Duplicate Rows:",

    df.duplicated().sum()

)



# ============================================================

# 11. REMOVE DUPLICATES

# ============================================================


df = df.drop_duplicates()


print(

    "\nShape After Duplicate Removal:"

)


print(

    df.shape

)



# ============================================================

# 12. STATISTICAL SUMMARY

# ============================================================


print(

    "\nStatistical Summary:"

)


display(

    df.describe()

)



# ============================================================

# 13. TARGET VARIABLE

# ============================================================


target = "MEDV"


print(

    "\nTarget Variable:",

    target

)



# ============================================================

# 14. TARGET DISTRIBUTION

# ============================================================


plt.figure(

    figsize=(9,5)

)


sns.histplot(

    df[target],

    kde=True

)


plt.title(

    "Distribution of Boston House Prices"

)


plt.xlabel(

    "MEDV"

)


plt.ylabel(

    "Frequency"

)


plt.show()



# ============================================================

# 15. OUTLIER DETECTION

# ============================================================


numeric_columns = (


    df.select_dtypes(

        include=np.number

    ).columns


)



plt.figure(

    figsize=(14,6)

)


sns.boxplot(

    data=df[

        numeric_columns

    ]

)


plt.title(

    "Boxplot of Boston Housing Features"

)


plt.xticks(

    rotation=45

)


plt.show()



# ============================================================

# 16. IQR OUTLIER ANALYSIS

# ============================================================


print(

    "\n============================================"

)


print(

    "          IQR OUTLIER ANALYSIS"

)


print(

    "============================================"

)



for column in numeric_columns:


    Q1 = df[column].quantile(

        0.25

    )


    Q3 = df[column].quantile(

        0.75

    )


    IQR = Q3 - Q1


    lower_bound = (

        Q1 - 1.5 * IQR

    )


    upper_bound = (

        Q3 + 1.5 * IQR

    )


    outliers = df[

        (df[column] < lower_bound)

        |

        (df[column] > upper_bound)

    ]


    print(

        column,

        "→",

        len(outliers),

        "outliers"

    )



# ============================================================

# 17. OUTLIER CAPPING

# ============================================================


df_capped = df.copy()


for column in numeric_columns:


    Q1 = df_capped[column].quantile(

        0.25

    )


    Q3 = df_capped[column].quantile(

        0.75

    )


    IQR = Q3 - Q1


    lower_bound = (

        Q1 - 1.5 * IQR

    )


    upper_bound = (

        Q3 + 1.5 * IQR

    )


    df_capped[column] = (


        df_capped[column]

        .clip(

            lower_bound,

            upper_bound

        )


    )



print(

    "\nIQR Outlier Capping Completed!"

)



# ============================================================

# 18. CORRELATION MATRIX

# ============================================================


correlation = df.corr(

    numeric_only=True

)


print(

    "\nCorrelation with MEDV:"

)


print(

    correlation[target]

    .sort_values(

        ascending=False

    )

)



# ============================================================

# 19. CORRELATION HEATMAP

# ============================================================


plt.figure(

    figsize=(12,9)

)


sns.heatmap(

    correlation,

    annot=True,

    fmt=".2f",

    cmap="coolwarm"

)


plt.title(

    "Boston Housing Correlation Matrix"

)


plt.show()



# ============================================================

# 20. CONVERT MEDV INTO CLASS

# ============================================================

#

# Original MEDV is continuous.

#

# KNN classification requires classes.

#

# 0 = Low Price

# 1 = High Price

#

# Median MEDV is used as threshold.

# ============================================================


median_price = df[target].median()


print(

    "\nMedian House Price:",

    median_price

)



df["Price_Class"] = (


    df[target] >= median_price


).astype(int)



# ============================================================

# 21. DISPLAY CLASS DISTRIBUTION

# ============================================================


print(

    "\nPrice Class Distribution:"

)


print(

    df["Price_Class"].value_counts()

)



# ============================================================

# 22. CLASS DISTRIBUTION GRAPH

# ============================================================


plt.figure(

    figsize=(7,5)

)


sns.countplot(

    x="Price_Class",

    data=df

)


plt.title(

    "Low Price vs High Price Houses"

)


plt.xlabel(

    "Price Class"

)


plt.ylabel(

    "Number of Houses"

)


plt.xticks(

    [0,1],

    [

        "Low Price",

        "High Price"

    ]

)


plt.show()



# ============================================================

# 23. FEATURES AND TARGET

# ============================================================


X = df.drop(

    columns=[

        "MEDV",

        "Price_Class"

    ]

)


y = df["Price_Class"]



print(

    "\nFeature Shape:",

    X.shape

)


print(

    "Target Shape:",

    y.shape

)



# ============================================================

# 24. IDENTIFY NUMERICAL FEATURES

# ============================================================


numeric_features = (


    X.select_dtypes(

        include=np.number

    )

    .columns

    .tolist()


)



categorical_features = (


    X.select_dtypes(

        exclude=np.number

    )

    .columns

    .tolist()


)



print(

    "\nNumerical Features:"

)


print(

    numeric_features

)


print(

    "\nCategorical Features:"

)


print(

    categorical_features

)



# ============================================================

# 25. TRAIN TEST SPLIT

# ============================================================


X_train, X_test, y_train, y_test = (


    train_test_split(


        X,

        y,


        test_size=0.20,


        random_state=42,


        stratify=y


    )


)



print(

    "\nTraining Shape:",

    X_train.shape

)


print(

    "Testing Shape:",

    X_test.shape

)



# ============================================================

# 26. NUMERICAL PREPROCESSING

# ============================================================


numeric_pipeline = Pipeline(


    steps=[


        (

            "imputer",


            SimpleImputer(

                strategy="median"

            )

        ),


        (

            "scaler",


            StandardScaler()

        )


    ]


)



# ============================================================

# 27. CATEGORICAL PREPROCESSING

# ============================================================


categorical_pipeline = Pipeline(


    steps=[


        (

            "imputer",


            SimpleImputer(

                strategy="most_frequent"

            )

        ),


        (

            "encoder",


            OneHotEncoder(

                handle_unknown="ignore"

            )

        )


    ]


)



# ============================================================

# 28. COMBINE PREPROCESSING

# ============================================================


preprocessor = ColumnTransformer(


    transformers=[


        (

            "numeric",


            numeric_pipeline,


            numeric_features

        ),


        (

            "categorical",


            categorical_pipeline,


            categorical_features

        )


    ]


)



# ============================================================

# 29. KNN MODEL

# ============================================================


knn_model = Pipeline(


    steps=[


        (

            "preprocessing",


            preprocessor

        ),


        (

            "model",


            KNeighborsClassifier(

                n_neighbors=5

            )

        )


    ]


)



# ============================================================

# 30. TRAIN KNN MODEL

# ============================================================


knn_model.fit(


    X_train,


    y_train


)



print(

    "\nKNN Model Trained Successfully!"

)



# ============================================================

# 31. PREDICTION

# ============================================================


y_pred = knn_model.predict(

    X_test

)



# ============================================================

# 32. PREDICT PROBABILITY

# ============================================================


y_probability = (


    knn_model

    .predict_proba(

        X_test

    )[:, 1]


)



# ============================================================

# 33. ACCURACY

# ============================================================


accuracy = accuracy_score(


    y_test,


    y_pred


)



# ============================================================

# 34. PRECISION

# ============================================================


precision = precision_score(


    y_test,


    y_pred,


    zero_division=0


)



# ============================================================

# 35. RECALL

# ============================================================


recall = recall_score(


    y_test,


    y_pred,


    zero_division=0


)



# ============================================================

# 36. F1 SCORE

# ============================================================


f1 = f1_score(


    y_test,


    y_pred,


    zero_division=0


)



# ============================================================

# 37. ROC-AUC

# ============================================================


roc_auc = roc_auc_score(


    y_test,


    y_probability


)



# ============================================================

# 38. DISPLAY MODEL RESULTS

# ============================================================


print(

    "\n============================================"

)


print(

    "             KNN RESULTS"

)


print(

    "============================================"

)


print(

    "Accuracy  :",

    round(accuracy, 4)

)


print(

    "Precision :",

    round(precision, 4)

)


print(

    "Recall    :",

    round(recall, 4)

)


print(

    "F1 Score  :",

    round(f1, 4)

)


print(

    "ROC-AUC   :",

    round(roc_auc, 4)

)


print(

    "============================================"

)



# ============================================================

# 39. CONFUSION MATRIX

# ============================================================


cm = confusion_matrix(


    y_test,


    y_pred


)



print(

    "\nConfusion Matrix:"

)


print(

    cm

)



# ============================================================

# 40. TRUE NEGATIVE / FALSE POSITIVE /

#     FALSE NEGATIVE / TRUE POSITIVE

# ============================================================


TN, FP, FN, TP = cm.ravel()



print(

    "\nTrue Negative (TN):",

    TN

)


print(

    "False Positive (FP):",

    FP

)


print(

    "False Negative (FN):",

    FN

)


print(

    "True Positive (TP):",

    TP

)



# ============================================================

# 41. CONFUSION MATRIX GRAPH

# ============================================================


plt.figure(

    figsize=(7,6)

)


sns.heatmap(


    cm,


    annot=True,


    fmt="d",


    cmap="Blues",


    xticklabels=[

        "Predicted Low",

        "Predicted High"

    ],


    yticklabels=[

        "Actual Low",

        "Actual High"

    ]


)


plt.title(

    "KNN Confusion Matrix"

)


plt.xlabel(

    "Predicted Class"

)


plt.ylabel(

    "Actual Class"

)


plt.show()



# ============================================================

# 42. CLASSIFICATION REPORT

# ============================================================


print(

    "\n============================================"

)


print(

    "          CLASSIFICATION REPORT"

)


print(

    "============================================"

)


print(


    classification_report(


        y_test,


        y_pred,


        target_names=[

            "Low Price",

            "High Price"

        ],


        zero_division=0


    )


)



# ============================================================

# 43. PERFORMANCE TABLE

# ============================================================


metrics_table = pd.DataFrame({


    "Metric": [


        "Accuracy",


        "Precision",


        "Recall",


        "F1 Score",


        "ROC-AUC"


    ],


    "Score": [


        accuracy,


        precision,


        recall,


        f1,


        roc_auc


    ]


})



metrics_table["Score"] = (


    metrics_table["Score"]

    .round(4)


)



print(

    "\nKNN Performance Metrics:"

)


display(

    metrics_table

)



# ============================================================

# 44. CROSS VALIDATION

# ============================================================


cv_scores = cross_val_score(


    knn_model,


    X,


    y,


    cv=5,


    scoring="accuracy"


)



print(

    "\n5-Fold Cross Validation Accuracy:"

)


print(

    cv_scores

)



print(

    "\nMean Cross Validation Accuracy:"

)


print(

    round(

        cv_scores.mean(),

        4

    )

)



# ============================================================

# 45. TEST DIFFERENT K VALUES

# ============================================================


k_values = range(

    1,

    21

)


k_accuracy = []



for k in k_values:


    temp_knn = Pipeline(


        steps=[


            (

                "preprocessing",


                preprocessor

            ),


            (

                "model",


                KNeighborsClassifier(

                    n_neighbors=k

                )


            )


        ]


    )



    temp_knn.fit(

        X_train,

        y_train

    )



    temp_prediction = temp_knn.predict(

        X_test

    )



    score = accuracy_score(


        y_test,


        temp_prediction


    )



    k_accuracy.append(

        score

    )



# ============================================================

# 46. K VALUE GRAPH

# ============================================================


plt.figure(

    figsize=(10,6)

)


plt.plot(


    list(k_values),


    k_accuracy,


    marker="o"


)


plt.xlabel(

    "Number of Neighbors (K)"

)


plt.ylabel(

    "Accuracy"

)


plt.title(

    "KNN Accuracy for Different K Values"

)


plt.xticks(

    list(k_values)

)


plt.grid(

    True

)


plt.show()



# ============================================================

# 47. BEST K VALUE

# ============================================================


best_k_index = np.argmax(

    k_accuracy

)


best_k = list(

    k_values

)[best_k_index]


best_k_accuracy = k_accuracy[

    best_k_index

]



print(

    "\n============================================"

)


print(

    "             BEST K VALUE"

)


print(

    "============================================"

)


print(

    "Best K:",

    best_k

)


print(

    "Accuracy:",

    round(

        best_k_accuracy,

        4

    )

)



# ============================================================

# 48. TRAIN FINAL KNN USING BEST K

# ============================================================


final_knn = Pipeline(


    steps=[


        (

            "preprocessing",


            preprocessor

        ),


        (

            "model",


            KNeighborsClassifier(


                n_neighbors=best_k


            )


        )


    ]


)



final_knn.fit(


    X_train,


    y_train


)



final_prediction = final_knn.predict(

    X_test

)



final_probability = (


    final_knn

    .predict_proba(

        X_test

    )[:, 1]


)



# ============================================================

# 49. FINAL METRICS

# ============================================================


final_accuracy = accuracy_score(


    y_test,


    final_prediction


)



final_precision = precision_score(


    y_test,


    final_prediction,


    zero_division=0


)



final_recall = recall_score(


    y_test,


    final_prediction,


    zero_division=0


)



final_f1 = f1_score(


    y_test,


    final_prediction,


    zero_division=0


)



final_auc = roc_auc_score(


    y_test,


    final_probability


)



print(

    "\n============================================"

)


print(

    "       FINAL KNN MODEL RESULTS"

)


print(

    "============================================"

)


print(

    "Best K       :",

    best_k

)


print(

    "Accuracy     :",

    round(

        final_accuracy,

        4

    )

)


print(

    "Precision    :",

    round(

        final_precision,

        4

    )

)


print(

    "Recall       :",

    round(

        final_recall,

        4

    )

)


print(

    "F1 Score     :",

    round(

        final_f1,

        4

    )

)


print(

    "ROC-AUC      :",

    round(

        final_auc,

        4

    )

)


print(

    "============================================"

)



# ============================================================

# 50. FINAL CONFUSION MATRIX

# ============================================================


final_cm = confusion_matrix(


    y_test,


    final_prediction


)



plt.figure(

    figsize=(7,6)

)


sns.heatmap(


    final_cm,


    annot=True,


    fmt="d",


    cmap="Blues",


    xticklabels=[

        "Predicted Low",

        "Predicted High"

    ],


    yticklabels=[

        "Actual Low",

        "Actual High"

    ]


)


plt.title(

    f"Final KNN Confusion Matrix (K={best_k})"

)


plt.xlabel(

    "Predicted Class"

)


plt.ylabel(

    "Actual Class"

)


plt.show()



# ============================================================

# 51. FINAL CLASSIFICATION REPORT

# ============================================================


print(

    "\nFinal Classification Report:"

)


print(


    classification_report(


        y_test,


        final_prediction,


        target_names=[

            "Low Price",

            "High Price"

        ],


        zero_division=0


    )


)



# ============================================================

# 52. ROC CURVE

# ============================================================


fpr, tpr, thresholds = roc_curve(


    y_test,


    final_probability


)



plt.figure(

    figsize=(8,6)

)


plt.plot(


    fpr,


    tpr,


    label=f"KNN ROC-AUC = {final_auc:.4f}"


)


plt.plot(


    [0,1],


    [0,1],


    linestyle="--"


)


plt.xlabel(

    "False Positive Rate"

)


plt.ylabel(

    "True Positive Rate"

)


plt.title(

    "KNN ROC Curve"

)


plt.legend()


plt.show()



# ============================================================

# 53. ACTUAL VS PREDICTED

# ============================================================


prediction_table = pd.DataFrame({


    "Actual Class":

        y_test.values,


    "Predicted Class":

        final_prediction,


    "High Price Probability":

        final_probability


})



prediction_table["Actual Class"] = (


    prediction_table[

        "Actual Class"

    ].map({


        0: "Low Price",


        1: "High Price"


    })


)



prediction_table["Predicted Class"] = (


    prediction_table[

        "Predicted Class"

    ].map({


        0: "Low Price",


        1: "High Price"


    })


)



prediction_table[

    "High Price Probability"

] = (


    prediction_table[

        "High Price Probability"

    ].round(4)


)



print(

    "\nActual vs Predicted:"

)


display(

    prediction_table.head(20)

)



# ============================================================

# 54. NEW HOUSE CLASSIFICATION

# ============================================================

#

# This version avoids the X.mean() categorical error.

# Numerical columns → median

# Categorical columns → mode

# ============================================================


new_house = pd.DataFrame(

    index=[0],

    columns=X.columns

)



# Numerical features


for column in numeric_features:


    new_house.loc[

        0,

        column

    ] = X[

        column

    ].median()



# Categorical features


for column in categorical_features:


    new_house.loc[

        0,

        column

    ] = X[

        column

    ].mode()[0]



# Restore category dtype if necessary


for column in categorical_features:


    if str(

        X[column].dtype

    ) == "category":


        new_house[column] = pd.Categorical(


            new_house[column],


            categories=X[

                column

            ].cat.categories


        )



# ============================================================

# 55. PREDICT NEW HOUSE

# ============================================================


new_prediction = final_knn.predict(


    new_house


)[0]



new_probability = final_knn.predict_proba(


    new_house


)[0][1]



if new_prediction == 1:


    new_result = "High Price"


else:


    new_result = "Low Price"



print(

    "\n============================================"

)


print(

    "        NEW HOUSE CLASSIFICATION"

)


print(

    "============================================"

)


print(

    "Predicted Class:",

    new_result

)


print(

    "Probability of High Price:",

    round(

        new_probability,

        4

    )

)


print(

    "Probability of Low Price:",

    round(

        1 - new_probability,

        4

    )

)



# ============================================================

# 56. FINAL KNN SUMMARY

# ============================================================


print(

    "\n============================================"

)


print(

    "             KNN SUMMARY"

)


print(

    "============================================"

)


print(

    "Algorithm: K-Nearest Neighbors"

)


print(

    "Number of Neighbors:",

    best_k

)


print(

    "Accuracy:",

    round(

        final_accuracy,

        4

    )

)


print(

    "Precision:",

    round(

        final_precision,

        4

    )

)


print(

    "Recall:",

    round(

        final_recall,

        4

    )

)


print(

    "F1 Score:",

    round(

        final_f1,

        4

    )

)


print(

    "ROC-AUC:",

    round(

        final_auc,

        4

    )

)


print(

    "Cross Validation Mean:",

    round(

        cv_scores.mean(),

        4

    )

)


print(

    "============================================"

)



# ============================================================

# 57. KNN THEORY

# ============================================================


print("""

============================================================

K-NEAREST NEIGHBORS (KNN) - THEORY

============================================================


KNN is a supervised machine learning algorithm used for

classification and regression.


For classification, KNN looks at the nearest K observations

and assigns the class that receives the majority vote.


Example:


K = 5


Nearest neighbors:


Neighbor 1 → High Price

Neighbor 2 → High Price

Neighbor 3 → Low Price

Neighbor 4 → High Price

Neighbor 5 → Low Price


High Price = 3

Low Price  = 2


Prediction = High Price


============================================================


WHY STANDARDIZATION IS IMPORTANT?


KNN uses distance calculations.


Features with large numerical values can dominate the distance.


StandardScaler converts features approximately to:


Mean = 0

Standard Deviation = 1


Therefore, scaling is particularly important for KNN.


============================================================


K VALUE


Small K:

- Sensitive to noise

- More flexible

- Can overfit


Large K:

- Smoother decision boundary

- Less sensitive to noise

- Can underfit


The code tests K = 1 to K = 20.


============================================================


BOSTON HOUSING CLASSIFICATION


Original target:


MEDV = Continuous house price


Converted target:


0 = Low Price

1 = High Price


Median MEDV is used as the classification threshold.


============================================================

""")