Total Pageviews

Thursday, October 1, 2026

📊 Clustering Algorithms Notebook

Four Major Clustering Techniques

K-Means • K-Medoids • DBSCAN • Hierarchical Clustering

Common Dataset Used by All Algorithms

Dataset: The same six two-dimensional points are used for all four clustering algorithms so that their results can be compared fairly.
import numpy as np

X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

print(X)
Dataset:
[[ 1  2]
 [ 2  2]
 [ 2  3]
 [ 8  7]
 [ 8  8]
 [25 80]]

1. K-Means Clustering

Concept: K-Means divides the dataset into a predefined number of clusters. It calculates the distance between data points and cluster centroids and repeatedly updates the centroids until the clusters become stable.
Important Parameter: n_clusters=2
This tells K-Means to create two clusters.

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create K-Means model
kmeans = KMeans(
    n_clusters=2,
    random_state=42,
    n_init=10
)

# 3. Fit model
kmeans.fit(X)

# 4. Get labels
labels = kmeans.labels_

print("K-Means Labels:", labels)

# 5. Get centroids
print("Centroids:")
print(kmeans.cluster_centers_)

# 6. Plot
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.scatter(
    kmeans.cluster_centers_[:,0],
    kmeans.cluster_centers_[:,1],
    marker='X',
    s=250,
    c='red',
    label='Centroids'
)

plt.title("K-Means Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.legend()
plt.show()

# 7. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)





2. K-Medoids Clustering

Concept: K-Medoids is similar to K-Means, but instead of calculating the average position of a cluster, it selects an actual data point as the representative center called a medoid. Unlike K-Means, the center of a cluster must be one of the existing data points.
Important Parameters:
n_clusters=2 → Number of clusters
random_state=42 → Makes the result reproducible

Python / Google Colab Code

!pip install scikit-learn-extra -q

import numpy as np
import matplotlib.pyplot as plt

from sklearn_extra.cluster import KMedoids
from sklearn.metrics import silhouette_score

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create K-Medoids model
kmedoids = KMedoids(
    n_clusters=2,
    random_state=42
)

# 3. Fit model
kmedoids.fit(X)

# 4. Get labels
labels = kmedoids.labels_

print("K-Medoids Labels:", labels)

# 5. Get medoids
print("Medoids:")
print(kmedoids.cluster_centers_)

# 6. Plot
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.scatter(
    kmedoids.cluster_centers_[:,0],
    kmedoids.cluster_centers_[:,1],
    marker='X',
    s=250,
    c='red',
    label='Medoids'
)

plt.title("K-Medoids Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.legend()
plt.show()

# 7. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)
Important Difference: K-Means uses calculated centroids, while K-Medoids uses actual data points as cluster centers. Therefore, K-Medoids is generally less sensitive to extreme values than K-Means.








3. DBSCAN Clustering

Concept: DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise. Instead of specifying the number of clusters, DBSCAN groups points according to their density. It can also identify isolated points as noise or anomalies. In this dataset, the point [25, 80] is far away from the other observations and can be identified as noise.
Important Parameters:
eps=3 → Maximum neighborhood distance
min_samples=2 → Minimum number of points required to form a dense region
label=-1 → Noise / anomaly

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import DBSCAN
from sklearn.metrics import silhouette_score

# 1. Our dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Setup and fit DBSCAN
dbscan = DBSCAN(
    eps=3,
    min_samples=2
)

dbscan.fit(X)

# 3. Get cluster labels
# -1 means noise/anomaly
labels = dbscan.labels_

print("Assigned Labels:", labels)

# 4. Plot results
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.colorbar(
    scatter,
    label='Cluster ID (-1 = Noise)'
)

plt.title("DBSCAN Clustering Results")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)

plt.show()

# 5. Silhouette Score
# Remove noise before calculating score
mask = labels != -1

if len(set(labels[mask])) >= 2:
    score = silhouette_score(
        X[mask],
        labels[mask]
    )
    print("Silhouette Score:", score)
else:
    print("Silhouette Score cannot be calculated.")
Important: DBSCAN is different from K-Means and K-Medoids because we do not need to specify the number of clusters beforehand. DBSCAN can also identify noise points using the label -1.



4. Hierarchical Clustering

Concept: Hierarchical clustering creates a hierarchy of clusters. In agglomerative hierarchical clustering, every data point initially forms its own cluster. The closest clusters are then repeatedly merged until the required number of clusters is obtained. A dendrogram can be used to visualize this merging process.
Important Parameters:
n_clusters=2 → Number of final clusters
linkage='ward' → Method used to calculate cluster distance

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import AgglomerativeClustering
from sklearn.metrics import silhouette_score
from scipy.cluster.hierarchy import dendrogram, linkage

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create hierarchical model
hierarchical = AgglomerativeClustering(
    n_clusters=2,
    linkage='ward'
)

# 3. Fit and predict
labels = hierarchical.fit_predict(X)

print("Hierarchical Labels:", labels)

# 4. Plot clusters
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.title("Hierarchical Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)

plt.show()

# 5. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)

# 6. Create dendrogram
linked = linkage(X, method='ward')

plt.figure(figsize=(8,5))

dendrogram(linked)

plt.title("Hierarchical Clustering Dendrogram")
plt.xlabel("Data Points")
plt.ylabel("Distance")

plt.show()







Comparison of Four Clustering Techniques

Technique Number of Clusters Required? Cluster Center Noise Detection Main Idea
K-Means Yes Centroid No Distance from centroid
K-Medoids Yes Actual data point No Distance from medoid
DBSCAN No Density Yes Density-based grouping
Hierarchical Yes No fixed center No Repeated cluster merging

Dataset Observation

The dataset contains two groups of nearby observations and one far-away point:

[1,2] [2,2] [2,3]

[8,7] [8,8]

[25,80]  ← isolated point

This makes the dataset useful for demonstrating an important difference between the algorithms, particularly DBSCAN's ability to identify an isolated observation as noise.


No comments:

Post a Comment