Four Major Clustering Techniques
K-Means • K-Medoids • DBSCAN • Hierarchical Clustering
Common Dataset Used by All Algorithms
Dataset:
The same six two-dimensional points are used for all four clustering
algorithms so that their results can be compared fairly.
import numpy as np
X = np.array([
[1, 2],
[2, 2],
[2, 3],
[8, 7],
[8, 8],
[25, 80]
])
print(X)
Dataset:
[[ 1 2]
[ 2 2]
[ 2 3]
[ 8 7]
[ 8 8]
[25 80]]
1. K-Means Clustering
Concept:
K-Means divides the dataset into a predefined number of clusters.
It calculates the distance between data points and cluster centroids
and repeatedly updates the centroids until the clusters become stable.
Important Parameter:
This tells K-Means to create two clusters.
n_clusters=2
This tells K-Means to create two clusters.
Python / Google Colab Code
import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score
# 1. Dataset
X = np.array([
[1, 2],
[2, 2],
[2, 3],
[8, 7],
[8, 8],
[25, 80]
])
# 2. Create K-Means model
kmeans = KMeans(
n_clusters=2,
random_state=42,
n_init=10
)
# 3. Fit model
kmeans.fit(X)
# 4. Get labels
labels = kmeans.labels_
print("K-Means Labels:", labels)
# 5. Get centroids
print("Centroids:")
print(kmeans.cluster_centers_)
# 6. Plot
plt.figure(figsize=(6,5))
scatter = plt.scatter(
X[:,0],
X[:,1],
c=labels,
cmap='viridis',
s=100,
edgecolors='black'
)
plt.scatter(
kmeans.cluster_centers_[:,0],
kmeans.cluster_centers_[:,1],
marker='X',
s=250,
c='red',
label='Centroids'
)
plt.title("K-Means Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.legend()
plt.show()
# 7. Silhouette Score
score = silhouette_score(X, labels)
print("Silhouette Score:", score)
2. K-Medoids Clustering
Concept:
K-Medoids is similar to K-Means, but instead of calculating the
average position of a cluster, it selects an actual data point as
the representative center called a medoid.
Unlike K-Means, the center of a cluster must be one of the existing
data points.
Important Parameters:
n_clusters=2 → Number of clusters
random_state=42 → Makes the result reproducible
Python / Google Colab Code
!pip install scikit-learn-extra -q
import numpy as np
import matplotlib.pyplot as plt
from sklearn_extra.cluster import KMedoids
from sklearn.metrics import silhouette_score
# 1. Dataset
X = np.array([
[1, 2],
[2, 2],
[2, 3],
[8, 7],
[8, 8],
[25, 80]
])
# 2. Create K-Medoids model
kmedoids = KMedoids(
n_clusters=2,
random_state=42
)
# 3. Fit model
kmedoids.fit(X)
# 4. Get labels
labels = kmedoids.labels_
print("K-Medoids Labels:", labels)
# 5. Get medoids
print("Medoids:")
print(kmedoids.cluster_centers_)
# 6. Plot
plt.figure(figsize=(6,5))
scatter = plt.scatter(
X[:,0],
X[:,1],
c=labels,
cmap='viridis',
s=100,
edgecolors='black'
)
plt.scatter(
kmedoids.cluster_centers_[:,0],
kmedoids.cluster_centers_[:,1],
marker='X',
s=250,
c='red',
label='Medoids'
)
plt.title("K-Medoids Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.legend()
plt.show()
# 7. Silhouette Score
score = silhouette_score(X, labels)
print("Silhouette Score:", score)
Important Difference:
K-Means uses calculated centroids, while K-Medoids uses actual
data points as cluster centers. Therefore, K-Medoids is generally
less sensitive to extreme values than K-Means.
3. DBSCAN Clustering
Concept:
DBSCAN stands for Density-Based Spatial Clustering of
Applications with Noise.
Instead of specifying the number of clusters, DBSCAN groups points
according to their density. It can also identify isolated points
as noise or anomalies.
In this dataset, the point [25, 80] is far away
from the other observations and can be identified as noise.
Important Parameters:
eps=3 → Maximum neighborhood distance
min_samples=2 → Minimum number of points required
to form a dense region
label=-1 → Noise / anomaly
Python / Google Colab Code
import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import DBSCAN
from sklearn.metrics import silhouette_score
# 1. Our dataset
X = np.array([
[1, 2],
[2, 2],
[2, 3],
[8, 7],
[8, 8],
[25, 80]
])
# 2. Setup and fit DBSCAN
dbscan = DBSCAN(
eps=3,
min_samples=2
)
dbscan.fit(X)
# 3. Get cluster labels
# -1 means noise/anomaly
labels = dbscan.labels_
print("Assigned Labels:", labels)
# 4. Plot results
plt.figure(figsize=(6,5))
scatter = plt.scatter(
X[:,0],
X[:,1],
c=labels,
cmap='viridis',
s=100,
edgecolors='black'
)
plt.colorbar(
scatter,
label='Cluster ID (-1 = Noise)'
)
plt.title("DBSCAN Clustering Results")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.show()
# 5. Silhouette Score
# Remove noise before calculating score
mask = labels != -1
if len(set(labels[mask])) >= 2:
score = silhouette_score(
X[mask],
labels[mask]
)
print("Silhouette Score:", score)
else:
print("Silhouette Score cannot be calculated.")
Important:
DBSCAN is different from K-Means and K-Medoids because we do not
need to specify the number of clusters beforehand. DBSCAN can also
identify noise points using the label
-1.
4. Hierarchical Clustering
Concept:
Hierarchical clustering creates a hierarchy of clusters.
In agglomerative hierarchical clustering, every
data point initially forms its own cluster. The closest clusters
are then repeatedly merged until the required number of clusters
is obtained.
A dendrogram can be used to visualize this merging process.
Important Parameters:
n_clusters=2 → Number of final clusters
linkage='ward' → Method used to calculate cluster
distance
Python / Google Colab Code
import numpy as np
import matplotlib.pyplot as plt
from sklearn.cluster import AgglomerativeClustering
from sklearn.metrics import silhouette_score
from scipy.cluster.hierarchy import dendrogram, linkage
# 1. Dataset
X = np.array([
[1, 2],
[2, 2],
[2, 3],
[8, 7],
[8, 8],
[25, 80]
])
# 2. Create hierarchical model
hierarchical = AgglomerativeClustering(
n_clusters=2,
linkage='ward'
)
# 3. Fit and predict
labels = hierarchical.fit_predict(X)
print("Hierarchical Labels:", labels)
# 4. Plot clusters
plt.figure(figsize=(6,5))
scatter = plt.scatter(
X[:,0],
X[:,1],
c=labels,
cmap='viridis',
s=100,
edgecolors='black'
)
plt.title("Hierarchical Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.show()
# 5. Silhouette Score
score = silhouette_score(X, labels)
print("Silhouette Score:", score)
# 6. Create dendrogram
linked = linkage(X, method='ward')
plt.figure(figsize=(8,5))
dendrogram(linked)
plt.title("Hierarchical Clustering Dendrogram")
plt.xlabel("Data Points")
plt.ylabel("Distance")
plt.show()
Comparison of Four Clustering Techniques
| Technique | Number of Clusters Required? | Cluster Center | Noise Detection | Main Idea |
|---|---|---|---|---|
| K-Means | Yes | Centroid | No | Distance from centroid |
| K-Medoids | Yes | Actual data point | No | Distance from medoid |
| DBSCAN | No | Density | Yes | Density-based grouping |
| Hierarchical | Yes | No fixed center | No | Repeated cluster merging |
Dataset Observation
The dataset contains two groups of nearby observations and one far-away point:
[1,2] [2,2] [2,3]
[8,7] [8,8]
[25,80] ← isolated point
This makes the dataset useful for demonstrating an important difference between the algorithms, particularly DBSCAN's ability to identify an isolated observation as noise.
No comments:
Post a Comment