Total Pageviews

Monday, October 5, 2026

📊 CREDIT CARD DEFAULT DATASET


📊 CREDIT CARD DEFAULT DATASET

Detailed Description of the Dataset and Its Variables

1 Dataset Name
Default of Credit Card Clients Dataset

This dataset contains information about credit-card customers and their payment history.

The dataset contains customer information, credit information, repayment history, bill amounts, previous payment amounts and the customer's payment status for the following month.

2 Basic Dataset Information
Property Details
Dataset Name Default of Credit Card Clients
Total Records 30,000 customers
Original Columns 25 columns including ID and target
Customer Information Demographic and financial information
Time Period April 2005 to September 2005
Currency New Taiwan Dollar (NT$)
Target Information Default payment status for the next month
3 What Does One Row Mean?

Each row in the dataset represents information about one credit-card customer.

Row 1 → Customer 1
Row 2 → Customer 2
Row 3 → Customer 3
...
Row 30,000 → Customer 30,000

Therefore, there are 30,000 customer records in the dataset.

4 ID – Customer Identification Number
Column Meaning
ID Unique identification number assigned to the customer
Example:

ID = 1 means one particular customer record.
ID = 2 means another customer record.

The ID is used to identify records; it does not describe the customer's financial condition.
5 Credit Limit
```
Column Meaning

📊 DATASETS USED IN MACHINE LEARNING

🤖 Introduction to Machine Learning Datasets

A dataset is a structured collection of data used by Machine Learning algorithms to learn patterns, relationships, and useful information from examples. Datasets are the foundation of almost every Machine Learning project.

In Machine Learning, datasets generally contain features (input variables) and a target variable (output). Features provide the information used by a model, while the target represents the value that a supervised learning model attempts to predict.

📊

Features

Input variables or attributes used by a Machine Learning model to learn patterns from the data.

🎯

Target

The output variable that a supervised Machine Learning model attempts to predict.

🧹

Data Preparation

Data may require cleaning, transformation, encoding, scaling, and feature selection before model training.

🧠

Model Training

Machine Learning algorithms use training data to identify patterns and build predictive models.

💡 Explore the datasets below to study different Machine Learning problems such as classification, regression, clustering, dimensionality reduction, and feature selection.

Thursday, October 1, 2026

🧠 SVM Kernel Mathematics & Kernel Matrices

🧠 SVM Kernel Mathematics & Kernel Matrices

A Support Vector Machine (SVM) uses a kernel function to calculate the similarity between two observations. Instead of explicitly transforming the data into a higher-dimensional feature space, the kernel computes:

K(x,z) = φ(x)ᵀφ(z)

For n observations, the kernel calculations produce an n × n Gram matrix.

1️⃣ Five-Point Dataset

Point x₁ x₂ Class y
A12-1
B21-1
C45+1
D54+1
E66+1

Example Points

We will frequently calculate the kernel between:

A = (1,2)
C = (4,5)

Dot product:

A · C = (1)(4) + (2)(5)
= 4 + 10
= 14

Squared Euclidean distance:

||A-C||² = (1-4)² + (2-5)²
= 9 + 9
= 18

L1 distance:

||A-C||₁ = |1-4| + |2-5|
= 3 + 3
= 6

2️⃣ Linear Kernel

K(x,z) = xᵀz
For A and C:

K(A,C) = (1)(4) + (2)(5)
= 4 + 10
= 14

5 × 5 Linear Kernel Matrix

[
54141318
45131418
1413414054
1314404154
1818545472
]

3️⃣ Polynomial Kernel

K(x,z) = (γxᵀz + r)ᵈ

Take:

γ = 1    r = 1    d = 2
K(A,C)
= (1 × 14 + 1)²
= 15²
= 225

Polynomial Degree 2 Matrix

[
3625225196361
2536196225361
225196176416813025
196225168117643025
361361302530255329
]

4️⃣ RBF / Gaussian Kernel

K(x,z) = exp(-γ||x-z||²)

Take:

γ = 0.1
For A and C:

||A-C||² = 18
K(A,C) = e-0.1(18)
= e-1.8
≈ 0.1653

RBF Kernel Matrix

[
1.00000.81870.16530.13530.0166
0.81871.00000.13530.16530.0166
0.16530.13531.00000.81870.6065
0.13530.16530.81871.00000.6065
0.01660.01660.60650.60651.0000
]

5️⃣ Sigmoid Kernel

K(x,z) = tanh(γxᵀz + r)

Take:

γ = 0.1    r = 0
K(A,C)
= tanh(0.1 × 14)
= tanh(1.4)
≈ 0.8854

Sigmoid Kernel Matrix

[
0.46210.37990.88540.86170.9468
0.37990.46210.86170.88540.9468
0.88540.86170.99950.99931.0000
0.86170.88540.99930.99951.0000
0.94680.94681.00001.00001.0000
]

6️⃣ Laplacian Kernel

K(x,z) = exp(-γ||x-z||₁)
For A and C:

||A-C||₁ = |1-4| + |2-5|
= 3 + 3
= 6

K(A,C) = e-0.1(6)
= e-0.6
≈ 0.5488

Laplacian Kernel Matrix

[
1.00000.81870.54880.54880.4066
0.81871.00000.54880.54880.4066
0.54880.54881.00000.81870.7408
0.54880.54880.81871.00000.7408
0.40660.40660.74080.74081.0000
]

7️⃣ ANOVA Kernel

K(x,z) = [ Σᵢ exp(-γ(xᵢ-zᵢ)²) ]ᵈ

Take:

γ = 0.1    d = 2
For A = (1,2) and C = (4,5):

K(A,C) = [e-0.1(1-4)² + e-0.1(2-5)²]²
= [e-0.9 + e-0.9]²
= [2(0.40657)]²
≈ 0.6612

ANOVA Kernel Matrix

[
4.00003.27490.66120.76080.0806
3.27494.00000.76080.66120.0806
0.66120.76084.00003.27492.4811
0.76080.66123.27494.00002.4811
0.08060.08062.48112.48114.0000
]

8️⃣ Chi-Square Kernel

K(x,z) = exp[ -γ Σᵢ (xᵢ-zᵢ)²/(xᵢ+zᵢ) ]
For A and C:

Term 1:
(1-4)²/(1+4) = 9/5 = 1.8

Term 2:
(2-5)²/(2+5) = 9/7 ≈ 1.2857

Total:
1.8 + 1.2857 = 3.0857

K(A,C) = e-0.1(3.0857)
≈ 0.7345

Chi-Square Kernel Matrix

[
1.00000.93550.73450.71650.5728
0.93551.00000.71650.73450.5728
0.73450.71651.00000.97800.9521
0.71650.73450.97801.00000.9521
0.57280.57280.95210.95211.0000
]

9️⃣ Histogram Intersection Kernel

K(x,z) = Σᵢ min(xᵢ,zᵢ)
For A and C:

K(A,C) = min(1,4) + min(2,5)
= 1 + 2
= 3

Histogram Intersection Matrix

[
32333
23333
33989
33899
339912
]

🔟 Kernel Matrix Concept

For five observations:

X = {x₁,x₂,x₃,x₄,x₅}

the kernel matrix is:

K = [ K(x₁,x₁)   K(x₁,x₂)   ...   K(x₁,x₅)
K(x₂,x₁)   K(x₂,x₂)   ...   K(x₂,x₅)
⋮            ⋱          ⋮
K(x₅,x₁)   K(x₅,x₂)   ...   K(x₅,x₅) ]

Each entry represents the similarity between two observations. For example:

K₂,₄ = K(B,D)

The diagonal contains self-similarity:

K(xᵢ,xᵢ)

📊 Kernel Comparison

Kernel Formula K(A,C)
Linear xᵀz 14
Polynomial d=2 (xᵀz+1)² 225
Polynomial d=3 (xᵀz+1)³ 3375
RBF e-0.1||x-z||² 0.1653
Sigmoid tanh(0.1xᵀz) 0.8854
Laplacian e-0.1||x-z||₁ 0.5488
ANOVA [Σe-0.1(xᵢ-zᵢ)²]² 0.6612
Chi-Square exp[-γΣ((xᵢ-zᵢ)²/(xᵢ+zᵢ))] 0.7345
Histogram Intersection Σmin(xᵢ,zᵢ) 3
📌 Important:

The kernel matrix is also called the Gram matrix. For five observations, every kernel produces a 5 × 5 matrix.

For the SVM, these kernel values are used inside the dual optimization formulation:

max α: Σᵢ αᵢ − 1/2 ΣᵢΣⱼ αᵢαⱼ yᵢyⱼ K(xᵢ,xⱼ)

subject to:

αᵢ ≥ 0

Σᵢ αᵢyᵢ = 0

🤖 Support Vector Machine (SVM)

🤖 Support Vector Machine (SVM)

Support Vector Machine (SVM) is a supervised machine-learning algorithm used mainly for classification. It can also be used for regression through Support Vector Regression (SVR).

The central idea is to find a decision boundary that separates classes while maintaining an appropriate margin from the nearest observations.

1. What is SVM?

SVM searches for a decision boundary that separates observations belonging to different classes.

Class A Class B Decision Boundary
SVM separates observations using a decision boundary.

2. Hyperplane

The hyperplane is the mathematical decision boundary.

Hyperplane x-axis y
w · x + b = 0

3. Margin

The margin represents the separation between the decision boundary and the closest observations.

Maximum Margin Margin Boundary
Margin = 2 / ||w||

4. Support Vectors

Support vectors are the observations closest to the decision boundary.

Support Vectors Hyperplane

5. Maximum-Margin Principle

Several boundaries may separate the classes. SVM searches for an appropriate boundary with maximum separation from the nearest examples.

Boundary A Maximum-margin boundary

6. Hard Margin SVM

Hard-margin SVM requires the training observations to satisfy the separation constraints.

No training violations

7. Soft Margin SVM

Soft-margin SVM allows some observations to violate the ideal margin.

Violation
Minimize ½ ||w||² + C Σ ξᵢ

8. C Parameter

C controls the penalty associated with margin violations.

Small C Large C Wider tolerance Stricter penalty

9. Linear SVM

Linear SVM uses a straight decision boundary.

Linear Decision Boundary

10. Non-Linear SVM

When a straight line cannot separate the classes effectively, SVM can use a non-linear kernel.

Curved Boundary

11. Kernel Trick

Original Data
→
Kernel
→
Higher Feature Space
→
Linear Separation
Original Space Kernel Feature Space

12. Linear Kernel

K(x,z) = x · z
Vector x Vector z Dot Product

13. Polynomial Kernel

K(x,z) = (γ x · z + r)d
Polynomial Decision Boundary

14. RBF / Gaussian Kernel

K(x,z) = exp(-γ ||x-z||²)
Radial Influence

15. Sigmoid Kernel

K(x,z) = tanh(γ x · z + r)
Sigmoid Curve

16. Gamma Parameter

Low Gamma High Gamma Broader influence Local influence

17. SVM Workflow

Data Preprocess Scale Train SVM Prediction

18. Feature Scaling

SVM models can be affected when features have very different numerical scales.

Before Scaling After Scaling
z = (x - μ) / σ

19. Spam Detection

Email
→
Text Processing
→
TF-IDF
→
SVM
→
Spam / Ham
EMAIL "Win free prize" TF-IDF SVM SPAM

20. Classification Metrics

Accuracy

Correct

Precision

Positive

Recall

21. Confusion Matrix

Predicted Actual True Negative False Positive False Negative True Positive

22. Iris Classification

Setosa Versicolor Virginica

23. Advantages of SVM

📐 Maximum Margin

Uses a margin-based decision principle.

📊 High Dimensions

Can work with high-dimensional feature representations.

🔄 Kernels

Can model non-linear relationships.

📝 Text Data

Useful with sparse text feature representations.

Margin Dimensions Kernels Text

24. Limitations of SVM

Large Data Tuning Scaling
  • Large datasets can increase computational cost.
  • Kernel and hyperparameter selection may require experimentation.
  • Feature scaling is often important.
  • Interpretability may be lower than simple rule-based models.

25. Applications of SVM

SVM Spam Detection Image Classification Medical Data Fraud Detection

26. SVM Classification vs Regression

SVC SVR

27. Important SVM Terminology

SVM Hyperplane Kernel Margin Support Vector

28. Kernel Comparison

Kernel Visual Boundary Main Idea
Linear ──────── Straight boundary
Polynomial ∿∿∿ Polynomial relationship
RBF ◉ Radial similarity
Sigmoid S Sigmoid-shaped relationship

29. Practical SVM Workflow

Dataset Clean Scale Tune Evaluate

30. SVM Parameters at a Glance

SVM C Gamma Degree Kernel

31. SVM Decision Process

Input SVM Model Class

32. Hard Margin vs Soft Margin

Property Hard Margin Soft Margin
Training violations Not allowed Allowed
Slack variables No Yes
Noise tolerance Low Higher
Parameter C Not the main formulation Important
Hard Margin Soft Margin

33. Classification Pipeline

Training Data
→
Feature Engineering
→
Scaling
→
SVM
→
Metrics

34. SVM vs Other Classification Models

SVM Decision Tree KNN Logistic Naive Bayes

35. SVM Mathematical View

w · x + b = 0 -1 boundary +1 boundary
yᵢ(w · xᵢ + b) ≥ 1

36. Support Vector Geometry

Support vector Margin

37. SVM Prediction

New Data Decision Boundary

38. Advantages and Limitations Summary

Advantages ✓ Maximum-margin learning ✓ Kernel support ✓ High-dimensional data Limitations ⚠ Parameter tuning ⚠ Scaling required ⚠ Large-data cost

39. Interview Concept Map

SVM Hyperplane Margin Support Vector Kernel C / Gamma Prediction

40. Complete SVM Concept Diagram

SVM Hyperplane Support Vectors Kernel Margin C / Gamma
Core SVM idea: SVM finds an appropriate separating boundary, identifies the observations that most strongly determine that boundary, and balances margin size with training violations through its optimization parameters.

41. Important Interview Questions

Q1. What is SVM?

A supervised machine-learning algorithm that constructs a decision boundary for classification and can also be extended to regression.

Q2. What is a hyperplane?

A mathematical decision boundary separating observations in feature space.

Q3. What are support vectors?

Observations closest to the decision boundary that strongly influence the fitted SVM boundary.

Q4. What is margin?

The separation between the decision boundary and the nearest observations.

Q5. What is the kernel trick?

A technique that allows SVM to model non-linear relationships using kernel similarity functions.

Q6. What does C do?

C controls the penalty associated with observations violating the desired margin constraints.

Q7. What does gamma do?

Gamma controls the locality of influence for kernels such as RBF.

Q8. Why is scaling important?

Because the geometry used by SVM can be affected when features have very different numerical scales.

Hierarchical Clustering

Hierarchical Clustering

Hierarchical Clustering is an unsupervised machine learning technique used to group similar data points into clusters. Instead of directly producing only one partition, it builds a hierarchy of clusters.

There are two main approaches:

  • Agglomerative Clustering – Bottom-Up approach
  • Divisive Clustering – Top-Down approach

1.

a )  Agglomerative Hierarchical Clustering

Agglomerative clustering starts by treating every data point as an individual cluster. The closest clusters are then repeatedly merged until all observations belong to one large cluster.

This is therefore called a Bottom-Up approach.


b )  Divisive Clustering

Divisive Clustering is known as top-down approach.

In this approach we take on huge cluster and starts breaking it up into smaller clusters until it reaches individual data points (or single point clusters).


2. Basic Working

  1. Initially, every data point is considered as a separate cluster.
  2. Calculate the distance between clusters.
  3. Find the two closest clusters.
  4. Merge those two clusters.
  5. Recalculate the distances between the new cluster and the remaining clusters.
  6. Continue the process until one large cluster remains.
  7. Represent the merging process using a Dendrogram.
  8. Choose the desired number of clusters by cutting the dendrogram at an appropriate height.

3. Distance Calculation

Euclidean Distance:

d(P₁,P₂) = √[(X₁ − X₂)² + (Y₁ − Y₂)²]

Example:
If P₁ = (170,56) and P₂ = (168,60)

d = √[(170−168)² + (56−60)²]
d = √(4 + 16)
d = √20 ≈ 4.47

4. Linkage Methods

Single Linkage

Distance between two clusters is based on the closest pair of points.

Complete Linkage

Distance is based on the farthest pair of points between two clusters.

Average Linkage

Distance is calculated using the average distance between points of two clusters.

Ward Linkage

Merges clusters while minimizing the increase in within-cluster variance.

5. Example Dataset

We will use a small Height and Weight dataset similar to the example described in the reference article.

Point Height Weight
P0 165 55
P1 170 56
P2 168 60
P3 178 70
P4 166 55
P5 180 72

6. Python Implementation

Google Colab / Python Code

# ============================================
# HIERARCHICAL CLUSTERING
# AGGLOMERATIVE APPROACH
# ============================================

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

from sklearn.cluster import AgglomerativeClustering
from scipy.cluster.hierarchy import dendrogram, linkage

# --------------------------------------------
# 1. Create Dataset
# --------------------------------------------

data = {
    "Height": [165, 170, 168, 178, 166, 180],
    "Weight": [55, 56, 60, 70, 55, 72]
}

df = pd.DataFrame(data)

print("Dataset:")
print(df)


# --------------------------------------------
# 2. Select Features
# --------------------------------------------

X = df[["Height", "Weight"]]


# --------------------------------------------
# 3. Create Dendrogram
# --------------------------------------------

linked = linkage(
    X,
    method="single",
    metric="euclidean"
)

plt.figure(figsize=(10,6))

dendrogram(
    linked,
    labels=["P0","P1","P2","P3","P4","P5"]
)

plt.title("Hierarchical Clustering Dendrogram")
plt.xlabel("Data Points")
plt.ylabel("Euclidean Distance")

plt.show()


# --------------------------------------------
# 4. Agglomerative Clustering
# --------------------------------------------

model = AgglomerativeClustering(
    n_clusters=2,
    metric="euclidean",
    linkage="single"
)

labels = model.fit_predict(X)


# --------------------------------------------
# 5. Add Cluster Labels
# --------------------------------------------

df["Cluster"] = labels

print("\nClustered Dataset:")
print(df)


# --------------------------------------------
# 6. Visualize Clusters
# --------------------------------------------

plt.figure(figsize=(8,6))

plt.scatter(
    df["Height"],
    df["Weight"],
    c=df["Cluster"],
    s=120
)

plt.xlabel("Height")
plt.ylabel("Weight")
plt.title("Agglomerative Hierarchical Clustering")

plt.grid(True)

plt.show()

7. Dendrogram

A dendrogram is a tree-like diagram that shows the sequence in which clusters are merged.

The vertical axis represents the distance at which clusters are merged. A horizontal cut through the dendrogram can be used to determine the number of clusters.

8. How to Select Number of Clusters?

  1. Generate the dendrogram.
  2. Look for the largest vertical gap that does not intersect another horizontal cluster line.
  3. Draw a horizontal line through that gap.
  4. Count the number of vertical branches crossed by the line.
  5. The number of branches represents the selected number of clusters.

9. Using Different Linkage Methods

Compare Single, Complete, Average and Ward Linkage

import matplotlib.pyplot as plt
from scipy.cluster.hierarchy import linkage, dendrogram

methods = [
    "single",
    "complete",
    "average",
    "ward"
]

plt.figure(figsize=(16,10))

for i, method in enumerate(methods):

    plt.subplot(2,2,i+1)

    linked = linkage(
        X,
        method=method,
        metric="euclidean"
    )

    dendrogram(
        linked,
        labels=["P0","P1","P2","P3","P4","P5"]
    )

    plt.title(method.capitalize() + " Linkage")
    plt.xlabel("Data Points")
    plt.ylabel("Distance")

plt.tight_layout()
plt.show()

10. Complete Agglomerative Clustering Code

Single Google Colab Code

# ============================================
# COMPLETE HIERARCHICAL CLUSTERING
# ============================================

import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

from sklearn.cluster import AgglomerativeClustering
from scipy.cluster.hierarchy import dendrogram, linkage

# Dataset
data = pd.DataFrame({
    "Height": [165,170,168,178,166,180],
    "Weight": [55,56,60,70,55,72]
})

print("Original Dataset")
print(data)


# Features
X = data[["Height","Weight"]]


# --------------------------------------------
# Dendrogram
# --------------------------------------------

Z = linkage(
    X,
    method="single",
    metric="euclidean"
)

plt.figure(figsize=(10,6))

dendrogram(
    Z,
    labels=["P0","P1","P2","P3","P4","P5"]
)

plt.title("Hierarchical Clustering Dendrogram")
plt.xlabel("Data Points")
plt.ylabel("Distance")

plt.show()


# --------------------------------------------
# Agglomerative Model
# --------------------------------------------

hc = AgglomerativeClustering(
    n_clusters=2,
    metric="euclidean",
    linkage="single"
)

data["Cluster"] = hc.fit_predict(X)


# --------------------------------------------
# Display Result
# --------------------------------------------

print("\nFinal Cluster Result")
print(data)


# --------------------------------------------
# Visualization
# --------------------------------------------

plt.figure(figsize=(8,6))

for cluster in sorted(data["Cluster"].unique()):

    subset = data[data["Cluster"] == cluster]

    plt.scatter(
        subset["Height"],
        subset["Weight"],
        s=150,
        label="Cluster " + str(cluster)
    )

    for index, row in subset.iterrows():

        plt.annotate(
            "P" + str(index),
            (row["Height"], row["Weight"]),
            xytext=(5,5),
            textcoords="offset points"
        )

plt.xlabel("Height")
plt.ylabel("Weight")

plt.title(
    "Agglomerative Hierarchical Clustering"
)

plt.legend()
plt.grid(True)

plt.show()

11. Advantages

No Need to Start With K

The hierarchy can be examined before deciding how many clusters to retain.

Dendrogram

The dendrogram provides a visual representation of the merging process.

Different Linkages

Single, complete, average and Ward linkage can be used depending on the data and objective.

Hierarchical Structure

The algorithm provides information about clusters at multiple levels.

12. Limitations

  • Computational cost can become high for very large datasets.
  • The result can depend strongly on the selected linkage method.
  • Once clusters are merged, the basic agglomerative process does not normally undo that merge.
  • Different distance measures can produce different clustering structures.

13. Summary

Hierarchical Clustering builds a hierarchy of clusters. In the Agglomerative approach, every observation starts as its own cluster and the closest clusters are repeatedly merged. A dendrogram records this merging process and can be used to select a desired number of clusters.

The reference article demonstrates this process using Euclidean distance and single linkage, including the formation of the distance matrix, successive merging of clusters, dendrogram construction and selection of two clusters.

📊 Clustering Algorithms Notebook

Four Major Clustering Techniques

K-Means • K-Medoids • DBSCAN • Hierarchical Clustering

Common Dataset Used by All Algorithms

Dataset: The same six two-dimensional points are used for all four clustering algorithms so that their results can be compared fairly.
import numpy as np

X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

print(X)
Dataset:
[[ 1  2]
 [ 2  2]
 [ 2  3]
 [ 8  7]
 [ 8  8]
 [25 80]]

1. K-Means Clustering

Concept: K-Means divides the dataset into a predefined number of clusters. It calculates the distance between data points and cluster centroids and repeatedly updates the centroids until the clusters become stable.
Important Parameter: n_clusters=2
This tells K-Means to create two clusters.

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import KMeans
from sklearn.metrics import silhouette_score

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create K-Means model
kmeans = KMeans(
    n_clusters=2,
    random_state=42,
    n_init=10
)

# 3. Fit model
kmeans.fit(X)

# 4. Get labels
labels = kmeans.labels_

print("K-Means Labels:", labels)

# 5. Get centroids
print("Centroids:")
print(kmeans.cluster_centers_)

# 6. Plot
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.scatter(
    kmeans.cluster_centers_[:,0],
    kmeans.cluster_centers_[:,1],
    marker='X',
    s=250,
    c='red',
    label='Centroids'
)

plt.title("K-Means Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.legend()
plt.show()

# 7. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)





2. K-Medoids Clustering

Concept: K-Medoids is similar to K-Means, but instead of calculating the average position of a cluster, it selects an actual data point as the representative center called a medoid. Unlike K-Means, the center of a cluster must be one of the existing data points.
Important Parameters:
n_clusters=2 → Number of clusters
random_state=42 → Makes the result reproducible

Python / Google Colab Code

!pip install scikit-learn-extra -q

import numpy as np
import matplotlib.pyplot as plt

from sklearn_extra.cluster import KMedoids
from sklearn.metrics import silhouette_score

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create K-Medoids model
kmedoids = KMedoids(
    n_clusters=2,
    random_state=42
)

# 3. Fit model
kmedoids.fit(X)

# 4. Get labels
labels = kmedoids.labels_

print("K-Medoids Labels:", labels)

# 5. Get medoids
print("Medoids:")
print(kmedoids.cluster_centers_)

# 6. Plot
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.scatter(
    kmedoids.cluster_centers_[:,0],
    kmedoids.cluster_centers_[:,1],
    marker='X',
    s=250,
    c='red',
    label='Medoids'
)

plt.title("K-Medoids Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)
plt.legend()
plt.show()

# 7. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)
Important Difference: K-Means uses calculated centroids, while K-Medoids uses actual data points as cluster centers. Therefore, K-Medoids is generally less sensitive to extreme values than K-Means.








3. DBSCAN Clustering

Concept: DBSCAN stands for Density-Based Spatial Clustering of Applications with Noise. Instead of specifying the number of clusters, DBSCAN groups points according to their density. It can also identify isolated points as noise or anomalies. In this dataset, the point [25, 80] is far away from the other observations and can be identified as noise.
Important Parameters:
eps=3 → Maximum neighborhood distance
min_samples=2 → Minimum number of points required to form a dense region
label=-1 → Noise / anomaly

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import DBSCAN
from sklearn.metrics import silhouette_score

# 1. Our dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Setup and fit DBSCAN
dbscan = DBSCAN(
    eps=3,
    min_samples=2
)

dbscan.fit(X)

# 3. Get cluster labels
# -1 means noise/anomaly
labels = dbscan.labels_

print("Assigned Labels:", labels)

# 4. Plot results
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.colorbar(
    scatter,
    label='Cluster ID (-1 = Noise)'
)

plt.title("DBSCAN Clustering Results")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)

plt.show()

# 5. Silhouette Score
# Remove noise before calculating score
mask = labels != -1

if len(set(labels[mask])) >= 2:
    score = silhouette_score(
        X[mask],
        labels[mask]
    )
    print("Silhouette Score:", score)
else:
    print("Silhouette Score cannot be calculated.")
Important: DBSCAN is different from K-Means and K-Medoids because we do not need to specify the number of clusters beforehand. DBSCAN can also identify noise points using the label -1.



4. Hierarchical Clustering

Concept: Hierarchical clustering creates a hierarchy of clusters. In agglomerative hierarchical clustering, every data point initially forms its own cluster. The closest clusters are then repeatedly merged until the required number of clusters is obtained. A dendrogram can be used to visualize this merging process.
Important Parameters:
n_clusters=2 → Number of final clusters
linkage='ward' → Method used to calculate cluster distance

Python / Google Colab Code

import numpy as np
import matplotlib.pyplot as plt

from sklearn.cluster import AgglomerativeClustering
from sklearn.metrics import silhouette_score
from scipy.cluster.hierarchy import dendrogram, linkage

# 1. Dataset
X = np.array([
    [1, 2],
    [2, 2],
    [2, 3],
    [8, 7],
    [8, 8],
    [25, 80]
])

# 2. Create hierarchical model
hierarchical = AgglomerativeClustering(
    n_clusters=2,
    linkage='ward'
)

# 3. Fit and predict
labels = hierarchical.fit_predict(X)

print("Hierarchical Labels:", labels)

# 4. Plot clusters
plt.figure(figsize=(6,5))

scatter = plt.scatter(
    X[:,0],
    X[:,1],
    c=labels,
    cmap='viridis',
    s=100,
    edgecolors='black'
)

plt.title("Hierarchical Clustering")
plt.xlabel("X-axis")
plt.ylabel("Y-axis")
plt.grid(True, linestyle='--', alpha=.5)

plt.show()

# 5. Silhouette Score
score = silhouette_score(X, labels)

print("Silhouette Score:", score)

# 6. Create dendrogram
linked = linkage(X, method='ward')

plt.figure(figsize=(8,5))

dendrogram(linked)

plt.title("Hierarchical Clustering Dendrogram")
plt.xlabel("Data Points")
plt.ylabel("Distance")

plt.show()







Comparison of Four Clustering Techniques

Technique Number of Clusters Required? Cluster Center Noise Detection Main Idea
K-Means Yes Centroid No Distance from centroid
K-Medoids Yes Actual data point No Distance from medoid
DBSCAN No Density Yes Density-based grouping
Hierarchical Yes No fixed center No Repeated cluster merging

Dataset Observation

The dataset contains two groups of nearby observations and one far-away point:

[1,2] [2,2] [2,3]

[8,7] [8,8]

[25,80]  ← isolated point

This makes the dataset useful for demonstrating an important difference between the algorithms, particularly DBSCAN's ability to identify an isolated observation as noise.