Total Pageviews

Tuesday, October 6, 2026

PRINCIPAL COMPONENT ANALYSIS

PRINCIPAL COMPONENT ANALYSIS

4-Feature Numerical Example
Mathematics • Physics • Chemistry • Biology
STEP 01 Get the Data
Consider the marks obtained by five students in four subjects: Mathematics, Physics, Chemistry and Biology.
Student Mathematics Physics Chemistry Biology
A2157
B4346
C6589
D8768
E1091011
Original Data Matrix
X = [ 2  1  5  7 ]
[ 4  3  4  6 ]
[ 6  5  8  9 ]
[ 8  7  6  8 ]
[10  9  10  11]
STEP 02 Compute the Mean Vector (μ)
μ = [ Mean(Math), Mean(Physics), Mean(Chemistry), Mean(Biology) ]
Mean(Math) = (2 + 4 + 6 + 8 + 10) / 5 = 6
Mean(Physics) = (1 + 3 + 5 + 7 + 9) / 5 = 5
Mean(Chemistry) = (5 + 4 + 8 + 6 + 10) / 5 = 6.6
Mean(Biology) = (7 + 6 + 9 + 8 + 11) / 5 = 8.2
Mean Vector:
μ = [ 6, 5, 6.6, 8.2 ]
STEP 03 Subtract Mean from the Given Data
Xcentered = X − μ
Student Math − 6 Physics − 5 Chemistry − 6.6 Biology − 8.2
A−4−4−1.6−1.2
B−2−2−2.6−2.2
C001.40.8
D22−0.6−0.2
E443.42.8
Xcentered = [ −4  −4  −1.6  −1.2 ]
[ −2  −2  −2.6  −2.2 ]
[ 0   0    1.4    0.8 ]
[ 2   2   −0.6  −0.2 ]
[ 4   4    3.4    2.8 ]
STEP 04 Calculate the Covariance Matrix
Following the Gate Vidyalay calculation style, covariance is calculated using 1/n.
C = (1/n) XcenteredTXcentered
C = [ 8.00  8.00  4.80  4.00 ]
[ 8.00  8.00  4.80  4.00 ]
[ 4.80  4.80  4.64  3.68 ]
[ 4.00  4.00  3.68  2.96 ]
Covariance Matrix obtained successfully.
STEP 05 Calculate Eigenvectors and Eigenvalues
Eigenvalues
Component Eigenvalue Variance Explained
PC1 21.5758 91.42%
PC2 2.0053 8.50%
PC3 0.0189 0.08%
PC4 0.0000 0.00%
Total Variance = 21.5758 + 2.0053 + 0.0189 + 0 = 23.6000
Principal Eigenvector — PC1
PC1 = [ −0.59800 ]
[ −0.59800 ]
[ −0.41254 ]
[ −0.33854 ]
The sign of an eigenvector is arbitrary. Therefore, the equivalent positive eigenvector can also be used:
PC1 = [ 0.59800 ]
[ 0.59800 ]
[ 0.41254 ]
[ 0.33854 ]
PC1 explains approximately 91.42% of the total variance. Therefore, PC1 is selected as the principal component.
Second Principal Component — PC2
PC2 ≈ [ −0.37660 ]
[ −0.37660 ]
[ 0.69244 ]
[ 0.48669 ]
STEP 06 Choosing Components and Forming Feature Vector
PC1 contains approximately 91.42% of the total variance. Therefore, the four original features can be reduced primarily to one principal component.
Feature Vector = [ PC1 ]
Feature Vector = [ 0.59800 ]
[ 0.59800 ]
[ 0.41254 ]
[ 0.33854 ]
Dimensionality Reduction: Original dimensions = 4
Reduced dimensions = 1
Variance retained ≈ 91.42%
STEP 07 Deriving the New Dataset
PC1 Score = Xcentered × PC1
Student PC1 Score
A −5.8503
B −4.2094
C 0.8484
D 2.0768
E 7.1345
FINAL PCA DATASET
Student Mathematics Physics Chemistry Biology PC1
A 2 1 5 7 −5.8503
B 4 3 4 6 −4.2094
C 6 5 8 9 0.8484
D 8 7 6 8 2.0768
E 10 9 10 11 7.1345
PCA Visualization
4-Feature Dataset — PC1 Projection
A B C D E Mathematics / Physics Chemistry / Biology PC1 Direction
Interpretation:
The four original features are transformed into principal components. PC1 captures approximately 91.42% of the total variation, making it the most important direction in the dataset.
Final PCA Summary
Item Result
Original Features 4
Mathematics Mean 6.0
Physics Mean 5.0
Chemistry Mean 6.6
Biology Mean 8.2
PC1 Eigenvalue 21.5758
PC2 Eigenvalue 2.0053
PC3 Eigenvalue 0.0189
PC4 Eigenvalue 0.0000
PC1 Variance Explained 91.42%
PC2 Variance Explained 8.50%
PC3 Variance Explained 0.08%
PC4 Variance Explained 0.00%
Reduced Dimension 4 → 1
4 FEATURES → 1 PRINCIPAL COMPONENT
PC1 retains approximately 91.42% of the total variance.

PRINCIPAL COMPONENT ANALYSIS

PRINCIPAL COMPONENT ANALYSIS

3-Feature Numerical Example
Mathematics • Physics • Chemistry
STEP 01 Get the Data
Consider the marks obtained by five students in three subjects: Mathematics, Physics and Chemistry.
Student Mathematics Physics Chemistry
A215
B434
C658
D876
E10910
Original Data Matrix
X = [ 2  1  5 ]
[ 4  3  4 ]
[ 6  5  8 ]
[ 8  7  6 ]
[10  9  10]
STEP 02 Compute the Mean Vector (μ)
μ = [ Mean(Math), Mean(Physics), Mean(Chemistry) ]
Mean(Math) = (2 + 4 + 6 + 8 + 10) / 5 = 6
Mean(Physics) = (1 + 3 + 5 + 7 + 9) / 5 = 5
Mean(Chemistry) = (5 + 4 + 8 + 6 + 10) / 5 = 6.6
Mean Vector: μ = [ 6, 5, 6.6 ]
STEP 03 Subtract Mean from the Given Data
Xcentered = X − μ
Student Math − 6 Physics − 5 Chemistry − 6.6
A−4−4−1.6
B−2−2−2.6
C001.4
D22−0.6
E443.4
Xcentered = [ −4  −4  −1.6 ]
[ −2  −2  −2.6 ]
[ 0   0    1.4 ]
[ 2   2   −0.6 ]
[ 4   4    3.4 ]
STEP 04 Calculate the Covariance Matrix
Following the calculation style used in the Gate Vidyalay example, the covariance matrix is calculated using 1/n.
C = (1/n) XcenteredTXcentered
C = 1/5 × [ 40  40  24 ]
       [ 40  40  24 ]
       [ 24  24  23.2 ]
Covariance Matrix:
C = [ 8.00  8.00  4.80 ]
[ 8.00  8.00  4.80 ]
[ 4.80  4.80  4.64 ]
STEP 05 Calculate Eigenvectors and Eigenvalues
Eigenvalues
Component Eigenvalue Variance Explained
PC1 19.1711 92.88%
PC2 1.4689 7.12%
PC3 0.0000 0.00%
Total Variance = 19.1711 + 1.4689 + 0 = 20.6400
Principal Eigenvector — PC1
PC1 = [ 0.64065 ]
[ 0.64065 ]
[ 0.42325 ]
PC1 explains approximately 92.88% of the total variance. Therefore, PC1 is selected as the principal component.
Second Principal Component — PC2
PC2 = [ −0.29928 ]
[ −0.29928 ]
[ 0.90602 ]
STEP 06 Choosing Components and Forming Feature Vector
Since PC1 explains 92.88% of the total variance, it can be selected as the main reduced feature.
Feature Vector = [ PC1 ]
Feature Vector = [ 0.64065 ]
[ 0.64065 ]
[ 0.42325 ]
Dimensionality Reduction: The original dataset contains 3 features. After PCA, the major information can be represented using 1 principal component with approximately 92.88% variance retention.
STEP 07 Deriving the New Dataset
PC1 Score = Xcentered × PC1
Student PC1 Calculation PC1 Value
A (−4)(0.64065)+(−4)(0.64065)+(−1.6)(0.42325) −5.8024
B (−2)(0.64065)+(−2)(0.64065)+(−2.6)(0.42325) −3.6630
C (0)(0.64065)+(0)(0.64065)+(1.4)(0.42325) 0.5925
D (2)(0.64065)+(2)(0.64065)+(−0.6)(0.42325) 2.3087
E (4)(0.64065)+(4)(0.64065)+(3.4)(0.42325) 6.5642
FINAL PCA DATASET
Student Mathematics Physics Chemistry PC1
A 2 1 5 −5.8024
B 4 3 4 −3.6630
C 6 5 8 0.5925
D 8 7 6 2.3087
E 10 9 10 6.5642
PCA Visualization
3-Feature Dataset — PC1 Projection
A B C D E Mathematics / Physics Direction Chemistry PC1 Direction
Interpretation:
PC1 combines Mathematics, Physics and Chemistry into a single new feature. Because PC1 explains approximately 92.88% of the variance, most of the information in the original three-dimensional dataset is retained.
Final PCA Summary
Item Result
Original Features 3
Mathematics Mean 6.0
Physics Mean 5.0
Chemistry Mean 6.6
PC1 Eigenvalue 19.1711
PC2 Eigenvalue 1.4689
PC3 Eigenvalue 0.0000
PC1 Variance Explained 92.88%
PC2 Variance Explained 7.12%
PC3 Variance Explained 0.00%
Reduced Dimension 3 → 1
3 FEATURES → 1 PRINCIPAL COMPONENT
PC1 retains approximately 92.88% of the total variance.

PRINCIPAL COMPONENT ANALYSIS (PCA)

PRINCIPAL COMPONENT ANALYSIS (PCA)

Two-Feature Numerical Example Using Gate Vidyalay PCA Steps
1. Get the Data
We have the marks of five students in two subjects: Mathematics and Physics. The objective is to transform the original two-dimensional data into a new coordinate system and represent it mainly using the first principal component.
Student Mathematics (X) Physics (Y)
A 2 1
B 4 3
C 6 5
D 8 7
E 10 9
X = [ 2   1 ]
[ 4   3 ]
[ 6   5 ]
[ 8   7 ]
[10   9 ]
STEP 01 Calculate the Mean Vector
According to the PCA procedure, first calculate the mean of each feature.
Mean of Mathematics
μX = (2 + 4 + 6 + 8 + 10) / 5 = 6
Mean of Physics
μY = (1 + 3 + 5 + 7 + 9) / 5 = 5
Mean Vector:

μ = [ 6   5 ]
STEP 02 Subtract Mean Vector from the Data
For every observation, subtract the mean vector:
xi − μ
Student Original X Original Y X − 6 Y − 5
A 2 1 -4 -4
B 4 3 -2 -2
C 6 5 0 0
D 8 7 2 2
E 10 9 4 4
X − μ = [ -4   -4 ]
[ -2   -2 ]
[ 0    0 ]
[ 2    2 ]
[ 4    4 ]
STEP 03 Calculate the Covariance Matrix
Following the calculation used in the Gate Vidyalay worked example, the covariance matrix is calculated using:
C = (1 / n) Σ (xi − μ)(xi − μ)T
Here, n = 5
Covariance of Mathematics with Mathematics
Cov(X,X) = [16 + 4 + 0 + 4 + 16] / 5 = 40 / 5 = 8
Covariance of Mathematics with Physics
Cov(X,Y) = [16 + 4 + 0 + 4 + 16] / 5 = 40 / 5 = 8
Covariance of Physics with Physics
Cov(Y,Y) = [16 + 4 + 0 + 4 + 16] / 5 = 40 / 5 = 8
Covariance Matrix:
C = [ 8   8 ]
[ 8   8 ]
STEP 04 Calculate Eigenvalues and Eigenvectors
An eigenvalue λ satisfies the characteristic equation:
|C − λI| = 0
| 8−λ   8 |
| 8   8−λ |
(8 − λ)(8 − λ) − 64 = 0
64 − 16λ + λ² − 64 = 0
λ² − 16λ = 0
λ(λ − 16) = 0
Eigenvalues:

λ1 = 16
λ2 = 0
Eigenvector corresponding to λ₁ = 16
CX = λX
[ 8   8 ] [X₁]    [16X₁]
[ 8   8 ] [X₂]    [16X₂]
8X₁ + 8X₂ = 16X₁
8X₂ = 8X₁
X₁ = X₂
Therefore, the eigenvector direction is:
[ 1 ]
[ 1 ]
Normalize the Eigenvector
√(1² + 1²) = √2
PC1 = [ 1/√2 ]
[ 1/√2 ] = [ 0.7071 ]
[ 0.7071 ]
Eigenvector corresponding to λ₂ = 0
PC2 = [ 1/√2 ]
[-1/√2 ] = [ 0.7071 ]
[-0.7071]
STEP 05 Choose the Principal Component
Component Eigenvalue Variance Contribution
PC1 16 100%
PC2 0 0%
Total Variance = 16 + 0 = 16
Explained Variance of PC1 = (16 / 16) × 100 = 100%
PC1 captures 100% of the variance.

Therefore, PC2 can be discarded without losing variance in this particular dataset.
STEP 06 Form the Feature Vector
Since PC1 has the largest eigenvalue, it is selected as the principal component.
Feature Vector = [ 0.7071 ]
[ 0.7071 ]
STEP 07 Derive the New Dataset
The new PCA value is obtained by projecting each centered data point onto the selected eigenvector:
Z = PC1T(X − μ)
Student A
Z₁ = (0.7071 × -4) + (0.7071 × -4)
Z₁ = -5.6569
Student B
Z₂ = (0.7071 × -2) + (0.7071 × -2)
Z₂ = -2.8284
Student C
Z₃ = (0.7071 × 0) + (0.7071 × 0)
Z₃ = 0
Student D
Z₄ = (0.7071 × 2) + (0.7071 × 2)
Z₄ = 2.8284
Student E
Z₅ = (0.7071 × 4) + (0.7071 × 4)
Z₅ = 5.6569
8. Original Data Plot and Principal Component
The red line represents the direction of the first principal component. The blue points represent the original observations.
Mathematics vs Physics with PC1 Direction
A (2,1) B (4,3) C (6,5) D (8,7) E (10,9) Mathematics Physics PC1 Direction
9. Final Dataset with PCA Values
The original two features are shown together with the new one-dimensional PCA representation.
Student Mathematics Physics PC1 Value
A 2 1 -5.6569
B 4 3 -2.8284
C 6 5 0.0000
D 8 7 2.8284
E 10 9 5.6569
10. Original Dataset vs Reduced Dataset
Original Dataset After PCA
Mathematics PC1
Physics
Dimensionality Reduction:

Original Dimension = 2
Reduced Dimension = 1

Variance Preserved = 100%
11. PCA Summary
Parameter Value
Number of Observations 5
Original Features 2
Mean Vector (6, 5)
Covariance Matrix 2 × 2
λ₁ 16
λ₂ 0
PC1 (0.7071, 0.7071)
PC1 Explained Variance 100%
PC2 Explained Variance 0%
Final Dimension 1
12. Final PCA Result
PCA DIMENSIONALITY REDUCTION
2 Original Features → 1 Principal Component
PC1 = 100%
The first principal component preserves all the variance of this dataset.
13. Conclusion
The original dataset contains two features: Mathematics and Physics.
The covariance matrix produces two eigenvalues:

λ₁ = 16
λ₂ = 0
Since λ₁ is much larger than λ₂, PC1 captures the complete variation of the dataset.
Therefore:

2-dimensional dataset → 1-dimensional dataset

with 100% variance preserved.
```

Monday, October 5, 2026

📊 CREDIT CARD DEFAULT DATASET


📊 CREDIT CARD DEFAULT DATASET

Detailed Description of the Dataset and Its Variables

1 Dataset Name
Default of Credit Card Clients Dataset

This dataset contains information about credit-card customers and their payment history.

The dataset contains customer information, credit information, repayment history, bill amounts, previous payment amounts and the customer's payment status for the following month.

2 Basic Dataset Information
Property Details
Dataset Name Default of Credit Card Clients
Total Records 30,000 customers
Original Columns 25 columns including ID and target
Customer Information Demographic and financial information
Time Period April 2005 to September 2005
Currency New Taiwan Dollar (NT$)
Target Information Default payment status for the next month
3 What Does One Row Mean?

Each row in the dataset represents information about one credit-card customer.

Row 1 → Customer 1
Row 2 → Customer 2
Row 3 → Customer 3
...
Row 30,000 → Customer 30,000

Therefore, there are 30,000 customer records in the dataset.

4 ID – Customer Identification Number
Column Meaning
ID Unique identification number assigned to the customer
Example:

ID = 1 means one particular customer record.
ID = 2 means another customer record.

The ID is used to identify records; it does not describe the customer's financial condition.
5 Credit Limit
```
Column Meaning

📊 DATASETS USED IN MACHINE LEARNING

🤖 Introduction to Machine Learning Datasets

A dataset is a structured collection of data used by Machine Learning algorithms to learn patterns, relationships, and useful information from examples. Datasets are the foundation of almost every Machine Learning project.

In Machine Learning, datasets generally contain features (input variables) and a target variable (output). Features provide the information used by a model, while the target represents the value that a supervised learning model attempts to predict.

📊

Features

Input variables or attributes used by a Machine Learning model to learn patterns from the data.

🎯

Target

The output variable that a supervised Machine Learning model attempts to predict.

🧹

Data Preparation

Data may require cleaning, transformation, encoding, scaling, and feature selection before model training.

🧠

Model Training

Machine Learning algorithms use training data to identify patterns and build predictive models.

💡 Explore the datasets below to study different Machine Learning problems such as classification, regression, clustering, dimensionality reduction, and feature selection.

Thursday, October 1, 2026

🧠 SVM Kernel Mathematics & Kernel Matrices

🧠 SVM Kernel Mathematics & Kernel Matrices

A Support Vector Machine (SVM) uses a kernel function to calculate the similarity between two observations. Instead of explicitly transforming the data into a higher-dimensional feature space, the kernel computes:

K(x,z) = φ(x)ᵀφ(z)

For n observations, the kernel calculations produce an n × n Gram matrix.

1️⃣ Five-Point Dataset

Point x₁ x₂ Class y
A12-1
B21-1
C45+1
D54+1
E66+1

Example Points

We will frequently calculate the kernel between:

A = (1,2)
C = (4,5)

Dot product:

A · C = (1)(4) + (2)(5)
= 4 + 10
= 14

Squared Euclidean distance:

||A-C||² = (1-4)² + (2-5)²
= 9 + 9
= 18

L1 distance:

||A-C||₁ = |1-4| + |2-5|
= 3 + 3
= 6

2️⃣ Linear Kernel

K(x,z) = xᵀz
For A and C:

K(A,C) = (1)(4) + (2)(5)
= 4 + 10
= 14

5 × 5 Linear Kernel Matrix

[
54141318
45131418
1413414054
1314404154
1818545472
]

3️⃣ Polynomial Kernel

K(x,z) = (γxᵀz + r)ᵈ

Take:

γ = 1    r = 1    d = 2
K(A,C)
= (1 × 14 + 1)²
= 15²
= 225

Polynomial Degree 2 Matrix

[
3625225196361
2536196225361
225196176416813025
196225168117643025
361361302530255329
]

4️⃣ RBF / Gaussian Kernel

K(x,z) = exp(-γ||x-z||²)

Take:

γ = 0.1
For A and C:

||A-C||² = 18
K(A,C) = e-0.1(18)
= e-1.8
≈ 0.1653

RBF Kernel Matrix

[
1.00000.81870.16530.13530.0166
0.81871.00000.13530.16530.0166
0.16530.13531.00000.81870.6065
0.13530.16530.81871.00000.6065
0.01660.01660.60650.60651.0000
]

5️⃣ Sigmoid Kernel

K(x,z) = tanh(γxᵀz + r)

Take:

γ = 0.1    r = 0
K(A,C)
= tanh(0.1 × 14)
= tanh(1.4)
≈ 0.8854

Sigmoid Kernel Matrix

[
0.46210.37990.88540.86170.9468
0.37990.46210.86170.88540.9468
0.88540.86170.99950.99931.0000
0.86170.88540.99930.99951.0000
0.94680.94681.00001.00001.0000
]

6️⃣ Laplacian Kernel

K(x,z) = exp(-γ||x-z||₁)
For A and C:

||A-C||₁ = |1-4| + |2-5|
= 3 + 3
= 6

K(A,C) = e-0.1(6)
= e-0.6
≈ 0.5488

Laplacian Kernel Matrix

[
1.00000.81870.54880.54880.4066
0.81871.00000.54880.54880.4066
0.54880.54881.00000.81870.7408
0.54880.54880.81871.00000.7408
0.40660.40660.74080.74081.0000
]

7️⃣ ANOVA Kernel

K(x,z) = [ Σᵢ exp(-γ(xᵢ-zᵢ)²) ]ᵈ

Take:

γ = 0.1    d = 2
For A = (1,2) and C = (4,5):

K(A,C) = [e-0.1(1-4)² + e-0.1(2-5)²]²
= [e-0.9 + e-0.9]²
= [2(0.40657)]²
≈ 0.6612

ANOVA Kernel Matrix

[
4.00003.27490.66120.76080.0806
3.27494.00000.76080.66120.0806
0.66120.76084.00003.27492.4811
0.76080.66123.27494.00002.4811
0.08060.08062.48112.48114.0000
]

8️⃣ Chi-Square Kernel

K(x,z) = exp[ -γ Σᵢ (xᵢ-zᵢ)²/(xᵢ+zᵢ) ]
For A and C:

Term 1:
(1-4)²/(1+4) = 9/5 = 1.8

Term 2:
(2-5)²/(2+5) = 9/7 ≈ 1.2857

Total:
1.8 + 1.2857 = 3.0857

K(A,C) = e-0.1(3.0857)
≈ 0.7345

Chi-Square Kernel Matrix

[
1.00000.93550.73450.71650.5728
0.93551.00000.71650.73450.5728
0.73450.71651.00000.97800.9521
0.71650.73450.97801.00000.9521
0.57280.57280.95210.95211.0000
]

9️⃣ Histogram Intersection Kernel

K(x,z) = Σᵢ min(xᵢ,zᵢ)
For A and C:

K(A,C) = min(1,4) + min(2,5)
= 1 + 2
= 3

Histogram Intersection Matrix

[
32333
23333
33989
33899
339912
]

🔟 Kernel Matrix Concept

For five observations:

X = {x₁,x₂,x₃,x₄,x₅}

the kernel matrix is:

K = [ K(x₁,x₁)   K(x₁,x₂)   ...   K(x₁,x₅)
K(x₂,x₁)   K(x₂,x₂)   ...   K(x₂,x₅)
⋮            ⋱          ⋮
K(x₅,x₁)   K(x₅,x₂)   ...   K(x₅,x₅) ]

Each entry represents the similarity between two observations. For example:

K₂,₄ = K(B,D)

The diagonal contains self-similarity:

K(xᵢ,xᵢ)

📊 Kernel Comparison

Kernel Formula K(A,C)
Linear xᵀz 14
Polynomial d=2 (xᵀz+1)² 225
Polynomial d=3 (xᵀz+1)³ 3375
RBF e-0.1||x-z||² 0.1653
Sigmoid tanh(0.1xᵀz) 0.8854
Laplacian e-0.1||x-z||₁ 0.5488
ANOVA [Σe-0.1(xᵢ-zᵢ)²]² 0.6612
Chi-Square exp[-γΣ((xᵢ-zᵢ)²/(xᵢ+zᵢ))] 0.7345
Histogram Intersection Σmin(xᵢ,zᵢ) 3
📌 Important:

The kernel matrix is also called the Gram matrix. For five observations, every kernel produces a 5 × 5 matrix.

For the SVM, these kernel values are used inside the dual optimization formulation:

max α: Σᵢ αᵢ − 1/2 ΣᵢΣⱼ αᵢαⱼ yᵢyⱼ K(xᵢ,xⱼ)

subject to:

αᵢ ≥ 0

Σᵢ αᵢyᵢ = 0

🤖 Support Vector Machine (SVM)

🤖 Support Vector Machine (SVM)

Support Vector Machine (SVM) is a supervised machine-learning algorithm used mainly for classification. It can also be used for regression through Support Vector Regression (SVR).

The central idea is to find a decision boundary that separates classes while maintaining an appropriate margin from the nearest observations.

1. What is SVM?

SVM searches for a decision boundary that separates observations belonging to different classes.

Class A Class B Decision Boundary
SVM separates observations using a decision boundary.

2. Hyperplane

The hyperplane is the mathematical decision boundary.

Hyperplane x-axis y
w · x + b = 0

3. Margin

The margin represents the separation between the decision boundary and the closest observations.

Maximum Margin Margin Boundary
Margin = 2 / ||w||

4. Support Vectors

Support vectors are the observations closest to the decision boundary.

Support Vectors Hyperplane

5. Maximum-Margin Principle

Several boundaries may separate the classes. SVM searches for an appropriate boundary with maximum separation from the nearest examples.

Boundary A Maximum-margin boundary

6. Hard Margin SVM

Hard-margin SVM requires the training observations to satisfy the separation constraints.

No training violations

7. Soft Margin SVM

Soft-margin SVM allows some observations to violate the ideal margin.

Violation
Minimize ½ ||w||² + C Σ ξᵢ

8. C Parameter

C controls the penalty associated with margin violations.

Small C Large C Wider tolerance Stricter penalty

9. Linear SVM

Linear SVM uses a straight decision boundary.

Linear Decision Boundary

10. Non-Linear SVM

When a straight line cannot separate the classes effectively, SVM can use a non-linear kernel.

Curved Boundary

11. Kernel Trick

Original Data
→
Kernel
→
Higher Feature Space
→
Linear Separation
Original Space Kernel Feature Space

12. Linear Kernel

K(x,z) = x · z
Vector x Vector z Dot Product

13. Polynomial Kernel

K(x,z) = (γ x · z + r)d
Polynomial Decision Boundary

14. RBF / Gaussian Kernel

K(x,z) = exp(-γ ||x-z||²)
Radial Influence

15. Sigmoid Kernel

K(x,z) = tanh(γ x · z + r)
Sigmoid Curve

16. Gamma Parameter

Low Gamma High Gamma Broader influence Local influence

17. SVM Workflow

Data Preprocess Scale Train SVM Prediction

18. Feature Scaling

SVM models can be affected when features have very different numerical scales.

Before Scaling After Scaling
z = (x - μ) / σ

19. Spam Detection

Email
→
Text Processing
→
TF-IDF
→
SVM
→
Spam / Ham
EMAIL "Win free prize" TF-IDF SVM SPAM

20. Classification Metrics

Accuracy

Correct

Precision

Positive

Recall

21. Confusion Matrix

Predicted Actual True Negative False Positive False Negative True Positive

22. Iris Classification

Setosa Versicolor Virginica

23. Advantages of SVM

📐 Maximum Margin

Uses a margin-based decision principle.

📊 High Dimensions

Can work with high-dimensional feature representations.

🔄 Kernels

Can model non-linear relationships.

📝 Text Data

Useful with sparse text feature representations.

Margin Dimensions Kernels Text

24. Limitations of SVM

Large Data Tuning Scaling
  • Large datasets can increase computational cost.
  • Kernel and hyperparameter selection may require experimentation.
  • Feature scaling is often important.
  • Interpretability may be lower than simple rule-based models.

25. Applications of SVM

SVM Spam Detection Image Classification Medical Data Fraud Detection

26. SVM Classification vs Regression

SVC SVR

27. Important SVM Terminology

SVM Hyperplane Kernel Margin Support Vector

28. Kernel Comparison

Kernel Visual Boundary Main Idea
Linear ──────── Straight boundary
Polynomial ∿∿∿ Polynomial relationship
RBF ◉ Radial similarity
Sigmoid S Sigmoid-shaped relationship

29. Practical SVM Workflow

Dataset Clean Scale Tune Evaluate

30. SVM Parameters at a Glance

SVM C Gamma Degree Kernel

31. SVM Decision Process

Input SVM Model Class

32. Hard Margin vs Soft Margin

Property Hard Margin Soft Margin
Training violations Not allowed Allowed
Slack variables No Yes
Noise tolerance Low Higher
Parameter C Not the main formulation Important
Hard Margin Soft Margin

33. Classification Pipeline

Training Data
→
Feature Engineering
→
Scaling
→
SVM
→
Metrics

34. SVM vs Other Classification Models

SVM Decision Tree KNN Logistic Naive Bayes

35. SVM Mathematical View

w · x + b = 0 -1 boundary +1 boundary
yᵢ(w · xᵢ + b) ≥ 1

36. Support Vector Geometry

Support vector Margin

37. SVM Prediction

New Data Decision Boundary

38. Advantages and Limitations Summary

Advantages ✓ Maximum-margin learning ✓ Kernel support ✓ High-dimensional data Limitations ⚠ Parameter tuning ⚠ Scaling required ⚠ Large-data cost

39. Interview Concept Map

SVM Hyperplane Margin Support Vector Kernel C / Gamma Prediction

40. Complete SVM Concept Diagram

SVM Hyperplane Support Vectors Kernel Margin C / Gamma
Core SVM idea: SVM finds an appropriate separating boundary, identifies the observations that most strongly determine that boundary, and balances margin size with training violations through its optimization parameters.

41. Important Interview Questions

Q1. What is SVM?

A supervised machine-learning algorithm that constructs a decision boundary for classification and can also be extended to regression.

Q2. What is a hyperplane?

A mathematical decision boundary separating observations in feature space.

Q3. What are support vectors?

Observations closest to the decision boundary that strongly influence the fitted SVM boundary.

Q4. What is margin?

The separation between the decision boundary and the nearest observations.

Q5. What is the kernel trick?

A technique that allows SVM to model non-linear relationships using kernel similarity functions.

Q6. What does C do?

C controls the penalty associated with observations violating the desired margin constraints.

Q7. What does gamma do?

Gamma controls the locality of influence for kernels such as RBF.

Q8. Why is scaling important?

Because the geometry used by SVM can be affected when features have very different numerical scales.