Interpretation:
The four original features are transformed into principal
components. PC1 captures approximately 91.42% of the
total variation, making it the most important direction in
the dataset.
Final PCA Summary
Item
Result
Original Features
4
Mathematics Mean
6.0
Physics Mean
5.0
Chemistry Mean
6.6
Biology Mean
8.2
PC1 Eigenvalue
21.5758
PC2 Eigenvalue
2.0053
PC3 Eigenvalue
0.0189
PC4 Eigenvalue
0.0000
PC1 Variance Explained
91.42%
PC2 Variance Explained
8.50%
PC3 Variance Explained
0.08%
PC4 Variance Explained
0.00%
Reduced Dimension
4 → 1
4 FEATURES → 1 PRINCIPAL COMPONENT
PC1 retains approximately 91.42% of the total variance.
Dimensionality Reduction:
The original dataset contains 3 features. After PCA, the major
information can be represented using 1 principal component with
approximately 92.88% variance retention.
STEP 07
Deriving the New Dataset
PC1 Score = Xcentered × PC1
Student
PC1 Calculation
PC1 Value
A
(−4)(0.64065)+(−4)(0.64065)+(−1.6)(0.42325)
−5.8024
B
(−2)(0.64065)+(−2)(0.64065)+(−2.6)(0.42325)
−3.6630
C
(0)(0.64065)+(0)(0.64065)+(1.4)(0.42325)
0.5925
D
(2)(0.64065)+(2)(0.64065)+(−0.6)(0.42325)
2.3087
E
(4)(0.64065)+(4)(0.64065)+(3.4)(0.42325)
6.5642
FINAL PCA DATASET
Student
Mathematics
Physics
Chemistry
PC1
A
2
1
5
−5.8024
B
4
3
4
−3.6630
C
6
5
8
0.5925
D
8
7
6
2.3087
E
10
9
10
6.5642
PCA Visualization
3-Feature Dataset — PC1 Projection
Interpretation:
PC1 combines Mathematics, Physics and Chemistry into a single
new feature. Because PC1 explains approximately 92.88%
of the variance, most of the information in the original
three-dimensional dataset is retained.
Final PCA Summary
Item
Result
Original Features
3
Mathematics Mean
6.0
Physics Mean
5.0
Chemistry Mean
6.6
PC1 Eigenvalue
19.1711
PC2 Eigenvalue
1.4689
PC3 Eigenvalue
0.0000
PC1 Variance Explained
92.88%
PC2 Variance Explained
7.12%
PC3 Variance Explained
0.00%
Reduced Dimension
3 → 1
3 FEATURES → 1 PRINCIPAL COMPONENT
PC1 retains approximately 92.88% of the total variance.
Two-Feature Numerical Example Using Gate Vidyalay PCA Steps
1. Get the Data
We have the marks of five students in two subjects:
Mathematics and Physics.
The objective is to transform the original two-dimensional data
into a new coordinate system and represent it mainly using the
first principal component.
Student
Mathematics (X)
Physics (Y)
A
2
1
B
4
3
C
6
5
D
8
7
E
10
9
X =
[ 2 1 ]
[ 4 3 ]
[ 6 5 ]
[ 8 7 ]
[10 9 ]
STEP 01
Calculate the Mean Vector
According to the PCA procedure, first calculate the mean of
each feature.
Detailed Description of the Dataset and Its Variables
1
Dataset Name
Default of Credit Card Clients Dataset
This dataset contains information about credit-card customers
and their payment history.
The dataset contains customer information, credit information,
repayment history, bill amounts, previous payment amounts and
the customer's payment status for the following month.
2
Basic Dataset Information
Property
Details
Dataset Name
Default of Credit Card Clients
Total Records
30,000 customers
Original Columns
25 columns including ID and target
Customer Information
Demographic and financial information
Time Period
April 2005 to September 2005
Currency
New Taiwan Dollar (NT$)
Target Information
Default payment status for the next month
3
What Does One Row Mean?
Each row in the dataset represents information about
one credit-card customer.
A dataset is a structured collection of data used by
Machine Learning algorithms to learn patterns, relationships, and
useful information from examples. Datasets are the foundation of
almost every Machine Learning project.
In Machine Learning, datasets generally contain
features (input variables) and a
target variable (output). Features provide the
information used by a model, while the target represents the value
that a supervised learning model attempts to predict.
📊
Features
Input variables or attributes used by a Machine Learning model
to learn patterns from the data.
🎯
Target
The output variable that a supervised Machine Learning model
attempts to predict.
🧹
Data Preparation
Data may require cleaning, transformation, encoding, scaling,
and feature selection before model training.
🧠
Model Training
Machine Learning algorithms use training data to identify
patterns and build predictive models.
💡 Explore the datasets below to study different
Machine Learning problems such as classification, regression,
clustering, dimensionality reduction, and feature selection.
A Support Vector Machine (SVM) uses a kernel function to calculate
the similarity between two observations.
Instead of explicitly transforming the data into a higher-dimensional
feature space, the kernel computes:
K(x,z) = φ(x)ᵀφ(z)
For n observations, the kernel calculations produce an
n × n Gram matrix.
Support Vector Machine (SVM) is a supervised machine-learning
algorithm used mainly for classification. It can also be used for
regression through Support Vector Regression (SVR).
The central idea is to find a decision boundary that separates classes
while maintaining an appropriate margin from the nearest observations.
1. What is SVM?
SVM searches for a decision boundary that separates observations belonging
to different classes.
SVM separates observations using a decision boundary.
2. Hyperplane
The hyperplane is the mathematical decision boundary.
w · x + b = 0
3. Margin
The margin represents the separation between the decision boundary and
the closest observations.
Margin = 2 / ||w||
4. Support Vectors
Support vectors are the observations closest to the decision boundary.
5. Maximum-Margin Principle
Several boundaries may separate the classes. SVM searches for an
appropriate boundary with maximum separation from the nearest examples.
6. Hard Margin SVM
Hard-margin SVM requires the training observations to satisfy the
separation constraints.
7. Soft Margin SVM
Soft-margin SVM allows some observations to violate the ideal margin.
Minimize ½ ||w||² + C Σ ξᵢ
8. C Parameter
C controls the penalty associated with margin violations.
9. Linear SVM
Linear SVM uses a straight decision boundary.
10. Non-Linear SVM
When a straight line cannot separate the classes effectively, SVM can
use a non-linear kernel.
11. Kernel Trick
Original Data
→
Kernel
→
Higher Feature Space
→
Linear Separation
12. Linear Kernel
K(x,z) = x · z
13. Polynomial Kernel
K(x,z) = (γ x · z + r)d
14. RBF / Gaussian Kernel
K(x,z) = exp(-γ ||x-z||²)
15. Sigmoid Kernel
K(x,z) = tanh(γ x · z + r)
16. Gamma Parameter
17. SVM Workflow
18. Feature Scaling
SVM models can be affected when features have very different numerical
scales.
z = (x - μ) / σ
19. Spam Detection
Email
→
Text Processing
→
TF-IDF
→
SVM
→
Spam / Ham
20. Classification Metrics
Accuracy
Precision
Recall
21. Confusion Matrix
22. Iris Classification
23. Advantages of SVM
📐 Maximum Margin
Uses a margin-based decision principle.
📊 High Dimensions
Can work with high-dimensional feature representations.
🔄 Kernels
Can model non-linear relationships.
📝 Text Data
Useful with sparse text feature representations.
24. Limitations of SVM
Large datasets can increase computational cost.
Kernel and hyperparameter selection may require experimentation.
Feature scaling is often important.
Interpretability may be lower than simple rule-based models.
25. Applications of SVM
26. SVM Classification vs Regression
27. Important SVM Terminology
28. Kernel Comparison
Kernel
Visual Boundary
Main Idea
Linear
────────
Straight boundary
Polynomial
∿∿∿
Polynomial relationship
RBF
◉
Radial similarity
Sigmoid
S
Sigmoid-shaped relationship
29. Practical SVM Workflow
30. SVM Parameters at a Glance
31. SVM Decision Process
32. Hard Margin vs Soft Margin
Property
Hard Margin
Soft Margin
Training violations
Not allowed
Allowed
Slack variables
No
Yes
Noise tolerance
Low
Higher
Parameter C
Not the main formulation
Important
33. Classification Pipeline
Training Data
→
Feature Engineering
→
Scaling
→
SVM
→
Metrics
34. SVM vs Other Classification Models
35. SVM Mathematical View
yᵢ(w · xᵢ + b) ≥ 1
36. Support Vector Geometry
37. SVM Prediction
38. Advantages and Limitations Summary
39. Interview Concept Map
40. Complete SVM Concept Diagram
Core SVM idea:
SVM finds an appropriate separating boundary, identifies the observations
that most strongly determine that boundary, and balances margin size with
training violations through its optimization parameters.
41. Important Interview Questions
Q1. What is SVM?
A supervised machine-learning algorithm that constructs a decision
boundary for classification and can also be extended to regression.
Q2. What is a hyperplane?
A mathematical decision boundary separating observations in feature space.
Q3. What are support vectors?
Observations closest to the decision boundary that strongly influence
the fitted SVM boundary.
Q4. What is margin?
The separation between the decision boundary and the nearest observations.
Q5. What is the kernel trick?
A technique that allows SVM to model non-linear relationships using
kernel similarity functions.
Q6. What does C do?
C controls the penalty associated with observations violating the
desired margin constraints.
Q7. What does gamma do?
Gamma controls the locality of influence for kernels such as RBF.
Q8. Why is scaling important?
Because the geometry used by SVM can be affected when features have
very different numerical scales.