PRINCIPAL COMPONENT ANALYSIS (PCA)
Two-Feature Numerical Example Using Gate Vidyalay PCA Steps
1. Get the Data
We have the marks of five students in two subjects:
Mathematics and Physics.
The objective is to transform the original two-dimensional data
into a new coordinate system and represent it mainly using the
first principal component.
| Student | Mathematics (X) | Physics (Y) |
|---|---|---|
| A | 2 | 1 |
| B | 4 | 3 |
| C | 6 | 5 |
| D | 8 | 7 |
| E | 10 | 9 |
X =
[ 2 1 ]
[ 4 3 ]
[ 6 5 ]
[ 8 7 ]
[10 9 ]
[ 4 3 ]
[ 6 5 ]
[ 8 7 ]
[10 9 ]
STEP 01
Calculate the Mean Vector
According to the PCA procedure, first calculate the mean of
each feature.
Mean of Mathematics
μX =
(2 + 4 + 6 + 8 + 10) / 5
= 6
Mean of Physics
μY =
(1 + 3 + 5 + 7 + 9) / 5
= 5
Mean Vector:
μ = [ 6 5 ]
μ = [ 6 5 ]
STEP 02
Subtract Mean Vector from the Data
For every observation, subtract the mean vector:
xi − μ
| Student | Original X | Original Y | X − 6 | Y − 5 |
|---|---|---|---|---|
| A | 2 | 1 | -4 | -4 |
| B | 4 | 3 | -2 | -2 |
| C | 6 | 5 | 0 | 0 |
| D | 8 | 7 | 2 | 2 |
| E | 10 | 9 | 4 | 4 |
X − μ =
[ -4 -4 ]
[ -2 -2 ]
[ 0 0 ]
[ 2 2 ]
[ 4 4 ]
[ -2 -2 ]
[ 0 0 ]
[ 2 2 ]
[ 4 4 ]
STEP 03
Calculate the Covariance Matrix
Following the calculation used in the Gate Vidyalay worked example,
the covariance matrix is calculated using:
C =
(1 / n)
Σ (xi − μ)(xi − μ)T
Here,
n = 5
Covariance of Mathematics with Mathematics
Cov(X,X) =
[16 + 4 + 0 + 4 + 16] / 5
= 40 / 5
= 8
Covariance of Mathematics with Physics
Cov(X,Y) =
[16 + 4 + 0 + 4 + 16] / 5
= 40 / 5
= 8
Covariance of Physics with Physics
Cov(Y,Y) =
[16 + 4 + 0 + 4 + 16] / 5
= 40 / 5
= 8
Covariance Matrix:
C =
[ 8 8 ]
[ 8 8 ]
[ 8 8 ]
STEP 04
Calculate Eigenvalues and Eigenvectors
An eigenvalue λ satisfies the characteristic equation:
|C − λI| = 0
| 8−λ 8 |
| 8 8−λ |
| 8 8−λ |
(8 − λ)(8 − λ) − 64 = 0
64 − 16λ + λ² − 64 = 0
λ² − 16λ = 0
λ(λ − 16) = 0
Eigenvalues:
λ1 = 16
λ2 = 0
λ1 = 16
λ2 = 0
Eigenvector corresponding to λ₁ = 16
CX = λX
[ 8 8 ] [X₁] [16X₁]
[ 8 8 ] [X₂] [16X₂]
[ 8 8 ] [X₂] [16X₂]
8X₁ + 8X₂ = 16X₁
8X₂ = 8X₁
X₁ = X₂
Therefore, the eigenvector direction is:
[ 1 ]
[ 1 ]
[ 1 ]
Normalize the Eigenvector
√(1² + 1²) = √2
PC1 =
[ 1/√2 ]
[ 1/√2 ] = [ 0.7071 ]
[ 0.7071 ]
[ 1/√2 ] = [ 0.7071 ]
[ 0.7071 ]
Eigenvector corresponding to λ₂ = 0
PC2 =
[ 1/√2 ]
[-1/√2 ] = [ 0.7071 ]
[-0.7071]
[-1/√2 ] = [ 0.7071 ]
[-0.7071]
STEP 05
Choose the Principal Component
| Component | Eigenvalue | Variance Contribution |
|---|---|---|
| PC1 | 16 | 100% |
| PC2 | 0 | 0% |
Total Variance = 16 + 0 = 16
Explained Variance of PC1
=
(16 / 16) × 100
=
100%
PC1 captures 100% of the variance.
Therefore, PC2 can be discarded without losing variance in this particular dataset.
Therefore, PC2 can be discarded without losing variance in this particular dataset.
STEP 06
Form the Feature Vector
Since PC1 has the largest eigenvalue, it is selected as the
principal component.
Feature Vector =
[ 0.7071 ]
[ 0.7071 ]
[ 0.7071 ]
STEP 07
Derive the New Dataset
The new PCA value is obtained by projecting each centered data point
onto the selected eigenvector:
Z = PC1T(X − μ)
Student A
Z₁ =
(0.7071 × -4)
+
(0.7071 × -4)
Z₁ = -5.6569
Z₁ = -5.6569
Student B
Z₂ =
(0.7071 × -2)
+
(0.7071 × -2)
Z₂ = -2.8284
Z₂ = -2.8284
Student C
Z₃ =
(0.7071 × 0)
+
(0.7071 × 0)
Z₃ = 0
Z₃ = 0
Student D
Z₄ =
(0.7071 × 2)
+
(0.7071 × 2)
Z₄ = 2.8284
Z₄ = 2.8284
Student E
Z₅ =
(0.7071 × 4)
+
(0.7071 × 4)
Z₅ = 5.6569
Z₅ = 5.6569
8. Original Data Plot and Principal Component
The red line represents the direction of the first principal component.
The blue points represent the original observations.
Mathematics vs Physics with PC1 Direction
9. Final Dataset with PCA Values
The original two features are shown together with the new
one-dimensional PCA representation.
| Student | Mathematics | Physics | PC1 Value |
|---|---|---|---|
| A | 2 | 1 | -5.6569 |
| B | 4 | 3 | -2.8284 |
| C | 6 | 5 | 0.0000 |
| D | 8 | 7 | 2.8284 |
| E | 10 | 9 | 5.6569 |
10. Original Dataset vs Reduced Dataset
| Original Dataset | After PCA |
|---|---|
| Mathematics | PC1 |
| Physics |
Dimensionality Reduction:
Original Dimension = 2
Reduced Dimension = 1
Variance Preserved = 100%
Original Dimension = 2
Reduced Dimension = 1
Variance Preserved = 100%
11. PCA Summary
| Parameter | Value |
|---|---|
| Number of Observations | 5 |
| Original Features | 2 |
| Mean Vector | (6, 5) |
| Covariance Matrix | 2 × 2 |
| λ₁ | 16 |
| λ₂ | 0 |
| PC1 | (0.7071, 0.7071) |
| PC1 Explained Variance | 100% |
| PC2 Explained Variance | 0% |
| Final Dimension | 1 |
12. Final PCA Result
PCA DIMENSIONALITY REDUCTION
2 Original Features → 1 Principal Component
PC1 = 100%
The first principal component preserves all the variance
of this dataset.
13. Conclusion
The original dataset contains two features:
Mathematics and Physics.
The covariance matrix produces two eigenvalues:
λ₁ = 16
λ₂ = 0
λ₁ = 16
λ₂ = 0
Since λ₁ is much larger than λ₂, PC1 captures the complete
variation of the dataset.
Therefore:
2-dimensional dataset → 1-dimensional dataset
with 100% variance preserved.
2-dimensional dataset → 1-dimensional dataset
with 100% variance preserved.
No comments:
Post a Comment