Total Pageviews

Sunday, September 27, 2026

๐ŸŽฏ RFE Estimator Model Guide

๐ŸŽฏ RFE Estimator Model Guide

Regression and Classification Estimators Used with Recursive Feature Elimination

Recursive Feature Elimination (RFE) is a wrapper-based feature selection technique. It repeatedly trains an estimator, evaluates feature importance, removes the least important features, and continues until the desired number of features remains.

RFE generally requires the estimator to expose either coefficients or feature importance values.

Feature Importance → |Coefficient| or Feature Importance

๐Ÿ“ˆ Regression Estimators for RFE

REGRESSION

Linear Regression

Models the relationship between independent variables and a continuous target using a linear equation.

ลท = ฮฒ₀ + ฮฒ₁X₁ + ฮฒ₂X₂ + ... + ฮฒโ‚šXโ‚š

RFE importance: Absolute coefficient |ฮฒ|

Strength: Simple and interpretable.

Limitation: Assumes a linear relationship.

RFE: Excellent

REGRESSION

Ridge Regression

Linear regression with L2 regularization to reduce the effect of large coefficients.

Loss = RSS + ฮฑฮฃฮฒ²

RFE importance: Absolute coefficient.

Strength: Handles multicollinearity better.

Limitation: Usually does not make coefficients exactly zero.

RFE: Excellent

REGRESSION

Lasso Regression

Linear regression using L1 regularization, which can shrink some coefficients to zero.

Loss = RSS + ฮฑฮฃ|ฮฒ|

RFE importance: Absolute coefficient.

Strength: Performs implicit feature selection.

Limitation: Can select one variable among highly correlated variables.

RFE: Excellent

REGRESSION

Elastic Net

Combines L1 and L2 regularization.

Loss = RSS + ฮฑ₁ฮฃ|ฮฒ| + ฮฑ₂ฮฃฮฒ²

RFE importance: Absolute coefficient.

Strength: Useful with correlated features.

Limitation: Requires tuning regularization parameters.

RFE: Excellent

REGRESSION

Decision Tree Regressor

Predicts continuous values by recursively splitting observations into groups.

RFE importance: Impurity reduction.

Strength: Captures nonlinear relationships.

Limitation: Individual trees can overfit.

RFE: Excellent

REGRESSION

Random Forest Regressor

Combines many decision trees and averages their predictions.

RFE importance: Feature importance from trees.

Strength: Handles nonlinear relationships and interactions.

Limitation: Feature importance can be biased toward some variables.

RFE: Excellent

REGRESSION

Extra Trees Regressor

Uses extremely randomized decision trees to model nonlinear relationships.

RFE importance: Tree-based feature importance.

Strength: Strong randomization and nonlinear modeling.

Limitation: Less directly interpretable than linear models.

RFE: Excellent

REGRESSION

Gradient Boosting Regressor

Builds trees sequentially, where each new tree attempts to improve the previous model.

RFE importance: Tree-based importance.

Strength: Powerful nonlinear prediction.

Limitation: More parameters and computational complexity.

RFE: Excellent

REGRESSION

SVR — Linear Kernel

Support Vector Regression finds a function that fits observations within an acceptable error margin.

Minimize model complexity while controlling ฮต-error

RFE importance: Coefficients for a linear kernel.

Strength: Effective for high-dimensional data.

Limitation: Requires careful parameter tuning.

RFE: Suitable with linear kernel

๐Ÿท️ Classification Estimators for RFE

CLASSIFICATION

Logistic Regression

Estimates the probability of belonging to one or more classes.

P(Y=1|X) = 1 / (1 + e⁻แถป)

RFE importance: Absolute coefficient.

Strength: Highly interpretable linear classifier.

Limitation: Limited for strongly nonlinear relationships.

RFE: Excellent

CLASSIFICATION

Decision Tree Classifier

Classifies observations using a sequence of decision-making rules.

RFE importance: Impurity reduction.

Strength: Easy to understand and handles nonlinear relationships.

Limitation: Can overfit without appropriate controls.

RFE: Excellent

CLASSIFICATION

Random Forest Classifier

Combines multiple classification trees using randomized sampling and feature selection.

RFE importance: Tree feature importance.

Strength: Captures nonlinear relationships and interactions.

Limitation: Less interpretable than a single tree.

RFE: Excellent

CLASSIFICATION

Extra Trees Classifier

Uses highly randomized decision trees to classify observations.

RFE importance: Tree feature importance.

Strength: Good nonlinear feature detection.

Limitation: Feature rankings may vary depending on data.

RFE: Excellent

CLASSIFICATION

Gradient Boosting Classifier

Builds a sequence of decision trees to progressively improve classification performance.

RFE importance: Tree-based importance.

Strength: Powerful nonlinear classification.

Limitation: Requires parameter tuning.

RFE: Excellent

CLASSIFICATION

AdaBoost Classifier

Combines several weak learners by giving additional emphasis to incorrectly classified observations.

RFE importance: Estimator-based feature importance.

Strength: Effective ensemble approach.

Limitation: Can be sensitive to noisy observations.

RFE: Suitable

CLASSIFICATION

Linear SVM

Finds a separating hyperplane that maximizes the margin between classes.

w · x + b = 0

RFE importance: Absolute coefficient |w|.

Strength: Effective in high-dimensional spaces.

Limitation: Linear SVM cannot directly model complex nonlinear boundaries.

RFE: Excellent

CLASSIFICATION

Ridge Classifier

A linear classifier using L2 regularization.

RFE importance: Absolute coefficient.

Strength: Useful when predictors are correlated.

Limitation: Produces a linear decision boundary.

RFE: Suitable

CLASSIFICATION

SGD Classifier

Uses stochastic gradient-based optimization to train linear classification models.

RFE importance: Coefficients.

Strength: Efficient for large datasets.

Limitation: Results depend on optimization and parameter settings.

RFE: Suitable

CLASSIFICATION

Perceptron

A basic linear classifier that learns a separating decision boundary.

RFE importance: Coefficients.

Strength: Simple and computationally efficient.

Limitation: Works best when classes are approximately linearly separable.

RFE: Suitable

CLASSIFICATION

Passive-Aggressive Classifier

An online learning algorithm that updates its model when a prediction violates the desired margin.

RFE importance: Coefficients.

Strength: Suitable for streaming and large datasets.

Limitation: Sensitive to parameter selection and noisy data.

RFE: Suitable

CLASSIFICATION

Linear Discriminant Analysis

Finds linear combinations of features that separate different classes.

RFE importance: Coefficient-based importance where supported.

Strength: Useful for relatively well-separated classes.

Limitation: Relies on assumptions about class distributions.

RFE: Suitable

๐Ÿ“Š RFE Estimator Comparison

Estimator Problem Importance Used by RFE Major Advantage Main Limitation
Linear Regression Regression |Coefficient| Simple and interpretable Linear assumption
Ridge Regression Regression |Coefficient| Handles multicollinearity Does not perform sparse selection by itself
Lasso Regression Regression |Coefficient| Sparse coefficients Can be unstable with correlated predictors
Elastic Net Regression |Coefficient| L1 + L2 regularization More parameters
Decision Tree Regressor Regression Feature importance Nonlinear relationships Can overfit
Random Forest Regressor Regression Tree importance Nonlinear + interactions Less interpretable
Extra Trees Regressor Regression Tree importance Strong randomization Less interpretable
Gradient Boosting Regressor Regression Tree importance Powerful nonlinear model Parameter tuning
Logistic Regression Classification |Coefficient| Interpretable Linear boundary
Decision Tree Classifier Classification Tree importance Easy to understand Can overfit
Random Forest Classifier Classification Tree importance Nonlinear modeling Less interpretable
Extra Trees Classifier Classification Tree importance Highly randomized trees Less interpretable
Gradient Boosting Classifier Classification Tree importance Strong predictive modeling Requires tuning
AdaBoost Classifier Classification Estimator importance Boosting approach Sensitive to noise
Linear SVM Classification |Coefficient| Good for high-dimensional data Linear boundary
Ridge Classifier Classification |Coefficient| Regularized linear model Linear boundary
SGD Classifier Classification Coefficients Scales to large datasets Optimization sensitive
Perceptron Classification Coefficients Simple and fast Linear separation
Passive-Aggressive Classifier Classification Coefficients Online learning Sensitive to noise
Linear Discriminant Analysis Classification Linear coefficients Dimensional discrimination Distribution assumptions

๐Ÿ”„ General RFE Process

1️⃣ Start with all features
2️⃣ Train estimator
3️⃣ Calculate importance
4️⃣ Rank features
5️⃣ Remove weakest features
6️⃣ Repeat until target number
Important: RFE is a wrapper method. The selected features depend on the estimator used, the data, preprocessing, hyperparameters, and the requested number of features. Therefore, two different estimators can produce different RFE rankings for the same dataset.

No comments:

Post a Comment