Feature Engineering and Data Preprocessing
Preparing Data for Learning
Why Preprocessing Matters
Raw data often contains missing values, inconsistent formats, outliers, and scales that make learning difficult. Preprocessing helps convert data into a form that supports reliable modeling.
Scaling and Encoding
Numerical features may be standardized using , while categorical features may be encoded as one-hot vectors or embeddings depending on the model.
Typical Data Pipeline
-
1
Step 1: Collect and inspect the raw dataset.
-
2
Step 2: Handle missing values and remove obvious errors.
-
3
Step 3: Encode categorical variables and scale numeric features.
-
4
Step 4: Split the data into training, validation, and test sets.
-
5
Step 5: Train and evaluate the model.
Why is feature scaling often used?
Scaling ensures that features with large numeric ranges do not dominate learning or distance calculations.
Correct answer: To put variables on comparable ranges
What is one-hot encoding?
One-hot encoding turns categorical values into machine-readable binary vectors.
Correct answer: A way of representing a category with a vector containing one 1 and the rest 0s.
Data leakage
Never use information from the validation or test set when preprocessing the training data in a way that would not be available at prediction time.
Which practice can cause data leakage?
Using the full dataset to compute preprocessing statistics can leak future information into training.
Correct answer: Computing preprocessing statistics using the full dataset before splitting
Preprocessing Goals