Unsupervised Learning and Clustering
Finding Structure Without Labels
Clustering
Clustering groups similar data points together without pre-existing labels. The goal is to discover natural structure, segments, or recurring patterns in the data.
K-Means Intuition
K-means tries to minimize within-cluster variance by repeatedly assigning points to the nearest centroid and then updating the centroids. It is simple but sensitive to initialization and the choice of .
Clustering Methods
K-means
- Fast and widely used
- Needs the number of clusters
- Best for roughly spherical clusters
Hierarchical clustering
- Builds nested groupings
- Does not require a fixed at the start
- Can be more expensive
What does unsupervised learning use?
Unsupervised learning finds structure in unlabeled data.
Correct answer: No labels or target outputs
What is the main objective of k-means?
K-means iteratively alternates between assignment and centroid updates.
Correct answer: To minimize within-cluster variance by assigning points to centroids.
Dimensionality Reduction
Dimensionality reduction compresses data into fewer variables while preserving as much useful structure as possible. Techniques such as principal component analysis can reveal hidden directions of variation.
A common use of dimensionality reduction is to:
Reducing dimensions can make data easier to inspect and sometimes easier to model.
Correct answer: Simplify data visualization and reduce noise
Why might clustering be useful in business?
Clustering helps identify groups with similar behavior or attributes.
Correct answer: It can reveal customer segments or patterns for targeted action.