Computer Vision and Perception
Seeing with Algorithms
Image Representation
An image can be treated as a grid of pixel values, often with channels such as red, green, and blue. Models learn patterns from local textures, edges, shapes, and object-level structure.
Convolutional Neural Networks
Convolutional neural networks use filters that slide over the image to detect features. Early layers often learn edges and textures, while deeper layers learn more abstract patterns.
Vision Tasks
| Task | Output | Example |
|---|---|---|
| Classification | One label | Cat vs dog |
| Detection | Bounding boxes and labels | Find all cars in a street scene |
| Segmentation | Pixel-level masks | Separate tumor from background |
| Tracking | Object identity over time | Follow a moving person in video |
What is a convolution filter used for?
Convolution filters scan the image to detect useful local features.
Correct answer: Detecting local visual patterns
What is the difference between classification and detection?
Detection adds localization, usually with bounding boxes or masks.
Correct answer: Classification assigns a single label, while detection identifies objects and their locations.
Visual bias
Vision systems can inherit dataset bias, such as poor performance across lighting conditions, skin tones, camera types, or geographic settings.
Which task produces pixel-level masks?
Segmentation labels each pixel or region, not just the entire image.
Correct answer: Segmentation
Why are deeper layers in CNNs useful?
This hierarchical feature learning is one reason CNNs work well for vision.
Correct answer: They can combine simpler features into more complex visual patterns.