Computer vision basics: how machines learn to see
How images become numbers, how convolutional networks learn edges, shapes and objects, the difference between classification, detection and segmentation, and where computer vision is used in India.
Computer vision is the branch of AI that lets machines understand images and video. A computer sees an image as a grid of numbers, and a neural network learns, from many labelled examples, which patterns of numbers mean an edge, a shape and finally an object. The same idea powers face unlock, number-plate readers and crop disease apps.
This guide explains how that works without heavy maths, walks through the key ideas and milestones, and shows where computer vision is being used in India today.
How does a computer see an image?
To a computer, a photo is a grid of tiny squares called pixels. Each pixel is stored as numbers. In a colour image, each pixel usually has three numbers for red, green and blue, each from 0 to 255. A single 1080p photo is roughly two million pixels, which means around six million numbers.
The challenge is enormous. The same cow looks completely different in bright sun, at dusk, from the side, half hidden behind a cart. The numbers change entirely, but a person still instantly says “cow”. Computer vision is about teaching machines that same robustness.
What are the main computer vision tasks?
How did computers get good at this?
For decades, engineers hand-designed rules and features, such as edge detectors and colour histograms, and fed them to simpler models. It worked for narrow tasks but broke easily.
Two things changed that. The first was data: ImageNet, a research dataset organised by concepts, now indexes over 14 million labelled images. The second was deep learning on GPUs. In 2012, a deep convolutional neural network, now known as AlexNet, won the ImageNet competition with a top-5 error rate of 15.3 percent, against 26.2 percent for the second-best entry. That gap convinced the field.
Progress continued with deeper networks such as ResNet (2015), which made very deep networks trainable, and later with vision transformers, which apply the same attention mechanism used in language models. Our explainer on how transformers work covers that side.
How does a convolutional neural network learn?
A convolutional neural network (CNN) slides small filters across the image, like a magnifying glass moving over a photo. Each filter responds to a particular pattern.
Nobody programs these filters. The network starts with random ones and adjusts them over thousands of labelled examples, keeping changes that reduce its mistakes. This is why data quality matters so much: the network learns whatever the examples teach, including their biases.
How do detection and segmentation work?
Classification gives one label per image. Real applications usually need more. Object detection finds each object and draws a box around it. The YOLO (“You Only Look Once”) approach made this fast enough for real-time video by predicting all boxes in a single pass through the network.
Segmentation goes further and labels every pixel. U-Net, originally designed for biomedical images, became a standard architecture for this, and variations are used well beyond medicine, for example to map roads and fields in satellite images.
Do you need millions of images to train a model?
Usually not, thanks to transfer learning. You start from a model already trained on a large dataset, which has learned general features such as edges and shapes, and fine-tune it on your own smaller dataset. A few hundred good, well-labelled images per class can be enough for a useful first version of many tasks.
Data augmentation helps too: flipping, rotating, cropping and changing the brightness of your images so the model learns to cope with variation instead of memorising your exact photos.
The most common failure in real projects is a mismatch between training images and real use. A model trained on clear, well-lit photos will struggle with blurry phone pictures taken in a dim shed. Collect training data in the conditions where the model will actually be used.
Where is computer vision used in India?
What does a first project look like?
A good first project is small, local and yours. Suppose you want a model that tells healthy tomato leaves from diseased ones.
That one project touches data collection, training, evaluation and deployment, which is exactly what interviewers ask about.
What are the limits and responsibilities?
Vision models can be biased if their training data under-represents certain lighting conditions, skin tones or regions. They can be fooled by small, deliberate changes to an image. And images of people are personal data: if you build anything involving faces or identifiable people, privacy law applies. Our guide to the DPDP Act for developers explains what that means in practice.
How can you start learning computer vision?
Start with Python and basic image handling using a library such as OpenCV. Then learn the fundamentals of neural networks, train a small CNN on a standard dataset, and finally fine-tune a pretrained model on images you collect yourself, such as leaves from your garden or products from a local shop.
Computer vision sits on top of programming, maths and machine learning, so the order matters. You can see where it fits in a complete path in the full Program Zero curriculum, where Phase 7 covers CNNs from LeNet to ResNet, transfer learning and deep learning in PyTorch, after the programming, data and maths phases that make it understandable.
Want to learn this live, with mentors?
Program Zero takes you from Python to deep learning, CNNs and language models over 18 months of live, mentor-led classes, with projects every quarter. ₹5,999 for all 18 months; the batch starts 9 January 2027.
Frequently asked questions
Related reading
We teach this, live
Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.