PROGRAM ZERO 18-month live AI, LLM & full-stack programme · ₹5,999 for all 18 months · Starts 9 January 2027
Explore Program Zero →
Careers Ninza — business and startup leadership training
JOIN ZERO NINZA KIDS
JOIN ZERO NINZA KIDS
Careers Ninza
AI & AUTOMATION 9 min read · Updated 24 September 2026

Computer vision basics: how machines learn to see

How images become numbers, how convolutional networks learn edges, shapes and objects, the difference between classification, detection and segmentation, and where computer vision is used in India.

CN
Careers Ninza AI faculty
Careers Ninza · Kolkata, India

Computer vision is the branch of AI that lets machines understand images and video. A computer sees an image as a grid of numbers, and a neural network learns, from many labelled examples, which patterns of numbers mean an edge, a shape and finally an object. The same idea powers face unlock, number-plate readers and crop disease apps.

This guide explains how that works without heavy maths, walks through the key ideas and milestones, and shows where computer vision is being used in India today.

How does a computer see an image?

To a computer, a photo is a grid of tiny squares called pixels. Each pixel is stored as numbers. In a colour image, each pixel usually has three numbers for red, green and blue, each from 0 to 255. A single 1080p photo is roughly two million pixels, which means around six million numbers.

The challenge is enormous. The same cow looks completely different in bright sun, at dusk, from the side, half hidden behind a cart. The numbers change entirely, but a person still instantly says “cow”. Computer vision is about teaching machines that same robustness.

What are the main computer vision tasks?

TaskQuestion it answersExample
ClassificationWhat is in this image?Is this leaf healthy or diseased?
Object detectionWhat objects are here, and where?Boxes around every vehicle in a traffic camera frame
SegmentationWhich exact pixels belong to what?Outlining a tumour region in a scan
OCRWhat text is written here?Reading a printed form in Hindi or Tamil
Pose estimationWhere are the body’s joints?Checking exercise form in a fitness app
TrackingWhere does this object move across frames?Counting footfall in a store

How did computers get good at this?

For decades, engineers hand-designed rules and features, such as edge detectors and colour histograms, and fed them to simpler models. It worked for narrow tasks but broke easily.

Two things changed that. The first was data: ImageNet, a research dataset organised by concepts, now indexes over 14 million labelled images. The second was deep learning on GPUs. In 2012, a deep convolutional neural network, now known as AlexNet, won the ImageNet competition with a top-5 error rate of 15.3 percent, against 26.2 percent for the second-best entry. That gap convinced the field.

Progress continued with deeper networks such as ResNet (2015), which made very deep networks trainable, and later with vision transformers, which apply the same attention mechanism used in language models. Our explainer on how transformers work covers that side.

How does a convolutional neural network learn?

A convolutional neural network (CNN) slides small filters across the image, like a magnifying glass moving over a photo. Each filter responds to a particular pattern.

—Early layers learn simple patterns: edges, corners, patches of colour.
—Middle layers combine those into textures and parts: an eye, a wheel, the vein pattern of a leaf.
—Later layers combine parts into whole objects: a face, a car, a diseased leaf.

Nobody programs these filters. The network starts with random ones and adjusts them over thousands of labelled examples, keeping changes that reduce its mistakes. This is why data quality matters so much: the network learns whatever the examples teach, including their biases.

How do detection and segmentation work?

Classification gives one label per image. Real applications usually need more. Object detection finds each object and draws a box around it. The YOLO (“You Only Look Once”) approach made this fast enough for real-time video by predicting all boxes in a single pass through the network.

Segmentation goes further and labels every pixel. U-Net, originally designed for biomedical images, became a standard architecture for this, and variations are used well beyond medicine, for example to map roads and fields in satellite images.

Do you need millions of images to train a model?

Usually not, thanks to transfer learning. You start from a model already trained on a large dataset, which has learned general features such as edges and shapes, and fine-tune it on your own smaller dataset. A few hundred good, well-labelled images per class can be enough for a useful first version of many tasks.

Data augmentation helps too: flipping, rotating, cropping and changing the brightness of your images so the model learns to cope with variation instead of memorising your exact photos.

The most common failure in real projects is a mismatch between training images and real use. A model trained on clear, well-lit photos will struggle with blurry phone pictures taken in a dim shed. Collect training data in the conditions where the model will actually be used.

Where is computer vision used in India?

—Agriculture. Apps that identify crop diseases or pests from a farmer’s phone photo.
—Traffic. Automatic number-plate recognition and vehicle counting.
—Manufacturing. Spotting defects on production lines faster than manual inspection.
—Healthcare. Tools that support doctors in screening medical images, alongside, not instead of, clinical judgement.
—Documents. Reading printed and handwritten forms in Indian scripts, which adds its own challenges.
—Retail. Shelf monitoring and footfall analysis.

What does a first project look like?

A good first project is small, local and yours. Suppose you want a model that tells healthy tomato leaves from diseased ones.

—Collect a few hundred photos of each kind with a phone, in the light and angles a farmer would actually use.
—Split them into training, validation and test sets, keeping the test set untouched until the end.
—Start from a pretrained CNN and fine-tune it on your images, with augmentation such as flips and brightness changes.
—Check accuracy on the test set, then look at the mistakes. They usually reveal a data problem, such as all diseased photos being taken in the shade.
—Wrap it in a simple web page where someone can upload a photo, and document everything on GitHub.

That one project touches data collection, training, evaluation and deployment, which is exactly what interviewers ask about.

What are the limits and responsibilities?

Vision models can be biased if their training data under-represents certain lighting conditions, skin tones or regions. They can be fooled by small, deliberate changes to an image. And images of people are personal data: if you build anything involving faces or identifiable people, privacy law applies. Our guide to the DPDP Act for developers explains what that means in practice.

How can you start learning computer vision?

Start with Python and basic image handling using a library such as OpenCV. Then learn the fundamentals of neural networks, train a small CNN on a standard dataset, and finally fine-tune a pretrained model on images you collect yourself, such as leaves from your garden or products from a local shop.

Computer vision sits on top of programming, maths and machine learning, so the order matters. You can see where it fits in a complete path in the full Program Zero curriculum, where Phase 7 covers CNNs from LeNet to ResNet, transfer learning and deep learning in PyTorch, after the programming, data and maths phases that make it understandable.

Want to learn this live, with mentors?

Program Zero takes you from Python to deep learning, CNNs and language models over 18 months of live, mentor-led classes, with projects every quarter. ₹5,999 for all 18 months; the batch starts 9 January 2027.

Frequently asked questions

What is computer vision in simple words?+

Computer vision is the part of AI that lets computers understand images and video. It treats an image as a grid of numbers and uses models, usually neural networks, trained on many labelled examples to recognise patterns such as edges, shapes and objects. It powers face unlock, number-plate readers, document scanning and crop disease apps.

What is the difference between object detection and segmentation?+

Object detection finds each object in an image and draws a box around it, such as every car in a traffic photo. Segmentation labels every pixel, giving the exact outline of each object or region, such as the boundary of a tumour in a scan. Segmentation is more precise but usually needs more detailed training labels.

Do I need a lot of data to train a computer vision model?+

Often not. With transfer learning, you start from a model pretrained on a large dataset and fine-tune it on your own images, and a few hundred good examples per class can be enough for a useful first version. Data augmentation and collecting images in real-world conditions matter more than sheer volume.

What should I learn before computer vision?+

Learn Python first, then basic linear algebra and probability, then the fundamentals of machine learning and neural networks. After that, study convolutional neural networks and transfer learning using a library such as PyTorch. OpenCV is useful early on for loading, resizing and processing images.

Related reading

We teach this, live

Every article here comes from something we teach. Sit in on a free masterclass and judge the mentors yourself.