
How to Be Found When the Query is a Picture
For most of the history of search, an image was something you published. It sat on the page, it made the page nicer, and if you were diligent, you gave it an alt attribute and moved on. The query — the thing that actually started the search — was always text.

That relationship has inverted. Today the image is frequently the query itself. A shopper points a camera at a chair, a pair of sunglasses or a box of miniatures, and the search begins from the pixels. No words are typed. And when words are typed, the machine answering them is increasingly one that reads your images directly rather than reading only the text you wrapped around them.
This changes what image optimization is for. It is no longer a finishing touch applied to a page that already ranks. It is one of the inputs that decides whether the page enters the retrieval set at all.
This guide follows the complete path from what a machine sees in a photograph to what eventually gets measured in search:
1. Why the Image Became the Query (you are here)
2. How Machines See Images
3. Google Images: How Image Search Works
4. Google Lens: Visual Search Explained
5. Images in Universal Search
6. Images in AI Search
7. Image & Product Structured Data
8. Technical SEO for Images
9. The Photography Brief for Machines
10. Visual Search Beyond Google
11. Measuring Visual Search
Why the image became the query
The two dilemmas visual search solves
Amazon once gave the clearest answer to the question of why anyone would search with a picture. Visual search, they said, is for shoppers facing one of two dilemmas:
I don’t know what I want, but I’ll know it when I see it.
I know what I want, but I don’t know what it’s called.
The first is the discovery problem. Nobody browsing a magazine spread of living rooms has a shopping list. They have a feeling, and they are waiting for something to trigger it. Marketing has always understood this; for instance, the iPhone sold to people who were not, before they saw it, in the market for an expensive phone to browse a 2G web with.
The second is the vocabulary problem, and it is more common than most ecommerce teams admit. Go into a hardware store and try to buy a screw. Cylindrical head, flared, flat? For wood, concrete, fixtures? Most of us solve this by walking in with the screw in our hand and saying: I need a box of these.
A camera is that outstretched hand. It lets a shopper skip the step where they have to know your vocabulary before they can find your product. That is the entire proposition, and it is why visual search grew fastest in exactly the categories where naming things is hard: furniture, fashion, plants, parts, and collectibles.
Both dilemmas were true a decade ago. What changed is that the tooling caught up, and the volume arrived.
The numbers that matter now
Google reported in October 2024 that Lens was being used for nearly 20 billion visual searches per month, and that 20% of those searches were shopping-related, matched against a Shopping Graph of more than 45 billion product listings. For scale: that was up from 12 billion a month cited at I/O 2023.

At Google I/O in May 2025, Sundar Pichai said Lens had grown 65% year over year, with more than 100 billion visual searches already that year.
The pattern is not confined to Google. Amazon reported a 70% year-over-year increase in visual searches worldwide in October 2024 and launched Amazon Lens Live in September 2025, real-time camera scanning with a swipeable product carousel, wired into its shopping assistant.
Set that against what is happening to the text query. SparkToro’s analysis with Datos found that 58.5% of US Google searches ended without a click in 2024, and their 2026 update puts that figure at 68%, with roughly 276 of every 1,000 US searches now reaching the open web, down from 374 two years earlier.
Read those two trends together and the strategic picture is unpleasant but clear. The text query is delivering fewer clicks to fewer sites. The visual query is growing fast and arrives with commercial intent already attached: someone photographing an object is much further down the funnel than someone typing a category name.
There is one more shift worth putting on the table early. In a survey of over 1,000 US consumers, Gen Z now discovers products on Instagram (30.4%) and TikTok (23.2%) more than on Google (18.8%). Google still leads comfortably with every older cohort. But the youngest buyers have already normalized a visual-first, feed-native way of finding things — and that is the behavior Lens, Pinterest, and AI Mode’s visual grid are all built to serve.
Why this is no longer an "easy opportunity" argument
For years the case for image SEO was made on neglect: nobody bothers with this, so there are cheap wins available. That argument has expired, and it is worth being clear about why.
It expired partly because it stopped being true: structured data adoption in ecommerce is now broad, and product badges in Images results are common rather than remarkable. But mostly it expired because the stakes changed shape. When images were a bonus channel, ignoring them cost you some incremental traffic. Now the systems that consume your images are the same systems deciding whether a human being ever sees your page.
An AI Mode answer that resolves a shopping question inside the result page is drawing on product data and imagery to do it. A Lens capture that returns six competitors and not you has not cost you a ranking but it has removed you from the shopper’s consideration set before a page was ever loaded. An agentic checkout flow that reads a product feed will not "see" a beautiful photograph that lacks the fields it requires.
So the argument is no longer about upside. It is about eligibility. Images are now part of how machines establish what your product is, and everything downstream — retrieval, matching, citation, purchase — depends on that identification being correct.
That is the premise this guide is built on, and it explains the order of what follows. Before covering any surface or tactic, the next section deals with how these systems actually see: what they extract from an image, what they read from the text around it, and where the two diverge. Almost every recommendation in this guide is derivable from those mechanics.
The Strategic Picture
The premise of this guide has been a single inversion: the image used to be something you published, and it is now something people search with.
Everything else follows from that. Structured data matters because it tells a machine what the object in the photograph is. Compression matters because artifacts are noise in the signal a recognition system depends on. Feeds matter because every surface that shops on a shopper’s behalf reaches for structured product data before it reaches for a page. Packaging photography matters because text rendered in pixels is text a machine can read. None of these are separate disciplines. They are one discipline seen from different angles.
For years, image SEO was treated as a finishing task — something applied to a page that already ranked. That framing no longer describes what is happening. Images are now part of how machines establish what your product is, and that identification determines whether the page enters consideration at all. The work has moved from decoration to eligibility.
What will age, and what will not
A guide claiming to be a reference should say which of its parts to re-check.
The mechanics in section 2 are durable. Embeddings, object detection, OCR and the blending of visual, object and annotation retrieval are the architecture of visual search across every platform in this guide. The models will improve. The structure will not change soon, and it is what makes the rest derivable rather than memorizable.
The photography brief in section 9 is durable. Sharp, well-lit, cleanly separated, honestly compressed, legibly printed, distinctively yours — none of that depends on a particular platform’s current rules. Better recognition systems reward the same qualities; they simply raise the ceiling on what a good photograph can achieve.
The interfaces will change, and fast. Section 3’s SERP anatomy, section 4’s filter tabs, section 6’s AI Mode layout... these are screenshots of a moment. Google shipped its Images tab in November 2025 and named visual search fan-out in September 2025. The next twelve months will produce equivalents.
The specifications will drift. Merchant Center’s 500 × 500 minimum enforces in January 2027. Structured data types have been retired at a steady rate since 2023. Verify the tables in sections 5 and 7 against live documentation before acting on them at scale.
The measurement will improve, which is the most likely development to change practice. When Search Console reports AI surface appearances or Lens-originated impressions, arguments that currently rest on eligibility will rest on numbers instead.
The two pipelines
If one idea from this guide survives contact with a real workday, make it this one.
There are two pipelines, and they fail independently.
The pixel pipeline is what a system extracts from the image itself: embeddings, detected objects, OCR-recovered text. It is what a shopper’s camera invokes, and no metadata participates in it. You win it with photography.
The text pipeline is everything you wrap around the image: alt text, captions, structured data, feeds, surrounding copy. It is what retrieval and indexing lean on, and it is what turns a picture into a product card with a price. You win it with markup and content discipline.
Optimizing one and neglecting the other produces exactly half the available visibility. Immaculate alt text on a dim, cluttered photograph fails the camera. A beautiful studio shot with an empty alt attribute and no Product markup appears as a picture rather than as a purchasable thing.
You cannot know in advance which pipeline any given shopper will invoke. So the strategy is not to choose. It is to cover both — which is affordable, because the two rarely conflict and often share the same underlying work.
A closing thought
There is a scene in The Silence of the Lambs in which Hannibal Lecter presses Clarice Starling on where desire begins and answers his own question: we start by coveting what we see every day, and our eyes seek out what we want before we have words for it.
It is an unsettling way to make a commercial point, but the point is sound and it predates every platform in this guide. Typing a query was always the unnatural step, an act of translation from something seen or imagined into a string of words that a machine would accept. Visual search removes the translation. The eye points, and the search begins.
Everything in this guide is, finally, about being present at that moment, when someone sees a thing they want, before they know what it is called, and reaches for a camera instead of a keyboard.
Be the answer they get.
Read the guide
Next → How Machines See Images
Article by
Gianluca Fiorelli
With almost 20 years of experience in web marketing, Gianluca Fiorelli is a Strategic and International SEO Consultant who helps businesses improve their visibility and performance on organic search. Gianluca collaborated with clients from various industries and regions, such as Glassdoor, Idealista, Rastreator.com, Outsystems, Chess.com, SIXT Ride, Vegetables by Bayer, Visit California, Gamepix, James Edition and many others.
A very active member of the SEO community, Gianluca daily shares his insights and best practices on SEO, content, Search marketing strategy and the evolution of Search on social media channels such as X, Bluesky and LinkedIn and through the blog on his website: IloveSEO.net.




