The Photography Brief for Machines

>

Images & Visual Search

The Photography Brief for Machines

10

min read

Art direction for machines: the photography brief

This section is the executional one. Everything in sections 2 through 8 describes how systems see; this converts that into instructions a photographer, a studio manager or a product-content team can work from.

The framing that makes it land with creative teams: you are no longer only art-directing for the shopper. You are art-directing for the system that decides whether the shopper ever sees the picture. Those two audiences agree far more often than they conflict — a clean, well-lit, honest product photograph is good for both — but where they diverge, the machine’s requirements are the ones that are testable.

Sharp, clean, frontal, and why compression is now a ranking issue

Recall the three parallel retrievals from section 2: visual, object and annotation. This subsection is about winning the first two, and the thing you are actually optimizing is detection confidence.

Separate the product from its background. Confidence scores rise when an object’s boundaries are unambiguous. This is why the plain-background catalogue shot remains the most reliably retrievable image type ever devised, and not because it looks good, but because it makes the object’s edges trivially computable. Google’s own structured data guidance says as much: images clearly showing the product, for example against a white background, are preferred.

Shoot frontal or near-frontal. A product photographed from a conventional angle produces an embedding close to the cluster where that category’s other images sit. An oblique, dramatic or unusual angle produces an embedding some distance away, which is to say, the system considers it a slightly different thing. Save the dramatic angles for supporting images; make the primary asset conventional.

Light evenly, and do not shift color. Filters that alter the natural color of light degrade both recognition and, more practically, the shopper’s trust when the product arrives. Color is also a variant attribute in structured data: a photograph whose colorway does not match its markup is a mismatch a system can detect.

Keep it sharp. Soft focus, motion blur and low resolution all reduce the information available for both embedding and detection.

Compress deliberately. This is the point most worth carrying into a studio conversation, because it reframes a decision usually made by whoever set the export preset three years ago. Over-compression does not make an image slightly worse-looking. It introduces information that was not in the original scene, aka artifacts the model cannot distinguish from real detail. The practical consequence is a higher likelihood of misidentification: the system fills ambiguity with plausible guesses.

Cheap, high-quality delivery is now available, so the old trade-off has largely dissolved. The remaining risk is a careless preset applied across a catalogue of ten thousand images.

A diagnostic worth adopting. Run a representative sample of your product photography through an object-detection tool — Cloud Vision’s demo interface is sufficient — and read the confidence scores.

Sharp, clean, frontal, and why compression is now a ranking issue

This is one of the few genuinely objective tests available in image SEO. If your hero product image scores poorly against its own category label, that is a photography brief, not a mystery. Establish a floor, and re-shoot below it.

Object co-occurrence: what else is in the frame

Before, I explained that Google’s visual search fan-out identifies "subtle details and secondary objects in addition to the primary subjects." The art-direction consequence is significant and runs counter to a long-standing instinct toward clean, single-product imagery.

A styled photograph is not one query. It is a dozen. A room shot containing your pendant lamp also generates retrievals for the mirror, the cushions, the rug and the curtains. Every recognizable object is an independent entry point.

For a retailer selling across a category, this is a strong argument for lifestyle photography, but with one hard constraint.

Eight clearly-separated products in a well-lit room are eight opportunities. The same eight in a dim, visually busy composition are eight low-confidence detections, which is worse than one good one. The failure mode is not "too many objects." It is too many objects rendered indistinctly. Clutter suppresses confidence across everything in frame, including your hero.

So the brief is specific: styled scenes, deliberately composed so that each merchandised object has clear separation, adequate light and an unbroken silhouette. That is a harder shot than either a plain-background catalogue image or a naturalistic room photograph, and it is worth briefing explicitly rather than hoping for.

Co-occurrence also signals positioning. What appears alongside your product tells a system what kind of product it is and who it is for. A chair photographed among mid-century furniture is classified differently from the same chair among industrial fittings. This is worth treating as a deliberate choice: the objects you place in frame are, in effect, an assertion about your brand’s category.

Then close the loop with interlinking. If a single photograph contains four products you sell, and those four product pages link to one another as related items, you have given the system both the visual co-occurrence and the explicit relationship. Recognition plus declared relationship is what makes cross-suggestion likely — the shopper who searches your lamp being shown your cushion.

Alt text should reflect this too. alt="Olive linen shirt worn with our wide-leg oatmeal trousers" describes the frame honestly, serves accessibility, and names two products instead of one.

OCR-legible packaging

Before, I also showed that text rendered in pixels is extracted and made available. For any product that ships in printed packaging, this is free structured product data, or a wasted opportunity.

Jessier’s companion piece on making products machine-readable for multimodal AI search (Search Engine Land, November 2025) makes the strategic case bluntly: ecommerce packaging now has to be engineered as a digital asset, because when a machine cannot read the packaging, the product becomes invisible at the moment of highest purchase intent.

The brief:

  • Include a back-of-pack or spec-panel shot in the standard shot list. Most catalogues photograph the front and stop.

  • Shoot text flat-on. Angled text degrades OCR sharply.

  • Resolution must support the smallest text you care about. Jessier’s working thresholds — roughly 30px character height and 40 grayscale contrast in the delivered image — are a usable specification to hand to a studio.

  • Do not crop the panel. A composition that trims the specification block for aesthetic balance has discarded machine-readable product data.

  • Check your compression against the smallest text, not against the product. Text is where artifacts do the most damage.

This applies well beyond boxed goods: care labels, ingredient panels, size charts, certification marks, model numbers on hardware. Anywhere a fact is printed rather than typed, a legible photograph makes it retrievable.

Turn visual search data into actionable insights

Advanced Web Ranking gives you a clearer view of visual search performance, from rankings in Google Images to image results in Universal SERPs and visibility across major AI engines.

👉 See how your images perform across search.

Stock photos and the uniqueness problem

Section 4 introduced Exact matches — the Lens tab that finds instances of this image across the web. It gives an old argument new force.

If your product page uses a supplier-provided image shared by forty other retailers, an exact-match retrieval cannot resolve to you specifically. You are one of forty identical results, differentiated only by whatever price and availability data you supplied, and probably not first.

The parallel with text is exact: republishing the distributor’s product description verbatim has never worked, for the same reason. Duplicate content cannot establish that you are the origin of anything.

The brief:

  • Shoot your own primary product imagery wherever volume allows. This is the single highest-value differentiator available on visual surfaces.

  • Where supplier images are unavoidable — long-tail catalogue, drop-ship, thousands of SKUs — modify them meaningfully: your own crop, your own background treatment, your own scale or context references, composited into your own scene.

  • Prioritize by revenue. Original photography for your top-performing and highest-margin lines; supplier imagery, modified, for the tail. This is a resourcing decision and framing it that way tends to get it funded.

For products whose form you share with many competitors — rattan lamps, generic homeware, white-label goods — you cannot out-design the shape. Distinctive photography is the available lever, and it is the one that decides which brand a system reaches for first.

Attribution that survives the journey

The previous subsection argued for owning your imagery. This one is about what happens to it afterwards, and – let’s be honest – what happens is that you no longer control whether your images travel.

They are scraped, screenshotted, republished, embedded in marketplace listings, pinned, reposted, and ingested by systems building visual indexes. Some of that you permit. Much of it you never hear about. The question that remains open to you is not whether your photographs circulate, but whether they remain attributable to you when they do.

That question has a better answer in 2026 than it did a few years ago, because attribution has moved from convention to infrastructure. It used to depend entirely on whoever republished your image choosing to credit you. Now credit can be embedded in the file, cryptographically signed, and carried wherever the image goes.

The layers, from simplest to strongest:

Layer

What it carries

Where it lives

IPTC metadata

Creator, credit, copyright, rights statement

Embedded in the file

`ImageObject` structured data

creditText, creator, copyrightNotice

On the page

Image License metadata

license, acquireLicensePage

On the page — enables the Licensable badge

C2PA Content Credentials

Cryptographically signed creation and edit history

Embedded and signed

The page-level signals establish your authority on your own site. The embedded signals are the ones that travel, and that distinction is the whole subsection.

Why this matters commercially, beyond the obvious rights argument:

It anchors the entity to you. Before, we established that an image’s meaning is partly established by the contexts it appears in. Distributed copies carrying your embedded credit data are contexts that point back, rather than contexts that dilute. Wide circulation without attribution builds an entity around a picture; wide circulation with attribution builds it around your brand.

It is the counterweight to the uncertainty problem we also saw before. When a system is guessing which brand a commodity-form product belongs to, signed provenance on the original is a real signal in a field of weak ones.

Provenance is becoming user-visible. When talking about technical foundations, I covered "About this image" and Google’s I/O 2026 extension letting shoppers ask directly whether an image was AI-generated. As that surface matures, images with verifiable histories and images with none become distinguishable to buyers, not just to systems.

Redistribution can be an asset rather than a leak, but only when it is attributed. Creative Commons licensing, where it suits your business, formalizes that: Attribution-ShareAlike where derivative works should stay open, Attribution-NoDerivatives where the image should be used unmodified.

The implementation, in order:

  1. Write provenance at capture. Configure cameras, DAM and photo pipelines to embed IPTC creator, credit and copyright. Add C2PA Content Credentials where your tooling supports it;  hardware support is arriving, including natively in recent phone cameras.

  2. Mirror it on-page. ImageObject credit properties plus Image License markup where licensing is relevant.

  3. Verify survival end to end. Run the delivered image — the file Googlebot actually fetches — through a metadata inspector. This is where most implementations fail.

That third step is the one to schedule rather than assume. As established in the ‘Technical Foundation’ section, optimization pipelines strip metadata by design — Cloudflare Polish documents that it does, and many CMS upload handlers and build-time compressors do it silently. Every layer of this can be implemented correctly at capture and discarded at delivery without anyone noticing.

The test is simple: download your own product image from your live site, inspect its metadata, and see what survived. Most teams will be surprised.

The brief, condensed

A single reference to hand to whoever actually shoots and processes your images.


Specification

Primary product shot

Frontal or near-frontal; product clearly separated from background; plain or neutral ground

Lighting

Even, natural; no color-shifting filters

Focus

Sharp throughout the product; no motion blur

Resolution

Minimum 500 × 500 for feeds; 1500 × 1500 recommended; structured data floor is 50,000 px (w × h)

Aspect ratios

Supply 16:9, 4:3 and 1:1

Product fill

75–90% of frame for feed images

Compression

Deliberate; verify against the smallest legible text, not the product

Format

AVIF or WebP with JPEG fallback via the picture element

Shot list

Front · detail · scale/in-use · back-of-pack or spec panel · one styled scene where relevant

Packaging text

Flat-on; 30px character height or more; 40 grayscale contrast or more in the delivered file

Styled scenes

Each merchandised object clearly separated and individually lit

Originality

Own photography for priority lines; meaningfully modified supplier assets for the tail

Variants

Distinct imagery per variant, at a distinct URL

Metadata

IPTC credit and copyright embedded; C2PA where relevant; verify survival through CDN

Alt text

Distinct per image; describes viewpoint and variant; no repetition across the page

Acceptance test

Object detection confidence above an agreed floor for the product’s own category label

That last row is the one that turns this from a style guide into a specification. Everything above it is advice; that is a test. A shot either passes or it does not, and the disagreement between a photographer and an SEO stops being a matter of taste.

Continue the guide

 Technical SEO for Images ← Previous · Next → Visual Search Beyond Google

Gianluca Fiorelli

Article by

Gianluca Fiorelli

With almost 20 years of experience in web marketing, Gianluca Fiorelli is a Strategic and International SEO Consultant who helps businesses improve their visibility and performance on organic search. Gianluca collaborated with clients from various industries and regions, such as Glassdoor, Idealista, Rastreator.com, Outsystems, Chess.com, SIXT Ride, Vegetables by Bayer, Visit California, Gamepix, James Edition and many others.

A very active member of the SEO community, Gianluca daily shares his insights and best practices on SEO, content, Search marketing strategy and the evolution of Search on social media channels such as X, Bluesky and LinkedIn and through the blog on his website: IloveSEO.net.