Retail image recognition is simultaneously oversold and genuinely useful, which makes it unusually hard to evaluate. Vendor demonstrations show a photograph converted instantly into a perfect shelf inventory. Real deployments show something more variable, dependent on conditions the demo controlled for.
Both pictures are accurate. The technology does work, and it works considerably better on some tasks than others. The gap between a successful deployment and a disappointing one usually comes down to whether the buyer understood which tasks those were.
This article is a realistic assessment for Gulf brands: what image recognition measures reliably, where accuracy degrades, what it costs to run beyond the licence, and how to test it on your own shelf rather than the vendor’s.
What it does
A rep photographs a shelf. The system identifies the products in the image and returns structured data: which SKUs are present, how many facings each has, their positions relative to each other, and by extension your share of shelf and your compliance against a planogram.
The value is not the identification itself but what it makes affordable. Counting facings manually across a large assortment takes a rep several minutes per bay and degrades badly under time pressure. A photograph takes seconds. That difference is what makes attribute-level compliance measurable across an estate rather than in a sample.
The comparison worth making is not image recognition against nothing, but image recognition against a rep with a clipboard. On presence checks the clipboard is competitive. On facings and share of shelf across forty SKUs, it is not. Our piece on the role of image recognition in retail execution covers the operational case in more depth.
What it delivers reliably
Facing counts and share of shelf
The strongest use case. Counting front-facing units of visually distinct products in a well-lit photograph is close to a solved problem, and share of shelf follows arithmetically. This is where the technology creates capability that manual auditing genuinely cannot match at scale.
Presence and absence
Reliable, though a checklist achieves the same thing. The advantage is that it comes free alongside the facing count rather than costing extra visit time.
Position and adjacency
Solid for planogram compliance, where the question is whether a SKU sits in its assigned position relative to its neighbours. This is the task that makes attribute-level planogram compliance measurement practical across a large estate.
Competitor capture
Underrated. Because the system processes the whole image, competitor facings arrive at no additional visit cost. In a contested category this is often the most commercially useful output, since your share of shelf falling is only interpretable against who gained.
An audit trail
Every measurement is backed by a timestamped, geolocated photograph. When a compliance figure is disputed by a retailer or a distributor, evidence settles it. This matters more in distributor-mediated markets than most buyers anticipate.
Where accuracy degrades
Vendors are generally honest about these when asked directly and rarely volunteer them.
Visually similar pack variants
The most significant real-world limitation. Flavour variants, size variants and multipack versions of the same product often differ only in a small area of pack graphics. Systems confuse them, and the confusion is systematic rather than random, which means it does not average out across a large sample. If your portfolio is variant-heavy, test this specifically before anything else.
Photo quality and capture discipline
Angle, lighting, distance, glare from chiller doors, and partial obstruction all degrade output. This is a field training problem more than a technology problem, and it is the most common reason a pilot underperforms. Capture standards need defining and enforcing.
Catalogue completeness
A system cannot identify a product it has never been shown. New listings, packaging refreshes and seasonal packs all require catalogue updates, and the lag between a pack change and catalogue coverage is a period of degraded accuracy. In markets with high assortment churn this is continuous overhead rather than a one-off setup task. It is fundamentally a retail product management discipline.
Imported and unfamiliar assortment
Gulf shelves — particularly in the UAE — carry a broader international assortment than most markets. Vendors trained predominantly on European or North American catalogues will have thinner coverage of regional brands and imports from Asia and the Levant. Ask specifically about regional catalogue depth.
Stacked, deep and irregular displays
Floor stacks, pallet displays, bins and hanging strips are much harder than a standard gondola shelf. Accuracy on secondary displays is generally well below accuracy on primary shelf, which matters because secondary displays are frequently what you funded and most want verified.
What it cannot see at all
Back-room stock, price accuracy on a label the camera did not capture, whether the product is findable from a shopper’s eye level, and why a gap exists. Image recognition tells you the state of the shelf face. It does not diagnose cause, which remains a human observation.
Like what your reading?
Take a moment to subscribe before continuing and never miss out on exclusive insights, news, and case studies.
The real cost of running it
The licence is rarely the largest line. Costs that surprise buyers:
- Per-image or per-scene pricing. A large field team photographing multiple bays per visit generates volume quickly. Model this on your actual visit frequency and bay count, not a per-user rate.
- Catalogue onboarding and maintenance. Initial training on your SKUs plus ongoing updates for new listings and pack changes. Ask explicitly whether maintenance is included or billable.
- Field training on capture standards. Non-optional, and in the Gulf it means training material in more than one language for a multilingual field team.
- Exception handling. Low-confidence recognitions need human review. Someone has to do that, and the volume is highest in the first months.
- Data and device cost. Image upload consumes mobile data and battery. On older devices this affects visit completion rates in ways that show up as adoption problems rather than technical ones.
None of these makes the technology a poor investment. They do mean a business case built on licence cost alone will be wrong, usually by a wide margin.
How to evaluate it properly
The only test that predicts deployment performance is one run on your conditions. A vendor demo predicts nothing.
- Use your own catalogue. Including your near-identical variants. If a vendor cannot onboard a representative subset for a pilot, that is informative in itself.
- Photograph real shelves in your real stores. A Riyadh hypermarket bay and a Jeddah grocery, in normal lighting, by an actual rep rather than a specialist.
- Count manually and compare. Take twenty bays, count facings by hand, run the same images through the system, and compare SKU by SKU. This is the whole evaluation. Everything else is secondary.
- Test the failure modes deliberately. A chiller with glare. A floor stack. A bay with a competitor’s product partly obscuring yours. A newly launched pack.
- Measure the exception rate. What proportion of images need human review, and how long does review take? This is your ongoing operational cost.
- Time the catalogue update. Add a new SKU and measure how long until it is recognised. In a high-churn assortment this determines steady-state accuracy.
Ask for accuracy stated as a percentage of SKU-facings correctly identified against a manual count, on your images. Vendors quote accuracy in various ways, some of which are considerably more flattering than others.
When it is worth it, and when it is not
Worth it when you need facing-level data across a wide estate, when share of shelf and planogram compliance drive commercial conversations with retailers, when competitor context matters, when your assortment is visually distinct enough for reliable recognition, and when you have the catalogue discipline to keep it current.
Not worth it yet when your primary question is presence and availability rather than facings, when your assortment is dominated by near-identical variants, when your outlet and product master data is not yet reliable, or when your field team’s basic visit completion is still inconsistent. In that last case the constraint is adoption, and adding a technology layer will not fix it.
A useful sequencing principle: image recognition amplifies a functioning field operation and does not create one. Brands that deploy it onto an unreliable field process generally conclude the technology failed. The AI and predictive analytics layer above it depends on the same foundation.
How image recognition fits alongside human observation
The framing that causes most disappointment is image recognition as a replacement for the field visit. It is better understood as a change in what the visit is for.
What moves to the camera
Counting. Facings, positions, adjacencies, share of shelf, competitor presence. These are mechanical measurements that a photograph captures faster and more consistently than a person working through a checklist, and consistency matters more than peak accuracy for trend analysis.
What stays with the rep
Judgement and cause. Why is there a gap — is there stock in the back room, has the SKU been delisted, has a competitor been given the space, is the store short-staffed? Is the price label correct? Is the product findable from shopper eye level? None of this is in the image.
This is the split that makes the technology worth its cost: the camera absorbs the measurement burden and the rep spends the recovered time on diagnosis and correction. A deployment that uses image recognition to shorten visits without redirecting the saved time captures only half the available value. We covered the field implications in using retail execution software to manage a field team.
The reporting consequence
Once facings arrive automatically and cause arrives from the rep, you can report share of shelf and compliance alongside the reason for every failure. That combination is what makes a category conversation possible, and neither half is sufficient alone.
Common deployment mistakes
- Rolling out before the catalogue is complete. Early accuracy will be poor, the field team will conclude the system does not work, and that impression is difficult to reverse. Onboard the catalogue properly first.
- No capture standard. Angle, distance and lighting drive accuracy more than the algorithm does. Define the standard, train it, and audit adherence in the first weeks.
- Photographing everything. Per-image costs and upload time both scale with volume. Photograph the bays that matter commercially, not every shelf in the store.
- No exception review process. Low-confidence recognitions accumulate silently and corrupt the data. Someone needs to own review from day one.
- Judging accuracy on a vendor benchmark. Accuracy is a property of your catalogue and your photos. A published figure from another market predicts nothing about yours.
- Ignoring device variation. Older phones produce lower-quality images and slower uploads. If part of your field team is on aging hardware, that segment’s data will be systematically worse.
- Treating it as a technology project. The determinants of success are catalogue discipline, capture training and what you do with the output — all organisational.
The pattern across these: image recognition failures are usually process failures. The algorithm is rarely the weakest component in a disappointing deployment.
What to expect in the first six months
A realistic trajectory, so that normal early friction does not get read as failure.
- Month 1. Catalogue onboarding and capture training. Accuracy below steady state. Exception review volume at its highest. Do not report externally from this data.
- Months 2-3. Accuracy improves as capture discipline settles and catalogue gaps are filled. First usable share-of-shelf and compliance reporting. Expect the numbers to be worse than assumed — that is the point.
- Months 4-5. Steady state on accuracy. Trend data becomes meaningful. Exception volume falls to a manageable baseline. Competitor data starts to be genuinely useful.
- Month 6. First category conversation supported by your own shelf data. This is the milestone that justifies the investment, and it is worth planning toward explicitly rather than waiting for it to happen.
The organisations that get least from image recognition tend to be those that expected month-six output in month one, concluded the accuracy was inadequate, and reverted to manual auditing before capture discipline had settled.
A note on what comes after recognition
Image recognition produces structured shelf data, which is an input rather than an answer. The question a brand manager actually has is why regional compliance slipped this month, and no recognition system answers that.
The layer above it — querying that data in plain language rather than through dashboard filters — is what makes shelf data usable by the people who need it, rather than by the analyst who maintains the reports. That is a different capability from recognition and worth evaluating separately. We covered it in getting the most out of your retail data with an AI assistant.
The practical point for a buyer: assess recognition accuracy and reporting usability as two separate things. A system with excellent recognition and reporting nobody opens delivers very little, and the reverse is equally true.
Where to start
Run the twenty-bay manual comparison described above before committing to anything. It costs a few days, produces a number specific to your catalogue and your stores, and is more informative than any reference call or vendor benchmark.
If the accuracy on your variants is acceptable, the case is usually straightforward. If it is not, the problem is unlikely to be solved by a different vendor — it is a property of your portfolio, and the sensible response is to use image recognition for the SKUs where it works and structured manual checks for the ones where it does not. A mixed approach is normal and considerably better than forcing either method across the whole assortment.
Shelvz runs image recognition alongside structured field checks in the same visit, so facings and share of shelf come from the image while cause, back-room stock and price accuracy come from the rep. To test accuracy on your own catalogue and your own shelves, book a walkthrough.


