Today's AI has learned to measure technique. It has not yet learned to measure art.

01 The question is: who is the judge?

There is one recognised way to name the best photographs: contests with a professional jury.

Not Instagram likes. Not even "panels of professionals". A photograph earns its standing only when truly accomplished photographers judge it.

This matters because of how today's AI photo-scorers were actually built. Every one of them learned from images labelled by ordinary people — amateurs, registered users, paid annotators — never by a professional photography jury.

Who annotates the data behind AI photo-scoring
SystemWho judges / annotatesExpert jury?
Top photography contestsSelected, recognised professionals — photographers, curators, editors, art professionalsYES
NIMA / AVARegistered DPChallenge users / amateur photographers; ~200 ratings per imageNO
LAION-Aesthetic V1Human raters via Simulacra, asked simply how much they liked the imageNO
LAION-Aesthetic V2Mixed human-rating datasets, including AVA / DPChallengeNO
EverypixelTrained on large stock-photo sets; algorithmic technical-aesthetic scoringNO
PickScoreReal users of Pick-a-PicNO
ImageRewardAnnotators trained for image-preference labellingNO
HPS v2Human annotators ranking generated images (~798k judgements)NO
VisionRewardHuman annotators giving multidimensional preference labelsNO

Almost every automatic scorer — Google's NIMA included — learned from a crowd (the AVA set), not from masters.2 A crowd feels "nice." A world jury sees what the crowd cannot. Researchers have shown models trained this way simply don't carry over to professional photography.3

"Models trained on the crowd-sourced AVA dataset behave differently on professional photographs, and do not generalize well to them."

— paraphrased from Chambe et al., Deep learning for assessing the aesthetics of professional photographs, 2022
02 How to solve it / Foundation

The right way is to learn from real contests, judged at the highest level.

Since 2018 we have run FOTOAWARD — an international non-profit photography contest judged in two jury rounds after an initial audience stage, whose defining feature is a jury of world-class experts.

Among the jury members are the finest photographers alive — repeat winners of the most prestigious awards in the world (World Press Photo, Wildlife Photographer of the Year, Bird Photographer of the Year, Trierenberg Super Circuit, Golden Turtle, GDT European, WPPI, etc.), National Geographic photographers, Nikon ambassadors and book authors, from ten countries. See the 2025 jury ↗

The contest itself draws winners of Hasselblad Masters, Smithsonian, HIPA and Xposure. From it comes a rare, structured record of expert taste — eight years of graded verdicts, layer by layer.

What we are looking for are the patterns in the winning work that bring these top experts to agreement. Independent agreement among judges is a measurable, scientific thing — statisticians call it inter-rater reliability.1

FOTOAWARD 2025 contest winners FOTOAWARD 2025 — winners ↗
Prizes from:
HOYA / Kenko Tokina
Topaz Labs
APE — Association of Photographers "Eurasia"

Our own tests point clearly in one direction: today, the best available judge of artistic worth is not these scoring services, but the leading frontier AI models. This is an illustration from our daily work, not a formal study. We took ten photographs that all reached the final of our contest — every one strong — and asked three of the world's leading AI models to score each on artistic worth, 0 to 100. Then we lined those scores up against how our jury actually voted, and against the technical scorer Everypixel.

Jury consensus vs. frontier AI vs. Everypixel — ten finalists
All bars share one 0–100 scale. Jury consensus (gold) is an index of how rare that level of agreement is — the gold standard. Then each of the three frontier models shown separately, and Everypixel (violet), a technical scorer. Watch the jury index fall in clear steps while the three models stay high — and often disagree sharply with each other.
Run: Aug 29, 2026 · Claude Opus 5 · GPT-5.5 · Gemini 3.5 · via OpenRouter API
Jury consensus (the gold standard)
Opus 5
GPT-5.5
Gemini 3.5
Everypixel (technical)
The jury index is plainly the fairer measure — because agreement this strong is extraordinarily rare. The AI scores barely move: nearly everything lands between 79 and 90, whether the jury almost crowned a frame or nearly rejected it — a work that earned a single vote can score higher from the machines than one the whole jury loved. The three models even disagree with each other by up to thirty points on one image.

About the jury index: a raw vote count understates the signal. In the spirit of Fleiss' kappa — the standard measure of inter-rater agreement — what matters is how statistically unlikely a given level of consensus is. Of 326 works, just one drew 8 of 11 votes; three drew six. Our index combines that rarity with how close the vote came to unanimous, on a 0–100 scale. It is deliberately conservative: 8 of 11 is near-miraculous agreement, but short of a unanimous 11, so it reads 85, not 100.

Even the best AI judge available today does not reach the level of a real world-class jury.

Our next ambitious undertaking is to study more than 35,000 photographs — invitation-only from the start, filtered by audience vote, then judged by a semifinal jury and finally by our world-class masters — and to represent each one as an embedding of hundreds of measurable traits, paired with the verdict it actually received. With enough such pairs, what brings judges to agreement stops being a matter of taste and becomes something a machine can learn.
03 AI is already useful

None of this means AI can't help photographers today.

Because we study how AI reads photographs every day, we have built real tools around it — free, and already in use on our contest. Load a photograph and the system can analyse it, write a critique, suggest titles, compose a poem or a short story from it, turn it into a 19th-century engraving, or check whether it was AI-generated.

Add a photograph to enable the tools. They run for real — text takes up to two minutes, technical checks are faster.
04 Our next ambition

FotoMatcher — the right contest for any photograph, found by reading the photograph itself.

We have built a unique, continuously updated database covering close to every photography contest in the world. But the real work is not matching by theme or contest name: the AI studies the photograph itself and weighs how well it truly fits each of the hundreds of open contests — even when genres and categories can't tell you — plus many other factors.

One AI agent keeps watch over 6,000+ FIAP salons and similar bodies, each on its own clock, ready for the moment an open call begins.
Another continuously parses the internet for newly announced contests, so the base stays current all year.
750+
contests profiled, 30 fields each
843+
category-level entries
45+
countries of organisers

It is a hard, separate project — and it is close to completion. fotomatcher.com ↗

FotoMatcher interface — profile, photo upload and live contest analysis
FotoMatcher — live interface · click to enlarge