Reading the GPHG: A prelude
How a throwaway prediction experiment led to a much deeper analysis of 25 years of watchmaking’s most important awards.
Note: If you just want to get to my 2026 GPHG predictions, scroll towards the bottom of this article and download the PDF. But you won’t be able to read them just yet – not until I reveal the password.
In 2025, I asked ChatGPT to predict the winners of the Grand Prix d’Horlogerie de Genève (GPHG) awards.
It was a bit of fun. Or so I thought.
I wanted to see what happened when I gave an AI model information about the competing watches and asked it to choose the likely winners.
It correctly selected five of the 15 category winners.
That’s a 33% hit rate, compared with the 17% that would be expected from randomly choosing one of six nominated watches in each category.
That result intrigued and irritated me in equal measure.
ChatGPT could not have somehow independently developed the taste, judgement and knowledge required to identify what judges might think the “best” watches were. Despite what Terminator 2 tells us.
Something else must have been happening. And I wanted to know what that was. Definitively.
Since then, things have gotten…a little out of hand. I have now remade the public record of the GPHG from 2001 to 2025 into a structured, machine-readable dataset and built a prediction model based on extensive historical backtests.
Why I’m talking about it now
I hadn’t intended to say anything about this analysis until a little closer to the GPHG revealing its nominated watches.
But then Chris Hall wrote about GPHG data in The Fourth Wheel.
Observing the hundreds of watches archived by the organisation each year, he suggested that somebody with “a little more coding ability” might scrape the GPHG archive and analyse trends in case size, materials and price. Which I had already done.
“I’d certainly publish it,” he added.
kingflum then went and did it.
His article, What 26 Years of GPHG Data Can Tell Us About the Watch Industry, assembles 5957 entries and uses them to examine changes in price, case size, materials, complications and participation. It’s an impressive piece of work and confirms what Chris had spotted: the GPHG has – possibly without meaning to – created one of the richest public datasets available anywhere in the watch industry.
kingflum’s work on this and my analysis overlap, but they are not the same.
His principal subject is the watches that brands chose to enter, and his article examines the archive as a record of what (a part of) the watch industry produced, valued and considered worth promoting over time. His piece also compares winners with the wider field of watches, and offers observations about price, established brands and familiarity.
The main focus of my analysis is the sequence of decisions made about the competing GPHG watches.
With that in mind, I reconstructed the awards’ selection stages, from competing watches to nominated watches to winners, over the competition’s history.
The main question I am asking is whether the successive decisions to nominate and reward watches contain persistent patterns – and whether those patterns have any predictive power.
My historical reconstruction and backtesting were completed at the beginning of this month, and by then the model that I had developed was used for the 2026 forecast I will share later in this article. The ScrewDownCrown piece did not initiate my analysis, but it has persuaded me to share it sooner than planned.
A remarkable public record
The GPHG has attracted criticism over the years because entry requires a fee and some of the watch industry’s largest brands do not regularly participate.
Some of that criticism seems reasonable to me. The GPHG archive is not representative of the whole watch industry. Brands that do enter are making a conscious marketing decision; brands that don’t participate are also making an intentional decision of sorts.
The pay-to-play model is clearly imperfect, but the alternatives are not obviously better to me. An awards system reliant on a handful of watch companies (be they conglomerates or independents), commercial sponsors or media partners might introduce different, potentially less visible biases.
The GPHG needs to be transparent to remain credible – that’s why there is such a rich public record available. The Academy is designed – as far as I can tell – to be a proxy for the entire industry, comprised of people from all parts of it.
Even if a fee is paid, entries tell us that somebody decided a particular watch was worth submitting. A watch moving from competing to nominated represents a collective, documented decision. Winning an award is the result of another collective decision. Across 25 years, those documented decisions create a detailed public history of how a significant part of the industry has marketed, edited and celebrated itself.
That must tell us something, surely.
Rebuilding the awards
It turns out that converting the GPHG archive into something capable of answering my question about why ChatGPT had been able to beat a random selection in 2025 involved more than just creating a spreadsheet and looking at it. The record is more complicated than that.
There are terminology changes over the years. Categories appear, disappear, split apart and then come back together. Some categories get renamed but the rules around entry are broadly the same, while others keep a name but the rules change. Then the underlying rules themselves change over time. Some parts of the archive have missing data while others contain inconsistencies and occasional mistakes – because these things happen with a manual record.
The 2026 Iconic Watch Prize is a good example of this moving feast.
In 2025, Iconic was a category that followed the normal entry and nomination process. This year it has become a separate prize, open to eligible watches whether or not they were officially entered by a brand, with its six candidates generated through Academy proposals. The name remains, but the rules have changed significantly. It could throw up some interesting results.
I’m glad, therefore, that I excluded the Iconic prize, alongside discretionary prizes such as Audacity, Revelation and Chronometry. These categories don’t offer up the same fixed, stable candidate universe as the other main categories.
I have now reconstructed the awards database to be able to look at three stages of decision-making across 25 years:
Competing – the watches entered into the competition.
Nominated – the watches selected to proceed to the final stage.
Winners – the watches that were awarded prizes.
Every watch in that dataset has attached to it metadata such as the category it was entered into, its broader category family (in order to smooth the volatile category data), what rules applied to it at the time, brand information associated with it and the historical results.
Ultimately, what I have tried to do is to model the decisions contained within the GPHG archive to see if there are persistent patterns present – not only describing what the data show (though my reconstruction can help me do that too).
kingflum’s finding that the median prize-winner was more expensive than the median entrant in every one of the 18 editions for which he had price data is striking. But comparing winners with the full field does not, by itself, tell us whether price necessarily helps to predict a winner – or at which stage any price effect arises.
It could be that expensive watches come from brands that tend to win more frequently, or that the price effect is largely contained during the nomination stage. It might be that price does contain useful predictive information even after those other factors are accounted for – but we should find a way to prove it either way.
Descriptive analysis is useful and interesting – what I seek to do is to answer the next questions that naturally arise in my mind from such observations.
To answer those next questions, historical backtests need to be performed, and those tests have to behave as if each award has not yet happened. Signals have to be confirmed as useful and eliminated from the model if they prove not to be.
With hindsight, following the award presentations in a given year, it is always possible to come up with some kind of convincing explanation for why something won. Of course it did.
I don’t want to do that. I want to see whether there are patterns in the data that can predict outcomes – and interpret what those patterns might mean.
Why it’s me doing this analysis
At this point I should be clear about the scope and limitations of this work – and its author.
I’m an independent, occasional watch writer – with a full-time, busy job outside of the industry – and definitely not an academic or a statistician.
My analysis has not been peer-reviewed by, well, anyone. It is not scientific research. And it definitely hasn’t been commissioned, approved or endorsed in any way by the GPHG – though I have informed them that I am doing it.
I do, however, spend my working life as a design leader in technology. Since my “fun” experiment last year, I’ve developed a bit more of an understanding about what AI models can do, when they are useful and, importantly, how confidently they frequently produce utter nonsense.
There are understandable reasons an actual academic researcher might not take this analysis on. It is a highly specialised subject with a narrow audience, the underlying data is awkward and the academic rewards are not obvious (at least to me, a non-academic). And it is also probably too labour-intensive for most watch journalism. I have undertaken this analysis in my spare time – albeit slightly obsessively – with no one telling me that I really shouldn’t bother.
I should also say that without AI this work would have been almost entirely impractical for me to complete.
Over the course of just the last year, AI models and tools have advanced to the point where they have changed the economics of this kind of analysis and made it possible for a person like me to do it. They helped with extraction, classification, comparison, testing and documentation at a scale I could not have managed by myself.
As A Watch Critic noted recently, using AI does not remove the need for judgement, particularly when it comes to watches. And there was a lot of checking and judgement needed here. Theoretically someone could undertake similar analyses to mine; but they might not make the same judgement calls I have made.
In summary, my analysis became possible because the GPHG has a well-maintained public archive, advances in AI made a previously impractical project possible, my professional life gave me some understanding of the tools – and my interest in what watches mean made me sufficiently obsessive to keep going when a much more sensible person would have surely done the right thing and stopped.
The historical analysis
My analysis looks at a handful of what I think are intuitive questions about what is associated with – and potentially predicts – GPHG success.
Questions such as: Does a brand’s previous awards record matter? Is recent success more important than older successes? Do brands develop particular strength in certain categories over time? Does the number of watches entered over multiple years or in a single year help? Are some categories easier to predict than others?
I also looked at whether data adjacent to the awards are signals that, over time, contain any useful information about the eventual results.
The distinction between “associated with winning” and “improves prediction” is something I have kept key to my analysis – and is the type of distinction my data-scientist colleagues have drummed into me during the course of my tech career. And even then it was hard to stay focused.
Two signals may be capturing the same underlying advantage. A pattern may look convincing when all 25 years are viewed together but disappear when each year is tested using only the past available to it.
A particular hypothesis may provide an excellent explanation for one striking result but might not establish a pattern that persists across years and award categories.
This is as much as I’m going to say right now about the model I’ve built and the signals I’ve tested, particularly while 2026 voting remains active.
A sealed test (in public)
The historical testing is the meat of the analysis I’ve done. Predicting the 2026 awards is less so, but still something I wanted to do. Maybe it will become an annual tradition. Let’s see how this year goes…
Whatever this year’s result, it can’t by itself validate or invalidate the 25 years of findings I’ve captured – those already exist and are verifiable.
But this year is particularly interesting as the jury has a new president, the category structure has changed fairly significantly and the awards – as they are every year – are ultimately decided through subjective human judgement rather than a static formula.
A strong forecast would be interesting, but then so would a poor one given the historical results that I will share down the line.
I am, however, conscious that any model that only explains past results is going to raise a little suspicion.
I therefore froze the methodology I landed on before applying it to the 2026 field of competing watches, recording the model version, data cut-off and decision rules in advance. I then used the model to create a forecast before the nominated watches were announced.
It includes:
The watches that the model predicts will progress from competing watches to nominated watches – that is, who the Academy votes for.
Winner predictions for each of the 14 standard competition categories.
A separate ranking for the Aiguille d’Or.
The forecast is linked below as a password-protected PDF.
You can download it now and save it locally. You just can’t read it yet.
Why? Because publishing the PDF now establishes that the predictions existed before the nominated watches and prize-winners were known, and, importantly for the integrity of this whole exercise, prevents me from changing them without someone who downloaded the PDF observing me doing so.
Once the nominated watches are announced, I will apply the same frozen methodology – without further adjustment – to that reduced field and publish a second sealed PDF. The first forecast tests both nomination and winner predictions from the full field of entrants. The second will test winner predictions once the nominated watches are known.
The passwords for both PDFs will remain secret while GPHG voting is active. I don’t want to inadvertently affect the voting, even in a small way.
When the voting has closed, I will share the passwords for both documents. Anyone will then be able to open them and see exactly what was predicted by the model before the results were known.
I have also extensively documented my methodology and process alongside capturing all the data so that someone could replicate exactly what I have done using the data I used. I won’t be making this widely available, but I may share more detail with a select group of people.
What comes next
This is the first in a series of Free Sprung articles on reading the GPHG archive.
The series will look at what those decisions can tell us about recognition, repetition and changing preferences within the GPHG – and even what that might tell us about the wider watch industry. It will also explore why several explanations that sound intuitively plausible for a watch winning an award become much less convincing when they are used to predict an unknown result.
The sealed PDF linked in this article will remain unchanged throughout that process, providing a record against which the eventual results can be judged.
I will share more historical data, some visualisations and details about the statistical analysis, explained as plainly as I can, showing how the model and its methodology work without necessarily giving away anything I consider proprietary.
Beyond that I will seek to answer this question: What patterns emerge when the GPHG decides what to nominate and what to reward?











Looking forward to seeing how this pans out. Massive kudos to you for the amount of effort put in here! 👏
Fantastic work Chris. I love the idea that genAI makes incredibly niche research like this accessible and possible, and I’m looking forward to seeing how your model predicts the results. Exciting times!