I mentioned the book The Formula: The Universal Laws of Success by the physicist and Northeastern University professor, Albert Laszlo-Barabasi in a recent post (see "Blind Taste Tests and the Universal Laws of Success"). Laszlo-Barabasi uses professional wine tasters as an example to help explain why there isn't always a perfect correlation between performance and success. In reality, there's often a fairly poor correlation between the two!
Given my growing obsession with wine, I was fascinated by his claim that professional wine tasters aren't always accurate in their assessments of fine wine! In his book, Laszlo-Barabasi references two studies performed by Robert Hodgson, an emeritus professor of oceanography and statistics at Cal Poly Humboldt. Hodgson is also a vintner who started making wine in 1976 at the Fieldbrook Winery, which he co-owns and operates with his wife Judy, in northern California. As the story goes, Hodgson started wondering how one of his wines could win a gold medal at one wine competition and "end up in the pooper" at another competition held in the same year. So, in order to better understand how professional wine tasters evaluated wines, he studied to become one himself. At some point, he became involved with the California State Fair wine competition, the oldest and most prestigious wine competition in North America. He convinced the other judges (after what he described as "lots of begging") to allow him to run a controlled scientific test of the wine tasting competition held there.
As he discussed in a podcast called Vintage Year, hosted by the Journal of Wine Economics, which published his study (see "An examination of judge reliability at a major U.S. wine competition"), Hodgson found that expert wine judges are rarely, if ever, consistent. Between 65-70 judges participated in the study at the California State Fair wine competition held every year from 2005 to 2008. Each panel of expert judges received a flight of 30 wines, which included four triplicate samples of the same wine poured from the same bottle (unbeknownst to the judges). Only ten percent of the judges were able to replicate their scores on the same wine. According to the 100-point wine rating system first popularized by the American wine critic Robert M. Parker, Jr (more on him shortly), the judges' ratings typically varied by plus or minus four points on scales ranging from 80 to 100. In other words, a wine rated 91 by one judge at the first tasting could be rating an 87 or even a 95 at the next tasting (again, each wine was scored three different times by the same judge in the same flight). Only 1 judge out of 10 could score the same wine with a two point range using the same scale.
In a follow-up study, again published in the Journal of Wine Economics (see "An analysis of the concordance among 13 U.S. wine competitions"), Hodgson analyzed over 4,000 wines entered into thirteen different U.S. wine competitions. Of the wines entered in more than three competitions, roughly 47% received a Gold medal at at least one competition. However, 84% of these exact same Gold-winning wines received no award at all in at least one other competition. When he analyzed all of the data, Hodgson concluded that the probability of receiving a Gold medal at any competition depended largely upon chance alone. Importantly, certain wines consistently received no award or Bronze medals across multiple competitions, suggesting that judges mainly agree on what they dislike rather than what they love.
Later that same year, the American physicist and author Leonard Mlodinow wrote an article for the Wall Street Journal reporting on Hodgson's findings, "A hint of hype, a taste of illusion" (I've posted about Mlodinow in the past, see "Drunkard's Walk"). Mlodinow references a number of other studies that shed further light on this issue. Apparently, as far back as 1963, a team of investigators from the University of California, Davis added flavorless food coloring to a dry white wine to simulate several different varieties of red wine. The professional tasters description of each wine reflected the type of wine that they thought they were tasting (e.g. sherry, rosé, Bordeaux, burgundy, etc) as opposed to what the wine actually was - a dry white wine.
Several years later, a French neurophysiologist named Frédéric Brochet served 57 French sommelier students the same Bordeaux wine served in two different bottles. One wine was served in an expensive Grand Cru bottle, while the other was served in the bottle of a cheap table wine. The sommelier students greatly preferred the Grand Cru wine to the cheap table wine, rating the Grand Cru as "excellent" and the table wine as "flat" or "unbalanced" (note that Brochet also published a study as the 1963 study in which food coloring was added to "turn" white wine into red wine, with similar results).
Several studies have replicated Brochet's method of presenting a cheap table wine as an expensive wine (through packaging and pricing information) and vice versa (see, for example, Robin Goldstein's working paper, "Do more expensive wines taste better? Evidence from a large sample of blind tasting"). One study even used functional MRI (which looks at brain activity) and showed that increasing the price of a wine increases subjective reports of wine quality as well as brain activity in the medial orbitofrontal cortex (one of the area's of the brain responsible for pleasantness). I guess that's nature's way of telling us that we can't tell a cheap wine from an expensive wine when it comes to flavor, but instead, we just need to look at the price on the bottle!
As you can imagine, Hodgson's results were fairly controversial when they were first released to the public. What's compelling is that the test subjects in his study were professional wine judges. The Goldstein study referenced in the paragraph above included both professionals and non-professionals. Non-professionals slightly preferred the cheaper table wines to the expensive wines in blind taste tests, while professionals did slightly better picking the more expensive wines correctly, though still leaving a lot to be desired. Similarly, Richard Wiseman, a professor of psychology at the University of Hertfordshire conducted a study in which over 400 members of the general public correctly classified their wines as expensive versus cheap about 50% of the time, which is about what would be predicted based on chance alone. So, the question is whether someone can be trained to taste wine accurately. Hodgson's studies suggest not.
While reporting on Hodgson's results, Leonard Mlodinow contacted the American wine critic Robert M. Parker, Jr. According to an article published in December 200 in The Atlantic magazine, "The Million-Dollar Nose" (the title of the article comes from the fact that Parker has a million dollar insurance policy for his ability to smell and taste fine wine), Parker tastes approximately 10,000 wines every year and claims that he can remember every wine he has tasted over the past 32 years and every score he has given, within a few points. Parker told Mlodinow that he didn't find Hodgson's results too surprising, and he further confessed that he had participated in a blind wine tasting of fifteen top wines from Bordeaux in 2005 (which he labeled, "the greatest vintage of my lifetime"). Parker apparently failed to correctly identify any of the wines, and confused so-called "left bank" wines and "right bank" wines several times during the tasting (note, in general "left bank" wines are grown in regions west of the Gironde Estuary, and "right bank" in the regions east). Parker had previously tasted and rated each wine, which was why these 15 wines were specifically selected. He told Mlodinow that his second rating (during the blind taste test) was always within 2-3 points of the first, unblinded tasting (which all ranged from 95-100).
While Parker did not do too bad, he is, after all, the world's expert at wine tasting. For all of us mere mortals, there appears to be no hope. Getting back to Albert Laszlo-Barabasi's argument in The Formula: The Universal Laws of Success, there are all kinds of reasons to explain why professional wine tasters fail to perform consistently or, at least in blind taste tests, accurately. He states, "Wine competition judges don't fail because they lack expertise, preparation, or thoroughness. They fail mainly because the wines they're judging are all excellent...And they have no 'stopwatch', no simple tool to determine which anonymous glass of Malbec, raised to the nose and swished in the mouth, is, definitively, a winner." More generally, "even though performance does drive success, the problem is that the differences among top contenders are so tiny that they're often nearly immeasurable."
