Showing posts with label Presenting data. Show all posts
Showing posts with label Presenting data. Show all posts

Monday, May 12, 2025

Issues when aggregating wine scores into an average

Wine competitions, and many web sites, involve summing assessors’ scores into a consensus “average” score for each wine. However, as an example for three assessors, the scores of:
5,5,5 have the same sum / average as 0,5,10
However, the former situation indicates complete agreement among the three assessors about the quality of the wine, while the latter situation is no different from random quality scores. Surely this difference matters?

This contradictory situation has long been ignored. Obviously, this issue does not matter when looking at a single critic’s review in a magazine, for example. However, it may matter enormously at sites like CellarTracker, which claim to represent the consensus of wine quality among many people. However, such sites seem not to have cared about this issue at all.

Recently, Jeffrey Bodington has looked at this situation in detail:
Wine Stars & Bars: the combinatorics of critic consensus
AAWE Working Paper no. 284

Here, I will summarize some of his ideas.

Judgment of Paris whites

First, however, let’s make the situation clear. The following lines indicate increasing agreement around the same average of 5, with 3 assessor scores having a maximum of 10 each:
0,5,10
1,5,9
2,5,8
3,5,7
4,5,6
5,5,5
Furthermore, there can also be clusters of scores (e.g. some scores indicating poor quality and some indicating good quality, with nothing in between), such as:
1,7,7
3,6,6

The basic mathematical issues here are that (potentially) many billions of combinations of scores have the same sum (and therefore average), and that uncertain ratings can have many different sums. Mathematically, an observed wine rating is one draw from a (latent) distribution of all possible ratings that is both wine-specific and judge-specific.

The further practical issues for wine assessments are that: (i) sample sizes (number of assessments) are often small (especially in competitions); (ii) some wine judges are more reliable or consistent assessors than are others; and (iii) clusters of scores can happen, for example in the case of stylistically distinctive wines. These three situations mean that the issue discussed by Bodington can potentially have a big effect.

Bodington proceeds mathematically:
A weighted sum of judges’ wine ratings is proposed and tested that (1) recognizes the uncertainty about a sum and (2) minimizes the disagreement among judges about that sum. A simple index of dispersion is [also] proposed and tested that measures a continuum from perfect consensus (dispersion = 1), to ratings that are indistinguishable from random assignments (dispersion = 0), and then to distant clusters of ratings when groups of judges disagree (dispersion is negative).
To make sense of this for you, he then illustrates his ideas with a straightforward example. This involves the 10 white wines from the 1976 Judgment of Paris comparative tasting of French and American wines, with 9 assessors per wine. The sums of the blind scores are shown in the graph above. The blue bars indicate the distribution of all possible sums of 9 scores of 20 each (ie. a minimum of 0 and a maximum of 180). The black lines represent the sums for each of the 10 Judgment wines (as labeled). As shown:
... the sums of points for the top two white wines, Chateau Montelena and Meursault Charmes are calculated to be the same at 130.5. The respective ranges of points assigned to those two wines were 3.0-to-18.5 and 12.0-to-16.0 ...
So, the overall mathematical assessment is the same for the two wines, but there is clearly much more consensus among the judges for Meursault Charmes (scores 12-to-16 out of 20) than for Chateau Montelena (scores 3-to-18.5). Bodington thinks that this difference should be dealt with, and that is the purpose of his weighted sum and his index of dispersion.

These calculations are shown in the next figure, with the weighted sum shown horizontally (ie. increasing assessed quality of the wine) and the index of dispersion vertically (ie. increasing agreement among the assessors), and each wine represented by a labeled point.

Bodington's calculations

Bodington notes:
Results for the dispersion index show that none are close to zero so they do not appear to be random results, and none are negative to indicate distant clusters. The weighted sum of points for highest-scoring Chateau Montelena has the second lowest dispersion index of any wine and the [other] wine Meursault Charmes has the highest index of any wine. Considering that finding, does it make sense to conclude that Montelena was better than Charmes?
In other words, the consistent critic judgements for Meursault Charmes should outweigh the relatively inconsistent ones for Chateau Montelena. This can be interpreted as indicating the “best” white wine at the Judgment of Paris.

This sort of situation can have a strong effect any time there is a relatively small number of wine assessments.

Monday, September 9, 2024

Why alcohol experiments are problematic

I recently published a post (Has the WHO lost its way regarding alcohol?) pointing out that what is recognized to be the best form of scientific experiment is what is called a “double-blind treatment–control” experiment (or sometimes a “randomized controlled trial”, or RCT). This procedure cannot usually be done ethically on people, and therefore no-one has ever admitted doing it in medical science. So, we will never have the best possible scientific evidence about the effects of alcohol on human health.

This does not mean that we do not have experiments about wine and health. What the scientists do is the best that they ethically can; and some of the pros and cons of this process is what I will discuss in this post. I cover several different but important topics.


What the researchers do is to follow groups of drinkers and non-drinkers through time, and see how these people get on — this is called a “descriptive” study rather than a “manipulative” one (as described above), of which there are several types as listed in the above picture. Here, we are concerned with the first one in the list, “observational”. In this type of study, health and behavior experts measure all of the consistent differences they can find between the studied people, to see what matches their patterns of their drinking and non-drinking. The results of the famous 1926 study by Raymond Pearl are shown in the next graph, as but one early example.

Also, we should ideally do all of this in such a manner that the people involved do not know which of the experimental groups they are in, and nor do the people evaluating their behavior (this is what “double blind” means). Is this actually feasible? Of course not. We can’t force people into the experimental groups, we can only ask them to volunteer to participate. That is, we rely on self–reporting of their alcohol consumption (often via detailed questionnaires filled in by the participants). Let’s look at the consequences of this now.

Here, is one useful recent discussion of how we might justifiably proceed (Causal inference about the effects of interventions from observational studies in medical journals):
Building on the extensive literature on causal inference across diverse disciplines, we suggest a framework for observational studies that aim to provide evidence about the causal effects of interventions based on 6 core questions: what is the causal question; what quantity would, if known, answer the causal question; what is the study design; what causal assumptions are being made; how can the observed data be used to answer the causal question in principle and in practice; and is a causal interpretation of the analyses tenable?

Raymond Pearl's 1926 observational study

This is all well and good, but the biggest recognized issue is how to choose the studied people. Basically, the choice should be literally random, but this is impossible. As a discussion of the problems with one example of this, we have:
In particular, people volunteer to take part in experiments, and this can never be described as “random”. Notably, people’s admissions regarding their own drinking may not be accurate (see also: Wine ratings involve both the accuracy and bias of the raters). This is discussed here:
As one specific example discussion (Reweighting UK Biobank corrects for pervasive selection bias due to volunteering):
Volunteers tend to be healthier and of higher socio-economic status than the population from which they were sampled ... Volunteer bias in all associations, as naively estimated in UKB, was substantial — in some cases so severe that unweighted estimates had the opposite sign of the association in the target population. For example, older individuals in UKB reported being in better health, in contrast to evidence from the UK Census.
So, we often end up with this rather jaundiced (but realistic) view:

The original of this cartoon hangs on my study wall

Moving on, another potential approach, to getting around the sampling problems discussed here, is to use animals as a substitute for humans (eg. mice or dogs). This generates a lot of emotional response from parts of the public, which I will not delve into here. Instead, I will focus on the actual experiments.

One useful discussion is (The flaws and human harms of animal experimentation):
Nonhuman animal (“animal”) experimentation is typically defended by arguments that it is reliable, that animals provide sufficiently good models of human biology and diseases to yield relevant information, and that, consequently, its use provides major human health benefits. I demonstrate that a growing body of scientific literature critically assessing the validity of animal experimentation generally (and animal modeling specifically) raises important concerns about its reliability and predictive value for human outcomes and for understanding human physiology ... The resulting evidence suggests that the collective harms and costs to humans from animal experimentation outweigh potential benefits and that resources would be better invested in developing human-based testing methods.
One important point is whether animal experiments actually lead to any benefit for human medical treatments. Sadly, it seems mostly not (Analysis of animal-to-human translation shows that only 5% of animal-tested therapeutic interventions obtain regulatory approval for human applications):
There is an ongoing debate about the value of animal experiments to inform medical practice, yet there are limited data on how well therapies developed in animal studies translate to humans. We aimed to assess 2 measures of translation across various biomedical fields: (1) The proportion of therapies which transition from animal studies to human application, including involved timeframes; and (2) the consistency between animal and human study results ... The overall proportion of therapies progressing from animal studies was 50% to human studies, 40% to RCTs, and 5% to regulatory approval.
Drink in moderation

I think that you can all see the bottom line here: things are not likely to get any better any time soon, regarding experiments of alcohol intake by humans. The medical scientists are doing the best that they ethically and practically can, in the real world. This, however, does not match what they would be doing in the theoretical world of scientific experiments, which is what would be the best for devising effective medical ideas.

Nevertheless, I am not the only one who has noted that there is: ‘No good evidence’ of risk from low-level alcohol consumption. Basically, the risks of one or two drinks per day are so low that they are very difficult to estimate; and drinking with meals also seems to be unproblematic (Drinking wine with meals linked to better health outcomes). Alternatively, there definitely are risks associated with heavy alcohol consumption, and people with known health problems related to alcohol may not have any safe level of consumption, as well as pregnant women. People with a family history of alcohol abuse also need to be careful.

Monday, January 22, 2024

Has WHO got it wrong with its new zero-alcohol policy? Probably.

A year ago, the World Health Organization (WHO) changed its attitude towards alcohol consumption, which it said it would recommend reducing as much as possible, because there is “no safe level of alcohol”, and that alcohol is associated with several different types of cancer. I wrote about this change in attitude in my previous post: Who started the current WHO completely negative attitude towards alcohol?. This is a follow-up post, so that previous one could be read as background information.

The important point of that post was that the previous (long-standing) evidence for possible beneficial effects of a small intake of alcohol on human mortality has recently been called into question. The previous evidence had been based on observing a so-called J-curve when plotting human mortality against alcohol intake, as shown in the first figure (below). Naturally, this graph might be right or wrong, and this distinction is the point at issue here.

J-curves of motality versus alcohol intake

Previous advice from WHO (ie. before last January) was based on comprehensive studies like this one (from which the above figure was taken): Di Castelnuovo A., Costanzo S., Bagnardi V., Donati M.B., Iacoviello L. and de Gaetano G. (2006) Alcohol dosing and total mortality in men and women: an updated meta-analysis of 34 prospective studies. Archives of Internal Medicine 166: 2437-2445. This 17-year-old paper concluded:
Low levels of alcohol intake (1-2 drinks per day for women and 2-4 drinks per day for men) are inversely associated with total mortality in both men and women. Our findings, while confirming the hazards of excess drinking, indicate potential windows of alcohol intake that may confer a net beneficial effect of moderate drinking, at least in terms of survival.
More recently, there are also summary papers like this one: Giovanni de Gaetano and Simona Costanzo (2017) Alcohol and health: praise of the J curves. Journal of the American College of Cardiology 70: 923–925. Clearly, this one supports the existence of the J-curves!

However, since then, this J-curve graph has been claimed to not be J-shaped after all, but to be monotonically increasing instead (ie. the more alcohol consumed then the greater the mortality), leading to the conclusion that the safest amount of alcohol is zero intake. That is, the J-curve was previously accepted as being correct, but it is now claimed to be wrong. This conclusion was clearly stated in a report by the Global Burden of Disease (GBD) collaborators; and the WHO has followed them.

My previous blog post called this new conclusion into question. I noted that, while this conclusion is literally true, it is not all of the truth — small amounts of alcohol were shown to be equally as safe as zero alcohol intake. I also claimed that I am appalled by this act of omission (leaving out part of the truth). I will continue my story here, pointing out some limitations of the study mentioned above, along with updated information from a more recent paper.

The GBD collaborators paper

The paper that I have been referring to above, by the Global Burden of Disease collaborators, is:
Alcohol use and burden for 195 countries and territories, 1990—2016: a systematic analysis for the Global Burden of Disease Study 2016. Lancet (2018) 392: 1015–1035.

Since I am questioning it, let’s look at what is in there. Their written summary in two sections of the paper is:

 Methods
Using 694 data sources of individual and population-level alcohol consumption, along with 592 prospective and retrospective studies on the risk of alcohol use, we produced estimates of the prevalence of current drinking, abstention, the distribution of alcohol consumption among current drinkers in standard drinks daily (defined as 10 g of pure ethyl alcohol), and alcohol-attributable deaths and DALYs [disability-adjusted life-years]. For our exposure estimates, we extracted 121,029 data points from 694 sources across all exposure indicators. For our relative risk estimates, we extracted 3,992 relative risk estimates across 592 studies. These relative risk estimates corresponded to a combined study population of 28 million individuals and 649,000 registered cases of respective outcomes.
 Findings
Globally, alcohol use was the seventh leading risk factor for both deaths and DALYs in 2016, accounting for 2.2% (95% uncertainty interval [UI] 1.5—3.0) of age-standardised female deaths and 6.8% (5.8—8.0) of age-standardised male deaths, [so that] the attributable burden for men around three times higher than that for women in 2016. The level of alcohol consumption that minimised harm across health outcomes was zero (95% UI 0·0—0·8) standard drinks per week.
There is no doubt that the contributors have performed an impressive study. They have collated a massive amount of data, and developed some innovative ways to analyze that data, accounting for previous limitations. However, there is still one basic limitation in this type of work — the authors compiled data from pre-existing sources, rather than doing an experiment of their own.

There are ways to grade what is called The Strength of Evidence of any published scientific paper. In this case, Lewis Perdue’s Stealth Syndromes Study grades this type of paper as only Strength C, with this comment:
Published pre-prints may be credible depending upon the study design (clinical, randomized, etc.), the investigators, methods, and institutional affiliations.
Basically, there are these possible Complicating Factors For Human Studies:
    C-SRD: Self reported / selected data
      C-SRDb: Social pressure / desirability approval bias

So, we do not have Grade A evidence, or even Grade B evidence. There are thus serious limitations to the conclusions from the study, and we should bear that in mind when evaluating them. Basically, these are what we call “observational” studies, which do not yield causal data, but merely offer potential connections or indications between observations and conclusions.

Updated data concerning mortality and alcohol intake

Follow-up paper

The paper discussed above is from the Global Burden of Disease (GBD) Study 2016. There has been another part of this series of studies that has been published since then, this time by the GBD 2020 Alcohol Collaborators:
Population-level risks of alcohol consumption by amount, geography, age, sex, and year: a systematic analysis for the Global Burden of Disease Study 2020. Lancet (2022) 400: 185–235

Their notes about their new work are:
For this analysis, we constructed burden-weighted dose–response relative risk curves across 22 health outcomes to estimate the theoretical minimum risk exposure level (TMREL) and non-drinker equivalence (NDE), the consumption level at which the health risk is equivalent to that of a non-drinker, using disease rates from the Global Burden of Diseases, Injuries, and Risk Factors Study (GBD) 2020 for 21 regions, including 204 countries and territories, by 5-year age group, sex, and year for individuals aged 15–95 years and older from 1990 to 2020.
This study thus focuses on variation among ages and countries, which is valuable in-depth information. Their results are:
The burden-weighted relative risk curves for alcohol use varied by region and age. Among individuals aged 15–39 years in 2020, the TMREL varied between 0 (95% uncertainty interval 0–0) and 0·603 (0·400–1·00) standard drinks per day, and the NDE varied between 0·002 (0–0) and 1·75 (0·698–4·30) standard drinks per day. Among individuals aged 40 years and older, the burden-weighted relative risk curve was J-shaped for all regions, with a 2020 TMREL that ranged from 0·114 (0–0·403) to 1·87 (0·500–3·30) standard drinks per day and an NDE that ranged between 0·193 (0–0·900) and 6·94 (3·40–8·30) standard drinks per day.
Note that they refer to the existence of J-curves for some age groups. So, J-curves do exist! Even using roughly the same data as in 2016! Furthermore, note that zero drinks is, indeed, the lower alcohol limit for safety, but that the authors also have a table updating the 2016 results to much less extreme levels. This table is shown as the second figure above.

This sort of apparent conflict among publications is the basic problem with what we call meta-analyses (where the results of multiple studies are considered together). It matters very much which studies are included in the meta-analysis, and what data analyses are done on the results (this is how you end up with Strength C evidence).

Conclusion

So, there you have it. The latest research (2020) is much less extreme in its conclusions about the mortality associated with low alcohol levels than is the previous one (2016). Imagine what might come next! The World Health Organization needs to take note, since it used the first one, but not the second one, even though the latter was published before WHO produced last year's recommendations. There is clearly no longer a consensus about alcohol — there is some sort of controversy, not a clear-cut solution. It seems to be far too early for WHO to make such a definitive (unambiguous) recommendation.

Monday, January 15, 2024

Who started the current WHO completely negative attitude towards alcohol?

I think that I have the answer to the title question, and that answer appalls me. The World Health Organization (WHO) used to accept the idea that small amounts of alcohol were not necessarily bad for you, and may actually have positive effects on some aspects of health. They no longer accept this — now, all amounts of alcohol are considered to be bad. Read on to see why they changed their minds.

I have written about this topic before, but I think that this is a very important one for the wine industry, as this seems to be one of the biggest global threats to that industry (along with local threats from neo-prohibitionists, etc), as also is global warming. This post is actually split into two halves, and will thus be continued next week.

As background, I've written several posts recently about wine and health:

Previously

The reason for the previous positive attitude towards alcohol is summarized by Mark Hicken (Don’t let anti-alcohol grinches ruin your holidays):

The science related to safe levels of [alcohol] consumption has not changed. Hundreds of studies, and decades of scientific research, have consistently shown that those who drink in moderation live about as long (or even slightly longer) than those who don’t drink at all. The reality is that moderate drinking provides some cardiovascular benefits while slightly increasing the risk of certain cancers, some of which are very rare ... For most people, there is little or no effect on overall health and mortality.

The so-called J-curve of mortailty and alcohol

This idea is usually pictured as a so-called J-curve, as shown above. It indicates that small amounts of alcohol (eg. one standard drink per day) actually reduce the risk of people dying (compared to zero alcohol), due to various medical causes. This particular picture is from: Giovanni de Gaetano and Simona Costanzo (2017) Alcohol and health: Praise of the J curves. Journal of the American College of Cardiology 70: 923–925.

Indeed, the Canadian Association for Responsible Drinkers hosts a whole web page covering this topic: Recent Studies on Alcohol + Health. It lists 12 science / medicine studies published from 2018—2023 confirming the health effects of moderate alcohol consumption. It also has links to pages containing both academic and medical commentary on the matter.

In spite of all of this, groups like the World Health Organization (WHO) would now have us believe that there is “no safe level of alcohol consumption and that alcohol causes cancer” (WHO shifts its alcohol narratives). Proactively, the WHO has suggested reducing consumption via global tax increases on “unhealthy products”, including wine (WHO demands tax increases on alcohol and sugar), as I recently discussed (WHO and the use of taxes to reduce alcohol consumption).

What caused WHO to change their tune?

The WHO makes it clear that their change of tune is based on accumulating medical evidence. This leads me to ask an obvious question: what is the first of the recent medical / scientific studies that made these new claims?

My research leads me to identify this published scientific paper, which appears to be a very important one of them, from 2018:

Alcohol use and burden for 195 countries and territories, 1990—2016: a systematic analysis for the Global Burden of Disease Study 2016. Lancet (2018) 392: 1015–1035. (Max G Griswold seems to be the senior author, and Emmanuela Gakidou is the corresponding author.)
This paper is actually a summary of a more detailed report:
Global Burden of Diseases, Injuries and Risk Factors Study 2016. This was authored by the MGBD 2016 Alcohol Collaborators (517 people are listed as the collaborators).

The revised J-curve

The conclusion from the published 2018 paper is:
Alcohol use is a leading risk factor for global disease burden and causes substantial health loss. We found that the risk of all-cause mortality, and of cancers specifically, rises with increasing levels of consumption, and the level of consumption that minimises health loss is zero.
Note that this is literally true, but that it is not all of the truth! Their Figure 5, as shown immediately above, is their revised version of the J-curve (as shown in the top figure). Note that the mortality curve does not drop below zero, which is the point that the authors are emphasizing. However, based on this graph, the authors could equally accurately have said that “the level of consumption that minimises health loss is one drink per day”. This is an act of omission, not commission — what they say is literally true, but it is only half of the story. That is why I am appalled!

An evaluation of the 2018 paper

Anyway, one can see why WHO might change their tack. The report is unambiguous, and actually concludes:

These results suggest that alcohol control policies might need to be revised worldwide, refocusing on efforts to lower overall population-level consumption ... In terms of reducing population-level alcohol use, WHO provides a set of best buys—policies that provide an individual year of healthy life at less than the cost of the average individual income. Governments should consider how these recommendations can be implemented within their local contexts and broader policy platforms, including excise taxes on alcohol, controlling the physical availability of alcohol and the hours of sale, and controlling alcohol advertising.
There is no doubt that the contributors have performed an impressive study. They have collated a massive amount of data, and developed some innovative ways to analyze that data, accounting for previous limitations. However, there is still one basic limitation in this type of work — the authors compiled data from pre-existing sources, rather than doing an experiment of their own. I will discuss this limitation in next week’s post.

Meanwhile, I noted above that previous studies found a positive effect on health of small amounts of alcohol (the so-called J-curve). When discussing these previous J-curves, the authors note:
Past findings subsequently suggested a persistent protective effect for some low or moderate levels of alcohol consumption on all-cause mortality. However, these studies were limited by small sample sizes, inadequate control for confounders, and non-optimal choices of a reference category for calculating relative risks. More recent research, which has used methodologies such as mendelian randomisation, pooling cohort studies, and multivariable adjusted meta-analyses, increasingly shows either a non-significant or no protective effect of drinking on all-cause mortality or cardiovascular outcomes. Our results on the weighted attributable risk are consistent with this body of work.
In estimating the weighted relative risk curve, we found that consuming zero (95% UI 0·0—0·8) standard drinks daily minimised the overall risk of all health loss (figure 5; shown above). The risk rose monotonically with increasing amounts of daily drinking. This weighted relative risk curve took into account the protective effects of alcohol use associated with ischaemic heart disease and diabetes in females. However, these protective effects were offset by the risks associated with cancers, which increased monotonically with consumption.
So, there you have it — the monotonic increase in mortality with increasing alcohol consumption is literally true, based on their data, but it is also true that lower levels of alcohol consumption have no notable difference in effects. The WHO have changed their mind for a very dubious reason. There is more to this topic, which I will cover in my next blog post.

Monday, July 12, 2021

Do online wine ratings and searches actually mean anything?

Social media sites like Vivino, Delectable and Cellar-Tracker collate wine ratings from their users, and Wine-Searcher does the same thing for critics, as well looking at the search popularity of wines. These sites sometimes present a compilation of their data as representing things like "the best wine in the world" (eg. Is this the best wine in the world? An app with 35 million subscribers says so) or "the world's most desired wines" (eg. The world's most wanted wines).


This seems like quite a radical conceptual leap, to me. It is one thing to note what wine is, in some sense, the most popular wine, on average, for the restricted set of users represented on any given online site. It is another thing altogether to present this as "the best" in any broader sense. To leap from a restricted user base to the entire world is a form of arrogance, at best, and complete and utter foolishness at worst.

After all, a highly rated wine (in the most general sense) may have little to do with "the best" for anyone other than the people doing the rating. First of all, ratings may have nothing to do with quality, but only with a desire to rate, based on any criteria you wish to name (eg. Should critics rate wine based on environmental impact?) — a rating is more like a popularity contest, rather than a quality evaluation. Second, the raters themselves are rarely representative of any group other than people who wish to provide ratings — they may have any motive at all for doing so. Third, the products being rated are rarely representative of the range of products available — they tend, for instance, to be associated with their snob value rather than their value for money.

So, does the highest average wine-rating on a site like Vivino represent the quality of the wines being rated, or does it represent their snob value? Vivino regularly tells us it is the former, while I suspect it is more likely to be the latter. Does the number of wine-label searches carried out on a site like Wine-Searcher represent the desirability of the wines being searched for, or does it represent their curiosity value? Once again, Wine-Searcher regularly tells us it is the former, while I suspect it is more likely to be the latter.

Social media and wine

What use, then, is the information provided by these types of sites? Sure, they are a valuable outlet for the modern penchant for social-media opinions, just like Facebook, Twitter, and YouTube in their respective domains. Given their extensive usage, I presume that these sorts of sites are providing a useful service for Millennials and Generation-Xers — my parents' generation nattered over the back fence or down at the pub, by my children's generation natters online. So, I doubt that we can take these services as anything other than opinions; and certainly not as a source of quantitative information about the actual goods and services, which are hiding firmly in the background. A Twitter storm, for example, has little to do with the topic at hand, but much more to do with human behavior when acting in groups — we learn much more about the people than we do about the topic.

These kinds of observations are a basic tenet of the social sciences. Getting reliable quantitative information out of human beings takes a lot of effort; and it is a topic that has been actively studied for more than a century, without any explicit resolution (see Best practices for survey research). Formally conducted surveys can be useful, although they have their own set of limitations, notably to do with what is known as sampling bias. Social media might seem like a quick and easy way to get at the same type of information; but I doubt that this works. It is more likely "quick and dirty", with the emphasis on the latter; because the "survey" respondents are self-selected (ie. they choose to use the site, and they choose to provide a rating) — this is known to be the worst form of sampling bias.

So, do not allow yourselves to be gulled by pronouncements from online social media sites. They have no more real information about the world than you do — they know only about the opinions of their own users, whoever they may be, and whatever motives they might have.


On a personal note, rarely are the wines commonly mentioned by most of the social-media sites of any relevance to me, in practice. Wine is not for bragging about, but for consuming with dinner (the benefits of which are explained here). As such, value for money is the main information of interest to me prior to a purchase. I will try any wine from anywhere, if there is evidence of it being good value for the money being asked. Most of the wines being bragged about and searched for online are therefore of no practical interest at all. It seems to be a pity that they get most of the attention, while "my type of wine" requires more effort to locate (see Calculating value for money wines). In this sense, the so-called "social media" is often anything but social (and, yes, I am quite well aware that a blog is a form of social media!).

Monday, August 3, 2020

The Judgment of Paris demonstrated nothing, statistically speaking

Way back in 1976 there was a comparative tasting of several fancy (and expensive) American and French wines, in which the US wines acquitted themselves quite well. This became known as The Judgment of Paris, a classical allusion that probably escapes most people.

The media made much of this outcome at the time, claiming that the American wines “won” some sort of contest, and should now be considered to be the best in the world, or at least the equal of the French stuff. Indeed, whole books have been written about this; and the volume of words in the media is horrendous to think of.


However, it seems to me that there is one thing missing here, and always has been. Since the quality scores assigned by each taster to each wine are available to us, we can judge for ourselves whether the scores actually show anything more than random variation. After all, the sample size was very small: 20 wines tasted by 11 people. I have already argued the case that the results are very variable in all the wrong ways, if one wishes to compare the wines of two countries (11 tasters and 20 wines, and very little consensus). Well, I have finally decided to bite the bullet, and apply some formal statistical analyses to the data — this blog is the only place likely to try this exercise!

Now, before you all jump up and down reminding everyone of the old saying that there are “Lies, damned lies, and statistics”, bear with me for a moment. That statement dates from the 1890s, long before the development of modern mathematical methods to study data involving probabilities. There are now many powerful methods for analyzing data in an objective and repeatable way, ones that generate no controversy about their outcome. Statistics is no longer whatever you want to make of it.

In the case at hand, there are several possible causes of variation in the quality scores assigned to the wines:
  1. the wines come from two countries — this is what we want to examine
  2. for each country, there were 10 reds and 10 whites selected out of all of the high-quality wines that could have been chosen for inclusion — the results might have been different if other wines had been chosen
  3. there were only 11 people selected as tasters, out of all of the wine critics that could have been chosen — the results might have been different if other tasters had been chosen
  4. these people were of different genders, and there were not equal numbers of males and females.
There are other factors that might be important, of course, but this list is all the information that we have. The point here is that we wish to study factor 1, and to do so we need to take into account the variation caused by factors 2, 3 and 4. This is what modern statistical analysis is all about — studying one source of variation if the face of variation cased by other factors. How do we do this in some objective and repeatable way?

The answer, as formalized by Ronald Fisher in the 1920s, is called Analysis of Variance. It does precisely what the name says — it tries to measure the amount of variation caused by each of the factors. If the variation for any particular factor has a strong pattern, compared to the others, then we can consider it to be an important one, and if the variation has no particular pattern then it is not important. This comparison is made with respect to the variation that is not accounted for by the specified factors.

The concept is fairly simple, but (of course) the mathematics is not. It can often be done by hand, if you are inclined to try that sort of thing, but I always use a computer program. I have used this type of analysis many times in my own research career; and I have even taught this analysis to undergraduate and postgraduate biology students. So, what happens if I apply it to the data from the Judgment of Paris?

Well, you already know the answer to that, from the title of this blog post. The formal details of the analysis are included at the bottom of the post, for those of you who might be interested in them. Here is a summary of the results:

Factor
Country
Wine
Gender
Person
F-value
1.37
5.64
0.01
3.63
Probability
0.256
0.000
0.908
0.000

Each row of the table refers to a formal statistical test of one of the four factors. The F-value (named after Fisher himself) is the result of the calculations — for purely random data this value will = 1. The probability is a measure of how likely it is that the F-value is not > 1. By convention, we might choose p < 0.05 as our criterion for concluding that a factor shows important variation.

As you can see, there is almost no variation between the two genders, in our sample, which may surprise none of you. There is some variation between the two countries, but it is not very large, and is nowhere near “statistical significance”. All of the fuss about the results is nonsense — the differences are nothing more than we would expect from random chance, 26% of the time.

This does not mean that there were no differences between the wines, because there surely were. However, the analysis indicates that these differences had nothing to with which country they came from. Simply put, some wines scored consistently much higher than others, meaning that the tasters agreed that they were better wines, irrespective of where they came from.

Unsurprisingly, the analysis also shows that there were big differences between the critics, with some of them giving much higher scores than others, irrespective of the wine. I produced a graph of this in my previous post on this topic, which I have reproduced here. It shows who scored highest (at the top) and who scored lowest (at the bottom).


Conclusion

The statistical analysis makes it clear that the variation in scores was no different from what we would expect from any single collection of wines and people at one place and one time. Some wines scored better than others, and some people gave higher scores than others, and the country of origin had nothing to do with it.

That is, we would not expect the results to be repeatable. Nor were they, as have I already pointed out in two previous blog posts (Was the Judgment of Paris repeatable? Did California wine-tasters agree with the results of the Judgment of Paris?). Different people tasting the same wines produced different results.

There were lots of differences in scores between the wines and between the people, but not much difference between the countries. So, what was all of the fuss about? Cultural politics, is my guess.




Analyses

I tried several different General Linear Models, using the Minitab package. The Country and Gender factors were both fixed; and the Wine and Person factors were nested within them.

Factor
Country
Wine(Country)
Gender
Person(Gender)  
Type
fixed
random
fixed
random
Levels
2
20
2
11

The simplest model (as reported above) has only the four main factors.

Factor
Country
Wine
Gender
Person
Residual  
DF
1
18
1
9
190
Sum Squares
69.323
908.359
0.456
292.376
1699.168
F-value
1.37
5.64
0.01
3.63
  
P-value
0.256
0.000
0.908
0.000
  

It is also possible to add several of the interactions between pairs of the four factors. However, in this case some of the F-tests will only be approximate, based on adjusted sums-of-squares (these are indicated with an asterisk). The most complex model has three 2-factor interactions; adding the fourth 2-factor interaction collapses the model.

Factor
Country
Wine
Gender
Person
Country*Gender
Country*Person
Gender*Wine
Residual  
DF
1
18
1
9
1
9
18
162
Adj. Sum Squares
50.464
576.866
0.767
250.450
1.146
104.767
67.162
1526.094
F-value
1.47
8.59
0.03
2.39
0.19
1.24
0.40
  
P-value
0.242 *
0.000
0.859 *
0.105
0.701 *
0.277
0.987
  

Note that in this analysis the Person factor is no longer statistically significant (due to the presence of the Person*Country interaction). Otherwise, the conclusions do not change.

Monday, September 30, 2019

How not to write a wine report

I do not usually do this sort of thing, but this time I am going to name names.

Let me explain. I used to teach university students about data analysis and presentation, and one of the exercises for those students was being given a series of published graphs and tables, and then being asked to find the inconsistencies (which I had previously found for myself). The idea was to show them real examples of all the ways one can go wrong when dealing with data, so that the students would (hopefully) be a bit more careful themselves, in their own future professional lives.

However, my examples were always kept anonymous — there is no point in fingering a few practitioners when the problem is actually widespread. We all make mistakes, and we don't necessarily expect to be pilloried for it, especially when most other people are also making the same sorts of mistakes.


In this post I am going to do the same exercise, but this time I am going to tell you the source of the errors, because they are all in the same wine report:
The Irish Wine Market Report 2018, produced by Drinks Ireland.
If you want to try the exercise for yourself, then stop reading here, and go read the report, instead. Make your own assessment of how many inconsistencies you can find; then come back here to see how well you did. Otherwise, read on now.

I have included a small part of each table or figure, to illustrate my points.

My comments

Page 5


The Wine Sales data are apparently in “millions”, but what are the units? Bottles? Liters? Euros? It turns out (as shown on page 8) that they are cases of wine. So, “0.2” ≈ 200,000 cases of wine.

Page 6


The Excise Receipts are apparently in “€”, but that cannot be right. Even the excise itself is €3.19 per wine bottle (the highest in the European Union!). Given the number of bottles per year (pages 5 and 8), I suspect that the numbers are actually each “million €”.

Page 8


The Country of Origin table has a duplicated column. The column is headed “2017”, but it contains a copy of the data from the “2018” column.


Furthermore, some of the columns do not sum to the “Total” given. The worst instance is the “2015” column, where the “Total” provided is 296 cases less than the sum of the numbers given. This, incidentally, is the error that used to most surprise my students — that, in the modern era of computer spreadsheets, people still can’t add up.

Even more inconsistent is that the data given on page 5 do not all agree with the Totals shown in this table. Notably, for the year 2000, page 5 shows 4.8 million (not 4.48 million, as shown here), and for the year 2013, page 5 shows 8.9 million (not 8.22 million, as shown here).

Page 9


The Percentage Share table contains some wrong numbers, based on the original count data shown on page 8. This occurs once in the “2016” column and four times in the “2000” column — the biggest difference is “2.9%” (shown here) instead of 3.6% (calculated from page 8).

Also, the number of decimal places for the bottom two rows is inconsistent, with 1 decimal place for most of the columns, but 2 places for the “2000” column and none at all for the “2013” column. Furthermore, the data for the bottom two rows shown in the “2018” column sum to 100.5% (not “100%”, as shown), which happens because both numbers are wrong (they should be 37.2% and 62.8%, not “37.8%” and “62.7%”). The same pair of numbers for column “2016” are also incorrectly rounded. Once again, all of this is based on the count data shown on page 8.

Page 10


First, what on earth is an “Excise: Tax on Tourism”? Excise taxes are placed on goods where the government doesn’t want you to overdose, like tobacco and alcohol. So, in Ireland each bottle of wine effectively has a minimum price, which will be equal to the excise duty — costs and profit are added to that price. So, what has tourism got to do with this?

Anyway, the reference to “HL Excise Rate”, does not make sense. The rates shown in the table are per bottle of wine. I suspect that the “HL” stands for “hectoliters”, which is how excise rates are usually quoted officially (see Excise duties in the EU), but it has no relevance to this table.

Page 13

Finally, I freely admit to being mystified by the High Excise Rate Tables. I have been unable to perform any calculations that will produce the numbers shown in this table. As far as I can determine, there is a constant €3.19 excise rate and a constant 23% value-added tax, but that information does not lead me to any of the numbers shown here. To me, there has been a complete failure to communicate with me, in this instance.

Conclusion

I always emphasized to my students that: (i) I would like them to get their reports right, and (ii) if they can't do that, then they should at least make them consistent. And when they are reading the literature, they should make sure that other authors have done the same — it is the first line of defense against false information. I still think that this is so.

Monday, April 29, 2019

How can you doubt global warming?

Every time we look at the viticulture news these days there is something about the current problems with grape harvests, whether it be drought, flood, or fire. These are all the result of climatic effects, and are therefore attributed to climate change.

Climate change is a big issue for all parts of the agriculture industry, not just grape-growing, because of the industry's almost complete reliance on the weather, which determines both the timing and the size of each harvest. Climate change thus leads inevitably to agriculture change. At the moment, the changes seem to consist of an increase in variability from year to year, from boom to bust. This is no way to try to make a living.

NASA forecasts of how global temperatures might change.

Global warming first became a big news story in 1988, which had the then hottest northern-hemisphere summer on record, with widespread droughts, and with fires from the Amazon rain forest to Yellowstone National Park. It is now 30 years since that time, and there are still climate-change skeptics running around, notably in Australia and the USA.

It is thus important to note that there are three separate issues related to the current concern about global climate change:
  • evidence that temperatures have increased recently
  • ideas about what might be causing this increase
  • and what, if anything, we might do about it.

The first of these issues seems to be unproblematic — every country on Earth has a Bureau of Meteorology of some sort, and as far as I know they have all been recording a slow but steady increase in world temperatures for at least 5 decades now. Furthermore, as discussed below, we have temperature records that go back 6 centuries that show the current temperatures to be unprecedented during that time.

The second point may be slightly more problematic, although the consensus is that the cause is increased carbon dioxide (CO2) in the atmosphere. Many people do not know that the idea that our climatic temperatures are determined by atmospheric gases actually dates back a century, now. So, I will look briefly at the history in the next section.

I guess that it is the third point that is the problematic one for the skeptics. They ask: "why should we do anything?" To me, the answer lies on our formal taxonomic name, Homo sapiens, which is Latin for "thinking person". Not only can I learn to use a screw-driver, which most other species cannot, I can think about my effect on the world, and about what sort of world I would like to live in. Maybe, just maybe, it is our effect on the world that is causing the sort of world we currently live in.

Atmospheric carbon

It was back in the 1820s that Joseph Fourier realized that some of sunlight's energy must be held within the atmosphere, helping to keep the Earth warm. It was thus apparently he who first concluded that Earth’s air layer acts like a greenhouse — energy enters through the transparent "walls" and is then trapped inside. It was 40 years later that John Tyndall tried to work out what kinds of atmospheric gases were most likely to play a role in absorbing sunlight, and demonstrated that CO2 is the principal culprit.


It was not until the 1890s that Svante Arrhenius started to wonder how much CO2 decrease was required to explain the Earth's cooling during past ice ages (as was apparent in the geological and fossil record), and what would be the effect of similar CO2 increases. It is ironic that he wrote:
By the influence of the increasing percentage of carbonic acid [CO2] in the atmosphere, we may hope to enjoy ages with more equable and better climates, especially as regards the colder regions of the Earth.
In the 1930s, Guy Stewart Callendar first noted that both the United States and the North Atlantic region had warmed significantly, on the heels of the Industrial Revolution. Sadly, this observed effect was then counter-balanced by the subsequent 30 years of global cooling, now attributed to aerosol pollutants blocking the entry of sunlight. So, it was thus not until the 1970s that the start of the current long steady increase in global temperatures became clearly evident in the world's many meteorological records.

Long-term temperature records

The problem with "long-term" records is that we haven't been accurately measuring anything for all that long, in the big scheme of things. Take, for example, our records of the amount of CO2 in the atmosphere, as shown in the first graph.


The Scripps Institute of Oceanology did not set up the Mauna Loa Observatory (on a mountain in Hawaii) until 1958, which is our main source of instrumental measurements of global carbon dioxide. So, we have to use indirect measure of atmospheric CO2 before that time, as in the graph. This makes it a bit tricky to show that CO2 has increased since the Industrial Revolution, which is a key part of connecting global warming to atmospheric gases.

Temperature, on the other hand is a bit easier to measure. Many of you will know that the international scale for measuring temperature is named after Anders Celsius (not that other scale, named after Daniel Fahrenheit, and now used by only a few countries). Celsius was professor of astronomy at Uppsala University, in Sweden, and so it should come as no surprise that the longest daily temperature record in the world is from the city of Uppsala.

Mind you, even that is not as simple as it seems. In the early years, for example, the temperature was recorded only on week days, not every day; and, of course, the instruments used were not as good as those used today. Nevertheless, the yearly averages from the year 1722 are shown in the next graph. Note that the numbers shown are the averages of the 365¼ daily average temperatures each year — if you want more details, then you can also look at the seasonal average temperatures here.

Average temperatures in Uppsala, Sweden, over the past 300 years.

Note that the interest is in temperature change, and so the graph is centered on the long-term average temperature, with blue years being below average and red years being above average. The recent cluster of warm decades is pretty obviously anomalous over the 3 centuries shown. Aside: being a Swede, Celsius was interested in how cold it gets, not how hot it gets, and so his original scale had water freezing at 100 °C and boiling at 0 °C — it was some time before it was inverted to the current form.

If we stick to the idea of yearly temperatures, then we can actually go back a bit further than the daily record. Gordon Manley has compiled the monthly mean temperatures for the Midlands region of England from the year 1659; and these are shown in the next graph, from the American Association of Wine Economists. Once again, the recent trend is clearly anomalous. [Note that the Midlands is on average 4°C warmer than the Uppsala region!]

Temperatures in central England over the past 350 years.

Finally, I have noted before (Grape harvest dates and the evidence for global warming) that grape harvest dates can be used as a reasonably accurate record of annual (summer) temperature variation, although obviously not as good as instrument measurements. There are formal records for the Burgundy region of France that go back to the late 1300s, which takes us back another 3 centuries before the above graphs. The data are shown in the final graph — as before, the recent trend is completely anomalous.

Grape harvest dates in Burgundy over the past 650 years.


Conclusion

Points 1 and 2 listed above do not seem to me to require much discussion, even by skeptics — we have clear temperature records, and we know how carbon dioxide works. Fortunately, for many people point 3 has moved beyond discussion, and into action.

Monday, December 31, 2018

Is there truth in (wine) numbers?

Everyone knows the expression in vino veritas (in wine there is truth), which (in one form or another) seems to date all the way back to the 6th century BCE. However, in this blog, I spend a lot of time looking at numbers. This immediately raises the oft-asked question of whether "truth" also lies in numbers. In this post I will look at four informative examples where there is truth in some wine numbers, but in each case all is not quite as it seems.


Introduction

Many people are wary of numbers. The issue is that truth lies not in the numbers themselves but in our ability to interpret them. Numbers cannot speak for themselves, and thus they can tell us nothing directly. We have to look at them and work out for ourselves what truth lies therein.

The same applies to words, of course. The same combination of letters can mean quite different things, in different contexts or in different places. Even in English the words "lead" (pronounced leed) and "lead" (pronounced led) look very similar but have different meanings. (Did you know that this is why the band Led Zeppelin spelled their name that way? That was how they wanted it to always be pronounced, as the name comes from the expression "going down like a lead zeppelin".)

We have to be aware of this sort of thing, if we are to make much sense of the world around us; and we all get it wrong more often than we would like. Having so many different languages only makes it much worse, of course.

It is the same with numbers, even though there is only one mathematical language. This is why Mark Twain famously referred to "Lies, damned lies, and statistics". The first two emphasize the problems with words, and the third one the problem with numbers. It is easy to fool ourselves when interpreting numbers, and to thereby intentionally or unintentionally mislead others.

I mention this because I recently encountered four different examples of misinterpreting numbers in the wine industry, in a way that lead to wrong conclusions, even though the numbers were (almost all) truthful. The first example comes from a book, the second from a blog post, the third from a press release, and the fourth comes from a research paper. Numbers are everywhere!

Example 1

An easy one to start with. This table is from a book about Madeira wine. It discusses the wine production from each of the main grape varieties. Back when I taught experimental design to university students, I used examples just like this one to drill into those students the importance of presenting numbers correctly in tables. Can you spot the error?


Any time you see a table where the numbers are supposed to add up to a given total, check whether they do — you might be surprised how often they don't (eg. see the Postscript.) In this case, the Production data for Other European Varieties cannot possibly be right, although the Percentage of total harvest is apparently correct. Working backwards from the Total given, the true Production should be 38,936.05 hL, not 39.04.

Note that this is similar to a typographical error, but of a somewhat complex type, and with important consequences.

Example 2

Now let's look at a slightly more tricky instance. Some years ago, a retailer blog post from Australia contained this comment:
I could not help but be struck by how many wines the tasters rated at 90 points or more on a scale where the maximum is 100. An analysis of the list of 365 wines (excluding the French champagnes - all of which were rated above 90) showed that the average score given to this range was 91.36 and the lowest score given to a wine was 83. Of the 211 reds listed, 83 were rated as 93 points or higher. Now under the 20 point system used in Australian wine shows, 18.5 points is gold medal standard. Multiply by five to get 92.5 and it seems that almost 40 per cent of all the reds (83 out of 211) are gold medal standard, and every red and white wine on their list is well above the minimum 15.5 out of 20 (or 77.5 out of 100) needed to gain a bronze medal.
The author's conclusion does, indeed, follow if his arithmetic is right; but it isn't right. The issue here is converting from one scale (20 points) to another (100 points). The arithmetic assumption that the author makes is that both scales start at 0, whereas the 100-point scale actually starts at 50. These different equivalences are compared in this table:
20-point
scale

0
0.5
1
1.5
2
2.5
3
3.5
4
4.5
5
5.5
6
6.5
7
7.5
8
8.5
9
9.5
10
10.5
11
11.5
12
12.5
13
13.5
14
14.5
15
15.5
16
16.5
17
17.5
18
18.5
19
19.5
20

0 = 0

0
2.5
5
7.5
10
12.5
15
17.5
20
22.5
25
27.5
30
32.5
35
37.5
40
42.5
45
47.5
50
52.5
55
57.5
60
62.5
65
67.5
70
72.5
75
77.5
80
82.5
85
87.5
90
92.5
95
97.5
100

   0 = 50
50
51.25
52.5
53.75
55
56.25
57.5
58.75
60
61.25
62.5
63.75
65
66.25
67.5
68.75
70
71.25
72.5
73.75
75
76.25
77.5
78.75
80
81.25
82.5
83.75
85
86.25
87.5
88.75
90
91.25
92.5
93.75
95
96.25
97.5
98.75
100
Show
medals
































Bronze


Silver


Gold

 

Allowing for the fact that 0 on the 20-point scale equals 50 on the 100-point scale does away with the author's concern about over-inflation of scores, because a Gold medal requires 96 points, not 93 points, and not all of the wines would get a Bronze medal (which requires 89 points not 77.5).

However, even this simple correction does not necessarily produce the "correct" conversion from 20 points to 100 points. For example, Australia's Winestate magazine has used a conversion where 15.5 points on the 20-point scale is equivalent to 90 points on the 100-point scale, not 89 points (as shown in the table).

Example 3

This example takes a lot of work to identify the source of the error.

In 2015, the climats and terroirs of the wine region of Burgundy were added to the UNESCO World Heritage List. To quote UNESCO: "The climats are precisely delimited vineyard parcels on the slopes of the Côte de Nuits and the Côte de Beaune south of the city of Dijon. They differ from one another due to specific natural conditions (geology and exposure) as well as vine types and have been shaped by human cultivation. Over time they came to be recognized by the wine they produce ... The site is an outstanding example of grape cultivation and wine production developed since the High Middle Ages."

The UNESCO documentation suggests that there are 1,247 climats in this World Heritage site. However, Paul Messerschmidt (an amateur wine researcher from the UK) (paulmess[at] gmail.com) noted a discrepancy between this number and the count of those actually listed in the UNESCO documentation.

After a lot of (tedious) work, he realized that, while there are 1,247 climat names in Burgundy, there are actually "1,628 separate, distinct, and precisely delimited vineyard parcels in the Côte d'Or". The difference appears to come from searching the database for "climat names" rather than "named climats". For example, "there are vineyards called Les Cras in Chambolle-Musigny, Vougeot, Aloxe-Corton, Pommard, and Meursault (ie. five "named climats"), but they share only one "climat name", as listed in UNESCO's count of 1,247". The difference of 381 vineyards is hardly trivial, especially if you happen to own part of one of them.

Paul is apparently now compiling the discrepancies between the UNESCO list and those of The Wines of Burgundy, by Sylvain Pitiot & Jean-Charles Servant, and Inside Burgundy, by Jasper Morris, if anyone wants to help him with his work.

Example 4

Let's return to the subject of scoring wines at a wine show, and awarding medals. This will illustrate a situation where we can easily be mislead when dealing with statistical summaries.

There are a number of research papers where judge scores have been compiled, and I will illustrate my point with a paper in the Journal of Wine Research (1996, 7:83-90). In this case, the judges evaluated 174 wines, and this graph shows the scores for three of the judges (each vertical bar represents the number of wines that received each of the scores shown on the horizontal axis):


One standard way to summarize data like this, and thus to compare the judges, is to calculate the mean score for each judge. In this case, the mean for Judge 3 is 11.8 and for Judge 5 it is 11.3. These two means are almost identical, suggesting that the judges are rather similar, and yet their scores, as shown in their graphs, are quite different. Indeed, Judge 3 seems to have two main groups of scores that are favored, with a score of 12 not commonly being used — and yet this is actually the mean score (11.8)! In this case, the mean does not help us understand the scoring behavior of this judge.

This is even more obvious when we look at the data for Judge 1. Once again, there are scores that the judge rarely uses, such as 9 and 10, and yet the mean score is 9.8. When the data have two distinct patterns, we call it "bi-modal", and in such a case the calculation of any sort of average score is going to mislead us badly. The data need to be clustered around the mean, if the mean is going to tell us anything useful.

Conclusion

So, remember that truth lies not in the words or numbers, but in our ability to interpret them. Compare this with the situation of a medical doctor diagnosing a disease based on the patient's symptoms. The symptoms really do indicate the disease, and hopefully the doctor extracts the truth most of the time. However, sometimes the doctor is unfamiliar with the disease, and sometimes the doctor misinterprets the symptoms, and sometimes the doctor fails to connect the symptoms with the disease. This is not good for the patient, or the doctor for that matter; but they both need to deal with it.

As a final word example, Swedes have an expression for couples living together, which is "samma boende", which they shorten to "sambo" (the first letters of each word). Americans do not introduce their partner as their sambo, but Swedes quite happily do so. This confuses Americans, but not Swedes.

Postscript

How many of you have ever noticed the error in a widely distributed description of the original UC Davis 20-point wine score card (as pointed out to me by Bob Henry)? The description is: "Appearance (2), Color (2), Aroma & Bouquet (4), Volatile Acidity (2), Total Acidity (2), Sugar (1), Body (1), Flavor (1), Astringency (1), and General Quality (2)". The true numbers should be: Flavor = 2, Astringency = 2, so that scores then correctly sum to 20. [I warned you to always check totals!] The written description also refers to a "fairy wine", which would be a very interesting thing, if it existed.