A month later, as a midsummer night stretched pleasantly ahead of them in Pasadena, California, dozens of NASA controllers helplessly watched a feed delayed by 19 minutes as the Viking 1 lander separated from the orbiter.一个月后,在加利福尼亚州帕萨迪纳一个惬意的仲夏夜,数十名美国国家航空航天局(NASA)的控制员们无助地看着延迟 19 分钟的信号,目睹了海盗 1 号着陆器与轨道器的分离。
Three hours later, the lander plunged into the Martian atmosphere at 10,000 miles per hour. Ten minutes later, following a fiendishly complex approach, and with zero help from JPL, the lander made a perfect touchdown on Mars.三小时后,着陆器以每小时 1 万英里的速度冲入火星大气层。十分钟后,在经历了极其复杂的进场过程,且在没有喷气推进实验室(JPL)任何帮助的情况下,着陆器在火星上实现了完美着陆。
For Mission Control, it hadn’t helped that, when Viking 1 entered Martian orbit and pointed its camera down at the surface, the originally selected landing site turned out to be “at least twice as rough as the Mars average” prompting a frantic search for a new site.对于任务控制中心来说,当海盗 1 号进入火星轨道并将摄像头对准表面时,最初选定的着陆点被发现“至少比火星平均水平粗糙两倍”,这迫使他们不得不疯狂寻找新的着陆点,这让情况变得更加棘手。
Fortunately, a safer site was found which, in the scientists’ words, was still “rougher than the Martian average,” but “near the Martian average for elevations accessible to Viking” and “near the Mars average in reflectivity.”幸运的是,他们找到了一个更安全的着陆点,用科学家的话说,虽然它仍然“比火星平均水平粗糙”,但“接近海盗号可到达区域的火星平均海拔”,并且“在反射率上接近火星平均水平”。
Even at 212 million miles from Earth, the humble mean turned out to be the statistical device guiding Viking 1 to a safe spot in an unfriendly world.即使在距离地球 2.12 亿英里的地方,谦逊的平均值也成为了引导海盗 1 号在这个不友好的世界中找到安全落脚点的统计工具。
The mean keeps making its usefulness felt in all sorts of situations, often in truly non-obvious ways. Consider, for instance, another happenstance.平均值在各种情况下不断展现其用途,且往往以真正出人意料的方式发挥作用。例如,让我们看看另一个巧合。
Five days after the Viking lander touched down on Mars, on July 25, 1976, the Viking orbiter transmitted to Earth, the following photograph.在海盗号着陆器在火星着陆五天后,即 1976 年 7 月 25 日,海盗号轨道器向地球传回了以下照片。

The photo was immediately dismissed as a mere play of light and shadow coupled with missing data (the dark “speckles” in the image). Nevertheless, it became one of the most famous examples of pareidolia: our tendency to detect patterns, especially faces, and ascribe meaning to them.这张照片立即被认为仅仅是光影效果加上数据缺失(图像中黑暗的“斑点”)造成的。然而,它成为了空想性错视(pareidolia)最著名的例子之一:即我们倾向于探测模式(尤其是人脸),并赋予它们意义。
And it is not just faces on Mars or the Moon.而且这不仅仅发生在火星或月球的人脸图案上。
People have spotted animal shapes and UFOs amid clouds. Down on Earth, people have spotted cups and saucers, saucepans with handles, flagpoles, hammers, people hanging from poles, triangles, rectangles, and above all, skyrocketing trends in charts and plots that have been comfortably out on little more than a random walk.人们在云层中发现了动物形状和不明飞行物。在地球上,人们曾看到杯碟、带柄的平底锅、旗杆、锤子、挂在杆子上的人、三角形、矩形,最重要的是,在图表和曲线中发现了实际上仅仅是随机游走的飙升趋势。
It’s not that these patterns do not exist. Often, they do exist, if only lightly and ephemerally, and when they are there, our vision system has evolved to detect their presence almost every single time. Sadly, the patterns often never mean what we think they do.并不是说这些模式不存在。它们通常确实存在,哪怕只是轻微且短暂的。当它们出现时,我们的视觉系统已经进化到几乎每次都能探测到它们。遗憾的是,这些模式往往并不意味着我们所认为的那样。
So, where in all this is the connection with the mean?那么,这一切与平均值有什么联系呢?
A Representative for Every Occasion适合各种场合的代表
The connection with the mean emerges from the counterfactual, i.e., from data in which you do not instantly see saucepans or people hanging from flag poles. In other words, data in which trends and patterns cannot be easily detected.与平均值的联系源于反事实,即在那些你无法一眼看出平底锅或挂在旗杆上的人的数据中。换句话说,在那些趋势和模式无法轻易探测到的数据中。
Often, trends in data aren’t expressed in ways that give you an evolutionary advantage to detect them. So, your eyes fail to see them. Your brain fails to interpret them. That is when you sorely need devices like the mean, median, mode, and percentile scores.通常,数据中的趋势并不会以让你具备进化优势的方式表现出来。因此,你的眼睛无法看到它们,大脑也无法解读它们。这时,你就会迫切需要平均值、中位数、众数和百分位数等工具。
Such statistical measures as the mean can compress the noise and pull out the pattern in a measurable and actionable form, provided it exists.只要模式存在,像平均值这样的统计指标就可以压缩噪声,并以可衡量且可操作的形式提取出模式。
Consider the following three charts showing daily absenteeism in NYC public schools in school years 2015-16, 2016-17, and 2017-18.考虑以下三张图表,它们显示了 2015-16、2016-17 和 2017-18 学年纽约市公立学校的每日缺勤情况。
The X-axis represents school days, and the Y-axis represents percentage daily absenteeism. Each bar represents the aggregate daily absenteeism rate across all NYC schools for a particular school day such October 22, 2015.X 轴代表上课天数,Y 轴代表每日缺勤百分比。每个柱状图代表特定上课日(例如 2015 年 10 月 22 日)所有纽约市学校的每日缺勤率总和。

This is a classic example of a panel data set in which the units, i, are school days, the time variable, t, is the school year, and the response variable, yit, is the aggregate daily percentage absenteeism across all schools on school day, i, in year, t.这是面板数据的一个经典示例,其中单位 i 是上课日,时间变量 t 是学年,响应变量 yit 是第 t 年第 i 个上课日所有学校的每日缺勤百分比总和。
From the charts alone, can you tell if absenteeism has worsened, improved, or simply meandered without a clear trend over the three years that constitute the data panel?仅从图表来看,你能判断出在这构成数据面板的三年里,缺勤情况是恶化了、改善了,还是仅仅在没有明确趋势的情况下徘徊不定吗?
Here is another example showing the state-level annual House Price Index (HPI) values in the United States from 2020 to 2025.这是另一个显示 2020 年至 2025 年美国各州年度房价指数(HPI)的例子。
The X-axis represents a U.S. state, and the Y-axis represents the annual HPI for that state. Each bar represents a state’s annual HPI value for the corresponding year, assuming a base of 100 in 1975.X 轴代表美国各州,Y 轴代表该州的年度 HPI。每个柱状图代表相应年份该州的年度 HPI 值,假设 1975 年的基数为 100。

What is your take on the house price trend expressed in this 6-year data panel? Have prices increased, decreased, or meandered listlessly?对于这个 6 年数据面板中表现出的房价趋势,你有什么看法?价格是上涨了、下跌了,还是漫无目的地徘徊?
It is hard for us to pick up trends and patterns in such data.我们很难在这样的数据中捕捉到趋势和模式。
So, what would be a good representative that would efficiently bring out the trend, if one exists?那么,什么样的代表性指标能有效地揭示趋势(如果存在的话)呢?
Would the maximum or the minimum value be a good candidate? How about the mode (the most frequently occurring value), or the median (the value around the 50th percentile)?最大值或最小值会是好的候选者吗?众数(最常出现的值)或中位数(第 50 百分位左右的值)又如何?
The statistic that you choose will depend entirely on how you intend to use it.你选择的统计指标完全取决于你打算如何使用它。
Let’s revisit the daily absenteeism data panel.让我们重温每日缺勤数据面板。
The U.S. Department of Education defines chronic absenteeism as students missing “10% or more of school”. Strictly speaking, chronic absenteeism is a student-level yearly measure. But if a school system had an additional operational goal of keeping the daily absenteeism rate also below 10%, then they’d want to monitor the maximum daily absenteeism rate.美国教育部将长期缺勤定义为学生“缺勤 10% 或以上”。严格来说,长期缺勤是一个学生层面的年度衡量指标。但如果一个学校系统还有一个将每日缺勤率也保持在 10% 以下的额外运营目标,那么他们就会想要监控每日最高缺勤率。
The following chart uses the data in the absenteeism data panel shown earlier. For each school year, it shows the maximum daily absenteeism across all schools.下表使用了前面显示的缺勤数据面板中的数据。它显示了每个学年所有学校的每日最高缺勤率。
It is immediately evident that plotting the daily maximum value paints a troubling picture of daily school attendance.显而易见,绘制每日最大值描绘了一幅令人担忧的每日学校出勤情况图景。

Now let’s revisit the House Price Index data panel. If federal economic policy requires the monitoring of the median annual state-level House Price Index, policymakers will want to keep an eye on a chart like the following.现在让我们重温房价指数数据面板。如果联邦经济政策要求监控年度州级房价指数的中位数,政策制定者们会希望关注像下面这样的图表。
The following chart uses the data from the HPI data panel and shows the median annual HPI across the 50-states-plus-DC panel from 2020 to 2025. Again, the trend is immediately visible when we use a suitable summary statistic.下表使用了房价指数数据面板中的数据,显示了 2020 年至 2025 年间 50 个州及哥伦比亚特区的年度房价指数中位数。同样,当我们使用合适的汇总统计指标时,趋势便一目了然。

How about using the sum as a summary statistic?那么使用“总和”作为汇总统计指标怎么样?
There are two problems with using the sum.使用总和存在两个问题。
First, its value is often proportional to the sample size. The sum confounds the “typical” (summary) value with the number of observations, making it tricky to compare samples of different sizes. 首先,它的值通常与样本量成正比。总和将“典型”(汇总)值与观测数量混为一谈,使得比较不同规模的样本变得棘手。
Second, often the sum doesn’t mean anything. Adding observations is sensible only if they are additive. Observations such as indexes, rates, and percentages are not additive.其次,总和往往毫无意义。只有当观测值具有可加性时,相加才有意义。像指数、比率和百分比这样的观测值是不可加的。
For example, the sum of House Price Indexes across all states is a real head-scratcher.例如,所有州的房价指数之和简直令人费解。
But let us not be so quick to dismiss the sum.但我们不要太快否定总和。
The Mean as a Probability-Weighted Sum作为概率加权和的平均值
Suppose we tweak the sum of the observed values so that the importance of each distinct value is weighted by just the right factor, a factor that reflects how frequently that value is observed.假设我们调整观测值的总和,使得每个不同值的重要性都乘以一个恰当的因子,这个因子反映了该值被观测到的频率。
Normally, when you multiply an integer A by another integer B, B acts as a numerical magnifying glass for A, and vice versa. You are magnifying the effect of A by B.通常,当你将一个整数 A 乘以另一个整数 B 时,B 充当了 A 的数值放大镜,反之亦然。你是在通过 B 来放大 A 的效果。
But here, we’d like to do the opposite. We would like to “weaken” the importance of each distinct value by a factor. The factor we choose is that value’s probability of occurrence, p.但在这里,我们想做相反的事情。我们想通过一个因子来“削弱”每个不同值的重要性。我们选择的因子就是该值出现的概率 p。
Let us develop this idea.让我们展开这个想法。
Suppose there are k distinct values x1, x2, …, xi, …, xk in a population of size N. Notice that k ≤ N since each xi can occur more than once in the population.假设在一个大小为 N 的总体中,有 k 个不同的值 x1, x2, …, xi, …, xk。注意 k ≤ N,因为每个 xi 在总体中可能出现不止一次。
Let Ni be the frequency of occurrence of xi. What is the probability of observing xi? It is the following ratio.设 Ni 为 xi 出现的频率。观测到 xi 的概率是多少?它是以下比率。

Using the probabilities, pi, we express the probability-weighted sum of the k distinct values in the population as follows.利用概率 pi,我们将总体中 k 个不同值的概率加权和表示如下。

Equation (1) is the formula for the population mean μ (“mu”), and it uses the population probabilities p1, p2, …, pi, …, pk.公式 (1) 是总体平均值 μ(“mu”)的公式,它使用了总体概率 p1, p2, …, pi, …, pk。
The population probability represents the true frequency of occurrence of the value in the entire population.总体概率代表了该值在整个总体中出现的真实频率。
Expected Value期望值
Eq. (1) is also the formula for the expected value or the unconditional expectation of the discrete random variable X which can take one of k distinct values x1, x2, …, xi, …, xk with corresponding probabilities p1, p2, …, pi, …, pk.公式 (1) 也是离散随机变量 X 的期望值或无条件期望的公式,该变量可以取 k 个不同值 x1, x2, …, xi, …, xk 中的一个,并对应相应的概率 p1, p2, …, pi, …, pk。

The Arithmetic Mean算术平均值
Suppose all distinct values x1 to xk in Eq. (1) occur at the same frequency, m. Notation-wise, we represent this situation as follows.假设公式 (1) 中所有不同的值 x1 到 xk 以相同的频率 m 出现。在符号表示上,我们将这种情况表示如下。

Then, each probability, pi, is the following constant value.那么,每个概率 pi 都是以下常数值。

In this case, the mean can be expressed as follows.在这种情况下,平均值可以表示如下。

Since m times xi is the same as (xi + xi + xi …m times), Eq. (1a) can be written as a simple sum over all N values in the population, as follows.由于 m 乘以 xi 等同于 (xi + xi + xi … m 次),公式 (1a) 可以写成对总体中所有 N 个值的简单求和,如下所示。

Eq. (1b) is the formula for the arithmetic mean of N values.公式 (1b) 是 N 个值的算术平均值公式。
In the simplest case, m = 1, that is, each distinct value in the population occurs exactly once. In other words, every value in the population is distinct. Then, we have the following situation.在最简单的情况下,m = 1,也就是说,总体中的每个不同值恰好出现一次。换句话说,总体中的每个值都是不同的。那么,我们有以下情况。

In this case, k = N, and the arithmetic mean of N values takes the following familiar form.在这种情况下,k = N,N 个值的算术平均值采用以下熟悉的格式。

Eq. (1c) is the familiar formula for the arithmetic mean of N values, where every value in the population has the same probability 1/N.公式 (1c) 是 N 个值算术平均值的常用公式,其中总体中的每个值都具有相同的概率 1/N。
The Sample and the Population样本与总体
In most cases, you wouldn’t have access to the entire population.在大多数情况下,你无法访问整个总体。
Often, the population size is unknown. Think about a country’s population: for most countries the exact population is unknown.通常,总体大小是未知的。想想一个国家的人口:对于大多数国家来说,确切的人口数量是未知的。
In practice, you would be working with a finite sample of size n in which you have observed q distinct values x1, x2, …, xi, …, xq where q ≤ k ≤ N, where k is the number of distinct values in the underlying population of size N.在实践中,你将使用一个大小为 n 的有限样本,其中你观测到了 q 个不同的值 x1, x2, …, xi, …, xq,其中 q ≤ k ≤ N,k 是大小为 N 的底层总体中不同值的数量。
In this finite sample of size n, you will note down how frequently each distinct value xi occurs. For the ith distinct value xi, this gives you the sample probability pi_hat (pronounced “p i hat” or “p hat i”).在这个大小为 n 的有限样本中,你将记录每个不同值 xi 出现的频率。对于第 i 个不同值 xi,这给了你样本概率 pi_hat(读作“p i hat”或“p hat i”)。
pi_hat is also called an empirical probability because it arises from an experiment involving the selection of an appropriate sampling strategy and taking measurements. Consequently, a statistic computed using empirical probabilities is also empirical in nature.pi_hat 也被称为经验概率,因为它源于一项涉及选择适当采样策略并进行测量的实验。因此,使用经验概率计算出的统计量本质上也是经验性的。
Using pi_hat, we write the formula for the sample mean as follows.使用 pi_hat,我们将样本平均值公式写为如下形式。

The n in X_bar_n denotes the sample size.X_bar_n 中的 n 表示样本量。
While the population mean in Eq. (1) is usually a theoretical quantity, the sample mean in Eq. (2) is practically computable.虽然公式 (1) 中的总体平均值通常是一个理论量,但公式 (2) 中的样本平均值在实践中是可以计算的。
The sample mean serves as your working estimate of the population mean.样本平均值作为你对总体平均值的有效估计。
All of us use the sample mean because we rarely, if ever, have access to the population. Consider restaurant ratings, for example. The 3.7-or-so rating that you saw on Google for the new restaurant in town was the mean of several 1-star, 2-star, 3-star, 4-star, and 5-star ratings left by a fraction of all patrons who dined there. Even Google cannot coax every patron into leaving a rating. That makes the star rating a sample mean and, therefore, a mere estimate of the true rating.我们所有人都使用样本平均值,因为我们很少(如果不是从未)能够访问总体。以餐厅评分为例。你在谷歌上看到的城里那家新餐厅 3.7 分左右的评分,是所有在那里用餐的顾客中一小部分人留下的 1 星、2 星、3 星、4 星和 5 星评分的平均值。即使是谷歌也无法诱导每一位顾客留下评分。这使得星级评分成为样本平均值,因此仅仅是对真实评分的估计。
In fact, things get slightly worse. The rating you saw was the mean of the ratings left by only those patrons who decided to give a rating. If you believe that people who had either an exceptionally fabulous or an exceptionally lousy experience are more likely to leave a rating, then 1-star, 2-star, and 5-star ratings will dominate the sample. This dominance will necessarily come at the expense of the 3- and 4-star ratings.事实上,情况会变得更糟。你看到的评分只是那些决定留下评分的顾客所留下的评分的平均值。如果你认为那些有极其美妙或极其糟糕体验的人更有可能留下评分,那么 1 星、2 星和 5 星的评分将在样本中占主导地位。这种主导地位必然会以牺牲 3 星和 4 星评分为代价。
Such over-representation of the extreme values can tilt the mean toward one or the other end of the scale, causing the mean to lose accuracy. The mean becomes a biased estimate of the “true” rating.这种极端值的过度代表会使平均值向量表的某一端倾斜,导致平均值失去准确性。平均值变成了对“真实”评分的有偏估计。
The following chart illustrates this situation. The sample (shown by the orange bars) contains an over-representation of 1- and 2-star ratings, causing the sample mean to become biased. In this case, the sample mean is biased downward.下图说明了这种情况。样本(由橙色柱表示)包含了过多的 1 星和 2 星评分,导致样本平均值产生偏差。在这种情况下,样本平均值向下偏移。

Platforms such as Google are surely aware of such biases, of course. And they almost certainly use every trick in the statistician’s handbook to tame them, though perhaps at the cost of increasing the variance of the estimated rating. In other words, the displayed rating becomes less biased, but it may also become a less precise estimate of the true rating. This tension between systematic error and random variation (the bias-variance tradeoff) influences the design of many statistical and neural-net models.当然,像谷歌这样的平台肯定意识到了这种偏差。他们几乎肯定会使用统计学家手册中的每一个技巧来控制它们,尽管这可能以增加估计评分的方差为代价。换句话说,显示的评分偏差较小,但它也可能成为对真实评分不太精确的估计。这种系统误差与随机变异之间的张力(偏差-方差权衡)影响了许多统计模型和神经网络模型的设计。
Biased or not, the sample is all you will have, and the sample value of the statistic, whether the sample mean, the sample median, the sample mode, or any other statistic, is what you will use as a working estimate for making decisions. In fact, you would be surprised how much data science, nay, science itself, is devoted to estimating the population mean, population median, population whatever, based on the sample value while beating the bias out of that estimate.无论是否有偏差,样本是你唯一拥有的东西,统计量的样本值(无论是样本平均值、样本中位数、样本众数还是任何其他统计量)都将作为你制定决策的有效估计。事实上,你会惊讶地发现,有多少数据科学,甚至科学本身,致力于基于样本值来估计总体平均值、总体中位数或总体其他指标,同时消除估计中的偏差。
The Role of Probability, p, in the Mean’s Formula概率 p 在平均值公式中的作用
In the probability-weighted sum formula for the mean, why should the weight be specifically p, and not some function of p like p2 or in general pr (r ≠ 1), or ep?在平均值的概率加权和公式中,为什么权重应该是特定的 p,而不是 p 的某种函数,比如 p2 或一般的 pr (r ≠ 1),或者是 ep?
There are two strong reasons why p is the perfect choice for the weight in the formula for the ordinary mean.有两个强有力的理由说明为什么 p 是普通平均值公式中权重的完美选择。
The Philosophical Reason哲学原因
The first reason is philosophical (or common-sensical, depending on your perspective), and is best explained with an example. Suppose your data contains only three distinct values. Suppose the first value is observed 70% of the time, the second one only 20% of the time, and the third the remaining 10% of the time. You would want a fair sample containing all three distinct values to give more weight to the first value, much less weight to the second, and even less weight to the third. Specifically, you would want the three values to carry weights in the ratio: 70 : 20 : 10, or 70/100 : 20/100 : 10/100. Thus, weighting each distinct value by its frequency, or standardized frequency, in other words, the probability, p, agrees pleasantly with our sense of fairness.第一个原因是哲学上的(或者取决于你的观点,是常识性的),最好用一个例子来解释。假设你的数据只包含三个不同的值。假设第一个值在 70% 的时间内被观测到,第二个值只在 20% 的时间内被观测到,第三个值在剩余的 10% 的时间内被观测到。你会希望一个包含所有这三个不同值的公平样本,能给予第一个值更多的权重,给予第二个值少得多的权重,给予第三个值更少的权重。具体来说,你会希望这三个值携带的权重比例为:70 : 20 : 10,即 70/100 : 20/100 : 10/100。因此,用每个不同值的频率(或标准化频率,换句话说,概率 p)来加权,与我们对公平的感知非常契合。
The Connection of the Mean with the Weak Law of Large Numbers平均值与弱大数定律的联系
The second reason for using p is based on something profound. To know what it is, we must go way back in time. Specifically, we must travel to the city-state of Basel in the old Swiss confederacy in the 1680s, when a little discovery, later to be known as the Weak Law of Large Numbers, was made by the one and only Jakob Bernoulli (1655-1705).使用 p 的第二个原因是基于一些深刻的东西。要了解它是什么,我们必须回到过去。具体来说,我们必须回到 1680 年代旧瑞士联邦的巴塞尔城邦,当时一个微小的发现——后来被称为弱大数定律——被雅各布·伯努利(Jakob Bernoulli,1655-1705)所发现。

I am going to indulge you in a slight detour here to give you a short explanation of the Weak LLN, just enough for us to address the question at hand: why is p the correct choice for the weight?我在这里要稍微绕个弯,给你简要解释一下弱大数定律(Weak LLN),足以让我们解决手头的问题:为什么 p 是权重的正确选择?
If you are eager to know the answer right away, it is this: If the target is the ordinary population mean μ, p is the only choice of weight in the formula for the mean that respects the Weak Law of Large Numbers.如果你急于立刻知道答案,那就是:如果目标是普通总体平均值 μ,p 是平均值公式中唯一符合弱大数定律的权重选择。
If you are happy to take that answer at face value, you don’t need to know how the Weak LLN works and why p, and only p, respects the Weak LLN. But in case you are curious, read on.如果你乐于接受这个答案,就不需要知道弱大数定律是如何工作的,以及为什么 p(且只有 p)符合弱大数定律。但如果你很好奇,请继续阅读。
A Short Introduction to the Weak Law of Large Numbers弱大数定律简介
The Weak Law of Large Numbers says, basically, the following.弱大数定律基本上说明了以下内容。
First, assume that your observations are independent and identically distributed, or i.i.d. “Independent” means the probability of observing a value isn’t influenced by the probability of observing any other value. “Identically distributed” means that your experimental setup doesn’t vary from one observation to the next, and all observations are generated by the same underlying probability distribution.首先,假设你的观测值是独立同分布的(i.i.d.)。“独立”意味着观测到一个值的概率不受观测到任何其他值的概率的影响。“同分布”意味着你的实验设置在每次观测时都不会改变,并且所有观测值都由相同的底层概率分布生成。
Also assume that the population mean exists and the population variance is finite.还要假设总体平均值存在且总体方差是有限的。
Second, even with i.i.d. observations, and no matter how hard you’ve worked to remove bias from your sample, your sample mean will generally contain an error with respect to the population mean. In other words, your sample mean will generally be an estimate of the population mean.其次,即使有 i.i.d. 观测值,无论你多么努力地消除样本中的偏差,你的样本平均值相对于总体平均值通常都会包含误差。换句话说,你的样本平均值通常只是总体平均值的估计。
Now comes the payoff.现在是重点部分。
Suppose you choose some positive threshold value, ϵ, for the error you can tolerate. Bernoulli’s big idea was that as you collect larger samples, it becomes increasingly hard to come across a sample sporting an absolute error larger than your tolerance level, ϵ. And this behavior does not depend on how vanishingly small an error threshold you select. This behavior is driven by a concept in statistics called convergence in probability.假设你为你能容忍的误差选择了一个正的阈值 ϵ。伯努利的一个伟大思想是,随着你收集的样本越来越大,遇到一个绝对误差大于你的容忍水平 ϵ 的样本变得越来越困难。这种行为并不取决于你选择的误差阈值有多小。这种行为是由统计学中一个称为“依概率收敛”(convergence in probability)的概念驱动的。
The bottom line is that the sample mean becomes progressively more reliable with larger samples. This is the Weak Law of Large Numbers, and Jakob Bernoulli quantified precisely how its machinery works.底线是,随着样本量的增加,样本平均值变得越来越可靠。这就是弱大数定律,雅各布·伯努利精确地量化了其机制是如何工作的。
The exact statement of the Weak LLN runs as follows.弱大数定律的精确陈述如下。
Consider a random sample made up of n independent, identically distributed random variables X1, X2, …, Xn. Let the population mean of the underlying distribution on which the random variables are defined be μ. Let Xn_bar be the sample mean. For any positive real number ϵ, the probability that the sample mean is more than ϵ away from μ approaches zero as the sample size approaches infinity.考虑由 n 个独立同分布随机变量 X1, X2, …, Xn 组成的随机样本。设定义随机变量的底层分布的总体平均值为 μ。设 Xn_bar 为样本平均值。对于任何正实数 ϵ,样本平均值与 μ 的距离超过 ϵ 的概率随着样本量趋于无穷大而趋于零。
Under these assumptions, the Weak LLN can be stated as follows.在这些假设下,弱大数定律可以陈述如下。

The Weak Law of Large Numbers implies that, under suitable assumptions, the sample mean converges in probability to the population mean.弱大数定律意味着,在适当的假设下,样本平均值依概率收敛于总体平均值。
It follows that the sample probability, p_hat, also converges in probability to the population probability, p.由此可见,样本概率 p_hat 也依概率收敛于总体概率 p。
To see why this is so, assume that X1, X2, …, Xn are n binary (1/0) random variables, each one denoting the outcome of a random trial. Let 1 denote the observation of a particular value, such as 5, and 0 denote any other observation, such as any value other than 5, in a single trial.为了理解为什么会这样,假设 X1, X2, …, Xn 是 n 个二元(1/0)随机变量,每一个都表示一次随机试验的结果。设 1 表示观测到一个特定值(例如 5),0 表示任何其他观测值(例如在单次试验中除 5 以外的任何值)。
In a random sample of size 10, i.e., 10 i.i.d. trials, suppose you see three 5s. So, your sample might look like the following set.在大小为 10 的随机样本中,即 10 次 i.i.d. 试验,假设你看到了三个 5。所以,你的样本可能看起来像下面这组。
[1, 0, 1, 0, 0, 0, 0, 0, 1, 0][1, 0, 1, 0, 0, 0, 0, 0, 1, 0]
The sample probability pi_hat of observing a 5 is 3/10. The sample mean also happens to work out to the same value 3/10, as follows.观测到 5 的样本概率 pi_hat 为 3/10。样本平均值也恰好是 3/10,如下所示。

Thus, in this case, the sample mean is also the sample probability, and the Weak LLN says that both converge to the population mean μ, whatever that might be, which for a binary indicator variable is exactly the population probability, p.因此,在这种情况下,样本平均值也是样本概率,而弱大数定律指出两者都收敛于总体平均值 μ(无论它可能是什么),对于二元指标变量,这恰好是总体概率 p。
With this background, let’s return to showing why probability p is the right choice for the weight in the formula for the mean.有了这个背景,让我们回到展示为什么概率 p 是平均值公式中权重的正确选择。
In an i.i.d. sample of size n containing q distinct values x1, x2, …, xq, suppose the ith distinct value xi is observed ni times. Then the sample probability pi_hat = ni /n, and the sample mean can be expressed as follows.在一个包含 q 个不同值 x1, x2, …, xq 的大小为 n 的 i.i.d. 样本中,假设第 i 个不同值 xi 被观测到 ni 次。那么样本概率 pi_hat = ni / n,样本平均值可以表示如下。

Next, we take identical limits on both sides of the above equation as follows.接下来,我们对上述方程的两边取相同的极限,如下所示。

Now let’s take the limit on the right-hand side inside the summation. 现在让我们将极限取到求和符号内。

We’ve seen how sample probabilities converge to their population counterparts.我们已经看到了样本概率如何收敛于它们的总体对应值。

Substituting (6) inside (5), we arrive at the following result.将 (6) 代入 (5),我们得到以下结果。

Thus, the choice of p as the weight causes the sample mean to converge to the population mean, thereby respecting the Weak LLN.因此,选择 p 作为权重会导致样本平均值收敛于总体平均值,从而符合弱大数定律。
Now, suppose we use p2 as the weight instead of p. The resulting statistic is obviously not the ordinary sample mean X_bar_n mentioned in Eq. (2). Let’s call it Tn.现在,假设我们使用 p2 作为权重而不是 p。由此产生的统计量显然不是公式 (2) 中提到的普通样本平均值 X_bar_n。我们称之为 Tn。

The summation of pj2over {1,…,q} in the denominator is essential to standardize the weights. Incidentally, we also performed this standardization while using pi via ∑ pj. But there, we did it invisibly because ∑ pj=1. That’s ’cause, in any random trial, exactly one of xi where i ϵ {1,…,q} is guaranteed to be observed. Therefore ∑ pj = 1 when summed over q is always 1. Therefore, mentioning this sum explicitly in the denominator wasn’t needed.分母中 pj2 在 {1,…,q} 上的求和对于标准化权重是必不可少的。顺便提一下,我们在使用 pi 时也通过 ∑ pj 进行了这种标准化。但在那里,我们做得不可见,因为 ∑ pj = 1。这是因为,在任何随机试验中,保证会观测到 i ϵ {1,…,q} 中的恰好一个 xi。因此,当在 q 上求和时,∑ pj 总是 1。因此,在分母中明确提到这个和是不需要的。
Let’s return to Eq. (7) and, as before, let’s take identical limits on both sides.让我们回到公式 (7),像以前一样,对两边取相同的极限。

Also as before, we’ll take the right-hand side limit inside the summation.同样像以前一样,我们将右侧极限取到求和符号内。

Remember that the Weak LLN causes the sample probability pi_hat to converge to the population value pi. So, a continuous transformation of pi_hat such as p2i_hat will also converge to the corresponding population value p2i. Thus, Eq. (8) will converge in probability to the corresponding population value as follows.记住,弱大数定律导致样本概率 pi_hat 收敛于总体值 pi。因此,pi_hat 的连续变换(如 p2i_hat)也将收敛于相应的总体值 p2i。因此,公式 (8) 将依概率收敛于相应的总体值,如下所示。

But the population value on the right-hand side of the above equation is not the ordinary mean μ. Thus, p2 weighting fails for estimating the ordinary mean as it converges to the wrong target. It estimates an entirely different statistic. The same issue arises with any other standardized transformation of p such as ep.但上述方程右侧的总体值不是普通平均值 μ。因此,p2 加权对于估计普通平均值是失败的,因为它收敛到了错误的目标。它估计的是一个完全不同的统计量。同样的问题也出现在 p 的任何其他标准化变换(如 ep)中。
So, You Want to Build Your Own Universe所以,你想建立你自己的宇宙
Let us not get ahead of ourselves. Suppose we still use p2, or ep as the weight. So what if the sample statistic Tn doesn’t converge to the ordinary population mean μ?我们不要操之过急。假设我们仍然使用 p2 或 ep 作为权重。那么如果样本统计量 Tn 不收敛于普通总体平均值 μ,那又怎样?
What’s so bad about that?那有什么不好的?
Well, to quote the line popularized by actor Tony Shalhoub in the detective series Monk:好吧,引用演员托尼·夏尔赫布(Tony Shalhoub)在侦探系列剧《神探阿蒙》(Monk)中普及的那句台词:
“Here’s the thing.”“事情是这样的。”

Suppose your stated target is some well-known population statistic such as the ordinary mean, μ. If your chosen sample statistic does not converge to your stated target, then you have made yourself a fractured promise. In doing so, you have also succeeded in constructing a teacup-universe, a make-believe world of your own liking, a Tolkienish existence if you like (no offense meant to Tolkien fans). A world that the fundamental laws of the universe, as we know them, barely bother to visit.假设你陈述的目标是某种众所周知的总体统计量,如普通平均值 μ。如果你选择的样本统计量不收敛于你陈述的目标,那么你就给自己许下了一个破裂的承诺。这样做,你也成功地构建了一个茶杯宇宙——一个你喜欢的虚构世界,一个托尔金式的存在(对托尔金粉丝并无冒犯之意)。一个我们所知的宇宙基本定律几乎懒得造访的世界。

There is a somber ending to this story.这个故事有一个凄凉的结局。
In 1705, Jakob Bernoulli died of tuberculosis when he was just 50 years old. This was an era when TB was rampant, yet doctors didn’t know what caused it or how to cure it. Diagnosis often amounted to a death sentence, and the decline could be long and painful, lasting from months to a few years. The patient gradually wasted away, “consumed” as it were by something mysterious inside the body. It would be nearly two more centuries before Robert Koch discovered the cause of TB: the microbe Mycobacterium tuberculosis.1705 年,雅各布·伯努利死于肺结核,享年仅 50 岁。那是一个肺结核猖獗的时代,然而医生们并不知道是什么导致了它,也不知道如何治愈它。诊断往往等同于死刑,而衰退的过程可能漫长而痛苦,持续数月到几年。病人逐渐消瘦,仿佛被体内某种神秘的东西所“吞噬”。直到近两个世纪后,罗伯特·科赫(Robert Koch)才发现了肺结核的病因:结核分枝杆菌(Mycobacterium tuberculosis)。
When Jakob Bernoulli died, he also left unfinished his magnum opus, Ars Conjectandi (“The Art of Conjecturing”), which contained his discovery of the Weak LLN. Fortunately for future generations, his nephew Nicolaus Bernoulli (1687–1759) managed to publish his uncle’s work in 1713, a full eight years after Jakob’s death.雅各布·伯努利去世时,也留下了未完成的巨著《猜度术》(Ars Conjectandi),其中包含了他在弱大数定律上的发现。幸运的是,他的侄子尼古拉斯·伯努利(Nicolaus Bernoulli,1687–1759)在 1713 年成功出版了他叔叔的作品,这距离雅各布去世整整八年。
The Mean’s Connection to Squared Loss平均值与平方损失的联系
There happens to be a third, indirect and non-obvious reason for using probability p as the weight. The role of p in the mean emerges from the mean’s deep historical connection with squared loss and the method of least squares estimation. To know about this connection and the circumstances that gave it shape, we must once again tunnel back into history, this time to a point roughly a century after Jakob’s discovery of the Weak LLN in the late 1680s.使用概率 p 作为权重还有一个第三种间接且不明显的理由。p 在平均值中的作用源于平均值与平方损失及最小二乘估计法之间深刻的历史联系。要了解这种联系以及塑造它的环境,我们必须再次深入历史,这次是回到雅各布在 1680 年代末发现弱大数定律后大约一个世纪的时间点。
We now enter Napoleonic Europe of the early 1800s, and the professional lives of two mathematicians on opposite sides of Napoleon Bonaparte’s conquests: the prolific Parisian mathematician Adrien-Marie Legendre (1752 – 1833), and the extraordinarily gifted polymath Johann Carl Friedrich Gauss (1777 – 1855) who had a knack of plucking brilliant ideas out of the ether, and who lived in the Duchy of Brunswick-Wolfenbüttel in what is now modern-day Germany.我们现在进入 19 世纪初拿破仑时代的欧洲,以及拿破仑·波拿巴征服行动对立双方的两位数学家的职业生涯:多产的巴黎数学家阿德里安-马里·勒让德(Adrien-Marie Legendre,1752 – 1833),以及才华横溢的天才博学家约翰·卡尔·弗里德里希·高斯(Johann Carl Friedrich Gauss,1777 – 1855),他有一种从以太中摘取天才思想的本领,生活在当时属于现代德国的布伦瑞克-沃尔芬比特尔公国。

Right: An 1803 portrait of Johann Carl Friedrich Gauss (1777 – 1855) by Johann Christian August Schwartz (Wikipedia, Public domain)左:朱利安-利奥波德·波伊利(Julien-Léopold Boilly)于 1820 年绘制的阿德里安-马里·勒让德(1777 – 1855)的极度夸张的漫画。这也是这位伟人唯一存世的真实肖像(Wikipedia,公有领域)。右:约翰·克里斯蒂安·奥古斯特·施瓦茨(Johann Christian August Schwartz)于 1803 年绘制的约翰·卡尔·弗里德里希·高斯(1777 – 1855)肖像(Wikipedia,公有领域)。
The Method of Least Squares and its Connection with the Arithmetic Mean最小二乘法及其与算术平均值的联系
In 1805, exactly 100 years after Bernoulli died in Basel leaving Ars Conjectandi unfinished, Legendre, one of mathematics’ all-time greatest heroes, published a book in Paris whose appendix would become the focal point of statistical discourse for the next 100 years.1805 年,即伯努利在巴塞尔去世并留下《猜度术》未完成的整整 100 年后,数学史上最伟大的英雄之一勒让德在巴黎出版了一本书,其附录成为了未来 100 年统计学讨论的焦点。
The book was Nouvelles Méthodes Pour La Détermination Des Orbites Des Comètes (New Methods for Determining Comet Orbits), and the appendix contained the method of least squares estimation.这本书是《确定彗星轨道的新方法》(Nouvelles Méthodes Pour La Détermination Des Orbites Des Comètes),附录中包含了最小二乘估计法。
It so happens that the humble mean owes a deep debt to the method of least squares. So deep is this debt that had it not been for the method of least squares, the mean would have lost quite a lot of its usefulness.谦逊的平均值恰好对最小二乘法有着深厚的亏欠。这种亏欠如此之深,以至于如果没有最小二乘法,平均值将失去很大一部分用途。
The following example will illustrate this deep connection.下面的例子将说明这种深刻的联系。
Suppose you wish to estimate an unknown quantity x. x could be the position of a new star you discovered in the night sky. To pin down x, you point your telescope at the presumed location of the star on three successive nights and make three error-ridden (a.k.a. “noisy”) observations x1, x2, and x3 such that x1, x2, x3 can be expressed as the true value x plus an error, as follows.假设你希望估计一个未知量 x。x 可能是你在夜空中发现的一颗新星的位置。为了确定 x,你连续三个晚上将望远镜对准这颗恒星的假定位置,并进行了三次充满误差(即“噪声”)的观测 x1, x2 和 x3,使得 x1, x2, x3 可以表示为真实值 x 加上一个误差,如下所示。
x1 = x + ε1
x2 = x + ε2
x3 = x + ε3x1 = x + ε1, x2 = x + ε2, x3 = x + ε3
Remember that the true x is unknown. The best you can do is compute a smashingly good estimate of the star’s true location using the three observations x1, x2, x3.记住,真实的 x 是未知的。你能做的最好事情就是使用这三次观测 x1, x2, x3 计算出这颗恒星真实位置的一个非常出色的估计。
In 1805, Legendre had a groundbreaking brainwave for creating such an estimate: he postulated that the best estimate of x would be one that heavily penalizes large errors, both negative and positive.1805 年,勒让德在创建这样一个估计时产生了一个突破性的灵感:他假设 x 的最佳估计将是那种对大误差(无论是负还是正)进行严厉惩罚的估计。
The following are Legendre’s words, translated into English.以下是勒让德的话(译成英文)。
“Among all the principles that can be proposed for this purpose, I think there is no one more general, more exact, and more easy to apply than that which we have made use of in the preceding researches, and which consists in making the sum of the squares of errors a minimum. In this way there is established a sort of equilibrium among the errors, which prevents the extremes to prevail and is well suited to make us know the state of the system most near to the truth.” (Legendre, 1805, pp. 72–73)“在所有可以为此目的提出的原则中,我认为没有比我们在前述研究中所使用的原则更普遍、更精确、更容易应用的了,该原则在于使误差平方和达到最小值。通过这种方式,在误差之间建立了一种平衡,防止了极值占主导地位,非常适合让我们了解最接近真理的系统状态。”(Legendre, 1805, pp. 72–73)
Unfortunately, it’s not possible to calculate these errors directly since each error εi is the difference between the observed xi and the actual x and the actual x is unknown. Therefore, in practice, we must use the residual ei which is the difference between the observed xi and the estimated value of x, that is, ei = (xi – x_hat).遗憾的是,不可能直接计算这些误差,因为每个误差 εi 都是观测值 xi 与实际值 x 之间的差值,而实际值 x 是未知的。因此,在实践中,我们必须使用残差 ei,它是观测值 xi 与 x 的估计值之间的差值,即 ei = (xi – x_hat)。
Note the difference in notation between the error and the residual. And note the difference in meaning between the two. εi denotes the unknown observation error, while ei denotes the residual which forms the estimate of the observation error.注意误差和残差在符号上的区别。还要注意两者在含义上的区别。εi 表示未知的观测误差,而 ei 表示构成观测误差估计的残差。
Effectively, Legendre had proposed that the best estimate of x would be one that minimizes the sum of squared residuals between the estimate and each observation xi.实际上,勒让德提出,x 的最佳估计将是使估计值与每个观测值 xi 之间的残差平方和最小化的估计。
Thus, the best possible estimate x_hat would be one that minimizes the following sum.因此,最佳估计 x_hat 将是使以下总和最小化的值。

In Eq. (9), S(x_hat) is a function of x_hat. To find the value of x_hat that minimizes the value of this function, we must take its derivative with respect to x_hat, set it to 0, and solve for x_hat as follows.在公式 (9) 中,S(x_hat) 是 x_hat 的函数。为了找到使该函数值最小的 x_hat,我们必须对其关于 x_hat 求导,令其等于 0,并求解 x_hat,具体如下。

By this time, you might have developed at least a vague premonition of where this derivation is headed. For look at what happens when we set the derivative in Eq. (10) to 0 and solve for x_hat.此时,你可能已经隐约预感到这个推导的方向了。让我们看看当我们将公式 (10) 中的导数设为 0 并求解 x_hat 时会发生什么。

Out pops the arithmetic mean of the three observations!结果正是这三个观测值的算术平均数!
Now, tell me, is that neat or what?现在告诉我,这难道不巧妙吗?
Also note that the second derivative of S(.) is 6 > 0 (see below), indicating that at the arithmetic mean of observations, the squared loss in the observations is minimized.还要注意,S(.) 的二阶导数为 6 > 0(见下文),这表明在观测值的算术平均数处,观测值的平方损失达到了最小值。

In general, given an unknown x, and a set of noisy observations x1, x2, x3, …, xn, of that quantity, the simple arithmetic mean of those observations is the estimate that minimizes the sum of squared residuals between the estimate and the observations. This arithmetic mean is the optimal estimate of x under least squares loss.通常情况下,给定一个未知量 x 以及该量的一组噪声观测值 x1, x2, x3, …, xn,这些观测值的简单算术平均数就是使估计值与观测值之间的残差平方和最小化的估计量。在最小二乘损失下,该算术平均数是 x 的最优估计。
Conversely, the sum of squared residuals between the estimate and the observations is minimized at a point where the estimate is the arithmetic mean of the observations.反之,当估计值等于观测值的算术平均数时,估计值与观测值之间的残差平方和达到最小。
Note once again that we are minimizing the sum of squared residuals, not errors, even though the method is usually called the method of least squared errors, or simply, least squares.再次提醒,我们是在最小化残差平方和,而不是误差平方和,尽管该方法通常被称为最小二乘法(method of least squared errors)或简称为最小二乘法(least squares)。
Now, let’s dial up our ambitions a bit.现在,让我们稍微提高一点目标。
Expected Squared Loss and its Connection with the Mean期望平方损失及其与均值的联系
Going beyond the special case of the arithmetic mean, if our goal is to minimize the expected value of the squared error (a.k.a. expected squared loss), then once again, the mean is the only representative of the observed sample that will meet this design goal. This can be proved as follows.超越算术平均数这一特殊情况,如果我们的目标是最小化平方误差的期望值(即期望平方损失),那么均值再次成为观测样本中唯一能满足这一设计目标的代表值。证明如下。
Let x1, x2, x3, …, xn be several observations of a random variable X. To illustrate, X could represent the air temperature recorded at 12 PM CEST on July 1 by the historic Parc Montsouris weather station in Paris. Then, x1, x2, x3, …, xn are the readings taken on July 1 in n randomly selected years.设 x1, x2, x3, …, xn 为随机变量 X 的若干次观测值。举例来说,X 可以代表巴黎历史悠久的蒙苏里公园气象站(Parc Montsouris weather station)在 7 月 1 日下午 12 点记录的气温。那么,x1, x2, x3, …, xn 就是在随机抽取的 n 个年份里 7 月 1 日当天测得的读数。
Suppose we wish to find a suitable constant representative c for X.假设我们希望为 X 找到一个合适的常数代表值 c。
The squared residual between xi and c is the following.xi 与 c 之间的残差平方如下。

Next, we’ll use the following formula for the expected value of a discrete random variable X.接下来,我们将使用以下公式来计算离散随机变量 X 的期望值。

Using the above formula and empirical probabilities pi_hat, the empirical expected squared loss E_hat(ei2) for a sample of size n containing q distinct observed values is given by the following formula.使用上述公式和经验概率 pi_hat,对于包含 q 个不同观测值的样本量为 n 的样本,其经验期望平方损失 E_hat(ei2) 由以下公式给出。

As before, to find the value of c that minimizes the empirical expected squared loss, 𝐿(c), we’ll differentiate 𝐿(c) with respect to c, set the derivative to 0, and solve for c as follows.像之前一样,为了找到使经验期望平方损失 𝐿(c) 最小化的 c 值,我们将 𝐿(c) 对 c 求导,令导数为 0,并求解 c,具体如下。

The second derivative of 𝐿(c) with respect to c is as follows.𝐿(c) 对 c 的二阶导数如下。

Since 𝐿”(c) is positive, setting Eq. (11) to zero will give us the value of c that minimizes the squared loss.由于 𝐿”(c) 为正,将公式 (11) 设为零将得到使平方损失最小化的 c 值。
Let us set Eq. (11) to zero and solve for c.让我们将公式 (11) 设为零并求解 c。

Thus, if our design goal is to find a constant representative for a random variable X that minimizes the empirical expected squared loss, that is, the following.因此,如果我们的设计目标是找到随机变量 X 的一个常数代表值,以最小化经验期望平方损失(即下式):

Then, the solution is the empirical expectation. In other words, the following.那么,其解就是经验期望值。换句话说,即下式:

In simple terms:简单来说:
- The empirical expectation E_hat(X) minimizes the empirical expected squared loss between the observed values of a random variable X and a constant representative c of that random variable.经验期望值 E_hat(X) 能够最小化随机变量 X 的观测值与该随机变量的常数代表值 c 之间的经验期望平方损失。
- Analogously, the population mean, μ, minimizes the population expected squared loss, E[(X – c)2].类似地,总体均值 μ 能够最小化总体期望平方损失 E[(X – c)2]。
Conversely,反之,
- The empirical expected squared loss between the observed values of X and a constant representative c of that random variable is minimized at a point where the estimate c is the empirical expectation E_hat(X) of X.当估计值 c 为 X 的经验期望值 E_hat(X) 时,X 的观测值与该随机变量的常数代表值 c 之间的经验期望平方损失达到最小。
- Analogously, the population expected squared loss, E[(X – c)2] is minimized at a point where the population representative is the expected value E(X) of X.类似地,当总体代表值为 X 的期望值 E(X) 时,总体期望平方损失 E[(X – c)2] 达到最小。
So far, we’ve been working in 1-dimensional space where we’ve seen the connection between the mean and the squared loss for a single random variable X.到目前为止,我们一直在 1 维空间中工作,并观察到了单个随机变量 X 的均值与平方损失之间的联系。
In 1805, when Legendre published his method of least squares, he went way beyond the single unknown. His goal was to show how minimizing the sum of squared errors could be used to estimate the optimal values of all the coefficients in a system of linear equations in which the number of observations exceeded the number of coefficients. Effectively, in the spring of 1805 when Legendre published “Nouvelles Méthodes…”, he had taken aim at nothing short of the Linear Regression model, and he presented a disarmingly simple least-squares based technique of estimating all the coefficients of such a model.1805 年,当勒让德(Legendre)发表他的最小二乘法时,他的研究远不止于单个未知量。他的目标是展示如何通过最小化误差平方和来估计线性方程组中所有系数的最优值,其中观测值的数量超过了系数的数量。实际上,在 1805 年春,当勒让德发表《彗星轨道测定新方法》(Nouvelles Méthodes…)时,他瞄准的正是线性回归模型,并提出了一种极其简单且基于最小二乘的技巧来估计该模型的所有系数。
Legendre, Gauss, and a Matter of Priority勒让德、高斯与优先权问题
Meanwhile, five hundred miles east of Paris, the young and extraordinarily talented Gauss was also busy working on a book: a tome on celestial mathematics.与此同时,在巴黎以东五百英里处,年轻且才华横溢的高斯(Gauss)也正忙于撰写一本关于天体数学的巨著。
In the early 1800s, Gauss could not have been unaffected in his work by the political climate of Europe. Napoleon Bonaparte’s armies were sweeping across the European continent with ever gathering speed. The Duchy of Brunswick-Wolfenbüttel he lived in was strongly aligned with the Kingdom of Prussia, and by Autumn of 1806, Napolean Bonaparte was at Gauss’s doorstep. For Gauss, 14 October 1806 was likely a life-changing day. On this day, Napolean’s armies engaged with the Kingdom of Prussia in two decisive battles at Jena and Auerstedt. The Duke of Brunswick commanded the troops at Auerstedt. It was an all-out disaster. The duke was shot through the eyes and left blind, helpless, and unable to lead. In subsequent weeks, he would die of his injuries. By the end of the day on 14 October, Prussia had lost both battles. It would soon loose half of its territory to France.19 世纪初,高斯的工作不可避免地受到了欧洲政治气候的影响。拿破仑·波拿巴的军队正以越来越快的速度横扫欧洲大陆。他所居住的不伦瑞克-沃尔芬比特尔公国(Duchy of Brunswick-Wolfenbüttel)与普鲁士王国关系紧密,到了 1806 年秋,拿破仑的军队已经兵临高斯家门口。对于高斯来说,1806 年 10 月 14 日很可能是改变人生的一天。这一天,拿破仑的军队在耶拿(Jena)和奥尔施泰特(Auerstedt)与普鲁士王国进行了两场决定性的战役。不伦瑞克公爵指挥了奥尔施泰特的军队,结果是一场彻底的灾难。公爵双眼受枪伤,失明且无助,无法继续指挥。在随后的几周里,他因伤去世。10 月 14 日当天结束时,普鲁士在两场战役中均告败北,并很快失去了其一半的领土。
With the Duke defeated and dead, Gauss’s homeland of Brunswick also came under Napoleonic occupation. In the following year, Napoleon dissolved the Duchy, incorporated its lands into a freshly minted territory called the Kingdom of Westphalia, and handed it over to his youngest brother, the wildly extravagant Jérôme Bonaparte who somehow managed to empty Westphalia’s treasury despite the punishing war taxes imposed by the Emperor on the occupied lands.随着公爵战败身亡,高斯的故乡不伦瑞克也陷入了拿破仑的占领之下。次年,拿破仑解散了该公国,将其领土并入一个新建立的“威斯特伐利亚王国”(Kingdom of Westphalia),并将其交给了他最年轻的弟弟——极其挥霍的热罗姆·波拿巴(Jérôme Bonaparte)。尽管皇帝对占领区征收了沉重的战争税,热罗姆还是设法掏空了威斯特伐利亚的国库。

Meanwhile, in 1806, Gauss had finished working on the German language version of his book but he found it hard to find a local publisher. Eventually, he found one who agreed to publish the work under the condition that it be translated into Latin – probably to reach a pan-European scientific audience. After all, given the heavy war levies coupled with Mr. J. Bonaparte’s spending stunts, it might be safe to assume that academic institutions and celestial observatories in Gauss’s homeland were not brimming with free cash to buy scientific tomes.与此同时,1806 年,高斯完成了他那本书的德语版本,但发现很难找到当地出版商。最终,他找到了一位愿意出版该书的出版商,条件是必须将其翻译成拉丁语——这可能是为了触及全欧洲的科学受众。毕竟,考虑到沉重的战争税再加上热罗姆·波拿巴先生的挥霍,可以肯定的是,高斯故乡的学术机构和天文台并没有多余的现金来购买科学巨著。
Finally in 1809 the Theoria motus corporum coelestium in sectionibus conicis solem ambientium (Theory of the Motion of the Heavenly Bodies Moving about the Sun in Conic Sections) saw the light of day. 最终,在 1809 年,《天体在圆锥曲线内绕日运动理论》(Theoria motus corporum coelestium in sectionibus conicis solem ambientium)问世了。
In his book, Gauss introduced his own version of the method of least squares. While Legendre’s exposition was purely algebraic, Gauss couched his method in the principles of probability and inverse probability which had been developed by Pierre-Simon Laplace in France.在高斯的书中,他介绍了自己版本的最小二乘法。虽然勒让德的阐述纯粹是代数性的,但高斯将他的方法建立在皮埃尔-西蒙·拉普拉斯(Pierre-Simon Laplace)在法国发展的概率论和逆概率论原则之上。
While introducing his own method, Gauss made a passing reference to Legendre’s 1805 paper: “…our principle, which we have made use of since the year 1795, has lately been published by LEGENDRE in the work Nouvelles méthodes pour la détermination des orbites des comètes, Paris 1806, where several other properties of this principle have been explained, which, for the sake of brevity, we here omit.” The emphasis is all mine.在介绍自己的方法时,高斯顺带提及了勒让德 1805 年的论文:“……我们的这一原则,自 1795 年以来我们一直在使用,最近由勒让德在 1806 年巴黎出版的著作《彗星轨道测定新方法》中发表,其中还解释了该原则的其他几种性质,为了简洁起见,我们在此省略。”(重点是我加的)。
It’s hardly surprising that when the above claim reached Legendre, it triggered an almost immediate response from the older, highly principled mathematician: a response that got the young Gauss into a lot of hot water.当上述声明传到勒让德耳中时,这位年长且原则性极强的数学家几乎立即做出了回应,这也就不足为奇了:这一回应让年轻的高斯陷入了极大的麻烦。
On May 31, 1809, Legendre shot off a letter to Gauss from Paris. After a polite preamble dripping with high praise for the younger mathematician’s achievements, Legendre came to the real point of his missive:1809 年 5 月 31 日,勒让德从巴黎给高斯寄去了一封信。在一段充满对这位年轻数学家成就高度赞扬的礼貌开场白之后,勒让德直奔主题:
“…I felt some regret on seeing that, while citing my memoir on page 221, you say principium nostrum quo jam inde ab anno 1795 usi sumus, etc. There is no discovery that one could not claim for oneself by saying that one had found the same thing some years earlier. But if one does not provide proof of this by citing the place where one published it, such an assertion becomes pointless and is nothing more than something hurtful to the true author of the discovery.…You have treasures enough of your own, Sir, to have no need to envy anyone;…” The emphasis is all mine.“……看到你在第 221 页引用我的回忆录时,写道‘principium nostrum quo jam inde ab anno 1795 usi sumus(我们自 1795 年以来一直使用的原则)’等,我感到十分遗憾。任何发现,只要声称自己几年前就已经发现,都可以据为己有。但如果不能通过引用发表地点来证明这一点,这种断言就毫无意义,只会伤害到发现的真正作者……先生,你拥有足够的财富,无需羡慕任何人……”(重点是我加的)。
What emotions lay behind that withering, reproachful prose, no one will ever know. 在那尖刻而责备的文字背后隐藏着怎样的情绪,无人知晓。
Although, one thing was clear: Legendre had pointed out to Gauss in no uncertain terms that one cannot simply claim priority for a discovery by saying that they thought of it first. They must be able to cite where and when they published it.尽管如此,有一点是明确的:勒让德毫不含糊地向高斯指出,一个人不能仅仅通过说自己先想到来声称对某项发现拥有优先权。他们必须能够引用发表的时间和地点。
In effect, the seemingly unanswerable question that Legendre also asked the scientific community was this: if Gauss had discovered the method of least squares in the 1790s, what was he doing keeping it a secret for over a decade? Why didn’t he publish it in the 1790s?实际上,勒让德向科学界提出的那个看似无法回答的问题是:如果高斯在 18 世纪 90 年代就发现了最小二乘法,为什么他要将其保密十多年?为什么他不在 18 世纪 90 年代发表?
Gauss’s correspondences with his friends and with the legendary Pierre-Simon Laplace reveal that his response to such questions was, in essence, the following: in Gauss’s opinion, the method of least squares was a rather simple concept that he had discovered while working on another problem. Besides, he thought it had probably already been used by others before him.高斯与朋友及传奇人物皮埃尔-西蒙·拉普拉斯的通信显示,他对这些问题的回应本质上是:在高斯看来,最小二乘法是一个相当简单的概念,他在研究另一个问题时偶然发现的。此外,他认为在自己之前可能已经有人使用过它。
In other words, it was no big deal, really, and Gauss would mention it in a future publication along with other discoveries (which, incidentally, is exactly what he did in his 1809 book “Theoria motus…”)换句话说,这真的没什么大不了,高斯会在未来的出版物中与其他发现一起提到它(顺便说一句,这正是他在 1809 年的著作《天体运动论》中所做的)。
Unfortunately for Gauss, Legendre took a very poor view of this line of logic, and he made Gauss the target of bitter rancor and name-calling for the next two decades. Laplace, a close contemporary of Legendre’s in Paris, and forever the astute and diplomatic scientist, did his part to douse the controversy by publicly giving both Legendre and Gauss credit for the discovery. But to no avail. Even after Legendre’s death in 1833, Gauss found it difficult to shake himself free of Legendre’s accusations.不幸的是,对于高斯来说,勒让德对这种逻辑非常反感,并在接下来的二十年里让高斯成为了他痛苦怨恨和辱骂的目标。拉普拉斯作为勒让德在巴黎的亲密同僚,一直是一位精明且圆滑的科学家,他试图通过公开将发现的功劳同时归于勒让德和高斯来平息这场争议。但无济于事。即使在 1833 年勒让德去世后,高斯仍然发现很难摆脱勒让德的指责。
Whatever Gauss thought of the method of least squares, it was to dominate statistical discourse throughout the 1800s.无论高斯如何看待最小二乘法,它都将在整个 19 世纪主导统计学的话语体系。
And the primacy that least squares gave to the mean propelled the mean to superpower-status.而最小二乘法赋予均值的首要地位,也推动了均值达到“超级大国”的地位。
Drawbacks of the Mean均值的缺点
For all its advantages, the mean also suffers from a few serious drawbacks.尽管均值有种种优点,但也存在一些严重的缺点。
Failure to Capture the Distributional Shape of the Data无法捕捉数据的分布形态
Look at the following set of distributions. Despite being remarkably different in shape, all of them have the same mean!看看下面这组分布。尽管它们的形状截然不同,但它们的均值却完全相同!

The mean, as a point estimate, has zero ability to capture distributional shape, in other words, the very essence of the data it represents. Thus, it loses much of the information contained in the shape of the data set. From an information-theoretic point of view, the mean is a lossy estimate. This has important implications.均值作为一种点估计,完全无法捕捉分布形态,换句话说,它无法捕捉数据所代表的本质。因此,它丢失了数据集中形状所包含的大部分信息。从信息论的角度来看,均值是一种有损估计。这具有重要的意义。
For example, suppose your data changes over time. Despite different groups within the data moving in opposite directions, the mean might remain the same thereby projecting a deceptive picture of stability.例如,假设你的数据随时间变化。尽管数据中的不同群体朝着相反的方向移动,但均值可能保持不变,从而投射出一种虚假的稳定感。

If your decisions depend on the shape of the distribution, the mean is a poor guide. Incidentally, the median, the mode, the max, and the min suffer from the same drawback, in that they too don’t capture the distributional shape of the data.如果你的决策取决于分布的形状,那么均值就是一个糟糕的指南。顺便提一下,中位数、众数、最大值和最小值都有同样的缺点,因为它们也无法捕捉数据的分布形态。
Unbounded Sensitivity to Outliers对异常值的无界敏感性
The second problem with the mean is that a single rogue outlier can have an outsized influence on the statistic. This sensitivity to outliers emerges directly from the formula for the mean which is a probability-weighted sum. In a sample of size n containing q distinct values x1, x2, …, xq with frequencies n1, n2, …, nq and sample probabilities p_hat1, p_hat2, …, p_hatq respectively, the sample mean is given by the following formula.均值的第二个问题是,单个异常值可能会对统计结果产生过大的影响。这种对异常值的敏感性直接源于均值的公式,即概率加权和。在包含 q 个不同值 x1, x2, …, xq 且频率分别为 n1, n2, …, nq,样本概率分别为 p_hat1, p_hat2, …, p_hatq 的样本量为 n 的样本中,样本均值由以下公式给出。

Now, suppose we extend this sample with an outlier, z. We will assume there is a single outlier.现在,假设我们在样本中增加一个异常值 z。我们假设只有一个异常值。
The new mean for the (n + 1)-sized sample containing (q+1) distinct values, i.e., the original set of q distinct values plus the new outlier, is the following.包含 (q+1) 个不同值(即原始 q 个不同值加上新异常值)的 (n + 1) 大小样本的新均值如下。

Inclusion of z has shifted the mean by the following quantity.加入 z 使得均值偏移了以下数量。

Simplifying, we get the following.简化后,我们得到以下结果。

If z is greater than the old mean, the shift is positive, and vice-versa. 如果 z 大于旧均值,则偏移为正,反之亦然。
Since z is an outlier, by definition, it lies far from the bulk of the sample, and often far from the old mean. So, the value in brackets in Eq. (12) will be large. Although the actual shift in the mean is the bracketed value scaled down by the new sample size (n + 1), for small samples (small n), the bracketed difference isn’t scaled down enough. So, the effect of an outlier on the old mean will be especially large in small samples.由于 z 是异常值,根据定义,它远离样本主体,通常也远离旧均值。因此,公式 (12) 中括号内的值会很大。尽管均值的实际偏移量是括号中的值除以新样本量 (n + 1),但对于小样本(n 很小)来说,括号中的差值并没有被充分缩小。因此,异常值对旧均值的影响在小样本中会特别大。
But even for large n, the shift in the mean can be substantial for extreme values of z.但即使对于较大的 n,对于极端值 z 来说,均值的偏移量也可能相当可观。
In fact, Eq. (12) shows that the mean has unbounded sensitivity to outliers. Even a single sufficiently extreme outlier is enough to thoroughly destabilize the statistic.事实上,公式 (12) 表明均值对异常值具有无界敏感性。即使单个足够极端的异常值也足以彻底破坏该统计量。
A Totally Synthetic Statistic完全合成的统计量
The third problem with the mean is that it’s a totally synthetic statistic. The mean is a computed value. The min, max, and mode are all observed values found in the sample. The median yields the observed value in an odd-sized sample, although it might need to be interpolated in an even-sized sample. The mean carries no such obligations and is little affected by such constraints. It is synthetic by design.均值的第三个问题是它是一个完全合成的统计量。均值是一个计算值。最小值、最大值和众数都是样本中观察到的实际值。中位数在奇数大小的样本中也是观察到的值,尽管在偶数大小的样本中可能需要进行插值。均值没有这样的义务,也不受此类约束的影响。它是人为设计的合成统计量。
The “Average Airman”“平均飞行员”
The synthetic nature of the mean, if not promptly called out, can lead to costly consequences, as illustrated by an anthropometric survey of considerable proportions that was carried out by the U.S. Air Force in 1950.如果均值的合成性质没有被及时指出,可能会导致代价高昂的后果,正如 1950 年美国空军进行的一项大规模人体测量调查所说明的那样。
The Air Force’s survey team visited 14 airbases in the U.S. and took 132 types of body measurements from more than 4,000 Air Force personnel in different roles and ranks. For each measurement such as height, weight, eye-height, arm length, head circumference and so on, the survey authors published percentile tables reporting the 1st to 99th percentile values, along with mean, median, and standard deviation. The stated goal of the survey was to “provide a basis for the design of clothing, equipment, and other aspects of the flight environment”.空军的调查小组走访了美国 14 个空军基地,对 4000 多名担任不同角色和军衔的空军人员进行了 132 种身体测量。对于身高、体重、眼高、臂长、头围等每一项测量指标,调查作者都发布了百分位表,报告了第 1 到第 99 百分位的值,以及均值、中位数和标准差。调查的既定目标是“为服装、设备和飞行环境的其他方面的设计提供基础”。
Unfortunately, the survey failed to adequately highlight the simple fact that the average value of a measurement such as the “sitting height”, or the “neck circumference” used either by itself or in combination with other measurements cannot form an adequate basis for designing any aspect of clothing, equipment, or operating environment.不幸的是,该调查未能充分强调一个简单的事实:像“坐高”或“颈围”这样的测量平均值,无论是单独使用还是与其他测量值结合使用,都不能作为设计服装、设备或操作环境任何方面的充分依据。
Fortunately, this omission was promptly corrected when one of the study authors, G. S. Daniels, published a follow-on report highlighting the dangers of designing anything based on the concept of the “average man”. Daniels took pains to point out that the “average man” did not exist.幸运的是,当研究作者之一 G. S. 丹尼尔斯(G. S. Daniels)发表了一份后续报告,强调了基于“平均人”概念进行设计的危险性时,这一疏忽得到了及时纠正。丹尼尔斯费尽心思指出,“平均人”根本不存在。
He drove home his point by showing how, for any measurement, if the middle 30% of the measured values are considered a rough indication of the average for that measurement, then in the 4063-strong sample in the original survey, not a single airman possessed all of the following: average stature, average circumferences of chest, torso, hip, neck, waist and thigh, average sleeve and crotch lengths, and average crotch height. And this was just a tiny subset of the 132 measurement types covered in the original survey.他通过展示对于任何测量指标,如果将测量值的中间 30% 视为该指标平均值的粗略指示,那么在原始调查的 4063 名样本中,没有一名飞行员同时具备以下所有平均值:平均身高、平均胸围、躯干围、臀围、颈围、腰围和大腿围、平均袖长和裆长,以及平均裆高。而这仅仅是原始调查中涵盖的 132 种测量类型中的一小部分。
To its credit, the Airforce survey did publish percentile scores for each measurement which provide the decision maker with a clear sense for the shape and the range of the measurements. 值得称赞的是,空军调查确实为每项测量发布了百分位分数,这为决策者提供了对测量形状和范围的清晰认识。
Unfortunately, the Airforce example wasn’t the first, and won’t be the last, to research anthropometric averages. The chimerical allure of the “Average Man” has attracted scientists for centuries.不幸的是,空军的例子并不是第一个,也不会是最后一个研究人体测量平均值的例子。“平均人”的虚幻魅力已经吸引了科学家几个世纪。
Adolphe Quetelet and the Myth of the “Average Man”阿道夫·凯特勒与“平均人”的神话
No discussion about the science of averages would be complete without mentioning Quetelet’s seminal contributions to the field.如果不提及凯特勒(Quetelet)对该领域的开创性贡献,关于平均值科学的讨论就不完整。
In the 1820s, the Belgian scientist and astronomer Adolphe Quetelet embarked upon quite possibly one of the largest systematic studies of averages in modern history. From his extensive labors spanning multiple years emerged (surprise, surprise!) the concept of the “Average Man”. In fact, he built a large collection of “average men” each one based on a different feature such as height or weight or a combination of features.19 世纪 20 年代,比利时科学家兼天文学家阿道夫·凯特勒开始了现代史上规模最大的平均值系统研究之一。经过多年广泛的劳动,他得出了(惊喜!)“平均人”的概念。事实上,他建立了一个庞大的“平均人”集合,每一个都基于不同的特征,如身高、体重或特征组合。
Quetelet used these fictional average beings to draw comparisons between groups. For example, he drew seemingly benign comparisons conscripts from different kingdoms in Europe based on aggregate anthropometric measurements such as the average height, weight, and chest measurements.凯特勒利用这些虚构的平均人来对群体进行比较。例如,他基于平均身高、体重和胸围等综合人体测量指标,对欧洲不同王国的应征者进行了看似无害的比较。
Then, Quetelet did something truly troubling.然后,凯特勒做了一件非常令人不安的事情。
He turned his analytical mind toward “moral” characteristics such as the aggregate rates of crime and drunkenness for the different groups.他将分析的目光转向了“道德”特征,如不同群体的犯罪率和酗酒率。
Quetelet noticed that some of these group-specific rates were remarkably stable from year to year and took it as a sign that they reflected persistent group-specific causes and tendencies. He associated such group-specific rates with the “average” individual representing that group.凯特勒注意到,其中一些群体特定的比率在逐年变化中非常稳定,他将其视为反映了持久的群体特定原因和倾向的迹象。他将这种群体特定的比率与代表该群体的“平均”个体联系起来。
In effect, Quetelet was anointing the average as the theoretically correct value, the value that nature intended for that group, and individual differences were simply random deviations from the average.实际上,凯特勒将平均值奉为理论上的正确值,即自然为该群体预定的值,而个体差异仅仅是偏离平均值的随机偏差。
Quetelet, who was an astronomer by profession, had constructed a false equivalence between noisy astronomical observations, and biological and socioeconomic diversity.凯特勒作为一名职业天文学家,在嘈杂的天文观测与生物及社会经济多样性之间构建了一种错误的等价关系。
A celestial object like a star can be said to have a theoretically true position in a specified coordinate system at a specified time. Multiple noisy observations taken through a telescope can be interpreted as deviations from this true value. But this approach doesn’t automatically transfer to a social group of living beings.像恒星这样的天体,可以说在特定时间、特定坐标系中具有理论上的真实位置。通过望远镜拍摄的多份嘈杂观测数据可以被解释为对该真实值的偏差。但这种方法并不能自动转移到生物构成的社会群体中。
When it comes to characteristics such as height, length, weight, intellect, or propensity for crime of living creatures, there is no such thing as a theoretically true, nature-intended value. The average value of such characteristics for a group of people is just that, an average.当涉及到生物的身高、长度、体重、智力或犯罪倾向等特征时,根本不存在什么理论上真实、自然预定的值。对于一群人来说,这些特征的平均值仅仅就是平均值而已。
Any kind of elevation of this statistic to the level of a nature-intended value for a social group, with individual variation being mere deviation from this true value constitutes a gravity-defying intellectual leap of startling proportions.将这一统计量提升到社会群体自然预定值的高度,并将个体差异视为对该真实值的纯粹偏离,这构成了令人震惊的、无视重力的智力飞跃。
On one hand, Quetelet had gifted future generations with a large and rigorously developed body of influential statistical work on averages. On the other hand, his work proved to be dangerously susceptible to misinterpretation.一方面,凯特勒为后代留下了大量且经过严格发展的、关于平均值的有影响力的统计工作。另一方面,他的工作被证明极易受到误解。
Without adequate contextualization, Quetelet’s work could be, and sadly was, extended and twisted into pseudoscientific narratives to elevate or condemn an entire group of people based purely on the notion of a fictitious average representative of that group and the average rates of various qualities attributed to that representative.如果没有充分的背景说明,凯特勒的工作可能会(遗憾的是也确实)被扩展并扭曲成伪科学叙事,仅基于虚构的群体平均代表和归因于该代表的平均素质,来提升或谴责整个群体。
It seems, to Quetelet, such deadly side-effects were never enough of a consideration to begin with.看来,对于凯特勒来说,这种致命的副作用从来都不是需要考虑的问题。
His own words betray the extent to which he was enamored of the deceptively effective representational power wielded by the “average”. At one point, he wrote the following: “If an individual at any given epoch of society possessed all the qualities of the average man, he would represent all that is great, good, or beautiful.” (Anonymous 1835, p. 661; Quetelet 1835, Vol. 2, p. 276; Quetelet 1842, p. 100)他自己的话暴露了他对“平均值”所拥有的、具有欺骗性的有效表征能力的迷恋程度。他曾写道:“如果社会在任何给定纪元的个体拥有平均人的所有品质,他将代表所有伟大、善良或美好的事物。”(Anonymous 1835, p. 661; Quetelet 1835, Vol. 2, p. 276; Quetelet 1842, p. 100)
The Mean Remains the First Among Equals均值依然是“平起平坐中的第一”
Despite all their shortcomings, the mean, the median, and the mode have become the indisputable superpowers of statistical science. But even among the three, the mean is numero uno: the first among equals. 尽管均值、中位数和众数都有其缺点,但它们已成为统计科学中无可争议的超级大国。但在三者之中,均值仍是当之无愧的“numero uno”:平起平坐中的第一。
The mean’s influence on data science runs deep and broad. From measures of dispersion such as the variance, to standardized observations such as the z-score; from moment-based statistics such as the mean and variance themselves, to shape statistics such as the skewness and excess kurtosis; from measures of association such as covariance and Pearson’s correlation, to statistical tests of significance such as the z-test,and economic statistics such as the Gini coefficient, there are literally hundreds of well-known statistical measures that use the mean in their formula.均值对数据科学的影响既深且广。从方差等离散度度量,到 z-score 等标准化观测值;从均值和方差等基于矩的统计量,到偏度和峰度等形状统计量;从协方差和皮尔逊相关系数等关联度量,到 z-test 等统计显著性检验,再到基尼系数等经济统计指标,字面上成百上千种著名的统计测度都在其公式中使用了均值。
In regression analysis, the mean occupies a place of unparalleled significance. A regression model that is trained to minimize the squared error will generally estimate the conditional mean of the response variable.在回归分析中,均值占据着无与伦比的重要地位。一个旨在最小化平方误差的回归模型通常会估计响应变量的条件均值。
The most well-known example of such a model is the linear regression model whose fitting procedure minimizes the squared error. The following equation expresses the mean response of a linear regression model as a linear function of the input variables x1, x2, …, xp.此类模型中最著名的例子是线性回归模型,其拟合过程就是最小化平方误差。以下方程将线性回归模型的平均响应表示为输入变量 x1, x2, …, xp 的线性函数。

In general, statistical models, including neural net models, that are trained to minimize the squared loss end up estimating the conditional mean of the response variable. For such models, Eq. (13) can be generalized as follows.通常,包括神经网络模型在内的旨在最小化平方损失的统计模型,最终都会估计响应变量的条件均值。对于此类模型,公式 (13) 可以推广如下。

In Eq. (14), the function fθ(.) captures the structure of the model. fθ(.) can range from a simple linear function as shown in Eq. (13) to a frightfully nonlinear and complex representation arising from a multilayer neural net model. The θ subscript denotes all the weights, biases, normalization parameters – everything that qualifies as a tunable or trainable value.在公式 (14) 中,函数 fθ(.) 捕捉了模型的结构。fθ(.) 的范围可以从公式 (13) 中所示的简单线性函数,到由多层神经网络模型产生的极其非线性和复杂的表征。下标 θ 表示所有的权重、偏置、归一化参数——所有符合可调或可训练值条件的内容。
From average commute times shown in Google Maps, to product ratings shown on Amazon, from HbA1c levels to average life expectancies that drive insurance premiums, from average standardized test scores that shape education policy, to average hourly earnings and consumer price indexes that inform economic policy, from average deal sizes, average retail order values, and average hotel room rates to the average amount of time audiences spend reading my work before moving on to other things, the list of areas in which the mean makes its presence felt is endless.从谷歌地图显示的平均通勤时间,到亚马逊上的产品评分;从 HbA1c 水平到驱动保险费的平均预期寿命;从塑造教育政策的平均标准化考试成绩,到为经济政策提供参考的平均时薪和消费者价格指数;从平均交易规模、平均零售订单价值和平均酒店房价,到读者在转向其他内容之前阅读我作品的平均时长,均值发挥作用的领域不胜枚举。
The mean informs and guides every aspect of human existence.均值告知并指导着人类存在的方方面面。
References参考文献
Anonymous (1835), “On Man, and the Developement of His Faculties, &c.—[Sur l’Homme et le Développement de ses Facultés, &c.], by A. Quetelet,” The Athenæum, August 8, 593–594; August 15, 611–613; August 29, 658–661 [URL]Anonymous (1835), “On Man, and the Developement of His Faculties, &c.—[Sur l’Homme et le Développement de ses Facultés, &c.], by A. Quetelet,” The Athenæum, August 8, 593–594; August 15, 611–613; August 29, 658–661 [URL]
Bernoulli, J. (2005 [1713]), On the Law of Large Numbers: Part Four of Ars Conjectandi, O. Sheynin, (translated to English)., Berlin: NG Verlag. [PDF]Bernoulli, J. (2005 [1713]), On the Law of Large Numbers: Part Four of Ars Conjectandi, O. Sheynin, (translated to English)., Berlin: NG Verlag. [PDF]
Bühler, W. K. (1981), Gauss: A Biographical Study, Berlin: Springer-Verlag, doi:10.1007/978-3-642-49207-5.Bühler, W. K. (1981), Gauss: A Biographical Study, Berlin: Springer-Verlag, doi:10.1007/978-3-642-49207-5.
Daniels, G. S. (1952), The “Average Man”?, WCRD Technical Note 53-7, Wright-Patterson Air Force Base, OH: Wright Air Development Center. [PDF]Daniels, G. S. (1952), The “Average Man”?, WCRD Technical Note 53-7, Wright-Patterson Air Force Base, OH: Wright Air Development Center. [PDF]
Gauss, C. F. (1809), Theoria Motus Corporum Coelestium in Sectionibus Conicis Solem Ambientium, Hamburg: F. Perthes and I. H. Besser.Gauss, C. F. (1809), Theoria Motus Corporum Coelestium in Sectionibus Conicis Solem Ambientium, Hamburg: F. Perthes and I. H. Besser.
Gauss, C. F. (1857), Theory of the Motion of the Heavenly Bodies Moving about the Sun in Conic Sections: A Translation of Gauss’s “Theoria Motus,” with an Appendix, C. H. Davis, trans., Boston: Little, Brown and Company.Gauss, C. F. (1857), Theory of the Motion of the Heavenly Bodies Moving about the Sun in Conic Sections: A Translation of Gauss’s “Theoria Motus,” with an Appendix, C. H. Davis, trans., Boston: Little, Brown and Company.
Hald, A. (2007), A History of Parametric Statistical Inference from Bernoulli to Fisher, 1713–1935, New York: Springer, doi:10.1007/978-0-387-46409-1.Hald, A. (2007), A History of Parametric Statistical Inference from Bernoulli to Fisher, 1713–1935, New York: Springer, doi:10.1007/978-0-387-46409-1.
Hertzberg, H. T. E., Churchill, E., and Daniels, G. S. (1954), Anthropometry of Flying Personnel—1950, WADC Technical Report 52-321, Wright-Patterson Air Force Base, OH: Aero Medical Laboratory, Wright Air Development Center, U.S. Air Force, and Antioch College, Contract AF 18(600)-30, PB111583, AD0047953. [PDF]Hertzberg, H. T. E., Churchill, E., and Daniels, G. S. (1954), Anthropometry of Flying Personnel—1950, WADC Technical Report 52-321, Wright-Patterson Air Force Base, OH: Aero Medical Laboratory, Wright Air Development Center, U.S. Air Force, and Antioch College, Contract AF 18(600)-30, PB111583, AD0047953. [PDF]
Legendre, A.-M. (1805), Nouvelles Méthodes pour la Détermination des Orbites des Comètes, Paris: Firmin Didot, pp. 72–80.Legendre, A.-M. (1805), Nouvelles Méthodes pour la Détermination des Orbites des Comètes, Paris: Firmin Didot, pp. 72–80.
Niedersächsische Akademie der Wissenschaften zu Göttingen (n.d.), Complete Correspondence of Carl Friedrich Gauß, online correspondence databaseNiedersächsische Akademie der Wissenschaften zu Göttingen (n.d.), Complete Correspondence of Carl Friedrich Gauß, online correspondence database
Plackett, R. L. (1972), “Studies in the History of Probability and Statistics. XXIX: The Discovery of the Method of Least Squares,” Biometrika, 59, 239–251, doi:10.1093/biomet/59.2.239. [PDF]Plackett, R. L. (1972), “Studies in the History of Probability and Statistics. XXIX: The Discovery of the Method of Least Squares,” Biometrika, 59, 239–251, doi:10.1093/biomet/59.2.239. [PDF]
Polasek, W. (2000), “The Bernoullis and the Origin of Probability Theory: Looking Back after 300 Years,” Resonance, 5, 26–42, doi:10.1007/BF02837935. [PDF]Polasek, W. (2000), “The Bernoullis and the Origin of Probability Theory: Looking Back after 300 Years,” Resonance, 5, 26–42, doi:10.1007/BF02837935. [PDF]
Quetelet, A. (1835), Sur l’homme et le développement de ses facultés, ou Essai de physique sociale, 2 vols., Paris: Bachelier.Quetelet, A. (1835), Sur l’homme et le développement de ses facultés, ou Essai de physique sociale, 2 vols., Paris: Bachelier.
Quetelet, A. (1842), A Treatise on Man and the Development of His Faculties, English translation supervised by R. Knox and edited by T. Smibert, Edinburgh: William and Robert Chambers.Quetelet, A. (1842), A Treatise on Man and the Development of His Faculties, English translation supervised by R. Knox and edited by T. Smibert, Edinburgh: William and Robert Chambers.
Seneta, E. (2013), “A Tricentenary History of the Law of Large Numbers,” Bernoulli, 19, 1088–1121, doi:10.3150/12-BEJSP12. [PDF]Seneta, E. (2013), “A Tricentenary History of the Law of Large Numbers,” Bernoulli, 19, 1088–1121, doi:10.3150/12-BEJSP12. [PDF]
Stigler, S. M. (1986), The History of Statistics: The Measurement of Uncertainty before 1900, Cambridge, MA: Belknap Press of Harvard University Press, doi:10.2307/2982057Stigler, S. M. (1986), The History of Statistics: The Measurement of Uncertainty before 1900, Cambridge, MA: Belknap Press of Harvard University Press, doi:10.2307/2982057
Copyrights版权
This article is © Sachin Date and is licensed under CC BY-SA.本文 © Sachin Date,采用 CC BY-SA 许可协议。
Unless otherwise indicated in an image caption, all images in this article are © Sachin Date and are licensed under CC BY-SA.除非图片说明中另有说明,本文中的所有图片均 © Sachin Date,并采用 CC BY-SA 许可协议。






