Preply — Study more efficiently by working with a personal tutor. Get 50% off.Affiliate

Wikipedia

E-values

In statistical hypothesis testing, e-values quantify the evidence in the data against a null hypothesis (e.g., "the coin is fair", or, in a medical context, "this new treatment has no effect"). They serve as a more robust alternative to p-values, addressing some shortcomings of the latter. In contrast to p-values, e-values can deal with optional continuation: e-values of subsequent experiments (e.g. clinical trials concerning the same treatment) may simply be multiplied to provide a new, "product" e-value that represents the evidence in the joint experiment. This works even if, as often happens in practice, the decision to perform later experiments may depend in vague, unknown ways on the data observed in earlier experiments, and it is not known beforehand how many trials will be conducted: the product e-value remains a meaningful quantity, leading to tests with Type-I error control. For this reason, e-values and their sequential extension, the e-process, are the fundamental building blocks for anytime-valid statistical methods (e.g. confidence sequences). Another advantage over p-values is that any weighted average of e-values remains an e-value, even if the individual e-values are arbitrarily dependent. This is one of the reasons why e-values have also turned out to be useful tools in multiple testing. E-values can be interpreted in a number of different ways: first, an e-value can be interpreted as rescaling of a test that is presented on a more appropriate scale that facilitates merging them. Second, the reciprocal of an e-value is a p-value, but not just any p-value: a special p-value for which a rejection `at level p' retains a generalized Type-I error guarantee. Third, they are broad generalizations of likelihood ratios and are also related to, yet distinct from, Bayes factors. Fourth, they have an interpretation as bets. Fifth, in a sequential context, they can also be interpreted as increments of nonnegative supermartingales. Interest in e-values has exploded since 2019, when the term 'e-value' was coined and a number of breakthrough results were achieved by several research groups. The first overview article appeared in 2023.

Definition and mathematical background Let the null hypothesis H 0 {\displaystyle H_{0}} be given as a set of distributions for data Y {\displaystyle Y} . Usually Y = ( X 1 , … , X τ ) {\displaystyle Y=(X_{1},\ldots ,X_{\tau })} with each X i {\displaystyle X_{i}} a single outcome and τ {\displaystyle \tau } a fixed sample size or some stopping time. We shall refer to such Y {\displaystyle Y} , which represent the full sequence of outcomes of a statistical experiment, as a sample or batch of outcomes. But in some cases Y {\displaystyle Y} may also be an unordered bag of outcomes or a single outcome. An e-variable or e-statistic is a nonnegative random variable E = E ( Y ) {\displaystyle E=E(Y)} such that under all P ∈ H 0 {\displaystyle P\in H_{0}} , its expected value is bounded by 1:

E P [ E ] ≤ 1 {\displaystyle {\mathbb {E} }_{P}[E]\leq 1} . The value taken by e-variable E {\displaystyle E} is called the e-value. In practice, the term e-value (a number) is often used when one is really referring to the underlying e-variable (a random variable, that is, a measurable function of the data).

Interpretations

As the continuous interpretation of a test A test for a null hypothesis H 0 {\displaystyle H_{0}} is traditionally modeled as a function ϕ {\displaystyle \phi } from the data to { not reject H 0 , reject H 0 } {\displaystyle \{{\text{not reject }}H_{0},{\text{ reject }}H_{0}\}} . A test ϕ α {\displaystyle \phi _{\alpha }} is said to be valid for level α {\displaystyle \alpha } if

P ( ϕ α = reject H 0 ) ≤ α , for every P ∈ H 0 . {\displaystyle P(\phi _{\alpha }={\text{reject }}H_{0})\leq \alpha ,{\text{ for every }}P\in H_{0}.}

This is classically conveniently summarized as a function ϕ α {\displaystyle \phi _{\alpha }} from the data to { 0 , 1 } {\displaystyle \{0,1\}} that satisfies

E P [ ϕ α ] ≤ α , for every P ∈ H 0 {\displaystyle \mathbb {E} ^{P}[\phi _{\alpha }]\leq \alpha ,{\text{ for every }}P\in H_{0}} . Moreover, this is sometimes generalized to permit external randomization by letting the test ϕ α {\displaystyle \phi _{\alpha }} take value in [ 0 , 1 ] {\displaystyle [0,1]} . Here, its value is interpreted as a probability with which one should subsequently reject the hypothesis. An issue with modelling a test in this manner, is that the traditional decision space { not reject H 0 , reject H 0 } {\displaystyle \{{\text{not reject }}H_{0},{\text{ reject }}H_{0}\}} or { 0 , 1 } {\displaystyle \{0,1\}} does not encode the level α {\displaystyle \alpha } at which the test ϕ α {\displaystyle \phi _{\alpha }} rejects. This is odd at best, because a rejection at level 1% is a much stronger claim than a rejection at level 10%. A more suitable decision space seems to be { not reject H 0 , reject H 0 at level α } {\displaystyle \{{\text{not reject }}H_{0},{\text{ reject }}H_{0}{\text{ at level }}\alpha \}} . The e-value can be interpreted as resolving this problem. Indeed, we can rescale from { 0 , 1 } {\displaystyle \{0,1\}} to { 0 , 1 / α } {\displaystyle \{0,1/\alpha \}} and [ 0 , 1 ] {\displaystyle [0,1]} to [ 0 , 1 / α ] {\displaystyle [0,1/\alpha ]} by rescaling the test by its level:

ε α = ϕ α / α {\displaystyle \varepsilon _{\alpha }=\phi _{\alpha }/\alpha } , where we denote a test on this evidence scale by ε α {\displaystyle \varepsilon _{\alpha }} to avoid confusion. Such a test is then valid if

E P [ ε α ] ≤ 1 , for every P ∈ H 0 {\displaystyle \mathbb {E} ^{P}[\varepsilon _{\alpha }]\leq 1,{\text{ for every }}P\in H_{0}} . That is: it is valid if it is an e-value. In fact, this reveals that e-values bounded to [ 0 , 1 / α ] {\displaystyle [0,1/\alpha ]} are rescaled randomized tests, that are continuously interpreted as evidence against the hypothesis. The standard e-value that takes value in [ 0 , ∞ ] {\displaystyle [0,\infty ]} appears as a generalization of a level 0 test. This interpretation shows that e-values are indeed fundamental to testing: they are equivalent to tests, thinly veiled by a rescaling. From this perspective, it may be surprising that typical e-values look very different from traditional tests: maximizing the objective

E Q [ ε α ] {\displaystyle \mathbb {E} ^{Q}[\varepsilon _{\alpha }]}

for an alternative hypothesis H 1 = { Q } {\displaystyle H_{1}=\{Q\}} would yield traditional Neyman-Pearson style tests. Indeed, this maximizes the probability under Q {\displaystyle Q} that ε α = 1 / α {\displaystyle \varepsilon _{\alpha }=1/\alpha } . But if we continuously interpret the value of the test ε α {\displaystyle \varepsilon _{\alpha }} as evidence against the hypothesis, then we may also be interested in maximizing different targets such as

E Q [ log ⁡ ε α ] {\displaystyle \mathbb {E} ^{Q}[\log \varepsilon _{\alpha }]} . This yields tests that are remarkably different from traditional Neyman-Pearson tests, and more suitable when merged through multiplication as they are positive with probability 1 under Q {\displaystyle Q} . From this angle, the main innovation of the e-value compared to traditional testing is to maximize a different power target.

As p-values with a stronger data-dependent-level Type-I error guarantee For any e-variable E {\displaystyle E} and any 0 < α ≤ 1 {\displaystyle 0<\alpha \leq 1} and all P ∈ H 0 {\displaystyle P\in H_{0}} , it holds that

P ( E ≥ 1 α ) = P ( 1 / E ≤ α ) ≤ ( ∗ ) α {\displaystyle P\left(E\geq {\frac {1}{\alpha }}\right)=P(1/E\leq \alpha )\ {\overset {(*)}{\leq }}\ \alpha } . This means p ′ = 1 / E {\displaystyle p^{\prime }=1/E} is a valid p-value. Moreover, the e-value based test with significance level α {\displaystyle \alpha } , which rejects P 0 {\displaystyle P_{0}} if p ′ ≤ α {\displaystyle p^{\prime }\leq \alpha } , has a Type-I error bounded by α {\displaystyle \alpha } . But, whereas with standard p-values the inequality (*) above is usually an equality (with continuous-valued data) or near-equality (with discrete data), this is not the case with e-variables. This makes e-value-based tests more conservative (less power) than those based on standard p-values. In exchange for this conservativeness, the p-value p ′ = 1 / E {\displaystyle p^{\prime }=1/E} comes with a stronger guarantee. In particular, for every possibly data-dependent significance level α ~ > 0 {\displaystyle {\widetilde {\alpha }}>0} , we have

E [ P ( p ′ ≤ α ~ ∣ α ~ ) α ~ ] ≤ 1 , {\displaystyle \mathbb {E} \left[{\frac {P(p^{\prime }\leq {\widetilde {\alpha }}\mid {\widetilde {\alpha }})}{\widetilde {\alpha }}}\right]\leq 1,}

if and only if E [ 1 / p ′ ] ≤ 1 {\displaystyle \mathbb {E} [1/p^{\prime }]\leq 1} . This means that a p-value satisfies this guarantee if and only if it is the reciprocal 1 / E {\displaystyle 1/E} of an e-variable E {\displaystyle E} . The interpretation of this guarantee is that, on average, the relative Type-I error distortion P ( p ′ ≤ α ~ ∣ α ~ ) / α ~ {\displaystyle P(p^{\prime }\leq {\widetilde {\alpha }}\mid {\widetilde {\alpha }})/{\widetilde {\alpha }}} caused by using a data-dependent level α ~ {\displaystyle {\widetilde {\alpha }}} is controlled for every choice of the data-dependent significance level. Traditional p-values only satisfy this guarantee for data-independent or pre-specified levels. This stronger guarantee is also called the post-hoc α {\displaystyle \alpha } Type-I error, as it allows one to choose the significance level after observing the data: post-hoc. A p-value that satisfies this guarantee is also called a post-hoc p-value. As p ′ {\displaystyle p^{\prime }} is a post-hoc p-value if and only if p ′ = 1 / E {\displaystyle p^{\prime }=1/E} for some e-value E {\displaystyle E} , it is possible to view this as an alternative definition of an e-value. Under this post-hoc Type-I error, the problem of choosing the significance level α {\displaystyle \alpha } vanishes: we can simply choose the smallest data-dependent level at which we reject the hypothesis by setting it equal to the post-hoc p-value: α ~ = p ′ {\displaystyle {\widetilde {\alpha }}=p^{\prime }} . Indeed, at this data-dependent level we have

E [ P ( p ′ ≤ p ′ ∣ p ′ ) p ′ ] = E [ 1 p ′ ] ≤ 1 , {\displaystyle \mathbb {E} \left[{\frac {P(p^{\prime }\leq p^{\prime }\mid p^{\prime })}{p^{\prime }}}\right]=\mathbb {E} \left[{\frac {1}{p^{\prime }}}\right]\leq 1,}

since 1 / p ′ {\displaystyle 1/p^{\prime }} is an e-variable. As a consequence, we can truly reject at level p ′ {\displaystyle p^{\prime }} and still retain the post-hoc Type-I error guarantee. For a traditional p-value p {\displaystyle p} , rejecting at level p comes with no such guarantee. Moreover, a post-hoc p-value inherits optional continuation and merging properties of e-values. But instead of an arithmetic weighted average, a weighted harmonic average of post-hoc p-values is still a post-hoc p-value.

As generalizations of likelihood ratios Let H 0 = { P 0 } {\displaystyle H_{0}=\{P_{0}\}} be a simple null hypothesis. Let Q {\displaystyle Q} be any other distribution on Y {\displaystyle Y} , and let

E := q ( Y ) p 0 ( Y ) {\displaystyle E:={\frac {q(Y)}{p_{0}(Y)}}}

be their likelihood ratio. Then E {\displaystyle E} is an e-variable. Conversely, any e-variable relative to a simple null H 0 = { P 0 } {\displaystyle H_{0}=\{P_{0}\}} can be written as a likelihood ratio with respect to some distribution Q {\displaystyle Q} . Thus, when the null is simple, e-variables coincide with likelihood ratios. E-variables exist for general composite nulls as well though, and they may then be thought of as generalizations of likelihood ratios. The two main ways of constructing e-variables, UI and RIPr (see below) both lead to expressions that are variations of likelihood ratios as well. Two other standard generalizations of the likelihood ratio are (a) the generalized likelihood ratio as used in the standard, classical likelihood ratio test and (b) the Bayes factor. Importantly, neither (a) nor (b) are e-variables in general: generalized likelihood ratios in sense (a) are not e-variables unless the alternative is simple (see below under "universal inference"). Bayes factors are e-variables if the null is simple. To see this, note that, if Q = { Q θ : θ ∈ Θ } {\displaystyle {\mathcal {Q}}=\{Q_{\theta }:\theta \in \Theta \}} represents a statistical model, and w {\displaystyle w} a prior density on Θ {\displaystyle \Theta } , then we can set Q {\displaystyle Q} as above to be the Bayes marginal distribution with density

q ( Y ) = ∫ q θ ( Y ) w ( θ ) d θ {\displaystyle q(Y)=\int q_{\theta }(Y)w(\theta )d\theta }

and then E = q ( Y ) / p 0 ( Y ) {\displaystyle E=q(Y)/p_{0}(Y)} is also a Bayes factor of H 0 {\displaystyle H_{0}} vs. H 1 := Q {\displaystyle H_{1}:={\mathcal {Q}}} . If the null is composite, then some special e-variables can be written as Bayes factors with some very special priors, but most Bayes factors one encounters in practice are not e-variables and many e-variables one encounters in practice are not Bayes factors.

As bets Suppose you can buy a ticket for 1 monetary unit, with nonnegative pay-off E = E ( Y ) {\displaystyle E=E(Y)} . The statements " E {\displaystyle E} is an e-variable" and "if the null hypothesis is true, you do not expect to gain any money if you engage in this bet" are logically equivalent. This is because E {\displaystyle E} being an e-variable means that the expected gain of buying the ticket is the pay-off minus the cost, i.e. E − 1 {\displaystyle E-1} , which has expectation ≤ 0 {\displaystyle \leq 0} . Based on this interpretation, the product e-value for a sequence of tests can be interpreted as the amount of money you have gained by sequentially betting with pay-offs given by the individual e-variables and always re-investing all your gains. The betting interpretation becomes particularly visible if we rewrite an e-variable as E := 1 + λ U {\displaystyle E:=1+\lambda U} where U {\displaystyle U} has expectation ≤ 0 {\displaystyle \leq 0} under all P ∈ H 0 {\displaystyle P\in H_{0}} and λ ∈ R {\displaystyle \lambda \in {\mathbb {R} }} is chosen so that E ≥ 0 {\displaystyle E\geq 0} a.s. Any e-variable can be written in the 1 + λ U {\displaystyle 1+\lambda U} form although with parametric nulls, writing it as a likelihood ratio is usually mathematically more convenient. The 1 + λ U {\displaystyle 1+\lambda U} form on the other hand is often more convenient in nonparametric settings. As a prototypical example, consider the case that Y = ( X 1 , … , X n ) {\displaystyle Y=(X_{1},\ldots ,X_{n})} with the X i {\displaystyle X_{i}} taking values in the bounded interval [ 0 , 1 ] {\displaystyle [0,1]} . According to H 0 {\displaystyle H_{0}} , the X i {\displaystyle X_{i}} are i.i.d. according to a distribution P {\displaystyle P} with mean μ {\displaystyle \mu } ; no other assumptions about P {\displaystyle P} are made. Then we may first construct a family of e-variables for single outcomes, E i , λ := 1 + λ ( X i − μ ) {\displaystyle E_{i,\lambda }:=1+\lambda (X_{i}-\mu )} , for any λ ∈ [ − 1 / ( 1 − μ ) , 1 / μ ] {\displaystyle \lambda \in [-1/(1-\mu ),1/\mu ]} (these are the λ {\displaystyle \lambda } for which E i , λ {\displaystyle E_{i,\lambda }} is guaranteed to be nonnegative). We may then define a new e-variable for the complete data vector Y {\displaystyle Y} by taking the product

E := ∏ i = 1 n E i , λ ˘ | X i − 1 {\displaystyle E:=\prod _{i=1}^{n}E_{i,{\breve {\lambda }}|X^{i-1}}} , where λ ˘ | X i − 1 {\displaystyle {\breve {\lambda }}|X^{i-1}} is an estimate for λ {\displaystyle {\lambda }} , based only on past data X i − 1 = ( X 1 , … , X i − 1 ) {\displaystyle X^{i-1}=(X_{1},\ldots ,X_{i-1})} , and designed to make E i , λ {\displaystyle E_{i,\lambda }} as large as possible in the "e-power" or "GRO" sense (see below). Waudby-Smith and Ramdas use this approach to construct "nonparametric" confidence intervals for the mean that tend to be significantly narrower than those based on more classical methods such as Chernoff, Hoeffding and Bernstein bounds.

A fundamental property: optional continuation E-values are more suitable than p-value when one expects follow-up tests involving the same null hypothesis with different data or experimental set-ups. This includes, for example, combining individual results in a meta-analysis. The advantage of e-values in this setting is that they allow for optional continuation. Indeed, they have been employed in what may be the world's first fully 'online' meta-analysis with explicit Type-I error control. Informally, optional continuation implies that the product of any number of e-values, E ( 1 ) , E ( 2 ) , … {\displaystyle E_{(1)},E_{(2)},\ldots } , defined on independent samples Y ( 1 ) , Y ( 2 ) , … {\displaystyle Y_{(1)},Y_{(2)},\ldots } , is itself an e-value, even if the definition of each e-value is allowed to depend on all previous outcomes, and no matter what rule is used to decide when to stop gathering new samples (e.g. to perform new trials). It follows that, for any significance level 0 < α < 1 {\displaystyle 0<\alpha <1} , if the null is true, then the probability that a product of e-values will ever become larger than 1 / α {\displaystyle 1/\alpha } is bounded by α {\displaystyle \alpha } . Thus if we decide to combine the samples observed so far and reject the null if the product e-value is larger than 1 / α {\displaystyle 1/\alpha } , then our Type-I error probability remains bounded by α {\displaystyle \alpha } . We say that testing based on e-values remains safe (Type-I valid) under optional continuation. Mathematically, this is shown by first showing that the product e-variables form a nonnegative discrete-time martingale in the filtration generated by Y ( 1 ) , Y ( 2 ) , … {\displaystyle Y_{(1)},Y_{(2)},\ldots } (the individual e-variables are then increments of this martingale). The results then follow as a consequence of Doob's optional stopping theorem and Ville's inequality. We already implicitly used product e-variables in the example above, where we defined e-variables on individual outcomes X i {\displaystyle X_{i}} and designed a new e-value by taking products. Thus, in the example, the individual outcomes X i {\displaystyle X_{i}} play the role of 'batches' (full samples) Y ( j ) {\displaystyle Y_{(j)}} above, and we can therefore even engage in optional stopping "within" the original batch Y {\displaystyle Y} : we may stop the data analysis at any individual outcome (not just "batch of outcomes") we like, for whatever reason, and reject if the product so far exceeds 1 / α {\displaystyle 1/\alpha } . Not all e-variables defined for batches of outcomes Y {\displaystyle Y} can be decomposed as a product of per-outcome e-values in this way though. If this is not possible, we cannot use them for optional stopping (within a sample Y {\displaystyle Y} ) but only for optional continuation (from one sample Y ( j ) {\displaystyle Y_{(j)}} to the next Y ( j + 1 ) {\displaystyle Y_{(j+1)}} and so on).

Construction and optimality If we set E := 1 {\displaystyle E:=1} independently of the data we get a trivial e-value: it is an e-variable by definition, but it will never allow us to reject the null hypothesis. This example shows that some e-variables may be better than others, in a sense to be defined below. Intuitively, a good e-variable is one that tends to be large (much larger than 1) if the alternative is true. This is analogous to the situation with p-values: both e-values and p-values can be defined without referring to an alternative, but if an alternative is available, we would like them to be small (p-values) or large (e-values) with high probability. In standard hypothesis tests, the quality of a valid test is formalized by the notion of statistical power but this notion has to be suitably modified in the context of e-values. The standard notion of quality of an e-variable relative to a given alternative H 1 {\displaystyle H_{1}} , used by most authors in the field, is a generalization of the Kelly criterion in economics and (since it does exhibit close relations to classical power) is sometimes called e-power; the optimal e-variable in this sense is known as log-optimal or growth-rate optimal (often abbreviated to GRO). In the case of a simple alternative H 1 = { Q } {\displaystyle H_{1}=\{Q\}} , the e-power of a given e-variable S {\displaystyle S} is simply defined as the expectation E Q [ log ⁡ E ] {\displaystyle {\mathbb {E} }_{Q}[\log E]} ; in case of composite alternatives, there are various versions (e.g. worst-case absolute, worst-case relative) of e-power and GRO.

Simple alternative, simple null: likelihood ratio Let H 0 = { P 0 } {\displaystyle H_{0}=\{P_{0}\}} and H 1 = { Q } {\displaystyle H_{1}=\{Q\}} both be simple. Then the likelihood ratio e-variable E = q ( Y ) / p 0 ( Y ) {\displaystyle E=q(Y)/p_{0}(Y)} has maximal e-power in the sense above, i.e. it is GRO.

Simple alternative, composite null: reverse information projection (RIPr) Let H 1 = { Q } {\displaystyle H_{1}=\{Q\}} be simple and H 0

Tags

  • Probability theory
  • Statistical concepts
  • Statistical hypothesis testing