In statistics, an effect size is a quantitative measure of the magnitude of a phenomenon. It can refer to the value of a statistic calculated from a sample of data, the value of one parameter for a hypothetical population, or the equation that operationalizes how statistics or parameters lead to the effect size value. Examples of effect sizes include the correlation between two variables, the regression coefficient in a regression, the mean difference, and the risk of a particular event (such as a heart attack). Effect sizes are a complementary tool for statistical hypothesis testing, and play an important role in statistical power analyses to assess the sample size required for new experiments. Effect size calculations are fundamental to meta-analysis, which aims to provide the combined effect size based on data from multiple studies. The group of data-analysis methods concerning effect sizes is referred to as estimation statistics. Effect size is an essential component in the evaluation of the strength of a statistical claim, and it is the first item (magnitude) in the MAGIC criteria. The standard deviation of the effect size is of critical importance, as it indicates how much uncertainty is included in the observed measurement. A standard deviation that is too large will make the measurement nearly meaningless. In meta-analysis, which aims to summarize multiple effect sizes into a single estimate, the uncertainty in studies' effect sizes is used to weight the contribution of each study, so larger studies are considered more important than smaller ones. The uncertainty in the effect size is calculated differently for each type of effect size, but generally only requires knowing the study's sample size (N), or the number of observations (n) in each group. Reporting effect sizes or estimates thereof (effect estimate [EE], estimate of effect) is considered good practice when presenting empirical research findings in many fields. The reporting of effect sizes facilitates the interpretation of the importance of a research result, in contrast to its statistical significance. Effect sizes are particularly significant in social science and medical research, with the latter emphasizing the importance of the magnitude of the average treatment effect. Effect sizes may be measured in relative or absolute terms. In relative effect sizes, two groups are directly compared with each other, as in odds ratios and relative risks. A larger absolute value always indicates a stronger effect for absolute effect sizes. Many types of measurements can be expressed as either absolute or relative, and these can be used together because they convey different information. A prominent task force in the psychology research community made the following recommendation:
Always present effect sizes for primary outcomes...If the units of measurement are meaningful on a practical level (e.g., number of cigarettes smoked per day), then we usually prefer an unstandardized measure (regression coefficient or mean difference) to a standardized measure (r or d).
Overview
Population and sample effect sizes As in statistical estimation, the true effect size is distinguished from the observed effect size. For example, to measure the risk of disease in a population (the population effect size) one can measure the risk within a sample of that population (the sample effect size). Conventions for describing true and observed effect sizes follow standard statistical practices—one common approach is to use Greek letters like ρ [rho] to denote population parameters and Latin letters like r to denote the corresponding statistic. Alternatively, a "hat" can be placed over the population parameter to denote the statistic, e.g. with ρ ^ {\displaystyle {\hat {\rho }}} being the estimate of the parameter ρ {\displaystyle \rho } . As in any statistical setting, effect sizes are estimated with sampling error, and may be biased unless the effect size estimator that is used is appropriate for the manner in which the data were sampled and the manner in which the measurements were made. An example of this is publication bias, which occurs when scientists report results only when the estimated effect sizes are large or are statistically significant. As a result, if many researchers carry out studies with low statistical power, the reported effect sizes will tend to be larger than the true (population) effects, if any. Another example where effect sizes may be distorted is in a multiple-trial experiment, where the effect size calculation is based on the averaged or aggregated response across the trials. Smaller studies sometimes show different, often larger, effect sizes than larger studies. This phenomenon is known as the small-study effect, which may signal publication bias.
Relationship to test statistics Sample-based effect sizes are distinguished from test statistics used in hypothesis testing, in that they estimate the strength (magnitude) of, for example, an apparent relationship, rather than assigning a significance level reflecting whether the magnitude of the relationship observed could be due to chance. The effect size does not directly determine the significance level, or vice versa. Given a sufficiently large sample size, a non-null statistical comparison will always show a statistically significant result unless the population effect size is exactly zero (and even there it will show statistical significance at the rate of the Type I error used). For example, a sample Pearson correlation coefficient of 0.01 is statistically significant if the sample size is 1000. Reporting only the significant p-value from this analysis could be misleading if a correlation of 0.01 is too small to be of interest in a particular application.
Standardized and unstandardized effect sizes The term effect size can refer to a standardized measure of effect (such as r, Cohen's d, or the odds ratio), or to an unstandardized measure (e.g., the difference between group means or the unstandardized regression coefficients). Standardized effect size measures are typically used when:
the metrics of variables being studied do not have intrinsic meaning (e.g., a score on a personality test on an arbitrary scale), results from multiple studies are being combined, some or all of the studies use different scales, or it is desired to convey the size of an effect relative to the variability in the population. In meta-analyses, standardized effect sizes are used as a common measure that can be calculated for different studies and then combined into an overall summary.
Interpretation The interpretation of an effect size of being small, medium, or large depends on its substantive context and its operational definition. Jacob Cohen suggested interpretation guidelines that are near ubiquitous across many fields. However, Cohen also cautioned:
"The terms 'small,' 'medium,' and 'large' are relative, not only to each other, but to the area of behavioral science or even more particularly to the specific content and research method being employed in any given investigation... In the face of this relativity, there is a certain risk inherent in offering conventional operational definitions for these terms for use in power analysis in as diverse a field of inquiry as behavioral science. This risk is nevertheless accepted in the belief that more is to be gained than lost by supplying a common conventional frame of reference which is recommended for use only when no better basis for estimating the ES index is available." (p. 25) Sawilowsky recommended that the rules of thumb for effect sizes should be revised, and expanded the descriptions to include very small, very large, and huge. Funder and Ozer suggested that effect sizes should be interpreted based on benchmarks and consequences of findings, resulting in adjustment of guideline recommendations. Lenth noted for a medium effect size, "you'll choose the same n regardless of the accuracy or reliability of your instrument, or the narrowness or diversity of your subjects. Clearly, important considerations are being ignored here. Researchers should interpret the substantive significance of their results by grounding them in a meaningful context or by quantifying their contribution to knowledge, and Cohen's effect size descriptions can be helpful as a starting point." Similarly, a U.S. Dept of Education sponsored report argued that the widespread indiscriminate use of Cohen's interpretation guidelines can be inappropriate and misleading. They instead suggested that norms should be based on distributions of effect sizes from comparable studies. Thus a small effect (in absolute numbers) could be considered large if the effect is larger than similar studies in the field. See Abelson's paradox and Sawilowsky's paradox for related points. The table below contains descriptors for various magnitudes of d, r, f and omega, as initially suggested by Jacob Cohen, and later expanded by Sawilowsky, and by Funder & Ozer.
Types About 50 to 100 different measures of effect size are known. Many effect sizes of different types can be converted to other types, as many estimate the separation of two distributions, so are mathematically related. For example, a correlation coefficient can be converted to a Cohen's d and vice versa.
Correlation family: Effect sizes based on "variance explained" These effect sizes estimate the amount of the variance within an experiment that is "explained" or "accounted for" by the experiment's model (Explained variation).
Pearson r or correlation coefficient Pearson's correlation, often denoted r and introduced by Karl Pearson, is widely used as an effect size when paired quantitative data are available; for instance if one were studying the relationship between birth weight and longevity. The correlation coefficient can also be used when the data are binary. Pearson's r can vary in magnitude from −1 to 1, with −1 indicating a perfect negative linear relation, 1 indicating a perfect positive linear relation, and 0 indicating no linear relation between two variables.
Coefficient of determination (r2 or R2) A related effect size is r2, the coefficient of determination (also referred to as R2 or "r-squared"), calculated as the square of the Pearson correlation r. In the case of paired data, this is a measure of the proportion of variance shared by the two variables, and varies from 0 to 1. For example, with an r of 0.21 the coefficient of determination is 0.0441, meaning that 4.4% of the variance of either variable is shared with the other variable. The r2 is always positive, so does not convey the direction of the correlation between the two variables.
Eta-squared (η2) Eta-squared describes the ratio of variance explained in the dependent variable by a predictor while controlling for other predictors, making it analogous to the r2. Eta-squared is a biased estimator of the variance explained by the model in the population (it estimates only the effect size in the sample). This estimate shares the weakness with r2 that each additional variable will automatically increase the value of η2. In addition, it measures the variance explained of the sample, not the population, meaning that it will always overestimate the effect size, although the bias grows smaller as the sample grows larger.
η 2 = S S Treatment S S Total . {\displaystyle \eta ^{2}={\frac {SS_{\text{Treatment}}}{SS_{\text{Total}}}}.}
Omega-squared (ω2)
A less biased estimator of the variance explained in the population is ω2
ω 2 = SS treatment − d f treatment ⋅ MS error SS total + MS error . {\displaystyle \omega ^{2}={\frac {{\text{SS}}_{\text{treatment}}-df_{\text{treatment}}\cdot {\text{MS}}_{\text{error}}}{{\text{SS}}_{\text{total}}+{\text{MS}}_{\text{error}}}}.}
This form of the formula is limited to between-subjects analysis with equal sample sizes in all cells. Since it is less biased (although not unbiased), ω2 is preferable to η2; however, it can be more inconvenient to calculate for complex analyses. A generalized form of the estimator has been published for between-subjects and within-subjects analysis, repeated measure, mixed design, and randomized block design experiments. In addition, methods to calculate partial ω2 for individual factors and combined factors in designs with up to three independent variables have been published.
Cohen's f2 Cohen's f2 is one of several effect size measures to use in the context of an F-test for ANOVA or multiple regression. Its amount of bias (overestimation of the effect size for the ANOVA) depends on the bias of its underlying measurement of variance explained (e.g., R2, η2, ω2). The f2 effect size measure for multiple regression is defined as:
f 2 = R 2 1 − R 2 . {\displaystyle f^{2}={R^{2} \over 1-R^{2}}.}
Likewise, f2 can be defined as:
f 2 = η 2 1 − η 2 {\displaystyle f^{2}={\eta ^{2} \over 1-\eta ^{2}}} or f 2 = ω 2 1 − ω 2 {\displaystyle f^{2}={\omega ^{2} \over 1-\omega ^{2}}}
for models described by those effect size measures. The f 2 {\displaystyle f^{2}} effect size measure for sequential multiple regression and also common for PLS modeling is defined as:
f 2 = R A B 2 − R A 2 1 − R A B 2 {\displaystyle f^{2}={R_{AB}^{2}-R_{A}^{2} \over 1-R_{AB}^{2}}}
where R2A is the variance accounted for by a set of one or more independent variables A, and R2AB is the combined variance accounted for by A and another set of one or more independent variables of interest B. By convention, f2 effect sizes of 0.1 2 {\displaystyle 0.1^{2}} , 0.25 2 {\displaystyle 0.25^{2}} , and 0.4 2 {\displaystyle 0.4^{2}} are termed small, medium, and large, respectively. Cohen's f ^ {\displaystyle {\hat {f}}} can also be found for factorial analysis of variance (ANOVA) working backwards, using:
f ^ effect = ( F effect d f effect / N ) . {\displaystyle {\hat {f}}_{\text{effect}}={\sqrt {(F_{\text{effect}}df_{\text{effect}}/N)}}.}
In a balanced design (equivalent sample sizes across groups) of ANOVA, the corresponding population parameter of f 2 {\displaystyle f^{2}} is
S S ( μ 1 , μ 2 , … , μ K ) K × σ 2 , {\displaystyle {SS(\mu _{1},\mu _{2},\dots ,\mu _{K})} \over {K\times \sigma ^{2}},}
wherein μj denotes the population mean within the jth group of the total K groups, and σ the equivalent population standard deviations within each groups. SS is the sum of squares in ANOVA.
Cohen's q Another measure that is used with correlation differences is Cohen's q. This is the difference between two Fisher transformed Pearson regression coefficients. In symbols this is
q = 1 2 log 1 + r 1 1 − r 1 − 1 2 log 1 + r 2 1 − r 2 {\displaystyle q={\frac {1}{2}}\log {\frac {1+r_{1}}{1-r_{1}}}-{\frac {1}{2}}\log {\frac {1+r_{2}}{1-r_{2}}}}
where r1 and r2 are the regressions being compared. The expected value of q is zero and its variance is
var ( q ) = 1 N 1 − 3 + 1 N 2 − 3 {\displaystyle \operatorname {var} (q)={\frac {1}{N_{1}-3}}+{\frac {1}{N_{2}-3}}}
where N1 and N2 are the number of data points in the first and second regression respectively.
Difference family: Effect sizes based on differences between means The raw effect size pertaining to a comparison of two groups is inherently calculated as the differences between the two means. However, to facilitate interpretation it is common to standardise the effect size; various conventions for statistical standardisation are presented below.
Standardized mean difference
A (population) effect size θ based on means usually considers the standardized mean difference (SMD) between two populations
θ = μ 1 − μ 2 σ , {\displaystyle \theta ={\frac {\mu _{1}-\mu _{2}}{\sigma }},}
where μ1 is the mean for one population, μ2 is the mean for the other population, and σ is a standard deviation based on either or both populations. In the practical setting the population values are typically not known and must be estimated from sample statistics. The several versions of effect sizes based on means differ with respect to which statistics are used. This form for the effect size resembles the computation for a t-test statistic, with the critical difference that the t-test statistic includes a factor of n {\displaystyle {\sqrt {n}}} . This means that for a given effect size, the significance level increases with the sample size. Unlike the t-test statistic, the effect size aims to estimate a population parameter and is not affected by the sample size. SMD values of 0.2 to 0.5 are considered small, 0.5 to 0.8 are considered medium, and greater than 0.8 are considered large.
Cohen's d Cohen's d is defined as the difference between two means divided by a standard deviation for the data, i.e.
d = x ¯ 1 − x ¯ 2 s . {\displaystyle d={\frac {{\bar {x}}_{1}-{\bar {x}}_{2}}{s}}.}
Jacob Cohen defined s, the pooled standard deviation, as (for two independent samples):
s = ( n 1 − 1 ) s 1 2 + ( n 2 − 1 ) s 2 2 n 1 + n 2 − 2 {\displaystyle s={\sqrt {\frac {(n_{1}-1)s_{1}^{2}+(n_{2}-1)s_{2}^{2}}{n_{1}+n_{2}-2}}}}
where the variance for one of the groups is defined as
s 1 2 = 1 n 1 − 1 ∑ i = 1 n 1 ( x 1 , i − x ¯ 1 ) 2 , {\displaystyle s_{1}^{2}={\frac {1}{n_{1}-1}}\sum _{i=1}^{n_{1}}(x_{1,i}-{\bar {x}}_{1})^{2},}
and similarly for the other group. Other authors choose a slightly different computation of the standard deviation when referring to "Cohen's d" where the denominator is without "-2"
s = ( n 1 − 1 ) s 1 2 + ( n 2 − 1 ) s 2 2 n 1 + n 2 {\displaystyle s={\sqrt {\frac {(n_{1}-1)s_{1}^{2}+(n_{2}-1)s_{2}^{2}}{n_{1}+n_{2}}}}}
This definition of "Cohen's d" is termed the maximum likelihood estimator by Hedges and Olkin, and it is related to Hedges' g by a scaling factor (see below). With two paired samples, an approach is to look at the distribution of the difference scores. In that case, s is the standard deviation of this distribution of difference scores (of note, the standard deviation of difference scores is dependent on the correlation between paired samples). This creates the following relationship between the t-statistic to test for a difference in the means of the two paired groups and Cohen's d' (computed with difference scores):
t = X ¯ 1 − X ¯ 2 SE d i f f = X ¯ 1 − X ¯ 2 SD d i f f N = N ( X ¯ 1 − X ¯ 2 ) S D d i f f {\displaystyle t={\frac {{\bar {X}}_{1}-{\bar {X}}_{2}}{{\text{SE}}_{diff}}}={\frac {{\bar {X}}_{1}-{\bar {X}}_{2}}{\frac {{\text{SD}}_{diff}}{\sqrt {N}}}}={\frac {{\sqrt {N}}({\bar {X}}_{1}-{\bar {X}}_{2})}{SD_{diff}}}}
and
d ′ = X ¯ 1 − X ¯ 2 SD d i f f = t N {\displaystyle d'={\frac {{\bar {X}}_{1}-{\bar {X}}_{2}}{{\text{SD}}_{diff}}}={\frac {t}{\sqrt {N}}}} However, for paired samples, Cohen states that d' does not provide the correct estimate to obtain the power of the test for d, and that before looking the values up in the tables provided for d, it should be corrected for r as in the following formula:
d ′ 1 − r . {\displaystyle {\frac {d'}{\sqrt {1-r}}}.} where r is the correlation between paired measurements. Given the same sample size, the higher r, the higher the power for a test of paired difference. Since d' depends on r, as a measure of effect size it is difficult to interpret; therefore, in the context of paired analyses, since it is possible to compute d' or d (estimated with a pooled standard deviation or that of a group or time-point), it is necessary to explicitly indicate which one is being reported. As a measure of effect size, d (estimated with a pooled standard deviation or that of a group or time-point) is more appropriate, for instance in meta-analysis. Cohen's d is frequently used in estimating sample sizes for statistical testing. A lower Cohen's d indicates the necessity of larger sample sizes, and vice versa, as can subsequently be determined together with the additional parameters of desired significance level and statistical power.
Glass' Δ In 1976, Gene V. Glass proposed an estimator of the effect size that uses only the standard deviation of the second group
Δ = x ¯ 1 − x ¯ 2 s 2 {\displaystyle \Delta ={\frac {{\bar {x}}_{1}-{\bar {x}}_{2}}{s_{2}}}}
The second group may be regarded as a control group, and Glass argued that if several treatments were compared to the control group it would be better to use just the standard deviation computed from the control group, so that effect sizes would not differ under equal means and different variances. Under a correct assumption of equal population variances a pooled estimate for σ is more precise.
Hedges' g Hedges' g, suggested by Larry Hedges in 1981, is like the other measures based on a standardized difference
g = x ¯ 1 − x ¯ 2 s ∗ {\displaystyle g={\frac {{\bar {x}}_{1}-{\bar {x}}_{2}}{s^{*}}}}
where the pooled standard deviation s ∗ {\displaystyle s^{*}} is computed as:
s ∗ = ( n 1 − 1 ) s 1 2 + ( n 2 − 1 ) s 2 2 n 1 + n 2 − 2 . {\displaystyle s^{*}={\sqrt {\frac {(n_{1}-1)s_{1}^{2}+(n_{2}-1)s_{2}^{2}}{n_{1}+n_{2}-2}}}.}
However, as an estimator for the population effect size θ it is biased. Nevertheless, this bias can be approximately corrected through multiplication by a factor
g ∗ = J ( n 1 + n 2 − 2 ) g ≈ ( 1 − 3 4 ( n 1 + n 2 ) − 9 ) g {\displaystyle g^{*}=J(n_{1}+n_{2}-2)\,\,g\,\approx \,\left(1-{\frac {3}{4(n_{1}+n_{2})-9}}\right)\,\,g}
Hedges and Olkin refer to this less-biased estimator g ∗ {\displaystyle g^{*}} as d, but it is not the same as Cohen's d. The exact form for the correction factor J() involves the gamma function
J ( a ) = Γ ( a / 2 ) a / 2 Γ ( ( a −
