Preply — Study more efficiently by working with a personal tutor. Get 50% off.Affiliate

Wikipedia

Design effect

In survey research, the design effect is a number that shows how well a sample of people may represent a larger group of people for a specific measure of interest (such as the mean). This is important when the sample comes from a sampling method that differs from simply picking people using a simple random sample. The design effect is a positive real number, represented by the symbol Deff {\displaystyle {\text{Deff}}} . If Deff = 1 {\displaystyle {\text{Deff}}=1} , then the sample was selected in a way that is just as good as if people were picked randomly. When Deff > 1 {\displaystyle {\text{Deff}}>1} , then inference from the data collected is not as accurate as it could have been if people were picked randomly. When researchers use complicated methods to pick their sample, they use the design effect to check and adjust their results. It may also be used when planning a study in order to determine the sample size.

Introduction In survey methodology, the design effect (generally denoted as Deff {\displaystyle {\text{Deff}}} , D eff {\displaystyle D_{\text{eff}}} , or D eft 2 {\displaystyle D_{\text{eft}}^{2}} ) is a measure of the expected impact of a sampling design on the variance of an estimator for some parameter of a population. It is calculated as the ratio of the variance of an estimator based on a sample from an (often) complex sampling design, to the variance of an alternative estimator based on a simple random sample (SRS) of the same number of elements. The Deff {\displaystyle {\text{Deff}}} (be it estimated, or known a priori) can be used to evaluate the variance of an estimator in cases where the sample is not drawn using simple random sampling. It may also be useful in sample size calculations and for quantifying the representativeness of samples collected with various sampling designs. The design effect is a positive real number that indicates an inflation ( Deff > 1 {\displaystyle {\text{Deff}}>1} ), or deflation ( Deff < 1 {\displaystyle {\text{Deff}}<1} ) in the variance of an estimator for some parameter, that is due to the study not using SRS (with Deff = 1 {\displaystyle {\text{Deff}}=1} , when the variances are identical). Intuitively we can get Deff < 1 {\displaystyle {\text{Deff}}<1} when we have some a-priori knowledge we can exploit during the sampling process (which is somewhat rare). And, in contrast, we often get Deff > 1 {\displaystyle {\text{Deff}}>1} when we need to compensate for some limitation in our ability to collect data (which is more common). Some sampling designs that could introduce Deff {\displaystyle {\text{Deff}}} generally greater than 1 include: cluster sampling (such as when there is correlation between observations), stratified sampling (with disproportionate allocation to the strata sizes), cluster randomized controlled trial, disproportional (unequal probability) sample (e.g. Poisson sampling), statistical adjustments of the data for non-coverage or non-response, and many others. Stratified sampling can yield Deff {\displaystyle {\text{Deff}}} that is smaller than 1 when using Proportionate allocation to strata sizes (when these are known a-priori, and correlated to the outcome of interest) or Optimum allocation (when the variance differs between strata and is known a-priori). Many calculations (and estimators) have been proposed in the literature for how a known sampling design influences the variance of estimators of interest, either increasing or decreasing it. Generally, the design effect varies among different statistics of interests, such as the total or ratio mean. It also matters if the sampling design is correlated with the outcome of interest. For example, a possible sampling design might be such that each element in the sample may have a different probability to be selected. In such cases, the level of correlation between the probability of selection for an element and its measured outcome can have a direct influence on the subsequent design effect. Lastly, the design effect can be influenced by the distribution of the outcome itself. All of these factors should be considered when estimating and using design effect in practice.

History The term "design effect" was coined by Leslie Kish in his 1965 book "Survey Sampling." In it, Kish proposed the general definition for the design effect, as well as formulas for the design effect of cluster sampling (with intraclass correlation); and the famous design effect formula for unequal probability sampling. These are often known as "Kish's design effect", and were later combined into a single formula. In a 1995 paper, Kish mentions that a similar concept, termed "Lexis ratio", was described at the end of the 19th century. The closely related Intraclass correlation was described by Fisher in 1950, while computations of ratios of variances were already published by Kish and others from the late 1940s to the 1950s. One of the precursors to Kish's definition was work done by Cornfield in 1951. In his 1995 paper, Kish proposed that considering the design effect is necessary when averaging the same measured quantity from multiple surveys conducted over a period of time. He also suggested that the design effect should be considered when extrapolating from the error of simple statistics (e.g. the mean) to more complex ones (e.g. regression coefficients). However, when analyzing data (e.g., using survey data to fit models), Deff {\displaystyle {\text{Deff}}} values are less useful nowadays due to the availability of specialized software for analyzing survey data. Prior to the development of software that computes standard errors for many types of designs and estimates, analysts would adjust standard errors produced by software that assumed all records in a dataset were i.i.d by multiplying them by a Deft {\displaystyle {\text{Deft}}} (see Deft definition below).

Definitions

Notations

Deff The design effect, commonly denoted by Deff {\displaystyle {\text{Deff}}} (or D eff {\displaystyle D_{\text{eff}}} , sometimes with additional subscripts), is the ratio of two theoretical variances for estimators of some parameter ( θ {\displaystyle \theta } ):

The numerator represents the actual variance for an estimator of a parameter ( θ ^ w {\displaystyle {\hat {\theta }}_{w}} ) under a given sampling design p {\displaystyle p} ; The denominator represents the variance assuming the same sample size, but if the sample were obtained using the estimator for simple random sampling without replacement ( θ ^ S R S W O R {\displaystyle {\hat {\theta }}_{SRSWOR}} ). So that:

Deff p ( θ ^ ) = v a r ( θ ^ w ) v a r ( θ ^ S R S W O R ) {\displaystyle {\text{Deff}}_{p}({\hat {\theta }})={\frac {var({\hat {\theta }}_{w})}{var({\hat {\theta }}_{SRSWOR})}}}

In other words, Deff {\displaystyle {\text{Deff}}} measures the extent to which the variance has increased (or, in some cases, decreased) because the sample was drawn and adjusted to a specific sampling design (e.g., using weights or other measures) compared to if the sample was from a simple random sample (without replacement). Notice how the definition of Deff {\displaystyle {\text{Deff}}} is based on parameters of the population that are often unknown, and that are hard to estimate directly. Specifically, the definition involves the variances of estimators under two different sampling designs, even though only a single sampling design is used in practice. For example, when estimating the population mean, the Deff {\displaystyle {\text{Deff}}} (for some sampling design p) is:

Deff p = v a r p ( y ¯ p ) ( 1 − f ) S y 2 / n {\displaystyle {\text{Deff}}_{p}={\frac {var_{p}({\bar {y}}_{p})}{(1-f)S_{y}^{2}/n}}}

Where n {\displaystyle n} is the sample size, f = n / N {\displaystyle f=n/N} is the fraction of the sample from the population, ( 1 − f ) {\displaystyle (1-f)} is the (squared) finite population correction (FPC), S y 2 {\displaystyle S_{y}^{2}} is the unbiassed sample variance, and v a r p ( y ¯ p ) {\displaystyle var_{p}({\bar {y}}_{p})} is some estimator of the variance of the mean under the sampling design. The issue with the above formula is that it is extremely rare to be able to directly estimate the variance of the estimated mean under two different sampling designs, since most studies rely on only a single sampling design. There are many ways of calculation Deff {\displaystyle {\text{Deff}}} , depending on the parameter of interest (e.g. population total, population mean, quantiles, ratio of quantities etc.), the estimator used, and the sampling design (e.g. clustered sampling, stratified sampling, post-stratification, multi-stage sampling, etc.). The process of estimating Deff {\displaystyle {\text{Deff}}} for specific designs will be described in the following section.

Deft A related quantity to Deff {\displaystyle {\text{Deff}}} , proposed by Kish in 1995, is the Design Effect Factor, abbreviated as Deft {\displaystyle {\text{Deft}}} (or also D eft {\displaystyle D_{\text{eft}}} ). It is defined as the square root of the variance ratios while also having the denominator use a simple random sample with replacement (SRSWR), instead of without replacement (SRSWOR):

Deft = var ( θ ^ w ) var ( θ ^ S R S W R ) {\displaystyle {\text{Deft}}={\sqrt {\frac {{\text{var}}({\hat {\theta }}_{w})}{{\text{var}}({\hat {\theta }}_{SRSWR})}}}}

In this later definition (proposed in 1995, vs 1965) Kish argued in favor of using Deft 2 {\displaystyle {\text{Deft}}^{2}} over Deff {\displaystyle {\text{Deff}}} for several reasons. It was argued that SRS "without replacement" (with its positive effect on the variance) should be captured in the denominator part in the definition of the design effect, since it is part of the sampling design. Also, since often the use of the factor is in confidence intervals), it was claimed that using Deft {\displaystyle {\text{Deft}}} will be simpler than writing Deff {\displaystyle {\sqrt {\text{Deff}}}} . It is also said that for many cases when the population is very large, Deft {\displaystyle {\text{Deft}}} is (almost) the square root of Deff {\displaystyle {\text{Deff}}} ( Deft ≈ Deff {\displaystyle {\text{Deft}}\approx {\sqrt {\text{Deff}}}} ), hence it is easier to use than exactly calculating the finite population correction (FPC). Even so, in various cases a researcher might approximate the Deft {\displaystyle {\text{Deft}}} by calculating the variance in the numerator while assuming SRS with replacement (SRSWR) instead of SRS without replacement (SRSWOR), even if it is not precise. For example, consider a multistage design with primary sampling units (PSUs) selected systematically with probability proportional to some measure of size from a list sorted in a particular way (say, by number of households in each PSU). Also, let it be combined with an estimator that uses raking to match the totals for several demographic variables. In such a design, the joint selection probabilities for the PSUs, which are needed for a without replacement variance estimator, are 0 for some pairs of PSUs - implying that an exact design-based (i.e., repeated sampling) variance estimator does not exist. Another example is when a public use file issued by some government agency is used for analysis. In such a case the information on joint selection probabilities of first-stage units is almost never released. As a result, an analyst cannot estimate a with replacement variance for the numerator even if desired. The standard workaround is to compute a variance estimator as if the PSUs were selected with replacement. This is the default choice in software packages such as Stata, the R survey package, and the SAS survey procedures.

Effective sample size The effective sample size, defined by Kish in 1965, is calculated by dividing the original sample size by the design effect. Namely:

n eff = n Deff {\displaystyle n_{\text{eff}}={\frac {n}{\text{Deff}}}}

This quantity reflects what would be the sample size that is needed to achieve the current variance of the estimator (for some parameter) with the existing design, if the sample design (and its relevant parameter estimator) were based on a simple random sample. A related quantity is the effective sample size ratio (ESSR), which can be calculated by simply taking the inverse of Deff {\displaystyle {\text{Deff}}} (i.e., n eff n = 1 Deff {\displaystyle {\frac {n_{\text{eff}}}{n}}={\frac {1}{\text{Deff}}}} ). For example, let the design effect, for estimating the population mean based on some sampling design, be 2. If the sample size is 1,000, then the effective sample size will be 500. It means that the variance of the weighted mean based on 1,000 samples will be the same as that of a simple mean based on 500 samples obtained using a simple random sample.

The design effect for well-known sampling designs

The design effect depends on sampling design and statistical adjustments Different sampling designs and statistical adjustments may have substantially different impact on the bias and variance of estimators (such as the mean). An example of a design which can lead to estimation efficiency, compared to simple random sampling, is Stratified sampling. This efficiency is gained by leveraging information about the composition of the population. For example, if it is known that gender is correlated with the outcome of interest, and also that the male-female ratio for some population is (say) 50%-50%, then sampling exactly half of the sample from each gender will reduce the variance of the outcome's estimator. Similarly, if a particular sub-population is of special interest, deliberately over-sampling from that sub-population will decrease the variance for estimations made about it. Improvement in variance efficiency might sometimes be sacrificed for convenience or cost. For example, in the cluster sampling case the units may have equal or unequal selection probabilities, irrespective of their intra-class correlation (and their negative effect of increasing the variance of the estimators). We might decide (for practical reasons) to collect responses from only 2 people of each household (i.e., a sampled cluster), which could lead to more complex post-sampling adjustment to deal with unequal selection probabilities. Also, such decisions could lead to less efficient estimators than just taking a fixed proportion of responses from a cluster. When the sampling design isn’t set in advance and needs to be figured out from the data we have, this can lead to an increase of both the variance and bias of the weighted estimator. This might happen when making adjustments for issues like non-coverage, non-response, or an unexpected strata split of the population that wasn’t available during the initial sampling stage. In these cases, we might use statistical procedures such as post-stratification, raking, or inverse propensity score weighting (where the propensity scores are estimated), among other methods. Using these methods requires assumptions about the initial design model. For example, when we use post-stratification based on age and gender, it is assumed that these variables can explain a significant portion of the bias in the sample. The quality of these estimators is closely tied to the quality of the additional information and the missing at random assumptions used when making them. Either way, even when estimators (like propensity score models) do a good job capturing most of the sampling design, using the weights can make a small or a large difference, depending on the specific data-set. Due to the large variety in sampling designs (with or without an effect on unequal selection probabilities), different formulas have been developed to capture the potential design effect, as well as to estimate the variance of estimators when accounting for the sampling designs. Sometimes, these different design effects can be compounded together (as in the case of unequal selection probability and cluster sampling, more details in the following sections). Whether or not to use these formulas, or just assume SRS, depends on the expected amount of bias reduction vs. the increase in estimator variance (and in the overhead of methodological and technical complexity).

Unequal selection probabilities

Sources of unequal selection probabilities

There are various ways to sample units so that each unit would have the exact same probability of selection. Such methods are called equal probability sampling (EPSEM) methods. Some of the more basic methods include simple random sampling (SRS, with or without replacement) and systematic sampling for getting a fixed sample size. There is also Bernoulli sampling with a random sample size. More advanced techniques such as stratified sampling and cluster sampling can also be designed to be EPSEM. For example, in cluster sampling we can use a two stage sampling in which we sample each cluster (which may be of different sizes) with equal probability, and then sample from each cluster at the second stage using SRS with a fixed proportion (e.g. sample half of the cluster, the whole cluster, etc.). This method will yield EPSEM, but the specific number of elements we end up with is stochastic (i.e., non deterministic). Another strategy for cluster sampling that leads to EPSEM is to sample clusters in a way that is proportional to their sizes, and then sample a fixed number of elements inside each cluster. In their works, Kish and others highlight several known reasons that lead to unequal selection probabilities:

Disproportional sampling due to selection frame or procedure. This happens when a researcher deliberately over- or under-samples specific sub-populations or clusters. For example: In stratified sampling when units from some strata are known to have a larger variance than other strata. In such cases, the intention of the researcher may be to use this prior knowledge about the variance between strata in order to reduce the overall variance of an estimator of some population level parameter of interest (e.g., the mean). This can be achieved by a strategy known as optimum allocation, in which a stratum h {\displaystyle h} is over sampled proportional to higher standard deviation and lower sampling cost (i.e., f h ∝ S h C h {\displaystyle f_{h}\propto {\frac {S_{h}}{\sqrt {C_{h}}}}} , where S h {\displaystyle S_{h}} is the standard deviation of the outcome in h {\displaystyle h} , and C h {\displaystyle C_{h}} relates to the cost of recruiting one element from h {\displaystyle h} ). An example of an optimum allocation is Neyman's optimal allocation which, when cost is fixed for recruiting people from each stratum, the sample size is: n h = n W h S U h ∑ h W h S U h {\displaystyle n_{h}=n{\frac {W_{h}S_{Uh}}{\sum _{h}W_{h}S_{Uh}}}} . Where the summation is over all strata: n is the total sample size; n h {\displaystyle n_{h}} is the sample size for stratum h; W h = N h N {\displaystyle W_{h}={\frac {N_{h}}{N}}} is the relative size of stratum h as compared to the entire population N; and S U h {\displaystyle S_{Uh}} is the standard error in stratum h. A related concept to optimum design is optimal experimental design. If there is interest in comparing two strata (e.g., people from two specific socio-demographic groups, or from two regions, etc.), in which case the smaller group may be over-sampled. This way, the variance of the estimator that compares the two groups is reduced. In cluster sampling there may be clusters of different sizes but the procedure samples from all clusters using SRS, and all elements in the cluster are measured (for example, if the cluster sizes are not known upfront at the stage of sampling). In some two-stage cluster sampling based cluster sizes. For example, when in the first stage the clusters are sampled proportionally to the estimation of their size (a.k.a.: PPS Probability Proportional to Size) and at the second stage a fixed proportion of elements are chosen (e.g., half, or all the elements in the cluster) - then the selection probabilities are different for elements from different clusters. A similar case is when the first stage attempts to sample the clusters using PPS, the second stage uses a fixed number of elements in each cluster - but the cluster sizes used for the first stage sampling were inaccurate (so that some smaller cluster may have a higher-than-it-should chance of being selected. And vice versa for larger clusters with too-small a chance of being sampled). In such cases, the larger the errors in the sampling probabilities used in the first stage, the larger the unequal selection probabilities for each element will be. When the frame used for sampling includes duplication of some of the items, thus leading some items to have a larger probability than others to be sampled (e.g., if the sampling frame was created by merging several lists. Or if recruiting users from several ad channels in which some of the users are available for recruitment from several of the channels, while others are available to be recruited from only one of the channels) so that different units would have different sampling probabilities, thus making this sampling procedure to not be EPSEM. When several different samples/frames are to be combined. For example, if running different ad campaigns for recruiting respondents. Or when combining results from several studies done by different researchers and/or at different times (i.e., Meta-analysis). When disproportional sampling happens, due to sampling design decisions, the researcher may (sometimes) be able to trace back the decision and accurately calculate the exact inclusion probability. When these selection probabilities are hard to trace back, they may be estimated using some propensity score model combined with information from auxiliary variables (e.g., age, gender, etc.). Non-coverage. This happens, for example, if people are sampled based on some pre-defined list that doesn't include all the people in the population (e.g., a phone book or using ads to recruit people to a survey). These missing units are missing due to some failure of creating the sampling frame, as opposed to deliberate exclusion of some people (e.g. minors, people who cannot vote, etc.). The effect of non-coverage on sampling probability is considered difficult to measure (and adjust for) in various survey situations, unless strong assumptions are made. Adjustments for non-coverage can lead to inadequate weights when the relevant covariates are not used for adjustment. If there are covariates that can be used to correct for non-coverage, they are expected to lead to unequal survey weights. Non-response. This refers to the failure of obtaining measurements on sampled units that are intended to be measured. Reasons for non-response are varied and depend on the context. A person may be temporarily unavailable, for example if they are not available to answer the phone when a telephone survey is done. A person may also refuse to answer the survey due to a variety of reasons, e.g. different tendencies of people from different ethnic/demographic/socio-economic groups to respond in general; insufficient incentive to spend the time or share data; the identity of the institution that is running the survey; inability to respond (e.g. due to illness, illiteracy, or a language barrier); respondent is not found (e.g. they moved); the response was lost/destroyed during encoding or transmission (i.e., measurement error). In the context of surveys, these reasons may be related to answering the entire survey or just specific questions. Statistical adjustments. These may include methods such as post-stratification, raking, or propensity score (estimation) models - used to perform an adjustment of the sample to some known (or estimated) strata sizes. These adjustments can be in addition of design weights, which aims to account for imbalances due to some known sampling design. Such procedures are used to mitigate issues in the sampling ranging from sampling error, under-coverage of the sampling frame to non-response. For example, these methods can be used to make the sample more similar to some target "controls" (i.e., population of interest), a process also called "standardization". In such cases, these adjustments help with providing unbiased estimators (often with the cost of increased variance, as seen in the following sections). If the original sample is a nonprobability sample, then post-stratification adjustments are just similar to quota sampling. Note that if a simple random sample is used, a post-stratification (using some auxiliary information) does not offer an estimator that is uniformly better than just an unweighted estimator. However, it can be viewed as a more "robust" estimator. Alternatively, when the sampling design is fully known (leading to some p h {\displaystyle p_{h}} probability of selection for some element from stratum h), and the non-response is measurable (i.e., we know that only r h {\displaystyle r_{h}} observations answered in stratum h), then an exactly known inverse probability weight can be calculated for each element i from stratum h using: w i = 1 p h r h {\displaystyle w_{i}={\frac {1}{p_{h}r_{h}}}} . Sometimes a statistical adjustment, such as post-stratification or raking, is used for estimating the selection probability. E.g., when comparing the sample we have with same target population, also known as matching to controls. The estimation process may be focused only on adjusting the existing population to an alternative population (for example, if trying to extrapolate from a panel drawn from several regions to an entire country). In such a case, the adjustment might be focused on some calibration factor c i {\displaystyle c_{i}} and the weights be calculated as w i = c i p h r h {\displaystyle w_{i}={\frac {c_{i}}{p_{h}r_{h}}}} . However, in other cases, both the under-coverage and non-response are all modeled as part of the statistical adjustment, which leads to an estimation of the overall sampling probability (lets say p i ′ {\displaystyle p_{i}'} ). In such a case, the weights are simply: w i = 1 p i ′ {\displaystyle w_{i}={\frac {1}{p_{i}'}}} . Notice that when statistical adjustments are used, w i {\displaystyle w_{i}} is often estimated based on some model. The formulation in the following sections assume this w i {\displaystyle w_{i}} is known, which is not true for statistical adjustments (since we only have w ^ i {\displaystyle {\widehat {w}}_{i}} ). However, if it is assumed that the estimation error of w ^ i {\displaystyle {\widehat {w}}_{i}} is very small then the following sections can be used as if it was known. Having this assumption be true depends on the size of the sample used for modeling, and is worth keeping in mind during analysis. When the selection probabilities may be different, the sample size is random, and the pairwise selection probabilities are independent, we call this Poisson sampling.

"Design based" vs "model based" for describing properties of estimators Adjusting for unequal probability selection through "individual case weights" (e.g. inverse probability weighting), yields various types of estimators for quantities of interest. Estimators such as Horvitz–Thompson estimator yield unbiased estimators (if the selection probabilities are indeed known, or approximately known), for total and the mean of the population. Deville and Särndal (1992) coined the term "calibration estimator" for estimators using weights such that they satisfy some condition, such as having the sum of weights equal the population size. And more generally, that the weighted sum of weights is equal some quantity of an auxiliary variable: ∑ w i x i = X {\displaystyle \sum w_{i}x_{i}=X} (e.g., that the sum of weighted ages of the respondents is equal to the population size in each age group). The two primary ways to argue about the properties of calibration estimators are:

randomization based (or, sampling design based) - in this case, the weights ( w i {\displaystyle w_{i}} ) and values of the outcome of interest y i {\displaystyle y_{i}} that are measured in the sample are all treated as known. In this framework, there is variability in the (known) values of the outcome (Y). However, the only randomness comes from which of the elements in the population were picked into the sample (often denoted as I i {\display

Tags

  • Design of experiments
  • Externally peer reviewed articles
  • Medical statistics
  • Wikipedia articles published in WikiJournal of Science
  • Wikipedia articles published in peer-reviewed literature
  • Wikipedia articles published in peer-reviewed literature (W2J)