Multilevel regression with poststratification (MRP) is a statistical technique used for correcting model estimates for known differences between a sample population (the population of the data one has), and a target population (a population one wishes to estimate for). The poststratification refers to the process of adjusting the estimates, essentially a weighted average of estimates from all possible combinations of attributes (for example age and sex). Each combination is sometimes called a "cell". The multilevel regression is the use of a multilevel model to smooth noisy estimates in the cells with too little data by using overall or nearby averages. One application is estimating preferences in sub-regions (e.g., states, individual constituencies) based on individual-level survey data gathered at other levels of aggregation (e.g., national surveys). Individual seat polls can struggle to have a high enough sample size, while MRPs have such large sample sizes that even smaller sub-demographics (eg grouping by age, or cultural background) will have a high enough sample size, which can then be used to adjust seat forecasts. Since the mid-2010s, MRP has seen rapid adoption by commercial pollsters and academic election forecasters as a means of producing seat-by-seat or district-by-district estimates from a single large national survey, particularly in the United Kingdom and United States.
Mathematical formulation Following the MRP model description, assume Y {\displaystyle Y} represents single outcome measurement and the population mean value of Y {\displaystyle Y} , μ Y {\displaystyle \mu _{Y}} , is the target parameter of interest. In the underlying population, each individual, i {\displaystyle i} , belongs to one of j = 1 , 2 , ⋯ , J {\displaystyle j=1,2,\cdots ,J} poststratification cells characterized by a unique set of covariates. The multilevel regression with poststratification model involves the following pair of steps: MRP step 1 (multilevel regression): The multilevel regression model specifies a linear predictor for the mean μ Y {\displaystyle \mu _{Y}} , or the logit transform of the mean in the case of a binary outcome, in poststratification cell j {\displaystyle j} ,
g ( μ j ) = g ( E [ Y j [ i ] ] ) = β 0 + X j T β + ∑ k = 1 K a l [ j ] k , {\displaystyle g{\left({\mathrm {\mu } }_{j}\right)}=g{\left(E{\left[Y_{j{\lbrack i\rbrack }}\right]}\right)}={\mathrm {\beta } }_{0}+{\boldsymbol {X}}_{j}^{T}\mathbf {\beta } +\sum _{k=1}^{K}a_{l{\lbrack j\rbrack }}^{k},}
where Y j [ i ] {\displaystyle Y_{j\lbrack i\rbrack }} is the outcome measurement for respondent i {\displaystyle i} in cell j {\displaystyle j} , β 0 {\displaystyle \beta _{0}} is the fixed intercept, X j {\displaystyle {\boldsymbol {X}}_{j}} is the unique covariate vector for cell j {\displaystyle j} , β {\displaystyle {\mathrm {\beta } }} is a vector of regression coefficients (fixed effects), a l [ j ] k {\displaystyle a_{l{\lbrack j\rbrack }}^{k}} is the varying coefficient (random effect), l [ j ] {\displaystyle l{\lbrack j\rbrack }} maps the j {\displaystyle j} cell index to the corresponding category index l {\displaystyle l} of variable k ∈ { 1 , 2 , ⋯ , K } {\displaystyle k\in \{1,2,\cdots ,K\}} . All varying coefficients are exchangeable batches with independent normal prior distributions
… excerpt ends here. Continue reading the full article.
