In epidemiology, Mendelian randomization (commonly abbreviated to MR) is a method using measured variation in genes to examine the causal effect of an exposure on an outcome. Under key assumptions (see below), the design reduces both reverse causation and confounding, which often substantially impede or mislead the interpretation of results from epidemiological studies.
The study design was first proposed in 1986 and subsequently described by Gray and Wheatley as a method for obtaining unbiased estimates of the effects of an assumed causal variable without conducting a traditional randomized controlled trial (the standard in epidemiology for establishing causality). These authors also coined the term Mendelian randomization.
Motivation One of the predominant aims of epidemiology is to identify modifiable causes of health outcomes and disease, especially those of public health concern. To ascertain whether modifying a particular trait (e.g. via an intervention, treatment or policy change) will convey a beneficial effect within a population, firm evidence that this trait causes the outcome of interest is required. However, many observational epidemiological study designs are limited in their ability to discern correlation from causation – specifically to distinguish whether a particular trait causes an outcome of interest, is simply related to that outcome (but does not cause it) or is a consequence of the disease processes leading up to the outcome, or of the outcome itself. Only the former will be beneficial within a public health setting where the aim is to modify that trait to reduce the burden of disease. Many epidemiological study designs aim to understand relationships between traits within a population sample, each with shared and unique advantages and limitations in terms of providing causal evidence, with the "gold standard" often being considered to be randomized controlled trials. Well-known successful demonstrations of causal evidence consistent across multiple studies with different designs include the identified causal links between smoking and lung cancer, and between blood pressure and stroke. However, there have also been notable failures when exposures hypothesized to be a causal risk factor for a particular outcome were later shown by well-conducted randomized controlled trials not to be causal. For instance, hormone replacement therapy was thought to prevent cardiovascular disease, but it is now known to have no such benefit. Another notable example is that of selenium and prostate cancer. Some observational studies found an association between higher circulating selenium levels (usually acquired through various foods and dietary supplements ) and lower risk of prostate cancer. However, the Selenium and Vitamin E Cancer Prevention Trial (SELECT) showed evidence that dietary selenium supplementation actually increased the risk of prostate and advanced prostate cancer and had an additional off-target effect on increasing type 2 diabetes risk. Mendelian randomization methods now support the view that high selenium status may not prevent cancer in the general population, and may even increase the risk of specific types. Such inconsistencies between observational epidemiological studies and randomized controlled trials are likely a function of social, behavioral or physiological confounding factors in many observational epidemiological designs, which are particularly difficult to measure accurately and difficult to control for. Moreover, randomized controlled trials (RCTs) are usually expensive, time-consuming, and laborious and many epidemiological findings cannot be ethically replicated in clinical trials. In some settings, Mendelian randomization studies appear capable of resolving questions of potential confounding more efficiently than RCTs
Definition Mendelian randomization (MR) uses the properties of germline genetic variation (usually in the form of single nucleotide polymorphisms or SNPs) strongly associated with a potential exposure, if those genetic variants are associated with the outcome then this adds strength to the conclusion that the exposure does have a causal effect on the outcome. The method is most commonly implemented using the instrumental variables estimation method hailing from econometrics. The genetic variants are then used as a "proxy" for that exposure to test for and estimate a causal effect of the exposure on an outcome of interest. The genetic variation used will have either well-understood effects on exposure patterns (e.g. propensity to smoke heavily) or effects that mimic those produced by modifiable exposures (e.g., raised blood cholesterol). Importantly, the genotype must only affect the disease status indirectly via its effect on the exposure of interest.
As genotypes are assigned randomly when passed from parents to offspring during meiosis, then groups of individuals defined by genetic variation associated with an exposure at a population level should be largely unrelated to the confounding factors that typically plague observational epidemiology studies. Given an individual's parents genotype, the genotype they inherit is truly random and so the method was initially proposed as being applied to data which included parents and their offspring. However, the number of datasets which include family data are limited and so Mendelian randomization is usually applied to data on unrelated individuals from a population. However, increasing availability of data is increasing the use of family based methods. Germline genetic variation (i.e. that which can be inherited) is fixed at conception and not modified by the onset of any outcome or disease, precluding reverse causation. Additionally, given improvements in modern genotyping technologies, measurement error and systematic misclassification is often low with genetic data. In this regard Mendelian randomization can be thought of as analogous to "nature's randomized controlled trial". Mendelian randomization requires three core instrumental variable assumptions. Namely that:
… excerpt ends here. Continue reading the full article.



