Expert Judgment (EJ) denotes a wide variety of techniques ranging from a single undocumented opinion, through preference surveys, to formal elicitation with external validation of expert probability assessments. Recent books are
. In the nuclear safety area, Rasmussen formalized EJ by documenting all steps in the expert elicitation process for scientific review. This made visible wide spreads in expert assessments and teed up questions regarding the validation and synthesis of expert judgments. The nuclear safety community later took onboard expert judgment techniques underpinned by external validation . Empirical validation is the hallmark of science, and forms the centerpiece of the classical model of probabilistic forecasting . A European Network coordinates workshops. Application areas include nuclear safety, investment banking, volcanology, public health, ecology, engineering, climate change and aeronautics/aerospace. For a survey of applications through 2006 see and give exhortatory overviews. A recent large scale implementation by the World Health Organization is described in . A long running application at the Montserrat Volcano Observatory is described in . The classical model scores expert performance in terms of statistical accuracy (sometimes called calibration) and informativeness . These terms should not be confused with "accuracy and precision". Accuracy "is a description of systematic errors" while precision "is a description of random errors". In the classical model statistical accuracy is measured as the p-value or probability with which one would falsely reject the hypotheses that an expert's probability assessments were statistically accurate. A low value (near zero) means it is very unlikely that the discrepancy between an expert's probability statements and observed outcomes should arise by chance. Informativeness is measured as Shannon relative information (or Kullback Leibler divergence) with respect to an analyst-supplied background measure. Shannon relative information is used because it is scale invariant, tail insensitive, slow, and familiar. Parenthetically, measures with physical dimensions, such as the standard deviation, or the width of prediction intervals, raise serious problems, as a change of units (meters to kilometers) would affect some variables but not others. The product of statistical accuracy and informativeness for each expert is their combined score. With an optimal choice of a statistical accuracy threshold beneath which experts are unweighted, the combined score is a long run "strictly proper scoring rule": an expert achieves his long run maximal expected score by and only by stating his true beliefs. The classical model derives Performance Weighted (PW) combinations. These are compared with Equally Weighted (EW) combinations, and recently with Harmonically Weighted (HW) combinations, as well as with individual expert assessments. While some mathematicians and decision analysts regard combining expert judgments as a mathematical problem, the classical model regards expert combination as more akin to an engineering problem. A bicycle obeys Newton's Laws but does not follow from them. It is designed to optimize performance under constraints. Similarly expert judgment combination is viewed as a tool for enabling rational consensus by optimizing performance measures under mathematical and decision theoretic constraints. The theory of rational consensus is summarized in . Real expert judgment studies differ in many ways from research or academic exercises. The experts are typically recruited in a traceable peer nomination process based on their knowledge of and engagement with the subject of the study; they may receive remuneration. In all cases, experts' reasoning is documented, and their names and affiliations are part of the reporting. However, to encourage candid judgments, individuals' responses are not exchanged within the group and association of names with assessments is not reported in the open literature, but is preserved to enable peer review by the problem owner. Elicitations typically last several hours; the elicitation protocol is formalized and is part of the public reporting. Elicitation styles differ among practitioners, including face-to-face interviews, with or without plenary briefing and training, and "supervised plenary". Remote elicitation is rarely used, but some recent studies use online face-to-face tools.
Why validate? Since experts are invoked when quantities of interest are uncertain, the goal of structured expert judgment is a defensible quantification of uncertainty. Confronted with uncertainty, society at large will always harken to prophets, oracles, pundits, blue ribbon panels, crowd wisdom reputed to have performed well in the past. Scientists and engineers, in contrast, are typically averse to any methodology which eschews empirical validation. Most invocations of expert judgment do not attempt any form of validation, as if the predicate "expert" were validation enough. The classical model's emphasis on validation is its distinguishing feature. Virtually all validation data with real experts and real applications (as opposed to academic exercises) has been generated by practitioners with the classical model. One of the first studies with experienced and inexperienced experts showed that expert performance on questions from their field of expertise was not predicted by their performance on "almanac questions". Experienced and inexperienced experts performed similarly on questions outside their field, but the experienced experts were much better on questions from their field. Hence, validation must be based on assessments of uncertain quantities from the experts' field, to which we know, or will know, the true values within the time frame of the study. Such quantities are called "calibration" or "seed" variables. Finding good calibration variables is difficult, and requires a deep dive into the subject matter at hand. The quality of the calibration, and the performance on calibration variables, buttresses the credibility of the whole study. At the end of the day, the problem owner will ask: "if expert A has very good performance on the calibration variables, whereas expert B has very poor performance, am I going to ignore that difference?" If the owner's answer is "yes" then the calibration variables have failed in their purpose and the effort has been for naught.
… excerpt ends here. Continue reading the full article.






