ArticleslgStudy

biology

Phylogenetic autocorrelation

Phylogenetic autocorrelation is a biology topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Phylogenetic autocorrelation rather than just read about it. In short: Phylogenetic autocorrelation, also known as Galton's problem after Sir Francis Galton who described it, is the problem of drawing inferences from cross-cultural data, due to the statistical phenomenon now called autocorrelation. The problem is now recognized as a general one that applies to all nonexperimental studies and to some experimental designs as well.

Phylogenetic autocorrelation — main illustration
Phylogenetic autocorrelation — illustration

Key takeaways

  • Phylogenetic autocorrelation belongs to biology; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Phylogenetic autocorrelation to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Phylogenetic autocorrelation from memory before moving on to harder problems.

Reference excerpt

Phylogenetic autocorrelation, also known as Galton's problem after Sir Francis Galton who described it, is the problem of drawing inferences from cross-cultural data, due to the statistical phenomenon now called autocorrelation. The problem is now recognized as a general one that applies to all nonexperimental studies and to some experimental designs as well. It is most simply described as the problem of external dependencies in making statistical estimates when the elements sampled are not statistically independent. Asking two people in the same household whether they watch TV, for example, does not give you statistically independent answers. The sample size, n, for independent observations in this case is one, not two. Once proper adjustments are made that deal with external dependencies, then the axioms of probability theory concerning statistical independence will apply. These axioms are important for deriving measures of variance, for example, or tests of statistical significance.

Origin In 1888, Galton was present when Sir Edward Tylor presented a paper at the Royal Anthropological Institute. Tylor had compiled information on institutions of marriage and descent for 350 cultures and examined the associations between these institutions and measures of societal complexity. Tylor interpreted his results as indications of a general evolutionary sequence, in which institutions change focus from the maternal line to the paternal line as societies become increasingly complex. Galton disagreed, pointing out that similarity between cultures could be due to borrowing, could be due to common descent, or could be due to evolutionary development; he maintained that without controlling for borrowing and common descent one cannot make valid inferences regarding evolutionary development. Galton's critique has become the eponymous Galton's Problem, as named by Raoul Naroll, who proposed the first statistical solutions. By the early 20th century unilineal evolutionism was abandoned and along with it the drawing of direct inferences from correlations to evolutionary sequences. Galton's criticisms proved equally valid, however, for inferring functional relations from correlations. The problem of autocorrelation remained.

Solutions Statistician William S. Gosset in 1914 developed methods of eliminating spurious correlation due to how position in time or space affects similarities. Today's election polls have a similar problem: the closer the poll to the election, the less individuals make up their mind independently, and the greater the unreliability of the polling results, especially the margin of error or confidence limits. The effective n of independent cases from their sample drops as the election nears. Statistical significance falls with lower effective sample size. The problem pops up in sample surveys when sociologists want to reduce the travel time to do their interviews, and hence they divide their population into local clusters and sample the clusters randomly, then sample again within the clusters. If they interview n people in clusters of size m the effective sample size (efs) would have a lower limit of 1 + (n − 1) / m if everyone in each cluster were identical. When there are only partial similarities within clusters, the m in this formula has to be lowered accordingly. A formula of this sort is 1 + d (n − 1) where d is the intraclass correlation for the statistic in question. In general, estimation of the appropriate efs depends on the statistic estimated, as for example, mean, chi-square, correlation, regression coefficient, and their variances. For cross-cultural studies, Murdock and White estimated the size of patches of similarities in their sample of 186 societies. The four variables they tested – language, economy, political integration, and descent – had patches of similarities that varied from size three to size ten. A very crude rule of thumb might be to divide the square root of the similarity-patch sizes into n, so that the effective sample sizes are 58 and 107 for these patches, respectively. Again, statistical significance falls with lower effective sample size. In modern analysis spatial lags have been modelled in order to estimate the degree of globalization on modern societies. Spatial dependency or auto-correlation is a fundamental concept in geography. Methods developed by geographers that measure and control for spatial autocorrelation do far more than reduce the effective n for tests of significance of a correlation. One example is the complicated hypothesis that "the presence of gambling in a society is directly proportional to the presence of a commercial money and to the presence of considerable socioeconomic differences and is inversely related to whether or not the society is a nomadic herding society."

Tests of this hypothesis in a sample of 60 societies failed to reject the null hypothesis. Autocorrelation analysis, however, showed a significant effect of socioeconomic differences. How prevalent is autocorrelation among the variables studied in cross-cultural research? A test by Anthon Eff on 1700 variables in the cumulative database for the Standard Cross-Cultural Sample, published in World Cultures, measured Moran's I for spatial autocorrelation (distance), linguistic autocorrelation (common descent), and autocorrelation in cultural complexity (mainline evolution). "The results suggest that ... it would be prudent to test for spatial and phylogenetic autoccorrelation when conducting regression analyses with the Standard Cross-Cultural Sample." The use of autocorrelation tests in exploratory data analysis is illustrated, showing how all variables in a given study can be evaluated for nonindependence of cases in terms of distance, language, and cultural complexity. The methods for estimating these autocorrelation effects are then explained and illustrated for ordinary least squares regression using again the Moran I significance measure of autocorrelation. When autocorrelation is present, it can often be removed to get unbiased estimates of regression coefficients and their variances by constructing a respecified dependent variable that is "lagged" by weightings on the dependent variable on other locations, where the weights are degree of relationship. This lagged dependent variable is endogenous, and estimation requires either two-stage least squares or maximum likelihood methods.

… excerpt ends here. Continue reading the full article.

Illustrations

Phylogenetic autocorrelation: The problem is sometimes termed after its progentior, Sir Francis Galton
The problem is sometimes termed after its progentior, Sir Francis Galton

Worked examples

Example 1 — a first encounter with Phylogenetic autocorrelation

Start with the simplest possible case. Write down what Phylogenetic autocorrelation claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In biology, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Phylogenetic autocorrelation before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Phylogenetic autocorrelation ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Phylogenetic autocorrelation

In research
Phylogenetic autocorrelation appears in biology research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Phylogenetic autocorrelation in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Phylogenetic autocorrelation is common in secondary-school and first-year university syllabi. It links to neighbouring topics Covariance and correlation, Cross-cultural studies, Regression with time series structure, so understanding it makes those chapters shorter.
In everyday life
Look for Phylogenetic autocorrelation outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “Phylogenetic autocorrelation” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Phylogenetic autocorrelation in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Phylogenetic autocorrelation means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Phylogenetic autocorrelation out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Phylogenetic autocorrelation in simple terms?

Phylogenetic autocorrelation, also known as Galton's problem after Sir Francis Galton who described it, is the problem of drawing inferences from cross-cultural data, due to the statistical phenomenon now called autocorrelation. The problem is now recognized as a general one that applies to all non…

Why does Phylogenetic autocorrelation matter?

Because it connects several biology ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Phylogenetic autocorrelation?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Phylogenetic autocorrelation.

Tags

  • Covariance and correlation
  • Cross-cultural studies
  • Regression with time series structure

Keep exploring