ArticleslgStudy

mathematics

Log transformation (statistics)

Log transformation (statistics) is a mathematics topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand Log transformation (statistics) rather than just read about it. In short: In statistics, the log transformation is the application of the logarithmic function to each point in a data set—that is, each data point zi is replaced with the transformed value yi = log(zi). The log transform is usually applied so that the data, after transformation, appear to more closely meet the assumptions of a statistical inference procedure that is to be applied, or to improve the interpretability or appear…

Log transformation (statistics) — main illustration
Log transformation (statistics) — illustration

Key takeaways

  • Log transformation (statistics) belongs to mathematics; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect Log transformation (statistics) to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of Log transformation (statistics) from memory before moving on to harder problems.

Reference excerpt

In statistics, the log transformation is the application of the logarithmic function to each point in a data set—that is, each data point zi is replaced with the transformed value yi = log(zi). The log transform is usually applied so that the data, after transformation, appear to more closely meet the assumptions of a statistical inference procedure that is to be applied, or to improve the interpretability or appearance of graphs. The log transform is invertible, continuous, and monotonic. The transformation is usually applied to a collection of comparable measurements. For example, if we are working with data on peoples' incomes in some currency unit, it would be common to transform each person's income value by the logarithm function.

Motivation

Guidance for how data should be transformed, or whether a transformation should be applied at all, should come from the particular statistical analysis to be performed. For example, a simple way to construct an approximate 95% confidence interval for the population mean is to take the sample mean plus or minus two standard error units. However, the constant factor 2 used here is particular to the normal distribution, and is only applicable if the sample mean varies approximately normally. The central limit theorem states that in many situations, the sample mean does vary normally if the sample size is reasonably large. However, if the population is substantially skewed and the sample size is at most moderate, the approximation provided by the central limit theorem can be poor, and the resulting confidence interval will likely have the wrong coverage probability. Thus, when there is evidence of substantial skew in the data, it is common to transform the data to a symmetric distribution before constructing a confidence interval. If desired, the confidence interval can be constructed for statistics in the original scale, such as the median or the mean, by transforming back to the original scale using exponent (with some adjustments for CI for the mean), the inverse of the log transformation that was applied to the data. it is possible to estimate a quantile using different methods, build a CI for it, and then transform these back to the original scale so to have a CI for the quantile in the original scale. For example, it's possible to estimate the location of the median, after the log transformation, using the arithmetic mean. Then build CI for the median using a CI for the mean and transform the CI back to the original scale using exponent. That transformed CI is then a CI for the median, not the mean. Data can also be transformed to make them easier to visualize. For example, suppose we have a scatterplot in which the points are the countries of the world, and the data values being plotted are the land area and population of each country. If the plot is made using untransformed data (e.g. square kilometers for area and the number of people for population), most of the countries would be plotted in tight cluster of points in the lower left corner of the graph. The few countries with very large areas and/or populations would be spread thinly around most of the graph's area. Simply rescaling units (e.g., to thousand square kilometers, or to millions of people) will not change this. However, following logarithmic transformations of both area and population, the points will be spread more uniformly in the graph. Another reason for applying the log data transformation is to improve interpretability, even if no formal statistical analysis or visualization is to be performed.

In regression

Data transformation may be used as a remedial measure to make data suitable for modeling with linear regression if the original data violates one or more assumptions of linear regression. For example, the simplest linear regression models assume a linear relationship between the expected value of Y (the response variable to be predicted) and each independent variable (when the other independent variables are held fixed). If linearity fails to hold, even approximately, it is sometimes possible to transform either the independent or dependent variables in the regression model to improve the linearity. For example, addition of quadratic functions of the original independent variables may lead to a linear relationship with expected value of Y, resulting in a polynomial regression model, a special case of linear regression. Another assumption of linear regression is homoscedasticity, that is the variance of errors must be the same regardless of the values of predictors. If this assumption is violated (i.e. if the data is heteroscedastic), it may be possible to find a transformation of Y alone, or transformations of both X (the predictor variables) and Y, such that the homoscedasticity assumption (in addition to the linearity assumption) holds true on the transformed variables and linear regression may therefore be applied on these. Yet another application of data transformation is to address the problem of lack of normality in error terms. Univariate normality is not needed for least squares estimates of the regression parameters to be meaningful (see Gauss–Markov theorem). However confidence intervals and hypothesis tests will have better statistical properties if the variables exhibit multivariate normality. Transformations that stabilize the variance of error terms (i.e. those that address heteroscedaticity) often also help make the error terms approximately normal.

Examples Equation:

Y = a + b X {\displaystyle Y=a+bX}

Meaning: A unit increase in X is associated with an average of b units increase in Y. Equation:

log ⁡ ( Y ) = a + b X {\displaystyle \log(Y)=a+bX}

… excerpt ends here. Continue reading the full article.

Illustrations

Log transformation (statistics): Fitted cumulative log-normal distribution to annually maximum 1-day rainfalls, see distribution fitting
Fitted cumulative log-normal distribution to annually maximum 1-day rainfalls, see distribution fitting

Worked examples

Example 1 — a first encounter with Log transformation (statistics)

Start with the simplest possible case. Write down what Log transformation (statistics) claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In mathematics, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to Log transformation (statistics) before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about Log transformation (statistics) ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of Log transformation (statistics)

In research
Log transformation (statistics) appears in mathematics research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses Log transformation (statistics) in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
Log transformation (statistics) is common in secondary-school and first-year university syllabi. It links to neighbouring topics Statistics, so understanding it makes those chapters shorter.
In everyday life
Look for Log transformation (statistics) outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study Log transformation (statistics) in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what Log transformation (statistics) means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain Log transformation (statistics) out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is Log transformation (statistics) in simple terms?

In statistics, the log transformation is the application of the logarithmic function to each point in a data set—that is, each data point zi is replaced with the transformed value yi = log(zi). The log transform is usually applied so that the data, after transformation, appear to more closely meet…

Why does Log transformation (statistics) matter?

Because it connects several mathematics ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study Log transformation (statistics)?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on Log transformation (statistics).

Tags

  • Statistics

Keep exploring