ArticleslgStudy

science

One in ten rule

One in ten rule is a science topic covered in the lgStudy science library. This page brings together a partial reference excerpt, illustrations, worked examples, real-world applications and a short study plan, so you can understand One in ten rule rather than just read about it. In short: In statistics, the one in ten rule is a rule of thumb for how many predictor parameters can be estimated from data when doing regression analysis (in particular proportional hazards models in survival analysis and logistic regression) while keeping the risk of overfitting and finding spurious correlations low. The rule states that one predictive variable can be studied for every ten events.

Key takeaways

  • One in ten rule belongs to science; place it in that map before memorising details.
  • Learn the definition first, then one example that makes the definition concrete.
  • Connect One in ten rule to a quantity you can measure, compute or draw — that is where exam questions come from.
  • Reproduce the core statement of One in ten rule from memory before moving on to harder problems.

Reference excerpt

In statistics, the one in ten rule is a rule of thumb for how many predictor parameters can be estimated from data when doing regression analysis (in particular proportional hazards models in survival analysis and logistic regression) while keeping the risk of overfitting and finding spurious correlations low. The rule states that one predictive variable can be studied for every ten events. For logistic regression the number of events is given by the size of the smallest of the outcome categories, and for survival analysis it is given by the number of uncensored events. In other words: for each feature we need 10 observations/labels. For example, if a sample of 200 patients is studied and 20 patients die during the study (so that 180 patients survive), the one in ten rule implies that two pre-specified predictors can reliably be fitted to the total data. Similarly, if 100 patients die during the study (so that 100 patients survive), ten pre-specified predictors can be fitted reliably. If more are fitted, the rule implies that overfitting is likely and the results will not predict well outside the training data. It is not uncommon to see the 1:10 rule violated in fields with many variables (e.g. gene expression studies in cancer), decreasing the confidence in reported findings.

Improvements A "one in 20 rule" has been suggested, indicating the need for shrinkage of regression coefficients, and a "one in 50 rule" for stepwise selection with the default p-value of 5%. Other studies, however, show that the one in ten rule may be too conservative as a general recommendation and that five to nine events per predictor can be enough, depending on the research question. More recently, a study has shown that the ratio of events per predictive variable is not a reliable statistic for estimating the minimum number of events for estimating a logistic prediction model. Instead, the number of predictor variables, the total sample size (events + non-events) and the events fraction (events / total sample size) can be used to calculate the expected prediction error of the model that is to be developed. One can then estimate the required sample size to achieve an expected prediction error that is smaller than a predetermined allowable prediction error value. Alternatively, three requirements for prediction model estimation have been suggested: the model should have a global shrinkage factor of ≥ 0.9, an absolute difference of ≤ 0.05 in the model's apparent and adjusted Nagelkerke R2, and a precise estimation of the overall risk or rate in the target population. The necessary sample size and number of events for model development are then given by the values that meet these requirements.

Other modalities For highly correlated input data the one-in-10 rule (10 observations or labels needed per feature) may not be directly applicable due to the high correlation of the features: For images there is a rule of thumb that per class 1000 examples are needed. This would mean that for a binary classification of images (with fictive 1000 pixel x 1000 pixel per image, i.e. 1 000 000 features per image), we would only require 2000 labels /1 000 0000 pixel = 0.002 labels per pixel or 0.002 labels per feature. This is however only due to the high (spatial) correlation of pixels.

Literature David A. Freedman (1983) "A Note on Screening Regression Equations," The American Statistician, 37:2, 152–155, doi:10.1080/00031305.1983.10482729

References

Worked examples

Example 1 — a first encounter with One in ten rule

Start with the simplest possible case. Write down what One in ten rule claims or describes in one sentence, then invent the smallest concrete situation in which that sentence is true. In science, the smallest case is usually a single object, a single equation or a single measurement. Check that every symbol or term in your sentence has a meaning in that case.

Example 2 — changing one variable

Take the situation from Example 1 and change exactly one quantity: double it, halve it, or set it to zero. Predict what should happen to One in ten rule before you calculate. Comparing your prediction with the result is the fastest way to find out whether you understand the idea or only the words.

Example 3 — an exam-style question

Typical questions about One in ten rule ask you to (a) state it precisely, (b) apply it to given data, and (c) explain a limitation. Practise writing all three answers in under five minutes; the third part is what separates a full-mark answer from an average one.

Applications of One in ten rule

In research
One in ten rule appears in science research whenever the underlying quantities have to be modelled precisely. Papers usually cite it as a starting assumption and then explore where it breaks down.
In technology and industry
Engineering practice reuses One in ten rule in design rules, simulations and safety margins. Knowing the idea lets you read a specification sheet and understand why the numbers look the way they do.
In the classroom
One in ten rule is common in secondary-school and first-year university syllabi. It links to neighbouring topics Regression variable selection, Rules of thumb, so understanding it makes those chapters shorter.
In everyday life
Look for One in ten rule outside the textbook — in sport, cooking, traffic, electronics or the sky above you. An example you found yourself is remembered far longer than one you were given.
Ask Teacher Smith questions about this articleOpens your AI tutor with a question about “One in ten rule” →

Affiliate

Preply — study more efficiently by working with a personal tutor. 50% off.

How to study One in ten rule in 20 minutes

  1. Read the reference excerpt below once, without taking notes.
  2. Close the page and write down what One in ten rule means in your own words.
  3. Compare your version with the excerpt and mark what you missed.
  4. Work through the three examples above with pen and paper.
  5. Explain One in ten rule out loud to somebody else — or to Teacher Smith in the lgStudy chat.

Frequently asked questions

What is One in ten rule in simple terms?

In statistics, the one in ten rule is a rule of thumb for how many predictor parameters can be estimated from data when doing regression analysis (in particular proportional hazards models in survival analysis and logistic regression) while keeping the risk of overfitting and finding spurious corre…

Why does One in ten rule matter?

Because it connects several science ideas at once: it gives you a definition you can apply, a quantity you can calculate, and a way to check whether a result is plausible.

How should I study One in ten rule?

Read the excerpt, restate it from memory, then work through the examples and applications listed on this page. The five-step study plan above takes about twenty minutes.

What does this page cover?

It gives you a compact reference excerpt plus original lgStudy explanations, examples, applications and study material on One in ten rule.

Tags

  • Regression variable selection
  • Rules of thumb

Keep exploring