Achievable logoAchievable logo
AP Statistics
Sign in
Sign up
Purchase
Textbook
Practice exams
Support
How it works
Resources
Exam catalog
Mountain with a flag at the peak
Textbook
Introduction
1. One variable data
2. Two variable data
3. Data collection
3.1 Introduction to collecting data
3.2 Experimental vs observational studies
3.3 Inferences and rules of generalizability
4. Probability and random variables
5. Sampling distributions
6. Categorical data
7. Quantitative data
8. Chi-square
9. Linear regression
Wrapping up
Achievable logoAchievable logo
3.3 Inferences and rules of generalizability
Achievable AP Statistics
3. Data collection
Our AP Statistics course is currently in development and is a work-in-progress.

Inferences and rules of generalizability

5 min read
Font
Discuss
Share
Feedback

Statistical significance and natural variation

After you collect results from an experimental study, the next step is to decide whether the differences you see are statistically significant or whether they could reasonably be explained by natural variation. Put another way: are the differences between subjects small enough that they might happen just by chance, or are they large enough that chance is an unlikely explanation?

Definitions
Replication
Having more than one experimental unit in each treatment group.
Natural variation
The inherent differences amongst individuals or measurements that occurs without anything systematic occurring.
Statistically significant
The results are unlikely to have occurred due to natural variation.

Replication

Example:

A research team wants to find out which of 5 medications works best. Instead of having only 5 patients total (one for each medication), the team recruits 100 participants and divides them into 5 groups of 20. That means there are 20 participants trying the first medication, another 20 trying the second medication, and so on. This is replication because there are 20 participants (not just one) in each treatment group.

Statistical significance

Example:

Suppose a professor wants to find out whether an exam jam session (an extra review session before the exam) increases exam scores.

  • In semester one, she did not offer an exam jam session and recorded students’ final exam grades.
  • In semester two, she offers an exam jam session and compares the semester two final exam scores to semester one.

If students score only 3% higher on average in semester two, the professor might conclude there isn’t enough evidence that the exam jam session had a real impact. A small increase like that could easily happen due to natural variation.

If students score 15% higher on average in semester two, the professor is more likely to attribute the increase to the exam jam session, because a jump that large is less likely to happen just by coincidence or natural variation alone.

A design that would make this evidence stronger would be for the professor to randomly assign half her students to attend the exam jam session and half not to attend, and then compare exam scores. Better yet, she could use blocking and ensure that the students selected come from a range of current grade levels. The challenge with these improvements is that it may be unfair for only some students to be invited to attend the exam jam while others are excluded.

Practice problems

Example:

A restaurant owner claims that when people eat his food, they report higher levels of happiness than people who do not eat at his restaurant. People were surveyed and asked to rate their happiness on a scale of 0–10 with 0 being very unhappy and 10 being extremely happy. Here are the results for people who had eaten at the restaurant within the last 3 months and people who had not done so:

People who ate at the restaurant: {1,2,3,3,4,6,7,7,8,8,9,10}

People who did not eat at the restaurant: {1,2,3,3,4,6,6,7,7,8,9,9}

Based on the results, does the data support the restaurant owner’s claim?

Solution:

(spoiler)

No, the data is not conclusive enough to support the restaurant owner’s claim. The mean and median in both data sets vary only slightly, so the difference could potentially be explained by natural variation amongst the happiness levels of different people and is not necessarily a result of eating at this particular restaurant.

Example:

The probability of rolling an even number on a fair 6-sided die is 50%. Kyle’s friend claims that a specific 6-sided die that they are using to play a game is fair, but Kyle claims that it is unfair because after Kyle rolled the die 100 times, he only got 23 even numbers out of 100. Kyle’s friend insists that Kyle’s experiment results can be explained by natural variation. Who is correct and why?

Solution:

(spoiler)

Kyle is correct because obtaining only 23 even numbers out of 100 when 50 out of 100 were expected is very unlikely.

Kyle’s friend would have more of a point if Kyle had rolled between 45 and 55 even numbers instead of the expected 50, but 23 is so much less than expected that it suggests the result is statistically significant. That makes it much more likely that the die is not fair, which is what Kyle is claiming.

Kyle is correct here, but it raises an important question: at what point would his friend be correct? What if 40 even numbers had been rolled? How about 41? Or 42? Or 43? Where exactly is the cutoff point where the results are statistically significant rather than just the result of natural variation?

That question is more complex and depends on factors such as sample size, theoretical probability, and the required confidence level. Later chapters will show how to answer it more precisely.

It’s also worth noting that the sampling method affects whether the results of an experiment can be generalized. With a random sampling method, you can be much more confident that the results apply to the entire population. If a convenience sample or voluntary sample is used, the results are less likely to be generalizable.

Statistical significance vs. natural variation

  • Statistical significance: results unlikely due to natural variation
  • Natural variation: inherent, random differences among individuals or measurements
  • Key question: are observed differences large enough to rule out chance?

Replication

  • Multiple experimental units per treatment group
  • Increases reliability and generalizability of results

Examples and interpretation

  • Small group differences may be due to natural variation, not treatment effect
  • Large, unexpected differences suggest statistical significance
  • Experimental design improvements:
    • Random assignment to groups
    • Blocking to control for confounding variables

Sampling methods and generalizability

  • Random sampling: results more generalizable to population
  • Convenience/voluntary samples: less reliable for generalization

Key questions when interpreting results

  • Is the result statistically significant or just natural variation?
  • To whom or what situations can the results be applied?

Sign up for free to take 5 quiz questions on this topic

Previous
Next  | 4.1 Law of large numbers
All rights reserved ©2016 - 2026 Achievable, Inc.

Inferences and rules of generalizability

Statistical significance and natural variation

After you collect results from an experimental study, the next step is to decide whether the differences you see are statistically significant or whether they could reasonably be explained by natural variation. Put another way: are the differences between subjects small enough that they might happen just by chance, or are they large enough that chance is an unlikely explanation?

Definitions
Replication
Having more than one experimental unit in each treatment group.
Natural variation
The inherent differences amongst individuals or measurements that occurs without anything systematic occurring.
Statistically significant
The results are unlikely to have occurred due to natural variation.

Replication

Example:

A research team wants to find out which of 5 medications works best. Instead of having only 5 patients total (one for each medication), the team recruits 100 participants and divides them into 5 groups of 20. That means there are 20 participants trying the first medication, another 20 trying the second medication, and so on. This is replication because there are 20 participants (not just one) in each treatment group.

Statistical significance

Example:

Suppose a professor wants to find out whether an exam jam session (an extra review session before the exam) increases exam scores.

  • In semester one, she did not offer an exam jam session and recorded students’ final exam grades.
  • In semester two, she offers an exam jam session and compares the semester two final exam scores to semester one.

If students score only 3% higher on average in semester two, the professor might conclude there isn’t enough evidence that the exam jam session had a real impact. A small increase like that could easily happen due to natural variation.

If students score 15% higher on average in semester two, the professor is more likely to attribute the increase to the exam jam session, because a jump that large is less likely to happen just by coincidence or natural variation alone.

A design that would make this evidence stronger would be for the professor to randomly assign half her students to attend the exam jam session and half not to attend, and then compare exam scores. Better yet, she could use blocking and ensure that the students selected come from a range of current grade levels. The challenge with these improvements is that it may be unfair for only some students to be invited to attend the exam jam while others are excluded.

Practice problems

Example:

A restaurant owner claims that when people eat his food, they report higher levels of happiness than people who do not eat at his restaurant. People were surveyed and asked to rate their happiness on a scale of 0–10 with 0 being very unhappy and 10 being extremely happy. Here are the results for people who had eaten at the restaurant within the last 3 months and people who had not done so:

People who ate at the restaurant: {1,2,3,3,4,6,7,7,8,8,9,10}

People who did not eat at the restaurant: {1,2,3,3,4,6,6,7,7,8,9,9}

Based on the results, does the data support the restaurant owner’s claim?

Solution:

(spoiler)

No, the data is not conclusive enough to support the restaurant owner’s claim. The mean and median in both data sets vary only slightly, so the difference could potentially be explained by natural variation amongst the happiness levels of different people and is not necessarily a result of eating at this particular restaurant.

Example:

The probability of rolling an even number on a fair 6-sided die is 50%. Kyle’s friend claims that a specific 6-sided die that they are using to play a game is fair, but Kyle claims that it is unfair because after Kyle rolled the die 100 times, he only got 23 even numbers out of 100. Kyle’s friend insists that Kyle’s experiment results can be explained by natural variation. Who is correct and why?

Solution:

(spoiler)

Kyle is correct because obtaining only 23 even numbers out of 100 when 50 out of 100 were expected is very unlikely.

Kyle’s friend would have more of a point if Kyle had rolled between 45 and 55 even numbers instead of the expected 50, but 23 is so much less than expected that it suggests the result is statistically significant. That makes it much more likely that the die is not fair, which is what Kyle is claiming.

Kyle is correct here, but it raises an important question: at what point would his friend be correct? What if 40 even numbers had been rolled? How about 41? Or 42? Or 43? Where exactly is the cutoff point where the results are statistically significant rather than just the result of natural variation?

That question is more complex and depends on factors such as sample size, theoretical probability, and the required confidence level. Later chapters will show how to answer it more precisely.

It’s also worth noting that the sampling method affects whether the results of an experiment can be generalized. With a random sampling method, you can be much more confident that the results apply to the entire population. If a convenience sample or voluntary sample is used, the results are less likely to be generalizable.

Key points

Statistical significance vs. natural variation

  • Statistical significance: results unlikely due to natural variation
  • Natural variation: inherent, random differences among individuals or measurements
  • Key question: are observed differences large enough to rule out chance?

Replication

  • Multiple experimental units per treatment group
  • Increases reliability and generalizability of results

Examples and interpretation

  • Small group differences may be due to natural variation, not treatment effect
  • Large, unexpected differences suggest statistical significance
  • Experimental design improvements:
    • Random assignment to groups
    • Blocking to control for confounding variables

Sampling methods and generalizability

  • Random sampling: results more generalizable to population
  • Convenience/voluntary samples: less reliable for generalization

Key questions when interpreting results

  • Is the result statistically significant or just natural variation?
  • To whom or what situations can the results be applied?

More from Data collection

  • Introduction to collecting data
  • Experimental vs observational studies