Inferences and rules of generalizability
Statistical significance and natural variation
After you collect results from an experimental study, the next step is to decide whether the differences you see are statistically significant or whether they could reasonably be explained by natural variation. Put another way: are the differences between subjects small enough that they might happen just by chance, or are they large enough that chance is an unlikely explanation?
Replication
Example:
A research team wants to find out which of medications works best. Instead of having only patients total (one for each medication), the team recruits participants and divides them into groups of . That means there are participants trying the first medication, another trying the second medication, and so on. This is replication because there are participants (not just one) in each treatment group.
Statistical significance
Example:
Suppose a professor wants to find out whether an exam jam session (an extra review session before the exam) increases exam scores.
- In semester one, she did not offer an exam jam session and recorded students’ final exam grades.
- In semester two, she offers an exam jam session and compares the semester two final exam scores to semester one.
If students score only higher on average in semester two, the professor might conclude there isn’t enough evidence that the exam jam session had a real impact. A small increase like that could easily happen due to natural variation.
If students score higher on average in semester two, the professor is more likely to attribute the increase to the exam jam session, because a jump that large is less likely to happen just by coincidence or natural variation alone.
A design that would make this evidence stronger would be for the professor to randomly assign half her students to attend the exam jam session and half not to attend, and then compare exam scores. Better yet, she could use blocking and ensure that the students selected come from a range of current grade levels. The challenge with these improvements is that it may be unfair for only some students to be invited to attend the exam jam while others are excluded.
Practice problems
Example:
A restaurant owner claims that when people eat his food, they report higher levels of happiness than people who do not eat at his restaurant. People were surveyed and asked to rate their happiness on a scale of – with being very unhappy and being extremely happy. Here are the results for people who had eaten at the restaurant within the last months and people who had not done so:
People who ate at the restaurant:
People who did not eat at the restaurant:
Based on the results, does the data support the restaurant owner’s claim?
Solution:
No, the data is not conclusive enough to support the restaurant owner’s claim. The mean and median in both data sets vary only slightly, so the difference could potentially be explained by natural variation amongst the happiness levels of different people and is not necessarily a result of eating at this particular restaurant.
Example:
The probability of rolling an even number on a fair -sided die is . Kyle’s friend claims that a specific -sided die that they are using to play a game is fair, but Kyle claims that it is unfair because after Kyle rolled the die times, he only got even numbers out of . Kyle’s friend insists that Kyle’s experiment results can be explained by natural variation. Who is correct and why?
Solution:
Kyle is correct because obtaining only even numbers out of when out of were expected is very unlikely.
Kyle’s friend would have more of a point if Kyle had rolled between and even numbers instead of the expected , but is so much less than expected that it suggests the result is statistically significant. That makes it much more likely that the die is not fair, which is what Kyle is claiming.
Kyle is correct here, but it raises an important question: at what point would his friend be correct? What if even numbers had been rolled? How about ? Or ? Or ? Where exactly is the cutoff point where the results are statistically significant rather than just the result of natural variation?
That question is more complex and depends on factors such as sample size, theoretical probability, and the required confidence level. Later chapters will show how to answer it more precisely.
It’s also worth noting that the sampling method affects whether the results of an experiment can be generalized. With a random sampling method, you can be much more confident that the results apply to the entire population. If a convenience sample or voluntary sample is used, the results are less likely to be generalizable.