Understanding central tendencies
The example below compares a data set without an extreme value to one with an extreme value, showing how each measure responds.
Example: Effect of an extreme value on center
Dataset A:
- Ordered: .
- Mean: .
- Median: the 3rd value in the ordered list .
- Mode: all values are unique, so there is no mode.
Dataset B: (replace 10 with an extreme value)
- Mean: .
- Median: the 3rd value is still .
- Mode: no mode.
Answer: Dataset A - Mean , Median , no mode. Dataset B - Mean , Median , no mode.
Measures of spread
Measures of spread describe how much the values in a data set vary around the center. For the Praxis, focus on range, interquartile range, and standard deviation.
Example: Computing range and IQR
- To find the interquartile range (IQR), start by ordering the data and identifying the median, which is 7. Then split the data into a lower half and an upper half around the median.
- Because is odd, we use the exclusive method: exclude the median (7) before splitting.
- The lower half is , so the first quartile is the average of those two values: .
- The upper half is , so the third quartile is .
- The IQR is then calculated as .
- To check for outliers, compute the fences: and . Any value below or above counts as an outlier - since all five values fall between and , this data set has none.
Answer: IQR ; no outliers (fences are and ).
Example: Computing standard deviation
- Mean . Squared deviations: .
- Variance: .
- Standard deviation: .
Answer:
Choosing the right measure
Now that you know how to compute both center and spread, here’s a guide for deciding which measure to use in a given situation.
This table matches each data situation to the best measure of center and spread. Use mode for categorical data, mean and standard deviation for symmetric data, and median and IQR for skewed data or data with outliers.
| Situation | Best measure of center | Best measure of spread | Why |
|---|---|---|---|
| Categorical data (e.g., favorite color) | Mode | Not applicable | Mean and median cannot be calculated, only mode makes sense. |
| Numerical data, no extreme values, symmetric | Mean | Standard deviation | Uses all values, accurate for well-behaved data. |
| Numerical data, no extreme values, skewed | Median | IQR | Median resists skew, IQR ignores extremes. |
| Numerical data with extreme values (outliers) | Median | IQR | Both are resistant to outliers, mean and standard deviation would be distorted. |
| Small data set | Median | Range or IQR | Range is quick to compute; IQR is still appropriate. Because range is sensitive to outliers, use IQR if any extreme values are present. |
Effects of transformations
When a constant is added to each value in a data set, the mean, median, and mode all increase by . Measures of spread such as the range, interquartile range (IQR), and standard deviation stay the same, because the distances between values don’t change.
When each data value is multiplied by a positive constant , the mean, median, and mode are all multiplied by . The range, IQR, and standard deviation are also multiplied by , because all distances are scaled by the same factor.
The next chapter, Understanding and representing data, covers how data is organized and displayed - including tables, graphs, and charts - before later chapters take up scatterplots and probability.