AP subjects/AP Statistics/Distribution Shape, Center & Spread Explorer
CED 1.6AP Statistics

Distribution Shape, Center & Spread Explorer

Use this free distribution shape, center and spread explorer to see how skewness pulls the mean away from the median, and how the SD, IQR and range measure variability. Drag two sliders and a histogram of 1,200 values and six summary statistics update together.

Controls
resetBtnskewspread

How to use the simulator

The simulator draws a histogram of 1,200 values in 34 bars. A fixed random seed means the same settings always give the same picture. The controls:
  • Skew runs from −1.00 to 1.00 in steps of 0.01 and starts at 0.00. Positive values stretch a long tail to the right, negative values to the left. The median stays near 50.
  • Spread (σ) runs from 4 to 20 in steps of 1 and starts at 10. It rescales the data so the standard deviation equals the slider value, changing spread but not shape.
  • Reset to symmetric puts Skew back to 0.00 and Spread back to 10.
On the histogram, a solid teal line marks the mean (x̄), a solid navy line the median, and a dashed gold line the mode (the middle of the tallest bar). The results box lists Mean (x̄), Median, Mode, SD (s), IQR (Q3−Q1) and Range to two decimal places. A verdict bar reads Skewed right → mean > median, Skewed left → mean < median or Roughly symmetric → mean ≈ median.
At Skew 1.00 and Spread 10 the box shows mean 53.02 and median 50.00: the right tail drags the mean about 3 units its way, while the mode (45.28) sits under the peak. At Skew −1.00 the order flips: mean 47.33, median 49.99, mode 53.64.
One caution: the verdict only says "skewed" when the mean and median differ by more than a fixed 1.2 units. A smaller Spread shrinks that gap, so at Spread 4 the verdict reads Roughly symmetric at nearly every Skew setting, even when the histogram is clearly lopsided. Judge shape from the picture, as you would on the exam.

The formulas

xˉ=∑xin\bar{x} = \frac{\sum x_i}{n}
  • The mean xˉ\bar{x} uses every value, so a few extreme values in a tail move it a long way.
  • The median is the middle of the ordered data. It is resistant to extreme values.
sx=∑(xi−xˉ)2n−1s_x = \sqrt{\frac{\sum (x_i - \bar{x})^2}{n - 1}}
  • The standard deviation sxs_x is roughly the typical distance of a value from the mean. It is not resistant.
  • IQR=Q3−Q1\text{IQR} = Q_3 - Q_1, the spread of the middle half of the data. It is resistant.
  • Range=max−min\text{Range} = \text{max} - \text{min}, set by just the two most extreme values.
  • The 1.5 × IQR rule: a value is an outlier if it is below Q1−1.5×IQRQ_1 - 1.5 \times \text{IQR} or above Q3+1.5×IQRQ_3 + 1.5 \times \text{IQR}.
At Skew 0.00 and Spread 10 the IQR is 13.72; at Skew 1.00 the SD is still 10.00 but the IQR falls to 9.58, because the tail inflates the SD and the IQR ignores it. For skewed data or outliers, report the median and IQR; for roughly symmetric data, the mean and SD.

Worked example

Problem: Ten students recorded their commute times in minutes: 5, 8, 10, 12, 12, 14, 15, 18, 25, 41. Describe the distribution.
Step 1: Center. The mean is xˉ=160/10=16.0\bar{x} = 160 / 10 = 16.0 minutes. With 10 values, the median is the average of the 5th and 6th: (12+14)/2=13(12 + 14)/2 = 13 minutes.
Step 2: Quartiles and IQR. Q1Q_1 is the median of the lower five values (5, 8, 10, 12, 12), so Q1=10Q_1 = 10. Q3Q_3 is the median of the upper five (14, 15, 18, 25, 41), so Q3=18Q_3 = 18. IQR=18−10=8\text{IQR} = 18 - 10 = 8 minutes.
Step 3: Outliers. 1.5×8=121.5 \times 8 = 12. The fences are 10−12=−210 - 12 = -2 and 18+12=3018 + 12 = 30. The value 41 is above 30, so it is an outlier.
Step 4: Other spread. sx≈10.37s_x \approx 10.37 minutes; range =41−5=36= 41 - 5 = 36 minutes.
Step 5: Describe. The distribution of commute times is skewed right with a high outlier at 41 minutes. Because of the skew and outlier, use resistant measures: median 13 minutes, IQR 8 minutes. The mean (16.0) is above the median, as the right tail predicts.
Compare in the simulator: set Skew to 1.00 for the same pattern on a larger data set: mode below median below mean (45.28, 50.00, 53.02), and SD (10.00) above IQR (9.58).

Common mistakes on the AP exam

  • Naming skew after the peak. "Skewed right" means the long tail points right. The peak is on the left.
  • Leaving out a feature or the context. Cover shape, center, spread and unusual features (outliers, gaps, clusters), with the variable and units: "median commute of 13 minutes."
  • Reporting the mean and SD for strongly skewed data. Both are pulled by the tail. Pair median with IQR, mean with SD.
  • Calling a point an outlier by eye. Show the 1.5 × IQR fence calculation, as in Step 3.
  • Using the range as the main measure of spread. It depends on only two values. Here, Skew −1.00 at Spread 10 gives a range of 101.87, even though the SD is still 10.00.

When the AP exam uses this

Describing the distribution of a quantitative variable is topic 1.6, the base for later Unit 1 work on boxplots and comparing distributions. Free-response questions often open by asking you to describe or compare distributions in a dotplot, histogram or boxplot; a full answer addresses shape, center, spread and outliers in context.
Embed this simulator on your class page

Free for classroom use. Paste this into your site, LMS page or blog; keep the credit link under it.