Use this free Type I and Type II errors and power simulator to see how α, β and power trade off in a significance test about a mean. Change the effect size, sample size, σ, significance level and tail direction, and the shaded areas and readouts update together.
Controls
effnsig
How to use the simulator
The graph shows two Normal sampling distributions of the sample mean. The solid navy curve labelled H0 is centred at the null value μ0; the dashed coral curve labelled Ha is centred at the true mean. The gold dashed line marked x* is the critical value: a sample mean beyond it leads you to reject H0. Both curves have standard deviation σ/n, so the simulator treats σ as known.
The controls:
Effect size |μtrue − μ0|: 0 to 4 in steps of 0.05 (default 1.50). This is how far the true mean sits from the null value.
Sample size (n): 5 to 200 in steps of 1 (default 30).
Population SD (σ): 0.5 to 3 in steps of 0.05 (default 1.00).
Significance level (α): buttons for 0.01, 0.05 (default) and 0.10.
Tail mode: Right tail (Ha: μ > μ0, the default) or Two tail (Ha: μ ≠ μ0), which splits α between two critical lines.
Three shaded regions match three boxes under the graph. Gold is α, the part of the H0 curve inside the rejection region. Coral is β, the part of the Ha curve on the fail-to-reject side. Teal is power, the rest of the Ha curve. The α box repeats your chosen level, β and power are given to three decimal places, and the power number turns teal once it reaches 0.80 (it is coral below that).
At the default settings the curves hardly overlap and power reads 1.000, so lower the effect size or n before comparing settings. Try an effect size of 0: β reads 0.950 and power reads 0.050 at α = 0.05. When H0 is actually true, the only way to reject it is by making a Type I error.
The key ideas
Type I error: rejecting H0 when H0 is true. Its probability is α.
Type II error: failing to reject H0 when a particular alternative is true. Its probability is β.
Power: the probability of correctly rejecting H0 when that alternative is true.
power=1−β
For a right-tailed test with known σ, which is what the simulator draws, the critical value and β are:
xˉ∗=μ0+z∗nσβ=P(Z<σ/nxˉ∗−μa)
Power rises when the effect size grows, n grows, σ shrinks or α grows. The first three push the curves apart or make them narrower. A larger α slides x* toward the centre of H0, which buys power at the cost of more Type I errors. Of the four, increasing n is the usual way to gain power without accepting more Type I errors.
Worked example
Problem: An engineer redesigns a battery and tests H0:μ=20 hours against Ha:μ>20 at α=0.05. Battery life has σ=2 hours. If the new design really averages 20.5 hours, what is the power of the test with 36 batteries? With 100?
Step 1: Standard deviation of xˉ.σ/n=2/36=0.333 hours.
Step 2: Critical value. For a right-tailed test at 0.05, z∗=1.645, so xˉ∗=20+1.645(0.333)=20.548 hours.
Step 3: Type II error probability.β=P(xˉ<20.548 when μ=20.5)=P(Z<0.33320.548−20.5)=P(Z<0.145)≈0.558. Power =1−0.558=0.442.
Step 4: Repeat with n = 100.σ/n=0.2, xˉ∗=20+1.645(0.2)=20.329, and β=P(Z<(20.329−20.5)/0.2)=P(Z<−0.855)≈0.196. Power ≈0.804.
Interpretation: with 36 batteries the test misses a real half-hour improvement more often than it detects it. About 100 batteries are needed; n = 99 is the smallest sample size that gets power to 0.80.
Check it in the simulator: set effect size to 0.50, σ to 2.00 and n to 36, with α = 0.05 and Right tail. The boxes read β = 0.558 and power = 0.442. Slide n to 100 and they change to 0.196 and 0.804, and the power number turns teal. Go back to n = 36 and press α = 0.10: power climbs to 0.586, but the gold α area doubles. Switch to Two tail and power at n = 36 drops to 0.323.
Common mistakes on the AP exam
Describing an error without context. "Rejecting a true null" is only the definition. Say what it means in the problem: concluding the new battery lasts longer on average when it does not.
Mixing up which error is possible. If you rejected H0, the only error you could have made is Type I. If you failed to reject, it could only be Type II.
Treating α as P(H0 is true). α is the probability of rejecting given thatH0 is true, not the probability that H0 is true.
Saying a bigger sample lowers α. The researcher sets α. A bigger n lowers β and raises power while α stays put.
Assuming a smaller α is always safer. Moving from 0.05 to 0.01 cut power from 0.442 to 0.204 in the example. Fewer Type I errors come at the price of more Type II errors.
Defining power as "the probability the null is false." Power is the probability of rejecting H0 when a specific alternative value is true.
Not explaining a consequence. When asked which error is worse, describe what would actually happen, such as a real improvement being shelved.
When the AP exam uses this
This is topic 3.8, Potential Errors When Performing Tests, in Unit 3 (Inference for Categorical Data: Proportions). The same reasoning carries over to the tests for means in Unit 4. Expect to identify and describe Type I and Type II errors in context, to state a consequence of each, and to explain how changing n or α affects power.
Embed this simulator on your class page
Free for classroom use. Paste this into your site, LMS page or blog; keep the credit link under it.
Sitting a paper sets one essential cookie so your attempt stays yours (paying or signing in sets the ones those need). Nothing identifying is set unless you accept: advertising cookies only help us measure our adverts, and the site works either way.
Sitting a paper sets one essential cookie so your attempt stays yours. Nothing identifying is set unless you accept; advertising cookies only measure our adverts.