Refresh KidLearning
LESSON 13 / 23 · TOPIC 3.8

How can a study become better at detecting a difference?

You will be able to: Explain power tradeoffs involving n, α and effect size.

Graphs, tables and mathematical reasoningFree study resourceReview editionTeacher review pending

How can a study become better at detecting a difference?

A small poll may miss a modest increase from 50% to 55%. A larger random poll can distinguish that small shift more reliably.

A useful starting point: What can go wrong with a testing decision? →

Words and symbols before equations

Effect size
The distance between a true alternative parameter and the null value.
Rejection cutoff
Statistic values that lead the pre-specified procedure to reject.
Detection probability
Power, evaluated under a specified alternative.
Rejection cutoff and true alternative share00.20.40.60.81Orange reference = 0.5Center 0.6; endpoints 0.59 to 1Horizontal scale: proportion / proportion difference
Read this model snapshot. Nominal α=0.05; actual false-alarm probability=0.04431. At true p=0.6, power=0.62253 and β=0.37747. Exact binomial evaluation of a greater-than normal z decision rule.
What this picture assumes

Greater-than normal z-test of p₀=.50. Exact binomial enumeration evaluates the finite-sample rejection rate of that approximate test rule under the null and selected alternative. Its actual false-alarm rate need not equal nominal α because counts are discrete.

Read the picture in three steps

  1. Read the axes and labels first. Identify what each symbol and line represents. Read the units and fixed conditions before comparing quantities.
  2. Nominal α=0.05; actual false-alarm probability=0.04431. At true p=0.6, power=0.62253 and β=0.37747. Exact binomial evaluation of a greater-than normal z decision rule.
  3. Check what the picture assumes below. Use the Explore task to predict one change before moving a control.

Connect the picture to the mathematics

For a fixed test direction and true effect, increasing n usually raises power because sampling variation becomes smaller. A larger true difference is easier to detect.

Increasing α makes rejection easier, increasing power but also allowing more false alarms. Lowering α generally reduces power if sample size and the alternative stay fixed.

Our investigation enumerates binomial counts for a greater-than normal z-test with p₀=.50. It shows both the exact finite-model false-alarm rate and power of that approximate decision rule. Discrete count thresholds can cause small steps or nonmonotonic changes even though the broad planning trend is upward with n.

A worked example, step by step

At a specified alternative, test A has power .70 and test B .90. Compare missed-effect probabilities.

  1. For A, β=1−.70=.30.
  2. For B, β=1−.90=.10.
  3. B misses that specified effect less often.
  4. Before preferring B, compare α, sample cost, design quality and which alternative the power calculation assumes.
Common mix-up

Power is a pre-study operating characteristic under a specified truth, not the probability that a rejected finding is real.

CHECK THE IDEA

Can more observations fix confounding or selection bias?

Compare with an explanation

No. Better power does not replace a valid design.

Now investigate one change Explore →

Predict. Change one thing. Explain.

Change n, the true alternative share and α one at a time. Explain the differences between the nominal α, actual false-alarm rate and power.

On narrow screens, swipe or scroll diagrams sideways to read all labels.

Rejection cutoff and true alternative share00.20.40.60.81Orange reference = 0.5Center 0.6; endpoints 0.59 to 1Horizontal scale: proportion / proportion difference

Nominal α=0.05; actual false-alarm probability=0.04431. At true p=0.6, power=0.62253 and β=0.37747. Exact binomial evaluation of a greater-than normal z decision rule.

QuantityProbability / count
Nominal α0.05
First rejecting success count59
Actual null rejection probability0.044313
Power at true p=0.60.622533
β at that alternative0.377467

Greater-than normal z-test of p₀=.50. Exact binomial enumeration evaluates the finite-sample rejection rate of that approximate test rule under the null and selected alternative. Its actual false-alarm rate need not equal nominal α because counts are discrete.

Explain what you noticed: Answer the investigation prompt above. State one observation and explain it using the probability values, reference groups, graph scales or model assumptions. Identify what the representation cannot tell you.

Apply the idea to a fresh problem Practice →

Show what you understand.

Two original questions are a starting check, not proof of mastery. Explain your choice before revealing the answer.

1. Holding other factors fixed, a larger true effect is generally…

Show answer and reasoning

Easier to detect. It separates the alternative from the null distribution.

2. Raising α generally…

Show answer and reasoning

Raises power and false-alarm allowance. The rejection rule becomes less demanding.

Original written challenge

4 points · self-check · not an official AP question

A planned study needs fewer missed effects without increasing α. Suggest a change and state its limits.

This response is not submitted or saved. Copy it before leaving.

Compare with the answer and four-point rubric
  1. 1 point: Increase sample size under the same appropriate random design.
  2. 1 point: This generally reduces SE and increases power for the specified effect.
  3. 1 point: It increases cost and does not eliminate all errors.
  4. 1 point: It does not repair biased sampling, confounding or an incorrect chance model.

Accept equivalent correct methods and explanations. This is a Refresh Kid teaching rubric, not an official AP scoring guideline.

Recall the ideas without notes Review →

Retrieve it before you reveal it.

RECALL 1What is the main cost of increasing α?

A higher allowed Type I error rate.

RECALL 2Why can larger n increase power?

It reduces sampling noise relative to a fixed effect.

RECALL 3Why can discrete power curves have steps?

Only integer success-count cutoffs are possible.

Revisit these tomorrow and a week later. Try a fresh problem and explain why the method applies.

How can a study become better at detecting a difference?

  • Power=1−β.
  • Larger n and larger effect generally increase power.
  • Larger α trades more false alarms for more sensitivity.

Remember: Power is a pre-study operating characteristic under a specified truth, not the probability that a rejected finding is real.

Conditions: Greater-than normal z-test of p₀=.50. Exact binomial enumeration evaluates the finite-sample rejection rate of that approximate test rule under the null and selected alternative. Its actual false-alarm rate need not equal nominal α because counts are discrete.

Refresh Kid · AP Statistics Unit 3 · Objectives 3.8.B, 3.8.C, 3.8.D · Review edition

Framework, scope and review status

Mapped to College Board, AP Statistics CED, Topic 3.8, objectives 3.8.B, 3.8.C, 3.8.D. Framework effective Fall 2026, checked September 17, 2026. Unit 3 includes inference for one and two population proportions, errors and power, and chi-square homogeneity/independence tests; it is part of the revised five-unit course.

Examples and datasets are synthetic, independently authored teaching material. Inference requires a justified design and appropriate counts. Intervals use observed proportions; null tests use their reference-model proportions. The revised CED states chi-square expected counts should be greater than 5; this unit follows that wording even though some companion texts use at least 5. Normal and chi-square inference are approximate. The chi-square goodness-of-fit test is not included in this unit’s official scope.

The Organic Chemistry Tutor companion title and destination were located; the full video was not reviewed. Khan Academy’s destination was checked, but its lesson content was not fully readable by the research tool. OpenStax provides optional reference reading. No provider scripts, questions or graphics were copied. Refresh Kid is not affiliated with these providers.

GitHub’s 3D website collection informed optional spatial inspection. Our original categorical count grid uses self-hosted Three.js with its MIT license. Two rows and three columns organize synthetic people into categorical cells. Each stacked block represents one person. Camera rotation changes only the view; exact counts, expectations and contributions are always available in the 2D table. Inference curves and intervals remain 2D; use the exact table rather than apparent 3D size for comparisons. Complete labeled diagrams, count tables and explanations remain available without 3D.

Independent teacher review and observation of students remain pending. Technical checks do not certify statistical accuracy, accessibility or learning effectiveness. This is a review edition.

Released AP Statistics questions and scoring guides are optional. Older exams use the earlier framework, so check alignment before selecting parts. All practice on this page is original, not official AP material.

Learn → Explore → Practice → Review is informed by the IES learning guide. This implementation has not yet been evaluated with learners.

OPTIONAL LIVE SUPPORT

Want to work through this with a tutor?

Bring your question about How can a study become better at detecting a difference? Your explanation and answers remain free to access.

Request a statistics tutor →Ask about this lesson on WhatsAppThe team can confirm teacher availability and next steps.