How variable is the difference between two sample shares?
You will be able to: Calculate and interpret the sampling distribution of p̂₁−p̂₂.
How variable is the difference between two sample shares?
Two large schools have true support shares .60 and .40. Independent random samples of 100 from each will produce a difference that varies around .20.
A useful starting point: How can a study become better at detecting a difference? →
Words and symbols before equations
- Difference statistic D
- p̂₁−p̂₂, with group order fixed.
- Independent groups
- The first sample’s outcomes do not determine the second’s.
- Variance
- Squared sampling spread; independent variances add.
- Percentage-point difference
- A difference of .20 in shares is 20 percentage points.
What this picture assumes
Synthetic study: independent SRSs from two separate populations, each N=100000. The chosen sizes satisfy both 10% conditions. Paired responses require a different method. Theoretical spread uses known population proportions. Four expected counts must be at least 10 for this normal approximation.
Read the picture in three steps
- Read the axes and labels first. Identify what each symbol and line represents. Read the units and fixed conditions before comparing quantities.
- Mean difference 0.2; SD 0.06928. Approximate P(D ≥ 0.3)=0.07446, without continuity correction.
- Check what the picture assumes below. Use the Explore task to predict one change before moving a control.
Connect the picture to the mathematics
The mean of D is p₁−p₂. Independence gives SD(D)=√[p₁(1−p₁)/n₁+p₂(1−p₂)/n₂]. Variances add even though the statistic subtracts.
For .60 and .40 with n₁=n₂=100, the center is .20 and SD=√.0048≈.06928. A normal approximation estimates P(D>.30)≈.0745, using z≈1.443.
Check independent random samples, the 10% rule for each population when sampling without replacement, and all four expected counts at least 10. A randomized experiment has a different design justification; random assignment supports a treatment comparison but does not itself establish population representativeness.
A worked example, step by step
For independent samples with p₁=.50,p₂=.50,n₁=n₂=200, find center and SD of the difference.
- Fix the order as group 1 minus group 2.
- Mean difference=.50−.50=0.
- Variance=.25/200+.25/200=.0025.
- SD=.05, so repeated differences fluctuate on a scale of 5 percentage points around zero.
Do not subtract variances, and do not use independent-group formulas for paired responses from the same people.
Why add variances for a difference?
Compare with an explanation
Independent fluctuations from either group can increase uncertainty in the difference.
Predict. Change one thing. Explain.
Change the two sample sizes separately. Explain why increasing only one sample leaves the other group’s variance contribution.
On narrow screens, swipe or scroll diagrams sideways to read all labels.
Mean difference 0.2; SD 0.06928. Approximate P(D ≥ 0.3)=0.07446, without continuity correction.
| Quantity | Value |
|---|---|
| Mean difference | 0.2 |
| SD difference | 0.06928 |
| Four expected counts | 60, 40, 40, 60 |
| Normal-count check | Pass |
Synthetic study: independent SRSs from two separate populations, each N=100000. The chosen sizes satisfy both 10% conditions. Paired responses require a different method. Theoretical spread uses known population proportions. Four expected counts must be at least 10 for this normal approximation.
Explain what you noticed: Answer the investigation prompt above. State one observation and explain it using the probability values, reference groups, graph scales or model assumptions. Identify what the representation cannot tell you.
Apply the idea to a fresh problem Practice →Show what you understand.
Two original questions are a starting check, not proof of mastery. Explain your choice before revealing the answer.
Original written challenge
4 points · self-check · not an official AP questionFor p₁=.60,p₂=.40 and n₁=n₂=100, calculate the difference’s mean and SD and check expected counts.
This response is not submitted or saved. Copy it before leaving.
Compare with the answer and four-point rubric
- 1 point: The mean is .20.
- 1 point: Variance=.0024+.0024=.0048.
- 1 point: SD≈.06928.
- 1 point: Expected counts are 60,40,40,60; all exceed 10. Independent random design and population-size checks still need justification.
Accept equivalent correct methods and explanations. This is a Refresh Kid teaching rubric, not an official AP scoring guideline.
Retrieve it before you reveal it.
RECALL 1What does a positive D mean?
Group 1’s sample share exceeds group 2’s.
RECALL 2Which counts matter for a known-population sampling model?
n₁p₁,n₁(1−p₁),n₂p₂,n₂(1−p₂).
RECALL 3Does random assignment alone support all-population generalization?
No; sampling and assignment support different claims.
Revisit these tomorrow and a week later. Try a fresh problem and explain why the method applies.
How variable is the difference between two sample shares?
- Mean(D)=p₁−p₂.
- SD(D)=√[p₁(1−p₁)/n₁+p₂(1−p₂)/n₂].
- Check four expected counts and independent groups.
Remember: Do not subtract variances, and do not use independent-group formulas for paired responses from the same people.
Conditions: Synthetic study: independent SRSs from two separate populations, each N=100000. The chosen sizes satisfy both 10% conditions. Paired responses require a different method. Theoretical spread uses known population proportions. Four expected counts must be at least 10 for this normal approximation.
Refresh Kid · AP Statistics Unit 3 · Objectives 3.9.A, 3.9.B, 3.9.C · Review edition
Framework, scope and review status
Mapped to College Board, AP Statistics CED, Topic 3.9, objectives 3.9.A, 3.9.B, 3.9.C. Framework effective Fall 2026, checked September 17, 2026. Unit 3 includes inference for one and two population proportions, errors and power, and chi-square homogeneity/independence tests; it is part of the revised five-unit course.
Examples and datasets are synthetic, independently authored teaching material. Inference requires a justified design and appropriate counts. Intervals use observed proportions; null tests use their reference-model proportions. The revised CED states chi-square expected counts should be greater than 5; this unit follows that wording even though some companion texts use at least 5. Normal and chi-square inference are approximate. The chi-square goodness-of-fit test is not included in this unit’s official scope.
The Organic Chemistry Tutor companion title and destination were located; the full video was not reviewed. Khan Academy’s destination was checked, but its lesson content was not fully readable by the research tool. OpenStax provides optional reference reading. No provider scripts, questions or graphics were copied. Refresh Kid is not affiliated with these providers.
GitHub’s 3D website collection informed optional spatial inspection. Our original categorical count grid uses self-hosted Three.js with its MIT license. Two rows and three columns organize synthetic people into categorical cells. Each stacked block represents one person. Camera rotation changes only the view; exact counts, expectations and contributions are always available in the 2D table. Inference curves and intervals remain 2D; use the exact table rather than apparent 3D size for comparisons. Complete labeled diagrams, count tables and explanations remain available without 3D.
Independent teacher review and observation of students remain pending. Technical checks do not certify statistical accuracy, accessibility or learning effectiveness. This is a review edition.
Released AP Statistics questions and scoring guides are optional. Older exams use the earlier framework, so check alignment before selecting parts. All practice on this page is original, not official AP material.
Learn → Explore → Practice → Review is informed by the IES learning guide. This implementation has not yet been evaluated with learners.
Want to work through this with a tutor?
Bring your question about How variable is the difference between two sample shares? Your explanation and answers remain free to access.
