Why can a huge survey still mislead?
You will be able to: Identify a bias mechanism and explain a plausible direction in context.
Why can a huge survey still mislead?
An online poll about school buses is posted only in the cycling club. A thousand responses could still miss the experiences of regular bus riders.
A useful starting point: Sample some from every group, or everyone from some groups? →
Words and symbols before equations
- Bias
- A systematic tendency for a procedure to overestimate or underestimate a target.
- Undercoverage
- Part of the target population is missing or less likely to enter the frame.
- Nonresponse
- Selected people do not supply usable responses.
- Response bias
- Recorded answers systematically differ from the truth.
- Voluntary response
- People select themselves into a survey.
What this picture assumes
Toy population: 10 early riders each wait 2 minutes and 10 late riders each wait 10 minutes. All early riders respond. This deliberately constructed example isolates differential nonresponse.
Read the picture in three steps
- Read the axes and labels first. Identify what each symbol and line represents. Read the units and fixed conditions before comparing quantities.
- 12 respondents: 10 early and 2 late. Mean 3.333 minutes versus population 6; difference -2.667 minutes.
- Check what the picture assumes below. Use the Explore task to predict one change before moving a control.
Connect the picture to the mathematics
Undercoverage occurs before or during selection: a cycling-club frame misses many bus users. Nonresponse occurs after selection: chosen bus users might not answer. These are different stages.
Leading wording, recall difficulty or pressure to appear favorable can cause response bias. A public link can attract people with unusually strong opinions, creating voluntary-response bias.
Explain direction only when the context supports it. If frequent bus users wait longer and are underrepresented, the mean wait may be too low. Increasing sample size within the same flawed frame reduces random variability but does not fix the mechanism.
A worked example, step by step
In a toy population, 10 early riders wait 2 minutes each and 10 late riders wait 10 minutes each. All early riders respond but only 2 late riders do. Find both means.
- Population total is 10×2+10×10=120 minutes across 20 people.
- The population mean is 6 minutes.
- Respondents contribute 10×2+2×10=40 minutes across 12 people, so their mean is 3.33 minutes.
- Differential nonresponse underrepresents long waits, pushing the respondent mean below the target mean.
Low response alone does not prove a specific direction; explain how nonrespondents differ on the variable being studied.
Does a random invitation guarantee unbiased responses?
Compare with an explanation
No. Coverage, nonresponse and measurement can still undermine a randomly selected sample.
Predict. Change one thing. Explain.
Change the number of late-rider respondents from 0 to 10. Compare respondent and population means. Explain why collecting more early riders would not solve the imbalance.
On narrow screens, swipe or scroll diagrams sideways to read all labels.
12 respondents: 10 early and 2 late. Mean 3.333 minutes versus population 6; difference -2.667 minutes.
| Group | Population | Respondents | Wait per person |
|---|---|---|---|
| Early | 10 | 10 | 2 min |
| Late | 10 | 2 | 10 min |
Toy population: 10 early riders each wait 2 minutes and 10 late riders each wait 10 minutes. All early riders respond. This deliberately constructed example isolates differential nonresponse.
Explain what you noticed: Answer the investigation prompt above. State one observation and explain it using the data values, graph scales, summary statistics or study-design conditions. Identify what the representation cannot tell you.
Apply the idea to a fresh problem Practice →Show what you understand.
Two original questions are a starting check, not proof of mastery. Explain your choice before revealing the answer.
Original written challenge
4 points · self-check · not an official AP questionA survey of all students uses only sports-team mailing lists. Athletes tend to exercise more. Identify the problem, likely direction and a better design.
This response is not submitted or saved. Copy it before leaving.
Compare with the answer and four-point rubric
- 1 point: The frame undercovers nonathletes.
- 1 point: Mean exercise time may be overestimated if athletes exercise more.
- 1 point: Use a complete student roster and an appropriate random selection procedure.
- 1 point: Follow up with selected nonrespondents and ask neutrally worded questions; a larger sports-list sample alone does not repair coverage.
Accept equivalent correct methods and explanations. This is a Refresh Kid teaching rubric, not an official AP scoring guideline.
Retrieve it before you reveal it.
RECALL 1What distinguishes nonresponse from undercoverage?
Selected but no answer versus missing or underrepresented in the frame.
RECALL 2Why can more responses fail to fix bias?
The same systematically unrepresentative mechanism can persist.
RECALL 3When should a bias direction be stated?
When contextual evidence supports why the estimate tends high or low.
Revisit these tomorrow and a week later. Try a fresh problem and explain why the method applies.
Why can a huge survey still mislead?
- Distinguish frame omissions, nonresponse and measurement bias.
- Larger n does not automatically remove bias.
- Support a predicted direction with a contextual mechanism.
Remember: Low response alone does not prove a specific direction; explain how nonrespondents differ on the variable being studied.
Conditions: Toy population: 10 early riders each wait 2 minutes and 10 late riders each wait 10 minutes. All early riders respond. This deliberately constructed example isolates differential nonresponse.
Refresh Kid · AP Statistics Unit 1 · Objectives 1.12.A · Review edition
Framework, scope and review status
Mapped to College Board, AP Statistics CED, Topic 1.12, objectives 1.12.A. Framework effective Fall 2026, checked September 17, 2026. Unit 1 includes one-variable data and data collection; it is part of the revised five-unit course.
Examples and datasets are synthetic, independently authored teaching material. Quartile calculations state a median-of-halves convention. Outlier screens identify values to investigate, not data to discard. Random selection and random assignment have different inferential roles. These lessons introduce design and descriptive reasoning; formal inference comes in later units.
The Organic Chemistry Tutor companion title and destination were checked; the full video was not reviewed. Khan Academy’s destination was checked, but its lesson content was not fully readable by the research tool. OpenStax provides optional reference reading. No provider scripts, questions or graphics were copied. Refresh Kid is not affiliated with these providers.
GitHub’s 3D website collection informed optional spatial inspection. Our original sampling model uses self-hosted Three.js with its MIT license. It shows labeled units in four groups; camera rotation does not change the sampling procedure. Quantitative graphs remain 2D to avoid perspective distortion. Complete labeled diagrams, selected IDs and explanations remain available without 3D.
Independent teacher review and observation of students remain pending. Technical checks do not certify statistical accuracy, accessibility or learning effectiveness. This is a review edition.
Released AP Statistics questions and scoring guides are optional. Older exams use the earlier framework, so check alignment before selecting parts. All practice on this page is original, not official AP material.
Learn → Explore → Practice → Review is informed by the IES learning guide. This implementation has not yet been evaluated with learners.
Want to work through this with a tutor?
Bring your question about Why can a huge survey still mislead? Your explanation and answers remain free to access.
