Learning
LESSON 10 / 20 · TOPIC 4.4

Which evidence makes a t procedure defensible?

You will be able to: Justify design and distribution conditions rather than ticking a sample-size box.

Graphs, tables and mathematical reasoningFree study resourceReview editionTeacher review pending

Which evidence makes a t procedure defensible?

A mean wait time comes from 12 volunteer responses with one extremely long wait. Calculating t is easy; defending it is harder.

A useful starting point: How does subtraction order control a paired claim? →

Words and symbols before equations

Randomization
Random sampling or randomized treatment assignment, supporting different scopes of inference.
Sample-data condition
A shape/size check supporting the t approximation.
Outlier
An observation far from the general pattern.
Scope of inference
The population or causal conclusion the design can justify.
Null t distribution: shaded p-value-5-4-3-2-10123450.00.10.20.30.4Density (fixed scale); teal reference, gray dashed normalStandardized value; numeric probability includes tails beyond ±5
Read this model snapshot. t=2, df=24, p=0.05694. At α=0.05, fail to reject H₀. Chosen design, sampling fraction where applicable, and shape/size assumptions support the procedure.
What this picture assumes

Synthetic summaries, with positive sample SDs. Random samples without replacement use source populations of N=100000 each, satisfying 10% for these sizes. A randomized experiment does not require that sampling-fraction check. Design and shape are assumptions chosen here, not conclusions that summary statistics can verify. Strong skewness/outlier selection conservatively withholds inference at all sizes; real data require examination.

Read the picture in three steps

  1. Read the axes and labels first. Identify what each symbol and line represents. Read the units and fixed conditions before comparing quantities.
  2. t=2, df=24, p=0.05694. At α=0.05, fail to reject H₀. Chosen design, sampling fraction where applicable, and shape/size assumptions support the procedure.
  3. Check what the picture assumes below. Use the Explore task to predict one change before moving a control.

Connect the picture to the mathematics

A one-sample or paired t procedure needs justified independent units. A random sample supports generalization; randomized assignment supports causal comparison in a suitable experiment. Convenience responses supply neither by themselves.

For sampling without replacement, n≤0.10N supports ignoring dependence. This is unnecessary as a sampling-fraction check in a randomized experiment that does not sample a population without replacement.

A normal population, a suitably large sample, or a small sample free of strong skewness and outliers can support the method. Extreme outliers or skewness warrant caution even beyond n=30. For paired data, apply shape and sample size to differences.

A worked example, step by step

Assess n=12 volunteer responses with one extreme outlier.

  1. The responses are voluntary, not a stated random sample.
  2. Population generalization may be biased.
  3. At n=12, the extreme outlier undermines the small-sample shape condition.
  4. A routine population t test is not justified; improve collection and examine the unusual observation rather than silently deleting it.
Common mix-up

A nonsignificant result cannot rescue an unsuitable design, and a significant result cannot validate it.

CHECK THE IDEA

Can an outlier be deleted just to make p smaller?

Compare with an explanation

No. Investigate its origin and use a justified, transparent analysis.

Now investigate one change Explore →

Predict. Change one thing. Explain.

Toggle the collection and shape assumptions. Find a case with n≥30 where the randomization check still fails and explain why inference is withheld.

On narrow screens, swipe or scroll diagrams sideways to read all labels.

Null t distribution: shaded p-value-5-4-3-2-10123450.00.10.20.30.4Density (fixed scale); teal reference, gray dashed normalStandardized value; numeric probability includes tails beyond ±5

t=2, df=24, p=0.05694. At α=0.05, fail to reject H₀. Chosen design, sampling fraction where applicable, and shape/size assumptions support the procedure.

QuantityValue
Estimate (minutes)22
SE (minutes)1
Degrees of freedom24

Synthetic summaries, with positive sample SDs. Random samples without replacement use source populations of N=100000 each, satisfying 10% for these sizes. A randomized experiment does not require that sampling-fraction check. Design and shape are assumptions chosen here, not conclusions that summary statistics can verify. Strong skewness/outlier selection conservatively withholds inference at all sizes; real data require examination.

Explain what you noticed: Answer the investigation prompt above. State one observation and explain it using the means, standard errors, pairing, graph scales or model assumptions. Identify what the representation cannot tell you.

Apply the idea to a fresh problem Practice →

Show what you understand.

Two original questions are a starting check, not proof of mastery. Explain your choice before revealing the answer.

1. A convenience sample with n=100 passes all conditions?

Show answer and reasoning

No. Sample size cannot establish random sampling.

2. For pairs, inspect…

Show answer and reasoning

Difference distribution. The difference is the analyzed response.

Original written challenge

4 points · self-check · not an official AP question

An SRS of 20 from N=500 has no strong skewness or outliers. Explain the conditions and one remaining limitation.

This response is not submitted or saved. Copy it before leaving.

Compare with the answer and four-point rubric
  1. 1 point: Random selection is stated.
  2. 1 point: 20≤50 satisfies 10%.
  3. 1 point: The small sample’s described shape supports a t approximation.
  4. 1 point: It remains an approximation and does not establish causation or remove all uncertainty.

Accept equivalent correct methods and explanations. This is a Refresh Kid teaching rubric, not an official AP scoring guideline.

Recall the ideas without notes Review →

Retrieve it before you reveal it.

RECALL 1What does random sampling support?

Population generalization.

RECALL 2What does random assignment support?

A causal comparison under a suitable experiment.

RECALL 3Does n≥30 guarantee good data?

No; design and unusual distributions still matter.

Revisit these tomorrow and a week later. Try a fresh problem and explain why the method applies.

Which evidence makes a t procedure defensible?

  • Randomization + justified independence + suitable shape/size.
  • Without replacement: n≤0.10N.
  • For pairs, n and shape refer to differences.

Remember: A nonsignificant result cannot rescue an unsuitable design, and a significant result cannot validate it.

Conditions: Synthetic summaries, with positive sample SDs. Random samples without replacement use source populations of N=100000 each, satisfying 10% for these sizes. A randomized experiment does not require that sampling-fraction check. Design and shape are assumptions chosen here, not conclusions that summary statistics can verify. Strong skewness/outlier selection conservatively withholds inference at all sizes; real data require examination.

Refresh Kid · AP Statistics Unit 4 · Objectives 4.4.C · Review edition

Framework, scope and review status

Mapped to College Board, AP Statistics CED, Topic 4.4, objectives 4.4.C. Framework effective Fall 2026, checked September 17, 2026. Unit 4 includes sampling distributions of means, one-sample and paired t inference, and independent two-sample t inference; it is part of the revised five-unit course.

Examples and datasets are synthetic, independently authored teaching material. Mean inference requires a justified design and suitable shape or sample size. Paired analysis uses one sample of differences. Independent two-sample inference uses separate variance estimates and technology-computed Welch degrees of freedom. Extreme skewness and influential observations need attention even in larger samples. This model conservatively withholds inference when those warnings are selected. Conclusions are limited by random sampling and/or assignment as appropriate.

The Organic Chemistry Tutor companion title and destination were located; the full video was not reviewed. Khan Academy’s destination was checked, but its lesson content was not fully readable by the research tool. OpenStax provides optional reference reading. No provider scripts, questions or graphics were copied. Refresh Kid is not affiliated with these providers.

GitHub’s 3D website collection informed optional spatial inspection. Our original paired-data display uses self-hosted Three.js with its MIT license. Two measurement columns are connected within each labeled student lane; horizontal position is before/after, vertical position is time, and depth separates identities rather than representing a numerical variable. Rotation can separate overlapping connectors. Exact values and differences always remain in the 2D table and labeled plot. The broken-matching option is an explicit counterexample, not a legitimate alternative analysis. No autoplay; complete teaching remains available without 3D.

Independent teacher review and observation of students remain pending. Technical checks do not certify statistical accuracy, accessibility or learning effectiveness. This is a review edition.

Released AP Statistics questions and scoring guides are optional. Older exams use the earlier framework, so check alignment before selecting parts. All practice on this page is original, not official AP material.

Learn → Explore → Practice → Review is informed by the IES learning guide. This implementation has not yet been evaluated with learners.

OPTIONAL LIVE SUPPORT

Want to work through this with a tutor?

Bring your question about Which evidence makes a t procedure defensible? Your explanation and answers remain free to access.

Request a statistics tutor →Ask about this lesson on WhatsAppThe team can confirm teacher availability and next steps.