Refresh KidLearning
LESSON 09 / 23 · TOPIC 3.6

What probability does a p-value measure?

You will be able to: Interpret tail probability conditional on the null model.

Graphs, tables and mathematical reasoningFree study resourceReview editionTeacher review pending

What probability does a p-value measure?

A poll result is far above a 50% benchmark. The question is how often a result at least that extreme would occur if the benchmark and sampling model were correct.

A useful starting point: Why does a test use expected counts under the null? →

Words and symbols before equations

p-value
Probability under H₀ of a statistic at least as extreme in the direction of Hₐ as the observed statistic.
Tail
A region of unusually low or high values.
Evidence
How incompatible the observed statistic is with the specified null model.
Null z distribution: shaded p-valueRelative density height; probability is area-4-2.67-1.3301.332.674Center 0; SD 1; finite display, full-tail calculation
Read this model snapshot. z=2; p-value=0.02275. At α=0.05, reject H₀. Alternative: p > 0.5. This does not give P(H₀).
What this picture assumes

Synthetic study: a simple random sample without replacement from N=100000 independent units. The chosen sizes satisfy the 10% condition. Real studies require their own design checks. Actual count x is rounded from the requested percentage. Normal z inference is withheld when either null expected count is below 10.

Read the picture in three steps

  1. Read the axes and labels first. Identify what each symbol and line represents. Read the units and fixed conditions before comparing quantities.
  2. z=2; p-value=0.02275. At α=0.05, reject H₀. Alternative: p > 0.5. This does not give P(H₀).
  3. Check what the picture assumes below. Use the Explore task to predict one change before moving a control.

Connect the picture to the mathematics

For Hₐ:p>p₀, use the upper tail beyond observed z. For Hₐ:p<p₀, use the lower tail. For a symmetric normal two-sided test use 2P(Z≥

z

).

If a greater-than test has z=2, the approximate p-value is .0228: under the null and assumptions, about 2.28% of repeated statistics would be at least this large.

A p-value is not P(H₀ is true), the probability results arose “by chance,” or an effect size. A large value means insufficient evidence against H₀, not proof of H₀. Small values must still be interpreted with design quality and practical consequences.

A worked example, step by step

A two-sided normal test has observed z=−2. Interpret its p-value.

  1. A two-sided alternative treats unusually positive and negative statistics as evidence.
  2. Use both tails beyond
  3. z
  4. =2.
  5. Each normal tail is about .02275, so p≈.0455.
  6. Assuming the null model, about 4.55% of repeated statistics are at least this far from zero in either direction.
Common mix-up

The condition “assuming H₀ is true” belongs in every interpretation.

CHECK THE IDEA

Does p=.30 establish H₀?

Compare with an explanation

No. It says the observed statistic is not sufficiently unusual under the null to provide strong evidence against it.

Now investigate one change Explore →

Predict. Change one thing. Explain.

Compare greater, less and different alternatives for the same positive z. Explain the shaded tails and why one-sided probabilities sum to 1.

On narrow screens, swipe or scroll diagrams sideways to read all labels.

Null z distribution: shaded p-valueRelative density height; probability is area-4-2.67-1.3301.332.674Center 0; SD 1; finite display, full-tail calculation

z=2; p-value=0.02275. At α=0.05, reject H₀. Alternative: p > 0.5. This does not give P(H₀).

QuantityValue
Observed successes x60
Observed failures n−x40
Actual p̂=x/n0.6
Null quantityValue
Expected successes50
Expected failures50
SE under H₀0.05

Synthetic study: a simple random sample without replacement from N=100000 independent units. The chosen sizes satisfy the 10% condition. Real studies require their own design checks. Actual count x is rounded from the requested percentage. Normal z inference is withheld when either null expected count is below 10.

Explain what you noticed: Answer the investigation prompt above. State one observation and explain it using the probability values, reference groups, graph scales or model assumptions. Identify what the representation cannot tell you.

Apply the idea to a fresh problem Practice →

Show what you understand.

Two original questions are a starting check, not proof of mastery. Explain your choice before revealing the answer.

1. A p-value is calculated assuming…

Show answer and reasoning

The null model. It is a conditional probability under H₀.

2. For z=2, a greater-tail p≈.0228 implies two-sided p≈…

Show answer and reasoning

.0456. The symmetric two-sided test includes both equal tails.

Original written challenge

4 points · self-check · not an official AP question

A test of p=.40 versus p>.40 gives p-value=.018. Interpret it without assigning a probability to H₀.

This response is not submitted or saved. Copy it before leaving.

Compare with the answer and four-point rubric
  1. 1 point: Assume the population share is .40 and sampling assumptions hold.
  2. 1 point: Consider repeated samples of the same size.
  3. 1 point: About 1.8% would yield a test statistic at least as large as observed.
  4. 1 point: This provides evidence for p>.40, not a 1.8% probability that H₀ is true.

Accept equivalent correct methods and explanations. This is a Refresh Kid teaching rubric, not an official AP scoring guideline.

Recall the ideas without notes Review →

Retrieve it before you reveal it.

RECALL 1What makes a result more extreme?

The direction and distance specified by the alternative.

RECALL 2What does a large p-value establish?

It does not establish H₀; it indicates limited evidence against it.

RECALL 3Does a p-value measure effect size?

No.

Revisit these tomorrow and a week later. Try a fresh problem and explain why the method applies.

What probability does a p-value measure?

  • Greater: 1−Φ(z). Less: Φ(z).
  • Two-sided normal: 2[1−Φ(
  • z
  • )].
  • A p-value is not the probability of a hypothesis.

Remember: The condition “assuming H₀ is true” belongs in every interpretation.

Conditions: Synthetic study: a simple random sample without replacement from N=100000 independent units. The chosen sizes satisfy the 10% condition. Real studies require their own design checks. Actual count x is rounded from the requested percentage. Normal z inference is withheld when either null expected count is below 10.

Refresh Kid · AP Statistics Unit 3 · Objectives 3.6.A · Review edition

Framework, scope and review status

Mapped to College Board, AP Statistics CED, Topic 3.6, objectives 3.6.A. Framework effective Fall 2026, checked September 17, 2026. Unit 3 includes inference for one and two population proportions, errors and power, and chi-square homogeneity/independence tests; it is part of the revised five-unit course.

Examples and datasets are synthetic, independently authored teaching material. Inference requires a justified design and appropriate counts. Intervals use observed proportions; null tests use their reference-model proportions. The revised CED states chi-square expected counts should be greater than 5; this unit follows that wording even though some companion texts use at least 5. Normal and chi-square inference are approximate. The chi-square goodness-of-fit test is not included in this unit’s official scope.

The Organic Chemistry Tutor companion title and destination were located; the full video was not reviewed. Khan Academy’s destination was checked, but its lesson content was not fully readable by the research tool. OpenStax provides optional reference reading. No provider scripts, questions or graphics were copied. Refresh Kid is not affiliated with these providers.

GitHub’s 3D website collection informed optional spatial inspection. Our original categorical count grid uses self-hosted Three.js with its MIT license. Two rows and three columns organize synthetic people into categorical cells. Each stacked block represents one person. Camera rotation changes only the view; exact counts, expectations and contributions are always available in the 2D table. Inference curves and intervals remain 2D; use the exact table rather than apparent 3D size for comparisons. Complete labeled diagrams, count tables and explanations remain available without 3D.

Independent teacher review and observation of students remain pending. Technical checks do not certify statistical accuracy, accessibility or learning effectiveness. This is a review edition.

Released AP Statistics questions and scoring guides are optional. Older exams use the earlier framework, so check alignment before selecting parts. All practice on this page is original, not official AP material.

Learn → Explore → Practice → Review is informed by the IES learning guide. This implementation has not yet been evaluated with learners.

OPTIONAL LIVE SUPPORT

Want to work through this with a tutor?

Bring your question about What probability does a p-value measure? Your explanation and answers remain free to access.

Request a statistics tutor →Ask about this lesson on WhatsAppThe team can confirm teacher availability and next steps.