What this deck covers
This deck focuses on Introducing Statistics Learning From Data, giving you a quick way to review the definitions, rules, and examples that matter most for AP Statistics.
Study Introducing Statistics Learning From Data in AP Statistics with focused flashcards that help you recognize the idea, recall the key rule, and apply it in practice-style prompts.
0% Complete
What does a correlation coefficient of 0 indicate?
Tap card or press Space to flip
A correlation coefficient of 0 indicates no linear relationship between variables. The variables move independently with no linear pattern.
How well did you know it?
Card 1 / 69
Space to flip · ← / → to move · once flipped, → Got it · ← Still learning
This deck focuses on Introducing Statistics Learning From Data, giving you a quick way to review the definitions, rules, and examples that matter most for AP Statistics.
Work through these flashcards in short sessions. Try to answer each prompt before flipping the card, then revisit any cards you miss until the explanation feels automatic.
Answer: A correlation coefficient of 0 indicates no linear relationship between variables. The variables move independently with no linear pattern.
Answer: The empirical rule states that 68%, 95%, and 99.7% of data fall within 1, 2, and 3 standard deviations from the mean, respectively. Applies specifically to normal distributions with predictable percentages.
Answer: Standard deviation is the square root of the variance. Provides a measure of spread in the same units as the data.
Answer: Variance = n−1sum of squared deviations from mean. Uses n−1 for sample variance to correct for bias.
Answer: A residual is the difference between an observed value and the value predicted by a model. Shows how far off the model's prediction was from reality.
Answer: Standard deviation is the square root of the variance. Provides a measure of spread in the same units as the data.
Answer: A z-score measures how many standard deviations an element is from the mean. Standardizes values for comparison across different datasets.
Answer: Inferential statistics make predictions or inferences about a population based on sample data. Uses sample data to draw conclusions about the entire population.
Answer: Two events are independent if the occurrence of one does not affect the probability of the other. Knowledge of one event provides no information about the other.
Answer: A normal distribution is a bell-shaped distribution that is symmetric about the mean. The classic bell curve with equal tails on both sides.
Answer: A confidence interval is a range of values that is likely to contain a population parameter. Provides a range estimate with a specified level of confidence.
Answer: Range measures the difference between the maximum and minimum values. Simple measure of spread calculated as max - min.
Answer: A sample is a subset of the population from which data is actually collected. A smaller group selected from the population for practical data collection.
Answer: A residual plot is used to assess the fit of a regression model. Helps identify patterns in errors and model appropriateness.
Answer: The significance level is the probability of a Type I error, typically denoted by alpha. Sets the threshold for rejecting the null hypothesis.
Answer: Variance = n−1sum of squared deviations from mean. Uses n−1 for sample variance to correct for bias.
Answer: A z-score measures how many standard deviations an element is from the mean. Standardizes values for comparison across different datasets.
Answer: Line of best fit: y=mx+b, where m is the slope, b is the y-intercept. Standard linear equation form for the regression line.
Answer: Two events are mutually exclusive if they cannot occur at the same time. No overlap possible; one event prevents the other.
Answer: A p-value indicates the probability of observing the test results under the null hypothesis. Measures how likely our results are if the null hypothesis is true.
Answer: The mode is the value that appears most frequently in a data set. Identifies the most common or repeated value in the dataset.
Answer: A confidence interval is a range of values that is likely to contain a population parameter. Provides a range estimate with a specified level of confidence.
Answer: A histogram is a graphical representation of the distribution of numerical data. Shows frequency of data values using bars or bins.
Answer: A box plot visually displays the distribution and identifies outliers of a data set. Shows five-number summary and highlights unusual values.
Answer: An outlier is a data point that differs significantly from other observations. An unusual value that stands apart from the typical pattern.
Answer: The mode is the value that appears most frequently in a data set. Identifies the most common or repeated value in the dataset.
Answer: Relative frequency is calculated as the frequency of an event divided by the total number of observations. Converts counts to proportions for easier comparison.
Answer: A sample is a subset of the population from which data is actually collected. A smaller group selected from the population for practical data collection.
Answer: A categorical variable is a variable that represents categories or groups. Takes on distinct labels or names rather than numerical values.
Answer: Extrapolation involves predicting beyond the range of observed data. Making predictions outside the observed data range.
Answer: To summarize or describe the characteristics of a data set. Focus is on organizing and presenting data, not making predictions.
Answer: Probability = P(A) × P(B) for independent events A and B. Multiply individual probabilities when events don't influence each other.
Answer: The central limit theorem states that the sampling distribution of the sample mean approaches a normal distribution as the sample size increases. Sample means become normally distributed regardless of population shape.
Answer: A continuous variable can take an infinite number of values within a given range. Can be measured to any desired precision within its range.
Answer: A Type I error occurs when the null hypothesis is incorrectly rejected. A false positive; rejecting a true null hypothesis.
Answer: A residual is the difference between an observed value and the value predicted by a model. Shows how far off the model's prediction was from reality.
Answer: The median is the middle value when data points are arranged in order. Half the values are above and half are below this central value.
Answer: To summarize or describe the characteristics of a data set. Focus is on organizing and presenting data, not making predictions.
Answer: The median is the middle value when data points are arranged in order. Half the values are above and half are below this central value.
Answer: Line of best fit: y=mx+b, where m is the slope, b is the y-intercept. Standard linear equation form for the regression line.
Answer: A histogram is a graphical representation of the distribution of numerical data. Shows frequency of data values using bars or bins.
Answer: A box plot visually displays the distribution and identifies outliers of a data set. Shows five-number summary and highlights unusual values.
Answer: A residual plot is used to assess the fit of a regression model. Helps identify patterns in errors and model appropriateness.
Answer: A contingency table displays the frequency distribution of variables. Cross-tabulates two categorical variables to show relationships.
Answer: Skewness describes asymmetry in the distribution of values in a data set. Indicates whether data leans left or right from center.
Answer: A null hypothesis is a statement that there is no effect or no difference, used as a starting point. Assumes no relationship exists; the default position to test against.
Answer: A contingency table displays the frequency distribution of variables. Cross-tabulates two categorical variables to show relationships.
Answer: A sampling distribution is the probability distribution of a statistic obtained from a large number of samples. Shows how sample statistics vary across repeated sampling.
Answer: Relative frequency is calculated as the frequency of an event divided by the total number of observations. Converts counts to proportions for easier comparison.
Answer: A null hypothesis is a statement that there is no effect or no difference, used as a starting point. Assumes no relationship exists; the default position to test against.
Answer: The correlation coefficient measures the strength and direction of a linear relationship. Ranges from -1 to +1, indicating weak to strong relationships.
Answer: Skewness describes asymmetry in the distribution of values in a data set. Indicates whether data leans left or right from center.
Answer: Range measures the difference between the maximum and minimum values. Simple measure of spread calculated as max - min.
Answer: A normal distribution is a bell-shaped distribution that is symmetric about the mean. The classic bell curve with equal tails on both sides.
Answer: The correlation coefficient measures the strength and direction of a linear relationship. Ranges from -1 to +1, indicating weak to strong relationships.
Answer: A p-value indicates the probability of observing the test results under the null hypothesis. Measures how likely our results are if the null hypothesis is true.
Answer: A scatter plot is used to show the relationship between two quantitative variables. Each point represents one observation with two measurements.
Answer: Mean = number of data pointssum of all data points. Add all values and divide by the count to find the average.
Answer: A categorical variable is a variable that represents categories or groups. Takes on distinct labels or names rather than numerical values.
Answer: A correlation coefficient of 0 indicates no linear relationship between variables. The variables move independently with no linear pattern.
Answer: An outlier is a data point that differs significantly from other observations. An unusual value that stands apart from the typical pattern.
Answer: The significance level is the probability of a Type I error, typically denoted by alpha. Sets the threshold for rejecting the null hypothesis.
Answer: Inferential statistics make predictions or inferences about a population based on sample data. Uses sample data to draw conclusions about the entire population.
Answer: Extrapolation involves predicting beyond the range of observed data. Making predictions outside the observed data range.
Answer: A population is the entire group of individuals or instances about whom we hope to learn. The complete group we want to study and make conclusions about.
Answer: A Type II error occurs when the null hypothesis is not rejected when it is false. A false negative; failing to reject a false null hypothesis.
Answer: A continuous variable can take an infinite number of values within a given range. Can be measured to any desired precision within its range.
Answer: The empirical rule states that 68%, 95%, and 99.7% of data fall within 1, 2, and 3 standard deviations from the mean, respectively. Applies specifically to normal distributions with predictable percentages.
Answer: A scatter plot is used to show the relationship between two quantitative variables. Each point represents one observation with two measurements.