Data, Statistics & Probability

Study Sheet

Data, Statistics & Probability

Measures of center and spread, reading and making graphs, interpreting data, and finding the probability of events as fractions, decimals, and percents

Mean, Median, and Mode (Measures of Center)

Concept
The Rule
3456789

A measure of center is one number that summarizes a whole list of numbers. There are three:

  • Mean (the average): add up all the values, then divide by how many values there are.
  • Median: the middle value once the numbers are put in order. If there are two middle values, take their mean (add them and divide by 22).
  • Mode: the value that appears most often. A data set can have one mode, several modes, or no mode at all.

One outlier drags the mean but not the median.

Example
Worked Example: all three centers

Find the mean, median, and mode of 7, 3, 9, 3, 87,\ 3,\ 9,\ 3,\ 8.

Order first: 3, 3, 7, 8, 93,\ 3,\ 7,\ 8,\ 9.

Mean: 3+3+7+8+95=305=6\dfrac{3+3+7+8+9}{5} = \dfrac{30}{5} = \mathbf{6}.

Median: the middle of five ordered numbers is the 33rd one, so median =7= \mathbf{7}.

Mode: 33 appears twice, more than any other, so mode =3= \mathbf{3}.

Example
Worked Example: even number of values

Find the median of 10, 4, 6, 210,\ 4,\ 6,\ 2. Order them: 2, 4, 6, 102,\ 4,\ 6,\ 10. There are two middle values, 44 and 66, so the median is their mean: 4+62=5\dfrac{4+6}{2} = \mathbf{5}. Notice the median need not be a value in the list.

Concept
When Is Each One Most Useful?
  • Use the mean when the data are fairly even, with no extreme values.
  • Use the median when there is an outlier (a value far from the rest), because the median is not pulled toward it.
  • Use the mode for things you count or categories, like the most common shoe size or favorite color.
Example
Worked Example: effect of an outlier

Four friends have $55, $66, $77, and $88. Mean =264=$6.50= \dfrac{26}{4} = \$6.50; median =6+72=$6.50= \dfrac{6+7}{2} = \$6.50. Now a fifth friend joins with $100100 (an outlier). New mean =1265=$25.20= \dfrac{126}{5} = \$25.20, but the ordered list 5,6,7,8,1005,6,7,8,100 has median =$7= \mathbf{\$7}. The single big value pulled the mean way up, while the median barely moved. That is why the median better describes a “typical” friend here.

Tip

Tip: Always put the numbers in order before you find the median. Forgetting to order the data is the most common mistake in this whole topic.

Range (A Measure of Spread)

Concept
The Rule
3456789

The range tells how spread out the data are. It is the largest value minus the smallest value:

range=maximumminimum.\text{range} = \text{maximum} - \text{minimum}.

A small range means the values are close together; a large range means they are spread far apart.

One outlier drags the mean but not the median.

Example
Worked Example: finding the range

Test scores were 88, 95, 72, 90, 8188,\ 95,\ 72,\ 90,\ 81. The maximum is 9595 and the minimum is 7272, so the range =9572=23= 95 - 72 = \mathbf{23}.

Tip

Tip: Range is a single number, not an interval. Say “the range is 2323,” not “the range is 7272 to 9595.”

Reading and Making Graphs

Concept
The Rule

Different graphs are good for different jobs:

  • Bar graph: compares separate categories (favorite fruits, sports). Taller bar means larger amount.
  • Line graph: shows how something changes over time (temperature through the day). Look at whether the line rises, falls, or stays flat.
  • Circle (pie) graph: shows parts of one whole. All the slices together make 100%100\% of the total.
  • Dot plot / line plot: an ×\times or dot stacked above a number line for each data value; great for seeing repeats and the shape of the data.
  • Frequency table: lists each value or category next to how many times it occurs (its frequency).
Example
Worked Example: frequency table and dot plot

A class recorded how many pets each student has:

The frequency column tells us 4+6+3+2=154+6+3+2 = 15 students answered. The mode is 11 pet (highest frequency, 66). A dot plot would stack 44 dots above 00, 66 dots above 11, and so on. To find the total number of pets, multiply and add: 0(4)+1(6)+2(3)+3(2)=0+6+6+6=180(4)+1(6)+2(3)+3(2) = 0+6+6+6 = 18 pets, so the mean is 1815=1.2\dfrac{18}{15} = 1.2 pets.

Example
Worked Example: reading a circle graph

A circle graph of how 6060 students get to school shows Walk =25%= 25\%, Bus =50%= 50\%, Car =25%= 25\%. Because the whole circle is 100%100\% of 6060 students, the bus slice stands for 50%50\% of 60=3060 = 30 students, and each 25%25\% slice stands for 1515 students. The slices 25%+50%+25%=100%25\% + 50\% + 25\% = 100\%, as they must.

Tip

Tip: On a circle graph the percents always add to 100%100\%. On any graph, first read the title and labels so you know what the numbers mean before you compare them.

Interpreting Data from Tables and Graphs

Concept
The Rule

To interpret data, read carefully and answer the exact question asked. Good questions to ask yourself:

  • Which category is largest or smallest?
  • What is the total, or the difference between two values?
  • Is the data increasing, decreasing, or staying the same over time?
Example
Worked Example: interpreting a line graph

A plant's height was measured each week: Week 11: 22 cm, Week 22: 55 cm, Week 33: 66 cm, Week 44: 1010 cm.

The plant grew every week, so the line rises the whole time. The greatest growth was between Weeks 33 and 44: 106=410 - 6 = 4 cm. Total growth over the whole period was 102=810 - 2 = 8 cm.

Tip

Tip: Watch the scale. If a bar graph counts by 55s, a bar reaching the third line means 1515, not 33.

Basic Probability

Concept
The Rule

Probability measures how likely an event is. An outcome is one possible result; a favorable outcome is a result you are hoping for. When every outcome is equally likely,

P(event)=number of favorable outcomestotal number of outcomes.P(\text{event}) = \frac{\text{number of favorable outcomes}}{\text{total number of outcomes}}.
Example
Worked Example: rolling a number cube

A number cube has faces 1,2,3,4,5,61,2,3,4,5,6, so there are 66 equally likely outcomes. What is P(even)P(\text{even})? The favorable outcomes are 2,4,62,4,6, so 33 of them.

P(even)=36=12.P(\text{even}) = \frac{3}{6} = \frac{1}{2}.

What is P(5)P(\text{a }5)? Only one face works, so P=16P = \dfrac{1}{6}.

Tip

Tip: Every probability is a number between 00 and 11. If your answer is negative or bigger than 11, you made a mistake. Count the total before you count the favorable outcomes.

Probability as a Fraction, Decimal, and Percent

Concept
The Rule

A probability can be written three equivalent ways. Start with the fraction, divide to get the decimal, then multiply by 100100 to get the percent. On the “likelihood” line:

  • P=0P = 0 means the event is impossible.
  • P=1P = 1 (that is, 100%100\%) means it is certain.
  • P=12P = \tfrac{1}{2} (50%50\%) means it is equally likely to happen or not. Larger than 12\tfrac12 is likely; smaller is unlikely.
Example
Worked Example: three forms

A bag has 2020 marbles, and 55 are red. Then

P(red)=520=14=0.25=25%.P(\text{red}) = \frac{5}{20} = \frac{1}{4} = 0.25 = 25\%.

Since 25%25\% is less than 50%50\%, drawing red is unlikely. Drawing “a marble” at all is certain: P=2020=1=100%P = \dfrac{20}{20} = 1 = 100\%. Drawing a green marble is impossible (there are none): P=0P = 0.

Tip

Tip: To turn a fraction into a percent, divide top by bottom, then move the decimal point two places right. 34=0.75=75%\dfrac{3}{4} = 0.75 = 75\%.

Compound Events and the Counting Principle

Concept
The Rule

A compound event involves more than one thing happening (such as flipping a coin and rolling a cube). The Counting Principle says: if one choice can happen in mm ways and a second choice in nn ways, then together they can happen in

m×n ways.m \times n \text{ ways}.

This gives the total number of outcomes, which you can then use in a probability fraction.

Example
Worked Example: counting outfits

You have 33 shirts and 44 pairs of shorts. How many different outfits? By the Counting Principle, 3×4=123 \times 4 = \mathbf{12} outfits.

Example
Worked Example: coin and cube together

Flip a coin (22 outcomes: H, T) and roll a cube (66 outcomes). Total outcomes =2×6=12= 2 \times 6 = 12. What is P(heads and a 3)P(\text{heads and a }3)? Only one pair works out of 1212, so

P(H and 3)=112.P(\text{H and }3) = \frac{1}{12}.
Tip

Tip: For “and” with the Counting Principle you multiply the numbers of ways. A tree diagram is a great way to list every outcome when the numbers are small.

Real-World Data and Probability Problems

Concept
The Rule

Word problems combine these ideas. Read slowly and decide what is being asked: a center (mean, median, mode), a spread (range), a reading from a graph, or a probability. Then set up the matching calculation and check that your answer makes sense.

Example
Worked Example: putting it together

In five games a player scored 12, 8, 15, 12, 312,\ 8,\ 15,\ 12,\ 3 points.

Mean: 12+8+15+12+35=505=10\dfrac{12+8+15+12+3}{5} = \dfrac{50}{5} = 10 points.

Median (order 3,8,12,12,153,8,12,12,15): the middle value is 1212.

Mode: 1212 (it appears twice).   Range: 153=1215 - 3 = 12.

If the coach picks one of these five games at random to rewatch, the probability it is a game where the player scored more than 1010 points (the 1212, 1515, and 1212 games =3=3 games) is 35=0.6=60%\dfrac{3}{5} = 0.6 = 60\% --- likely.

Tip

Tip: After solving, reread the question. Answers should have the right units (points, students, %) and land in a sensible range --- probabilities between 00 and 11, and a mean somewhere between the smallest and largest data values.

Going Deeper: Advanced Statistics & Probability

Concept
The Weighted Mean
3456789

A plain mean treats every value equally. A weighted mean lets some values count more than others. If each value xx has a weight ww (how many times it counts), then

weighted mean=(xw)w=x1w1+x2w2+w1+w2+.\text{weighted mean} = \frac{\sum (x \cdot w)}{\sum w} = \frac{x_1 w_1 + x_2 w_2 + \cdots}{w_1 + w_2 + \cdots}.

This is exactly how a course grade works when tests count more than homework.

One outlier drags the mean but not the median.

Example
Worked Example: a weighted class grade

Homework is worth 20%20\%, quizzes 30%30\%, and the final exam 50%50\%. A student earns 9090, 8080, and 7070 on these. Multiply each score by its weight (as a decimal) and add:

grade=90(0.20)+80(0.30)+70(0.50)=18+24+35=77.\text{grade} = 90(0.20) + 80(0.30) + 70(0.50) = 18 + 24 + 35 = \mathbf{77}.

The plain average would be 90+80+703=80\dfrac{90+80+70}{3} = 80, but because the heavily weighted final was lowest, the weighted grade drops to 7777. The weights 0.20+0.30+0.50=10.20+0.30+0.50 = 1 must add to 11 (that is 100%100\%).

Concept
Measuring Spread: MAD and Standard Deviation

Range uses only two values. Two better measures of spread use every value's distance from the mean.

  • Mean Absolute Deviation (MAD): average how far each value is from the mean, using positive distances. @@BLOCK0@@
  • Standard deviation (σ\sigma): instead of absolute values, square each distance, average the squares (this average is the variance), then take the square root. @@BLOCK1@@

For both, a larger value means the data are more spread out. Standard deviation is the one you will meet most in later courses.

Example
Worked Example: MAD and standard deviation

Find the MAD and standard deviation of 2, 4, 6, 8, 102,\ 4,\ 6,\ 8,\ 10. First the mean: xˉ=2+4+6+8+105=305=6\bar{x} = \dfrac{2+4+6+8+10}{5} = \dfrac{30}{5} = 6. Now each distance from 66:

MAD: 4+2+0+2+45=125=2.4\dfrac{4+2+0+2+4}{5} = \dfrac{12}{5} = \mathbf{2.4}.

Standard deviation: the squared distances sum to 16+4+0+4+16=4016+4+0+4+16 = 40, so the variance is 405=8\dfrac{40}{5} = 8 and σ=82.83\sigma = \sqrt{8} \approx \mathbf{2.83}. Both numbers describe the “typical” distance of a value from the mean.

Concept
Quartiles, IQR, and Outliers

The median splits ordered data in half. The quartiles split it into quarters:

  • Q1Q_1 (lower quartile) is the median of the lower half.
  • Q2Q_2 is the overall median.
  • Q3Q_3 (upper quartile) is the median of the upper half.

The interquartile range is IQR=Q3Q1\text{IQR} = Q_3 - Q_1; it measures the spread of the middle 50%50\% of the data and ignores extremes. A value is usually called an outlier if it lies below Q11.5(IQR)Q_1 - 1.5(\text{IQR}) or above Q3+1.5(IQR)Q_3 + 1.5(\text{IQR}).

Example
Worked Example: quartiles, IQR, and an outlier check

For 4, 7, 8, 9, 11, 12, 204,\ 7,\ 8,\ 9,\ 11,\ 12,\ 20 (already in order, n=7n=7), the median is the 44th value, Q2=9Q_2 = 9. The lower half is 4,7,84,7,8, so Q1=7Q_1 = 7; the upper half is 11,12,2011,12,20, so Q3=12Q_3 = 12. Then

IQR=Q3Q1=127=5.\text{IQR} = Q_3 - Q_1 = 12 - 7 = 5.

Outlier fences: Q11.5(5)=77.5=0.5Q_1 - 1.5(5) = 7 - 7.5 = -0.5 and Q3+1.5(5)=12+7.5=19.5Q_3 + 1.5(5) = 12 + 7.5 = 19.5. Since 20>19.520 > 19.5, the value 20\mathbf{20} is an outlier.

Concept
Permutations and Combinations

When the Counting Principle is used to arrange or choose from one group, two special cases appear. Write n!=n(n1)(n2)1n! = n(n-1)(n-2)\cdots 1 (read “nn factorial”).

  • A permutation counts arrangements where order matters (like finishing 11st, 22nd, 33rd): @@BLOCK0@@
  • A combination counts selections where order does not matter (like choosing a team): @@BLOCK1@@

Because order matters in more cases, nPr{}_nP_r is always at least as large as nCr{}_nC_r.

Example
Worked Example: order matters vs. order does not

From 55 runners, how many ways can they take gold, silver, and bronze? Order matters, so

5P3=5!(53)!=5!2!=1202=60.{}_5P_3 = \frac{5!}{(5-3)!} = \frac{5!}{2!} = \frac{120}{2} = \mathbf{60}.

From those same 55 runners, how many ways can we pick 33 to form a relay team (no ordering)? Now

5C3=5!3!2!=12062=12012=10.{}_5C_3 = \frac{5!}{3!\,2!} = \frac{120}{6 \cdot 2} = \frac{120}{12} = \mathbf{10}.

Each unordered team of 33 can be ordered in 3!=63! = 6 ways, and indeed 10×6=6010 \times 6 = 60, matching the permutation count.

Concept
Expected Value

The expected value is the long-run average outcome of a random situation. Multiply each outcome's value by its probability, then add:

E=(value×probability).E = \sum (\text{value} \times \text{probability}).

It need not be a possible single outcome; it is what you would average over many, many repeats.

Example
Worked Example: is the game fair?

A game costs $22 to play. You roll a number cube: rolling a 66 wins you $1010; anything else wins nothing. The expected winnings are

E=1016+056=106$1.67.E = 10 \cdot \frac{1}{6} + 0 \cdot \frac{5}{6} = \frac{10}{6} \approx \$1.67.

Since the expected win $1.671.67 is less than the $22 cost, the game favors the house: on average you lose about $0.330.33 each play.

Concept
Independent vs. Dependent Events and Conditional Probability

Two events are independent if one happening does not change the other's probability (like two separate coin flips). They are dependent if it does (like drawing cards without replacing them).

  • Independent “and”: P(A and B)=P(A)P(B)P(A \text{ and } B) = P(A) \cdot P(B).
  • Dependent “and”: P(A and B)=P(A)P(B given A)P(A \text{ and } B) = P(A) \cdot P(B \text{ given } A), where P(B given A)P(B \text{ given } A) is the conditional probability of BB once AA has happened.
Example
Worked Example: drawing without replacement

A bag has 33 red and 22 blue marbles (55 total). You draw two marbles without putting the first back. Find P(both red)P(\text{both red}). The first draw is red with probability 35\dfrac{3}{5}. Now only 22 red remain out of 44 marbles, so the second is red with probability 24\dfrac{2}{4}. These are dependent, so multiply:

P(both red)=3524=620=310.P(\text{both red}) = \frac{3}{5} \cdot \frac{2}{4} = \frac{6}{20} = \frac{3}{10}.

If instead you replaced the first marble, the draws would be independent and P=3535=925P = \dfrac{3}{5} \cdot \dfrac{3}{5} = \dfrac{9}{25}.

Tip

Tip: The word “and” tells you to multiply probabilities; the word “or” (for events that cannot both happen) tells you to add them. Always ask whether the second event's probability changed --- if it did, the events are dependent.

Concept
“At Least One” via the Complement

The complement of an event is everything except that event, and P(not A)=1P(A)P(\text{not }A) = 1 - P(A). Problems that ask for the probability of “at least one” success are usually far easier through the complement, because “at least one” is the opposite of “none”:

P(at least one)=1P(none).P(\text{at least one}) = 1 - P(\text{none}).
Example
Worked Example: at least one head

Flip a fair coin 33 times. Find P(at least one head)P(\text{at least one head}). Listing all the winning cases is tedious, so use the complement. The only way to get no heads is all tails: P(TTT)=121212=18P(\text{TTT}) = \dfrac{1}{2} \cdot \dfrac{1}{2} \cdot \dfrac{1}{2} = \dfrac{1}{8}. Therefore

P(at least one head)=118=78.P(\text{at least one head}) = 1 - \frac{1}{8} = \frac{7}{8}.
Tip

Tip: Whenever a probability question contains the phrase “at least one,” try the complement first: find P(none)P(\text{none}) and subtract from 11. It almost always saves work.

Formulas, Proofs & Tips

Tip
Mean, median, mode, range
xˉ=x1+x2++xnn\bar{x}=\frac{x_1+x_2+\cdots+x_n}{n}

What it means. The mean is the balancing point; the median is the middle value once sorted; the mode is the most frequent value; the range is largest minus smallest.

Example. For 3,5,5,93,5,5,9: mean =224=5.5=\tfrac{22}{4}=5.5, median =5=5, mode =5=5, range =93=6=9-3=6.

Why it works. The mean shares the total equally among the nn values: if everyone had xˉ\bar{x}, the total would still be nxˉ=xin\bar{x}=\sum x_i.

Tip. Sort the list before taking a median. With an even count the median is the average of the two middle values. One extreme outlier drags the mean but barely moves the median.

Tip
Standard deviation
σ=(xixˉ)2n\sigma=\sqrt{\frac{\sum (x_i-\bar{x})^{2}}{n}}

What it means. The typical distance of a value from the mean.

Example. For 2,4,62,4,6: mean 44, so σ=4+0+43=831.63\sigma=\sqrt{\tfrac{4+0+4}{3}}=\sqrt{\tfrac83}\approx1.63.

Why it works. Raw deviations xixˉx_i-\bar x sum to zero, so they are squared to stop cancellation, averaged to get a typical squared distance, then square-rooted to return to the original units.

Tip. Adding a constant to every value leaves σ\sigma unchanged; multiplying every value by kk multiplies σ\sigma by k|k|.