
Probability Distributions
Some special distributions and visualizing probabilities
- definition of a probability distribution
- visual representations of a known probability distribution
- visual representations of an unknown probability distribution (empirical histogram)
- definition of a Binomial distribution and a Hypergeometric distribution
Concept Acquisition
- Probability distributions
- Probability histograms
- Empirical histograms
- Distribution tables
Tool Acquisition
- How to write down the distribution of the probabilities of outcomes
- What a probability histogram represents
- Empirical histograms vs probability histograms
geom_col(), R script files,replicate()
Concept Application
- Drawing probability histograms
- Using R to simulate probabilities
- Drawing empirical histograms
So far we have seen examples of outcome spaces, and descriptions of how we might compute probabilities, along with explicitly writing out the probabilities. In this set of notes, we are going to see how to visualize probabilities using tables and histograms, as well as how to visualize simulations of outcomes from actions such as tossing coins or rolling dice.
Probability distributions and histograms
Probability distributions
Recall the example in which we drew a ticket from a box with 5 tickets in it:

If we draw one ticket at random from this box, we know that the probabilities of the four distinct outcomes can be listed in a table as:
| Outcome | \(1\) | \(2\) | \(3\) | \(4\) |
|---|---|---|---|---|
| Probability | \(\displaystyle \frac{1}{5}\) | \(\displaystyle \frac{2}{5}\) | \(\displaystyle \frac{1}{5}\) | \(\displaystyle \frac{1}{5}\) |
What we have described in the table above is a probability distribution. We have shown how the total probability of 1 (100%) is distributed among all the possible outcomes. Since the ticket \(\fbox{2}\) is twice as likely as any of the other outcomes, it gets twice as much of the probability.
Probability histograms
A table is nice, but maybe a visual representation would be even better.
We have represented the distribution in the form of a histogram, with the areas of the bars representing probabilities. Notice that this histogram is different from the ones we have seen before, since we didn’t collect any data. We just defined the probabilities based on the outcomes, and then drew bars with the heights being the probabilities. This type of theoretical histogram is called a probability histogram. Note that the heights add up to 1 (or 100%), and the width of each bar is 1, making the area of each bar the probability of seeing the value at the center.
Empirical histograms
What about if we don’t know the probability distribution of the outcomes of an experiment? For example, what if we didn’t know how to compute the probability distribution above? What could we do to get an idea of what the probabilities might be? Well, we could keep drawing tickets over and over again from the box, with replacement (that is, we put the selected tickets back before choosing again), keep track of the tickets we draw, and make a histogram of our results. This kind of histogram, which is the kind we have seen before, is a visual representation of data, and is called an empirical histogram.

On the x-axis of this histogram, we have the ticket values; on the y-axis, we have the proportion of times that this ticket was selected out of the 50 with-replacement draws we took. We can see that the sample proportions look similar to the values given by the probability distribution, but there are some differences. For example, we appear to have drawn more \(3\)s and fewer \(4\)s than what was to be expected. It turns out that the counts and proportions of the drawn tickets are:
| Ticket | Number of times drawn | Proportion of times drawn |
|---|---|---|
| \(\fbox{1}\) | 10 | 0.2 |
| \(\fbox{2}\) | 24 | 0.48 |
| \(\fbox{3}\) | 10 | 0.2 |
| \(\fbox{4}\) | 6 | 0.12 |
What we have seen here is how when we draw at random, we get a sample that resembles the population, that is, a representative sample, but it isn’t exactly the true probabilities. If we increase our sample, however, say to 500, we will get something that more closely aligns with the truth (just like we did with the proportion of heads when we toss a coin over and over again in the last chapter).
Examples
Tossing a coin
Let’s think about the probabilities when we toss a coin. If we toss a fair coin once, we have two equally likely outcomes that are possible, “Heads” and “Tails”. In order to plot a histogram that visualizes the probability of the coin landing heads, we need to record each toss that lands heads as \(1\) and each toss that lands tails as \(0\). This is like drawing a ticket from a box that has two tickets marked \(0\) and \(1\).
What if the coin is not fair, and the chance of landing heads is \(\dfrac{2}{3}\) or even \(\dfrac{3}{4}\)! In these cases, the histogram’s bar over \(1\) has to reflect this probability. Since the area of the bar is the probability, if the width was \(1\), the height of the bar over \(1\) would be \(3/4\). Here are three probability histograms, one for a fair coin, and the other two for biased coins.

Suppose we wanted to represent tossing a biased coin by drawing tickets from a box. What box would we use?
For example, if the probability of the coin landing heads is \(\dfrac{3}{4}\), and we wanted to represent this by drawing from a box of tickets marked \(\fbox{0}\) or \(\fbox{1}\), then we would need three tickets marked \(\fbox{1}\), and one ticket marked \(\fbox{0}\).
What about if the probability of the coin landing heads was \(\dfrac{2}{3}\)? What tickets would the corresponding box contain?
What about if the probability of the coin landing heads was \(0.3\)?
Check your answer
If the probability of heads is \(2/3\), then we need the probability of drawing a \(\fbox{1}\) to be \(2/3\), so we would need two tickets marked \(\fbox{1}\) and one marked \(\fbox{0}\).
If the probability of the coin landing heads is \(0.3\), then we need three out of ten tickets in the box to be marked \(\fbox{1}\) and the other seven to be \(\fbox{0}\).
Rolling a pair of dice and summing the spots
The outcomes are already numbers, so we don’t need to represent them differently. We know that there are \(36\) total possible equally likely outcomes when we roll a pair of dice, but when we add the spots, we have only 11 possible outcomes, which are not equally likely (the chance of seeing a \(2\) is \(1/36\), but \(P(6)=5/36\)).
Again, the probability histogram will have the possible outcomes listed on the x-axis, and bars of width \(1\) over each possible outcome. The height of these bars will be the probability, so that the areas of the bars represent the probability of the value under the bar. The height, which is the probability, is written on the top of each bar.

What about the probability distribution? Make a table showing the probability distribution for rolling a pair of dice and summing the spots.
Check your answer
| Outcome | \(2\) | \(3\) | \(4\) | \(5\) | \(6\) | \(7\) | \(8\) | \(9\) | \(10\) | \(11\) | \(12\) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| Probability | \(\displaystyle \frac{1}{36}\) | \(\displaystyle \frac{2}{36}\) | \(\displaystyle \frac{3}{36}\) | \(\displaystyle \frac{4}{36}\) | \(\displaystyle \frac{5}{36}\) | \(\displaystyle \frac{6}{36}\) | \(\displaystyle\frac{5}{36}\) | \(\displaystyle \frac{4}{36}\) | \(\displaystyle \frac{3}{36}\) | \(\displaystyle \frac{2}{36}\) | \(\displaystyle \frac{1}{36}\) |
Tossing a fair coin 3 times and counting the number of heads
We have seen that there are 8 equally likely outcomes from tossing a fair coin three times: \(\{HHH, HHT, HTH, THH, TTH, THT, HTT, TTT\}\). If we count the number of \(H\) in each outcome, and write down the probability distribution of the number of heads, we get:
| Outcome | \(0\) | \(1\) | \(2\) | \(3\) |
|---|---|---|---|---|
| Probability | \(\displaystyle \frac{1}{8}\) | \(\displaystyle \frac{3}{8}\) | \(\displaystyle \frac{3}{8}\) | \(\displaystyle \frac{1}{8}\) |
What would the probability histogram look like?
Check your answer

We are going to introduce some special distributions. We have seen most of these distributions, but will introduce some names and definitions. Before we do this, let’s recall how to count the number of outcomes for various experiments such as tossing coins or drawing tickets from a box (both with and without replacement).
Basic rule of counting
Recall that if we have multiple steps (say \(n\)) of some action, such that the \(j^{\text{th}}\) step has \(k_j\) outcomes; then the total number of outcomes is \(k_1 \times k_2 \times \ldots \times k_n\) and is obtained by multiplying the number of outcomes at each step. This principle is illustrated in the picture below: Geni the Gentoo penguin1 is trying to count how many outfits they have, if each outfit consists of a t-shirt and a pair of pants. The tree diagram below shows the number of possible outfits Geni can wear. In this example, \(n=2\), \(k_1 = 3\), and \(k_2 = 2\), since Geni has three t-shirts to choose from, and for each t-shirt, they have two pairs of pants, leading to a total of \(3 \times 2 = 6\) outfits.

This example seems trivial, but it illustrates the basic principle of counting: we get the total number of possible outcomes of an action that has multiple steps, by multiplying together the number of outcomes for each step. This is called the fundamental rule of counting, and you can check it for yourself by drawing a tree with the possible outcomes as shown for Geni above.
All the counting that follows in our notes applies this rule.
For example, let’s suppose we are drawing tickets from a box which has tickets marked with the letters \(\fbox{C}\), \(\fbox{R}\), \(\fbox{A}\), \(\fbox{T}\), \(\fbox{E}\), and say we draw three letters with replacement (that means that we put each drawn ticket back, so the box is the same for each draw). How many possible sequences of three letters can we get? Using the counting rule, since we have \(5\) choices for the first letter, \(5\) for the second, and \(5\) for the third, we will have a total of \(5\times 5 \times 5 = 125\) possible words, allowing for repeated letters (that means that \(CCC\) is a possible outcome).
How many possible words are there if we draw without replacement? That is, we don’t put the drawn ticket back?
Check your answer
We have \(5\) choices for the first letter, \(4\) for the second, and \(3\) for the third, leading to \(5 \times 4 \times 3 = 60\) possible outcomes or words. Note that here we count the word \(CRA\) as different from the word \(CAR\). That is, the order in which we draw the letters matters.
We usually write the quantity \(5 \times 4 \times 3\) as \(\displaystyle \frac{5!}{2!}\).Talking about factorials
Let’s briefly digress to admire the quantity we have just encountered. We won’t test you on this (you just need to know how to compute a factorial for not very large \(n\)), but these numbers can get mind-bendingly large.
\(n! = n\times(n-1)\times(n-2)\times \ldots \times 3 \times 2 \times 1\). We read \(n!\) as “n-factorial”, and define \(0\)-factorial to be 1 (\(0!=1\)). \(n!\) is a quantity that is denoted with a very apt symbol, actually. One of surprise! And perhaps this is because it blows up very quickly. You can try it out in R by using the function factorial(). For small numbers, there is no surprise: \(1! = 1,\, 2! = 2,\, 3! = 6,\, 4! = 24, \, 5! = 120\). But \(6!=720\) gets close to \(1,000\), and \(7!\) is over \(5,000\)!! We begin to see why using an exclamation point might be well-deserved. Factorials appear in mathematics in many situations, and especially in counting. How big would you guess \(20!\) to be? What about \(100!\)?
Check your answer
factorial(20) = 2.432902e+18, which is scientific notation for about \(2,432,902,000,000,000,000\), or roughly \(2.43\) quintillion (\(2.43\) billion billion). The “e+18” stands for \(10^{18}\).
If you think about it, \(100\) isn’t such a big number. But what about \(100!\)?
factorial(100) = 9.332622e+157. Just to put this number in perspective, the number of atoms in the observable universe is estimated to be between \(10^{78}\) to \(10^{82}\).
Counting the number of ways to select a subset of outcomes
Selecting a subset means that we only care about which tickets we get, not the order in which we get them. For example, while playing standard five card poker, all that matters is which cards you have in your hand, not in what order you got them.
What if, in this example of selecting \(3\) letters without replacement from \(\fbox{C}\), \(\fbox{R}\), \(\fbox{A}\), \(\fbox{T}\), \(\fbox{E}\), the order does not matter - we don’t count the order in which the letters are selected, just which letters were selected, that is, only the letters themselves matter. For example, the words \(CRA,\,CAR,\,ARC,\,ACR,\,RAC,\,RCA\) all count as the same word, so we will count all \(6\) words as the same subset of letters. We have to take the number that we got from earlier and divide it by the number of words that can be made from \(3\) letters (number of rearrangements), which is \(3 \times 2\times 1\). This gives us the number of ways that we can choose \(3\) letters out of \(5\), which is \[ \frac{\left(5!/2!\right)}{3!} = \frac{5!}{2!\; 3!}, \]
In general, the number of ways we can choose a subset of \(k\) things out of \(n\) possible things (when we only care about which things is chosen, not their order) is given by \[ \binom{n}{k} = \frac{n!}{k!\times (n-k)!} \] and is read as “n choose k”.
In this note, we see how to compute the number of ways we can draw tickets from a box when we are sampling without replacement. We will use these numbers later, but you do not need to know how to derive them. Please read this only if you are interested in how we get these numbers.
The number we computed above is called the number of combinations of \(5\) things taken \(3\) at a time.
To recap: when we draw \(k\) items from \(n\) items without replacement, we have two cases: either we care in what order we draw the \(k\) items (so the different arrangements of the same set of \(k\) items have to be counted separately), or we don’t, and we count all possible orders of drawing as one.
In the first case, the number of such arrangements is called the permutations of \(n\) things taken \(k\) at a time. In the example above, we have \(5 \times 4\times 3 = 60\) ways of arranging \(3\) things out of \(5\) when we count every sequence as different.
- Permutations
- The number of possible arrangements or sequences of \(n\) things taken \(k\) at a time which is given by (the ordering matters): \[ \frac{n!}{(n-k)!} \]
- Combinations
- Number of ways to choose a subset of \(k\) things out of \(n\) possible things which is given by \[ \frac{n!}{k!\; (n-k)!} \] This number is just the number of distinct arrangements or permutations of \(n\) things taken \(k\) at a time divided by the number of arrangements of \(k\) things. It is denoted by \(\displaystyle \binom{n}{k}\), which is read as “n choose k”.
Example How many ways can I deal \(5\) cards from a standard deck of 52 cards?
Check your answer
When we deal cards, order does not matter, so this number is \(\displaystyle \binom{52}{5} = \frac{52!}{(52-5)!5!} = 2,598,960\).
Special distributions
There are some important special distributions that every student of probability must know. Here are a few, and we will learn some more later in the course. These should all be familiar as we have seen them before. All we are doing now is identifying their names. First, we need a vocabulary term:
- Parameter of a probability distribution
- This is some constant(s) number associated with the distribution. If you know the parameter(s) of a probability distribution, then you can compute the probabilities of all the possible outcomes.
Each of the distributions we list below has a parameter(s) associated with it.
Discrete uniform distribution
This is one of the simplest probability distributions. We say that this distribution is over the numbers \(1, 2, 3 \ldots, n\) and assigns the same probability (\(\dfrac{1}{n}\)) to each of the numbers. We have seen it for the experiment of rolling a fair six-sided die in previous chapters, or for drawing a ticket from a box with \(n\) tickets labeled \(\fbox{1}, \fbox{2}, \fbox{3}, \ldots\) etc.
This probability distribution is called the discrete uniform probability distribution, since each possible outcome has the same probability of \(1/n\). Here \(n\) is the parameter of the discrete uniform distribution.
Bernoulli distribution
This is a probability distribution describing the probabilities associated with binary outcomes that result from one action, such as one coin toss that can either land Heads or Tails. We can represent the action as drawing one ticket from a box with tickets marked \(\fbox{1}\) or \(\fbox{0}\), where the probability of \(\fbox{1}\) is \(p\), and therefore, the probability of \(\fbox{0}\) is \((1-p)\). We have already seen some examples of probability histograms for this distribution. We usually think of the possible outcomes of a Bernoulli distribution as success and failure, and represent a success by \(\fbox{1}\) and a failure by \(\fbox{0}\).

For the Bernoulli distribution, our parameter is \(p = P\left(\fbox{1}\right)\). If we know \(p\), we also know the probability of drawing a ticket marked \(\fbox{0}\).
In the figure above, the first histogram is for a Bernoulli distribution with parameter \(p = 1/2\), the second \(p=3/4\), and the third has \(p = 2/3\).
Binomial Distribution
The binomial distribution, which describes the total number of successes in a sequence of \(n\) independent Bernoulli trials, is one of the most important probability distributions. For example, consider the outcomes from tossing a coin \(n\) times and counting the total number of heads across all \(n\) tosses, where the probability of heads on each toss is \(p\). Each toss is one Bernoulli trial, where a success would be the coin landing heads.
That is, the binomial distribution describes the probabilities of the total number of heads in \(n\) tosses. We saw what this distribution looked like in the case where \(n = 3\) for three tosses of a fair coin.
What would the probability distribution and histogram for the number of heads in three tosses of a biased coin look like, where \(P(H) = 2/3\)? Make sure you know how the probabilities are computed. For example, \(P(HHH) = (2/3)^3 = 8/27\).
Check your answer

Note that the outcomes \(HHT, HTH, THH\) all have the same probability, as do the outcomes \(TTH, THT, HTT\), since the probability only depends on how many heads and how many tails we see in three tosses, not the order in which we see them. We get the probability of 2 heads in 3 tosses by adding the probabilities of all three outcomes \(HHT, HTH, THH\), since they are mutually exclusive. (In three tosses, we can see exactly one of the possible 8 sequences listed above.)
More generally, suppose that we have \(n\) independent trials, where each trial can either result in a “success” (like drawing a ticket marked \(\fbox{1}\)) with probability \(p\); or a “failure” (like drawing \(\fbox{0}\)) with probability \(1-p\).
The probability of \(k\) successes in \(n\) trials is given by the following formula: \[ \binom{n}{k} \times p^k \times (1-p)^{n-k} \]
The multiplication rule for independent events tells us how to compute the probability of a sequence that consisted of the first \(k\) trials being successes and the rest of the \(n-k\) trials being failures. The probability of this particular sequence of \(k\) successes followed by \(n-k\) failures is (by multiplying their probabilities) given by: \[ p^k \times (1-p)^{n-k} \] Now this is the probability of one particular sequence: \(SSS\ldots SSFF \ldots FFF\), but as we saw in the example above, only the number of successes and failures matter, not the particular order. So every sequence of \(n\) trials in which we have \(k\) successes and \(n-k\) failures has the same probability.
How many such sequences are there? We can count them using our rules above. We have \(n\) spots in the sequence, of which \(k\) have to be successes. The number of such sequences of length \(n\) consisting of \(k\) \(S\)’s and \(n-k\) \(F\)’s) is given by \(\displaystyle \binom{n}{k}\). Each such sequence has probability \(\displaystyle p^k \times (1-p)^{n-k}\). Adding up all these \(\displaystyle \binom{n}{k}\) probabilities (of each such sequence) gives us the formula for the probability of \(k\) successes in \(n\) trials: \[ \binom{n}{k} \times p^k \times (1-p)^{n-k} \]
The probability distribution described by the above formula is called the binomial distribution. It is named after \(\displaystyle \binom{n}{k}\), which is called the binomial coefficient and is defined by: \[ \binom{n}{k} = \frac{n!}{k!\times (n-k)!} \]
The binomial distribution has two parameters: the number of trials \(n\) and the probability of success on each trial, \(p\).
Example
Toss a weighted coin five times, where \(P(\text{heads}) = 0.7\). What is the probability that you see exactly four heads across these five tosses?
Check your answer
Using the Binomial formula from above, \(n = 5\). We define \(k = 4\) to be the number of successes whose probability we want to compute. Plugging these values into the Binomial formula using \(p = 0.7\) and \(1-p = 0.3\), we get the probability of exactly 4 heads in 5 tosses to be: \[ \binom{5}{4} (0.7)^4 \times (0.3)^1 \approx 0.36 \]Hypergeometric distribution
In the binomial scenario described above, we had \(n\) independent trials, where each trial resulted in a success or a failure. This is like drawing \(n\) times with replacement from a box of tickets marked \(0\) or \(1\).
Now consider the situation when we have a box with \(N\) tickets marked with either \(\fbox{0}\) or \(\fbox{1}\), and we draw \(n\) times from this box without replacement.
As usual, the ticket marked \(\fbox{1}\) represents a success. Say the box has \(G\) tickets marked \(\fbox{1}\) (and therefore \(N-G\) tickets marked \(\fbox{0}\) representing failures). Suppose we draw \(n\) tickets without replacement from this box. The \(n\) tickets we have drawn (so we must have \(n \le N\)) form a “sample” which is called a “simple random sample”. A simple random sample is just the result of a bunch of draws without replacement such that on each draw, every ticket is equally likely to be selected from among the remaining tickets. Thus, the probability of drawing a ticket marked \(\fbox{1}\) changes from draw to draw.
What is the probability that we will have exactly \(k\) successes among these \(n\) draws? The formula for this probability is: \[ \frac{\binom{G}{k}\times \binom{N-G}{n-k}}{\binom{N}{n}} \]
The distribution of the probabilities defined by this formula is called the Hypergeometric distribution. It has three parameters, \(n\), \(N\) and \(G\).
This is a wacky formula, so let’s explain it piece by piece, starting with the numerator and then moving to the denominator!
Numerator
We count the number of samples drawn without replacement that have \(k\) tickets marked \(\fbox{1}\). Since there are \(G\) tickets marked \(\fbox{1}\) in the box, and \(\displaystyle \binom{G}{k}\) ways to choose exactly \(k\) of them. Similarly, there are \(N-G\) tickets marked \(\fbox{0}\) in the box, and \(\displaystyle \binom{N-G}{n-k}\) ways to choose exactly \(n-k\) of them. The total number of ways to have \(k\) \(\fbox{1}\)s and \(n-k\) \(\fbox{0}\)s is therefore (by multiplication):
\[ \binom{G}{k}\times \binom{N-G}{n-k} \]
Denominator
We count the total number of simple random samples of size \(n\) that can be drawn from a pool of \(N\) observations. This is given by \(\displaystyle \binom{N}{n}\).
Example
Say we have a box of \(10\) tickets consisting of \(4\) tickets marked \(\fbox{0}\) and \(6\) tickets marked \(\fbox{1}\), and draw three times from this box without replacement.
What is the probability that exactly two of the tickets drawn are marked \(\fbox{1}\)?
Check your answer
We draw \(3\) tickets without replacement. We need two of these tickets to be marked \(\fbox{1}\) and there are six in total to choose from; we need one of them to be marked \(\fbox{0}\) and there are four in total to choose from. Therefore, there are
\[ \binom{6}{2}\times \binom{4}{1} = 60 \] different ways to pick three tickets in this manner.
How many total ways are there to draw \(n=3\) tickets from a box of \(N=10\)? This is given by \(\binom{10}{3} = 120\).
Therefore, our final answer is given by
\[ \frac{\binom{6}{2}\times \binom{4}{1}}{\binom{10}{3}} = \frac{60}{120} = \frac{1}{2} \]
Binomial vs Hypergeometric distributions
Both these distributions deal with:
A fixed number of trials, or instances of the random experiment;
Two kinds of outcomes, one of which is defined to be a “success”, and the other is defined to be a “failure”.
The difference is that for a binomial distribution, we model drawing n times with replacement from a box with all its tickets marked \(\fbox{0}\) or \(\fbox{1}\) and counting the number \(\fbox{1}\)’s we drew.
For a hypergeometric distribution, we also draw n times from a box with tickets marked \(\fbox{0}\) or \(\fbox{1}\) and counting the number \(\fbox{1}\)’s we drew, but we draw without replacement.
This means that for a binomial distribution, the probability of a success stays the same for each trial, while for a hypergeometric distribution, the probability of a success changes for each trial.
Summary
- We defined probability distributions
- Visualized probability distributions using a special kind of plot called a probability histogram
- Stated the fundamental rule of counting
- Defined some famous named distributions (Bernoulli, discrete uniform, binomial, hypergeometric)
- Looked at the relationship between empirical histograms and probability histograms.
Footnotes
Penguin taken from art by @allison_horst↩︎