Probability theory studies uncertainty, but “uncertain” doesn’t mean unanalyzable. A probability model first lists possible outcomes, then describes how likely different outcomes are with numbers. Statistics, in turn, works backward from already-observed data to infer a model, estimate unknown parameters, or judge whether some hypothesis is credible.

Beginners often memorize combinatorics, probability distributions, and expectation formulas directly, without first distinguishing sample points, events, random variables, and observed values. Mixing these concepts together makes conditional probability, probability density, expectation, statistical inference, and stochastic processes much harder to understand. This article first builds a vocabulary that can be reused repeatedly, then uses examples such as dice rolls, height, quality inspection, and server requests to clarify how the pieces fit together.

1. From Random Experiments to a Probability Space

1.1 Random experiments and sample spaces

An operation whose result can’t be determined before it’s carried out can be treated as a random experiment. Typical examples include rolling a die and recording the number, drawing a product and checking whether it passes inspection, recording the number of requests a server receives in the next minute, and measuring an electronic component’s lifespan. A probability model needs to clearly define “one trial” and “the outcome to record” — a vaguely defined experiment can’t produce a well-defined probability.

The set of every possible outcome of a random experiment is called the sample space, usually written Ω\Omega; a single outcome in the sample space is written ω\omega. The sample space of rolling a six-sided die is Ω={1,2,3,4,5,6}\Omega=\{1,2,3,4,5,6\}. When rolling two distinguishable dice, Ω={(i,j):i,j∈{1,2,3,4,5,6}}\Omega=\{(i,j):i,j\in\{1,2,3,4,5,6\}\} represents 36 ordered outcomes. (1,6)(1,6) and (6,1)(6,1) are different outcomes, even though both sum to 7.

A sample space can also be an infinite set. Repeatedly flipping a coin until the first heads can be represented with {1,2,3,… }\{1,2,3,\dots\} for the waiting count; an electronic component’s lifespan can be represented with [0,∞)[0,\infty); a temperature curve over a period of time needs to be represented with a set of functions instead. A finite sample space is suited to counting methods, while a continuous or function-valued sample space needs tools like density, integration, and measure.

1.2 Events and set operations

An event is a subset of the sample space. When rolling a die, let A={2,4,6}A=\{2,4,6\} represent “the number is even,” and let B={4,5,6}B=\{4,5,6\} represent “the number is greater than 3.” The observed outcome falling within the event’s set means the event occurred.

Set notationEvent meaning
A∩BA\cap BAA and BB both occur
A∪BA\cup Bat least one of AA or BB occurs
AcA^cAA does not occur
A∖BA\setminus BAA occurs but BB does not
A∩B=∅A\cap B=\varnothingAA and BB cannot occur simultaneously

In the example above, A∩B={4,6}A\cap B=\{4,6\}, so “even and greater than 3” contains the numbers 4 and 6.

For a finite sample space, every subset can generally be treated as an event. In an infinite sample space, pathological sets exist that can’t be consistently assigned a probability, so a rigorous probability model specifies a family of event sets F\mathcal F. F\mathcal F is closed under complementation and countable union, and is called a σ\sigma-algebra. At an introductory level, F\mathcal F can be understood as the collection of events the model allows you to ask about and assign a probability to.

1.3 Probability measures and the equally-likely assumption

A probability measure PP maps an event A∈FA\in\mathcal F to a value within [0,1][0,1]. Kolmogorov’s axioms require nonnegativity P(A)≥0P(A)\ge 0, normalization P(Ω)=1P(\Omega)=1, and countable additivity over mutually exclusive events. The triple (Ω,F,P)(\Omega,\mathcal F,P) is called a probability space.

The axioms let us derive P(Ac)=1−P(A)P(A^c)=1-P(A) and P(A∪B)=P(A)+P(B)−P(A∩B)P(A\cup B)=P(A)+P(B)-P(A\cap B). The union formula subtracts the intersection because directly adding P(A)P(A) and P(B)P(B) would double-count the intersection.

If the elementary outcomes of a finite sample space are equally likely, P(A)=∣A∣/∣Ω∣P(A)=|A|/|\Omega| can be used. The 36 ordered outcomes of two fair dice are equally likely, but the sums of the two dice are not. A sum of 7 has six outcomes — (1,6),(2,5),(3,4),(4,3),(5,2),(6,1)(1,6),(2,5),(3,4),(4,3),(5,2),(6,1) — so P(D1+D2=7)=6/36=1/6P(D_1+D_2=7)=6/36=1/6; a sum of 2 has only one outcome, (1,1)(1,1). Treating the sums from 2 to 12 as 11 equally likely outcomes gives the wrong answer.

Counting formulas can’t be applied to a biased die, components with different failure rates, or continuous measurements. Before using combinatorics, you must first confirm that the finiteness and equally-likely assumptions actually hold.

2. Random Variables and Distributions

2.1 A random variable is a function

A random variable isn’t a symbol that changes randomly — it’s a function that maps from the sample space to a set of numbers, written X:Ω→RX:\Omega\rightarrow\mathbb R. In the two-dice experiment, we can define X(i,j)=i+jX(i,j)=i+j. The elementary outcome is (2,5)(2,5), and the random variable’s observed value is X(2,5)=7X(2,5)=7.

Multiple random variables can be defined on the same sample space — for example, the sum of the dice, the maximum of the two, the difference between them, or whether the first die is even. A random variable lets a researcher ignore irrelevant details of the trial and keep only the quantity the question actually cares about.

2.2 Discrete random variables and the PMF

A discrete random variable has finitely or countably many possible values, such as a die roll, the number of server requests in one minute, the number of defective products in a batch, or the number of failures in a time interval.

A discrete random variable can assign a probability directly to each possible value. This function is called the probability mass function (PMF):

pX(x)=P(X=x).p_X(x)=P(X=x).

For example, if XX is the sum of two fair dice, then P(X=2)=1/36P(X=2)=1/36, while P(X=7)=6/36=1/6P(X=7)=6/36=1/6. Because many different sample points can map to the same value of XX, computing the PMF requires adding the probabilities of all elementary outcomes that produce that value:

P(X=x)=∑ω:X(ω)=xP({ω}).P(X=x)=\sum_{\omega:X(\omega)=x}P(\{\omega\}).

For a discrete random variable, a single concrete value can have positive probability, such as P(X=7)=1/6>0P(X=7)=1/6>0, and the probabilities of all possible values must sum to 1:

∑xpX(x)=1.\sum_x p_X(x)=1.

Continuous random variables behave differently.

2.3 Continuous random variables: why a single-point probability is 0

Consider a person’s height XX. Intuitively, a person does have a concrete height, such as a measured value of X=170.2X=170.2 cm. So it is natural to ask: if the final observation is a specific number, why does a continuous random variable satisfy P(X=170.2)=0P(X=170.2)=0?

The key is to separate two statements: one experiment will indeed produce some value, and the probability of hitting one exact real number specified in advance.

If height is idealized as a continuous quantity, then there are infinitely many real numbers between 170 cm and 171 cm. More importantly, between any two distinct real numbers, no matter how close, there are still infinitely many real numbers. A continuous distribution therefore cannot allocate total probability 1 to isolated points in the same way a discrete distribution does.

For a continuous random variable with a probability density function,

P(X=x)=0P(X=x)=0

for every single point xx. But probability 0 does not mean impossible. A trial will still end with some concrete real value. It only means that if we specify an infinitely precise real number before the trial, such as 170.234817291…170.234817291\dots, the probability of landing exactly on that number is 0.

One way to see this is by shrinking intervals. The probability

P(170≤X≤171)P(170\le X\le171)

is attached to an interval with positive width, so it may be positive. If we shrink the range to P(170.2≤X≤170.3)P(170.2\le X\le170.3), then to P(170.20≤X≤170.21)P(170.20\le X\le170.21), the probability usually shrinks with the interval width. When the interval width tends to 0 and only one point remains, the probability tends to 0 as well.

Real measurements also have finite precision. If a height scale displays 170.2170.2 cm, it does not mean the true height is exactly the mathematical value 170.200000…170.200000\dots cm. If the instrument rounds to 0.10.1 cm, then the reading 170.2170.2 cm usually represents a true height in a small interval such as

170.15≤X<170.25.170.15\le X<170.25.

So a “specific measurement” in the real world usually already hides a small interval created by finite measurement precision.

2.4 PDF: density is not probability

Continuous random variables are usually described with a probability density function (PDF) fX(x)f_X(x). The most important point is

fX(x)≠P(X=x).f_X(x)\neq P(X=x).

The PDF value itself is not the probability of taking the value xx; it describes how densely probability is concentrated near xx. Actual probability comes from integrating over an interval:

P(a≤X≤b)=∫abfX(x) dx.P(a\le X\le b)=\int_a^b f_X(x)\,dx.

For a very small Δx\Delta x, a narrow interval probability can be approximated by

P(x≤X≤x+Δx)≈fX(x)Δx.P(x\le X\le x+\Delta x)\approx f_X(x)\Delta x.

Thus fX(x)f_X(x) can be read as probability density per unit length. A higher density means more probability is concentrated nearby; a wider interval usually accumulates more probability.

To compute P(a≤X≤b)P(a\le X\le b), divide [a,b][a,b] into many narrow subintervals of width Δx\Delta x. The probability in the ii-th subinterval is approximately fX(xi)Δxf_X(x_i)\Delta x, so the total interval probability is approximated by

∑ifX(xi)Δx.\sum_i f_X(x_i)\Delta x.

This is a Riemann sum. As the partition becomes finer and Δx→0\Delta x\rightarrow0, the sum converges to the integral:

P(a≤X≤b)=lim⁡Δx→0∑ifX(xi)Δx=∫abfX(x) dx.P(a\le X\le b) = \lim_{\Delta x\to0}\sum_i f_X(x_i)\Delta x = \int_a^b f_X(x)\,dx.

The area under the PDF curve is the probability.

This also explains why a single-point probability is 0. If an interval degenerates to a point [x,x][x,x], its width is 0, so

P(X=x)=∫xxfX(u) du=0.P(X=x)=\int_x^x f_X(u)\,du=0.

For continuous random variables with a PDF,

P(a<X<b)=P(a≤X<b)=P(a<X≤b)=P(a≤X≤b).P(a<X<b) = P(a\le X<b) = P(a<X\le b) = P(a\le X\le b).

Including or excluding endpoints does not change the probability, because each endpoint has probability 0. This is very different from a discrete random variable: for a die roll, P(X=3)=1/6P(X=3)=1/6, so whether an endpoint is included can change the event probability.

Another common misunderstanding is that PDF values must lie between 0 and 1. In fact, a PDF value can be greater than 1. For example,

fX(x)={2,0≤x≤0.5,0,otherwise,f_X(x)= \begin{cases} 2,&0\le x\le0.5,\\ 0,&\text{otherwise}, \end{cases}

has density value 2, but total probability

∫00.52 dx=1.\int_0^{0.5}2\,dx=1.

What must lie in [0,1][0,1] is the probability obtained by integration, not the density value itself. The two basic requirements for a PDF are

fX(x)≥0f_X(x)\ge0

and

∫−∞∞fX(x) dx=1.\int_{-\infty}^{\infty}f_X(x)\,dx=1.

2.5 CDF: a common language for discrete and continuous variables

Any real-valued random variable can use a cumulative distribution function (CDF), FX(x)=P(X≤x)F_X(x)=P(X\le x). The CDF is monotonically nondecreasing, and satisfies lim⁡x→−∞FX(x)=0\lim_{x\to-\infty}F_X(x)=0 and lim⁡x→∞FX(x)=1\lim_{x\to\infty}F_X(x)=1. A discrete variable’s CDF looks like a staircase; if a continuous variable has a PDF, then FX(x)=∫−∞xfX(u) duF_X(x)=\int_{-\infty}^{x}f_X(u)\,du.

For a discrete random variable, the CDF usually has jumps. If a point xx has positive probability, then the CDF jumps upward at xx, and the jump height is exactly that point probability:

P(X=x)=FX(x)−FX(x−).P(X=x)=F_X(x)-F_X(x^-).

If FXF_X is differentiable for a continuous random variable with a PDF, the density can be recovered from the CDF:

fX(x)=FX′(x).f_X(x)=F_X'(x).

The roles of PMF, PDF, and CDF can be summarized as follows:

FunctionIntuition
PMF pX(x)p_X(x)how much probability sits at the discrete point xx
PDF fX(x)f_X(x)how dense probability is near xx
CDF FX(x)F_X(x)how much probability has accumulated up to xx

Discrete probabilities are usually computed by summation:

P(X∈A)=∑x∈ApX(x).P(X\in A)=\sum_{x\in A}p_X(x).

Continuous probabilities with a PDF are computed by integration:

P(X∈A)=∫AfX(x) dx.P(X\in A)=\int_A f_X(x)\,dx.

The forms differ, but both describe the same idea: how probability is distributed across the possible values of a random variable.

3. Conditional Probability, Independence, and Bayes’ Theorem

3.1 Conditional probability narrows the sample range

Given that event BB has already occurred, the conditional probability of event AA is defined as P(A∣B)=P(A∩B)/P(B)P(A\mid B)=P(A\cap B)/P(B), where P(B)>0P(B)>0. Conditional probability restricts the analysis to within BB, then renormalizes the probability using P(B)P(B).

Using two fair dice as an example: let AA represent “the sum is 8,” and let BB represent “the first die is even.” Event BB contains 18 equally likely outcomes; the outcomes satisfying both AA and BB are (2,6),(4,4),(6,2)(2,6),(4,4),(6,2), so P(A∣B)=3/18=1/6P(A\mid B)=3/18=1/6. Without knowing the parity of the first die, a sum of 8 has five outcomes, giving P(A)=5/36P(A)=5/36. The extra piece of information changed the comparable set of outcomes, and therefore changed the event’s probability too.

The multiplication rule P(A∩B)=P(A∣B)P(B)P(A\cap B)=P(A\mid B)P(B) can be extended to multiple events. For a sequence of events A1,…,AnA_1,\dots,A_n, their joint probability can be decomposed into a product of successive conditional probabilities. Autoregressive models in machine learning use the same chain rule, breaking a token sequence’s joint probability into the probability of each token conditioned on the tokens before it.

3.2 Independence is a condition on the joint distribution

Events AA and BB are independent when P(A∩B)=P(A)P(B)P(A\cap B)=P(A)P(B). If P(B)>0P(B)>0, an equivalent condition is P(A∣B)=P(A)P(A\mid B)=P(A): learning that BB occurred doesn’t change the probability of AA.

Mutual exclusivity and independence have different meanings. Two mutually exclusive events, each with positive probability, cannot occur together, so P(A∩B)=0P(A\cap B)=0, yet P(A)P(B)>0P(A)P(B)>0 — mutually exclusive events are therefore not independent. Independence also can’t just be declared by intuition; it must come from the sampling design, the generating mechanism, or a testable assumption about the joint distribution.

For multiple events, pairwise independence only requires that every pair’s joint probability factors; mutual independence additionally requires that the joint probability of any subset factors. Mutual independence always implies pairwise independence, but the converse generally doesn’t hold.

3.3 The law of total probability and Bayes’ theorem

If B1,…,BmB_1,\dots,B_m are pairwise mutually exclusive and their union is the entire sample space, then the probability of event AA can be split by source: P(A)=∑iP(A∣Bi)P(Bi)P(A)=\sum_iP(A\mid B_i)P(B_i). This is called the law of total probability.

Bayes’ theorem reverses the direction of conditioning: P(Bj∣A)=P(A∣Bj)P(Bj)/P(A)P(B_j\mid A)=P(A\mid B_j)P(B_j)/P(A). The numerator is made up of the likelihood P(A∣Bj)P(A\mid B_j) and the prior probability P(Bj)P(B_j); the denominator sums the joint probability over every possible source, making the posterior probabilities sum to 1.

Suppose a disease has a 1% prevalence, a test has a 95% true-positive rate for patients, and a 5% false-positive rate for non-patients. Let DD represent having the disease, and ++ represent a positive test. The overall probability of a positive result is P(+)=0.95×0.01+0.05×0.99=0.059P(+)=0.95\times0.01+0.05\times0.99=0.059, so the probability of actually having the disease given a positive result is P(D∣+)=0.95×0.01/0.059≈0.161P(D\mid +)=0.95\times0.01/0.059\approx0.161. Even with a highly sensitive test, someone testing positive only has about a 16.1% chance of actually being sick — because there are far more non-patients than patients, and a 5% false-positive rate accumulates into a large number of positive results.

A Bayes calculation needs to keep both the base rate and the test’s conditional probability in view at the same time. Looking only at P(+∣D)P(+\mid D) and mistaking that 95% figure for P(D∣+)P(D\mid +) is a common mix-up of the direction of conditioning.

4. Summary: A Probability Space Is the Shared Language for Later Topics

A complete probability model can be written as (Ω,F,P)(\Omega,\mathcal F,P). Here Ω\Omega describes all possible elementary outcomes, F\mathcal F describes which sets of outcomes count as events, and PP assigns probabilities to those events. On top of this structure, a random variable X:Ω→RX:\Omega\rightarrow\mathbb R maps the full random outcome into the numerical quantity we actually want to analyze.

Random experiment

Sample space Ω

Sample point ω

Event A ⊆ Ω

Probability P(A)

Random variable X(ω)

Discrete random variable

Continuous random variable

PMF p(x)

PDF f(x)

CDF F(x)

Random experiment

Sample space Ω

Sample point ω

Event A ⊆ Ω

Probability P(A)

Random variable X(ω)

Discrete random variable

Continuous random variable

PMF p(x)

PDF f(x)

CDF F(x)

Discrete random variables use a PMF to describe probability at individual values:

pX(x)=P(X=x).p_X(x)=P(X=x).

Continuous random variables with a density use a PDF to describe how concentrated probability is near different locations:

P(a≤X≤b)=∫abfX(x) dx.P(a\le X\le b)=\int_a^b f_X(x)\,dx.

For a continuous random variable with a PDF, P(X=x)=0P(X=x)=0, but this does not mean a concrete value cannot be observed. It only means a single mathematical point has no width and therefore contains no positive probability area. The CDF

FX(x)=P(X≤x)F_X(x)=P(X\le x)

then puts discrete and continuous distributions into one shared language.

Conditional probability describes how to recompute probabilities after receiving new information; independence describes whether one event changes the probability of another; the law of total probability and Bayes’ theorem provide ways to move between different conditions and information states. Once this vocabulary is clear, later topics such as expectation, variance, joint distributions, statistical inference, and stochastic processes are all extensions built on the same foundation.

From here, you can continue on to Stochastic Processes 1: What Is a Stochastic Process?, which extends a random variable into an indexed family of random variables; for a counting-based application, see Stochastic Processes 9: The Poisson Process.