Quantitative Essentials — CFA Level I

Easy

Find each statistical concept hidden in the grid. Selecting a word reveals its definition and a link to study it in depth.

10 terms · Choose how you want to study

New to the CFA Level I exam? Read our how-to-pass guide →

Study modes

Terms in this set

Mean

The arithmetic average of a set of numerical observations.

The exam rarely just asks you to add and divide — it tests which mean to use. The “tell” is the verb in the stem: averaging a sequence of prices paid per share (dollar-cost averaging) signals the harmonic mean; equal-weighting rates, ratios, or P/E multiples across a portfolio often does too (a value-weighted index average uses the weighted harmonic mean), while compounded growth over time signals geometric. A reliable check is the ordering harmonic ≤ geometric ≤ arithmetic, with equality only when all positive observations are identical — so the wrong choice usually lands too high. The weighted mean appears as expected return, where weights are portfolio allocations (or probabilities) that sum to one.

The classic trap is conflating center with spread: the mean feeds variance (every squared deviation is measured from it), but it is not itself a dispersion measure. A second trap is robustness — unlike the median, the arithmetic mean is pulled toward outliers and the skewed tail, which is exactly why mean-versus-median comparisons diagnose skew direction (mean > median signals right skew). Memory hook: harmonic for what you paid, geometric for what you earned.

Median

The middle value when observations are arranged in order; robust to outliers and skew.

The exam rarely asks you to just compute a median; the harder version hands you a skewed data set (incomes, returns, drawdowns) and asks which measure of central tendency is most appropriate, where the answer hinges on median because outliers distort the mean. Watch for the positional tell: the median sits at the (n + 1)/2-th position of the ordered observations, so a vignette handing you that location index is steering you toward it. The median is also the 50th percentile (second quartile, Q2), the center of the five-number summary, so quartile and percentile items quietly test the same skill.

The classic trap is reaching for the mean out of habit on tail-heavy data, or misreading the skew direction: positive skew pulls the mean above the median (the examTip rule), so a long right tail gives mean > median > mode. Another snare is forgetting to sort first before picking the middle value. Memory hook: the median is the “middle child” — it ignores how loud the extremes shout and just holds the center position.

Variance

The expected squared deviation of a random variable from its mean — a fundamental measure of dispersion.

Item-writers love to make you pick the divisor: read the stem for “a sample of…” versus the full population, because the wrong choice is the single most common variance error. The other classic pattern hands you a two-asset portfolio variancew₁²σ₁² + w₂²σ₂² + 2·w₁·w₂·Cov(1,2) — and tests whether you square the weights and keep that factor-of-2 covariance term. Since Cov(1,2) = ρ·σ₁·σ₂, the cross-term’s contribution flips sign with the correlation, so lower correlation shrinks total variance (the engine of diversification); note variance itself never goes negative. Watch units, too: variance is in squared units (e.g., %²), so any answer needing comparable, real-world units points to standard deviation.

Don’t conflate variance with correlation. Variance and covariance are unstandardized and unbounded (they run to ±∞), while correlation is covariance scaled to [−1, +1]. And variance measures spread around the mean, not the mean’s level — two datasets with identical means can have wildly different variances. Memory hook: variance squares, deviation roots it back.

Skewness

A measure of asymmetry in a distribution; positive skew has a long right tail, negative skew a long left tail.

The classic item gives you a mean, median, and mode and asks for the skew direction, so memorize the ordering: a positively skewed (right-tailed) distribution has mean > median > mode; reverse all three for negative skew. The tell is that the mean is dragged toward the long tail because it alone uses every observation’s magnitude, while the median depends only on rank/position. A second pattern asks for the sign of sample skewness—but watch the trap that a symmetric distribution can still be leptokurtic (fat-tailed), because skewness and kurtosis are independent.

Don’t confuse the two: skewness is the standardized third moment (asymmetry); kurtosis is the standardized fourth moment (tail thickness), so a symmetric fat-tailed series has zero skew but high positive excess kurtosis. The frequent error is assuming mean > median means the mean is the better center—it signals the opposite, that the median is more robust. (Mean > median > mode is an empirical rule for unimodal distributions, not an exact identity.) Hook: the mean chases the tail.

Kurtosis

A measure of tail heaviness; leptokurtic distributions have fatter tails (more extreme observations) than the normal distribution.

The exam’s favorite trap is the kurtosis-vs-excess-kurtosis swap: a question reports “kurtosis = 4” and asks if the distribution is leptokurtic — the answer turns on the excess value, 4 − 3 = 1 > 0, so yes. The “tell” is any value near 3; that’s mesokurtic, normal-like. Watch for platykurtic distributions (excess kurtosis < 0, thinner tails) as the wrong-direction distractor. Both measures come from the fourth moment about the mean.

The classic confusion is kurtosis vs skewness: skewness measures asymmetry (which tail is longer), while kurtosis measures tail weight, not direction — both tails fatten together (the curriculum and most question banks also describe leptokurtic shapes as more peaked, though modern statistics treats kurtosis as a tail-extremity measure, not peakedness). Don’t conflate fat tails with high variance either: two series can share the same variance yet differ in kurtosis, the leptokurtic one carrying far more extreme-outcome risk. Memory hook: “lepto = leaping” tails jump out further; platy = “plateau,” flat and squat.

Correlation

A standardized measure of linear association between two variables, bounded by −1 and +1.

Correlation equals covariance divided by the product of the two standard deviations, which is what gives it its bounded scale. In portfolio construction, lower correlation across holdings is the engine of diversification — adding an asset with correlation below 1 to an existing portfolio reduces portfolio risk for any given expected return.

Regression

A statistical technique for estimating the relationship between a dependent variable and one or more independent variables.

Simple linear regression fits a line minimizing the sum of squared residuals. The slope is interpreted as the expected change in Y for a one-unit change in X, holding everything else constant. R² measures the share of variance in Y explained by the regression — but a high R² does not validate the model’s assumptions or its causal direction.

Sampling

The process of selecting a subset of observations from a population for analysis.

Exam items rarely ask “what is sampling” — they hand you a scenario and make you name the bias or pick the sampling method. The classic trap is sampling error versus sampling bias: sampling error is the random gap between a statistic and the parameter, shrinks as n grows, and is not a mistake; bias is a systematic flaw that more data won’t fix. Know the named biases cold — survivorship bias (dead funds dropped from a database inflate returns), data-snooping/data-mining bias, look-ahead bias, and time-period bias. The central limit theorem is the workhorse: for n ≥ 30 the sampling distribution of the mean is approximately normal regardless of the population’s shape, which is what lets you build a confidence interval.

Distinguish stratified random sampling (divide into strata, sample within each — greater precision, used for bond-index replication) from cluster sampling (sample whole clusters — cheaper, usually less precise). Don’t confuse a sampling distribution with a probability distribution of one variable, and remember the standard error (of the mean) is not the standard deviation of the data.

Quantile

A value that divides a distribution into equal-sized groups — quartiles, deciles, and percentiles are the common cases.

The classic item gives you a small ordered data set and asks for a specific percentile, so the whole problem is the position formula: the location of the yth percentile is L_y = (n + 1)(y/100), where n is the number of observations. The “tell” is a non-integer answer — say L = 3.75 — which means you linearly interpolate between the 3rd and 4th ordered values, adding 0.75 of the gap to the 3rd value. The trap is grabbing the value at position 3.75 instead of interpolating, or forgetting the +1 (it’s n+1, not n). Order the data smallest-to-largest first; an unsorted list is the setup for a wrong answer.

Distinguish the quantile (a position-based cut point) from the median, which is just the 50th percentile (the second quartile), and from probability, a 0-to-1 likelihood measure rather than a data value. Students conflate the quartile number with the quartile value — Q1 is the observation at the 25th percentile, not “25.” Memory hook: q is for “quarter-style cuts,” any equal slice you choose.

Probability

A numerical measure of the likelihood that an event will occur, bounded between 0 and 1.

The exam loves to hand you the general addition rule, P(A or B) = P(A) + P(B) − P(A and B), then trick you with non-mutually-exclusive events where students forget to subtract the overlap and overstate the union. Watch for the total probability rule, where an unconditional P(A) is rebuilt by weighting conditional probabilities across mutually exclusive, exhaustive scenarios — that setup is almost always the lead-in to a Bayes’ update. Another staple is converting between probabilities and odds: odds for an event equal P/(1 − P), odds against are the reciprocal (1 − P)/P, and a common slip is reporting one when the item asks for the other.

A deeper trap is treating independence and mutual exclusivity as near-equivalent when they are nearly opposites: two events with positive probability that are mutually exclusive cannot be independent, since one occurring forces the other’s probability to zero. Distinguish probability from sampling — probability assumes the model is known and computes outcome likelihoods, whereas sampling estimates unknown parameters from data — and from regression, which fits a relationship between variables rather than quantifying uncertainty.

More Quantitative Methods study sets

This is the only Quantitative Methods set so far.

All Quantitative Methods sets and terms → · All CFA Level I study games → · Not sure where to start? Take the CFA Level I diagnostic →