VRM 3: Measuring and Monitoring Volatility
A volatility that never moved would be easy to pin down. Feed a long run of past returns into the usual standard deviation formula, and the answer would serve for every day that follows. Markets refuse to cooperate. Volatility shifts through time, sometimes by degrees and sometimes in a jump, and the collected returns then stop looking like draws from one fixed normal distribution. What appears instead carries fatter tails than a normal distribution allows, which matters because value at risk and expected shortfall are statements about the far end of the loss distribution.
Three ways a return distribution can depart from normality
The departures sort into three groups.
First, returns can arrive with tails fatter than those a normal distribution allows. Second, the distribution can be non-symmetrical, so a loss and a gain of equal size do not carry equal probability. Third, it can be unstable, meaning the parameters describing it vary through time rather than holding still.
The third case drives the other two. Let the volatility parameter wander from one period to the next, and the returns gathered across those periods stack up into a distribution with fat tails. Why it wanders is no puzzle: stress pushes volatility up, and a calm market lets it settle back down.
Conditioning on volatility rather than fixing it
One repair keeps the normal distribution and attaches it to a volatility treated as known for the day in question. On a turbulent day the return is drawn from a normal distribution with a large standard deviation, and on a quiet day from one with a small standard deviation. This conditionally normal model is not flawless, and it is a clear improvement on the constant volatility model.
Working with it demands what the constant volatility model never asked for: a current estimate of volatility, refreshed as each return arrives. Two updating models do most of that work, the exponentially weighted moving average model and the GARCH(1,1) model. What they produce feeds the delta-normal approach for value at risk and expected shortfall, extends to longer horizons, and works on correlations too.
The clearest way to see fat tails being manufactured is to build them deliberately. Take a variable X whose mean is zero in either state, with a 50% probability that its standard deviation equals sigma 1 and a 50% probability that it equals sigma 2. Each component alone is an ordinary normal distribution.
What X follows is called a mixture distribution, and its probability density function is simply the weighted average of the two component density functions, with the weights being the probabilities of the two states.
What the mixture looks like next to a normal distribution
Set the two standard deviations at 0.5 and 1.5 and compare the result against a normal distribution of the same standard deviation. The variance of that comparison distribution is the average of 0.5 squared and 1.5 squared, which is 1.25, so its standard deviation is roughly 1.118. Any difference in shape is then not a difference in scale.
Probability mass has been drawn away from the region around the one standard deviation point. Some of it lands in the tails, which is the fat tail result, and the rest lands in the centre, which is why the mixture is visibly more peaked. Extra weight in the tails has to be paid for out of the shoulders, and a symmetrical fat-tailed distribution generally behaves this way.
Volatility is not the only parameter that moves. Means move as well, and for reasons that are easy to name. The expected return on equities is a risk-free rate plus a risk premium. Both components shift through time in practice, so their sum shifts too.
An unstable mean can be modelled with the same device. Mix a normal distribution whose mean is 0.3 and whose standard deviation is 0.5 with one whose mean is 1.5 and whose standard deviation is 2.0, and the result is skewed, because the components no longer sit on top of one another. A mixture sharing a mean but differing in standard deviation gives fat tails; a mixture differing in both gives a non-symmetrical distribution.
Why this lesson concentrates on volatility
Both parameters matter, yet not equally at a one-day horizon. Over a single day the standard deviation of the return dominates the mean return, and setting the mean daily return at zero does little damage. Over a year the same move would be indefensible, since a stock might carry an estimated mean return of 12% against a volatility of 20%.
What fat tails do to the risk measures
Value at risk at a high confidence level and expected shortfall both live in the tail, which is where a fat-tailed distribution departs most from a normal one. Fit a normal distribution to returns that are truly fat-tailed and the extreme quantiles come out too small, so the figures understate the exposure at exactly the confidence levels boards and regulators care about. Expected shortfall suffers more, since it averages the losses beyond the threshold.
A second implication shapes the rest of this lesson. If those fat tails come from a moving volatility rather than from a genuinely strange distribution, the cure is not an exotic distributional assumption but a better estimate of today’s volatility.
Fat tails produced by a stochastic volatility, meaning a volatility that changes through time in a way nobody can predict, are easier to reason about once two models are held apart.
Where returns are unconditionally normal, every day’s return is drawn from one and the same normal distribution, with a single standard deviation applying throughout. In a model where returns are conditionally normal, each day still delivers a normal return, but the standard deviation attached to that day varies: high in some stretches, low in others.
The consequence is the fat tail result again, arriving by a slightly different door. Pooling normal distributions of differing standard deviations produces an unconditional distribution with fat tails.
Why the data show fat tails
This resolves the apparent contradiction between a normal model and non-normal data. When daily returns are collected and plotted, what is observed is the unconditional distribution. Each observation came from its own conditional distribution, and once they are mixed together those individual standard deviations are no longer visible, so fat tails in the historical record are exactly what a conditionally normal world should generate.
What conditioning buys in practice
Suppose the standard deviation of returns on some asset has averaged 1% per day across its history, while a monitoring procedure indicates that volatility right now is 2%. Assuming normality with a standard deviation of 2% yields more accurate value at risk and expected shortfall than fitting a fat-tailed distribution to the historical sample, and much more accurate figures than assuming normality with a standard deviation of 1%. That last option is wrong about the scale of today’s distribution by a factor of two.
Slow changes in volatility versus regime switching
Volatility moves in two distinguishable ways, and they call for different responses. The first pattern is gradual drift. Volatility today resembles volatility yesterday, and change accumulates over weeks rather than arriving at once. Volatility clustering of this kind is what Benoit Mandelbrot described when he wrote that “large changes tend to be followed by large changes” and that small changes are similarly followed by small ones. A high volatility stretch shows large price moves in both directions, a low volatility stretch small ones, and nothing in the observation concerns the direction of returns.
Gradual movement is the friendly case. If today’s volatility is close to yesterday’s, recent returns carry usable information about the current level, and every method in this lesson rests on that premise.
The second pattern breaks the premise. Volatility sometimes moves abruptly, so that the clustering description fails. An unexpected event or a government action can lift volatility from 1% to 3% per day almost overnight, and when markets settle it can drop back to 1% per day just as quickly. Sudden shifts of this kind between distinct volatility levels are called regime switching.
Regime switches are awkward for risk managers because they are normally unanticipated. A model calibrated in one regime carries no information about when the next begins, so a method that has been working can stop working without warning, not because the arithmetic broke but because the economic environment underneath it changed. Any rule based on recent returns also lags a jump, since the observations it needs only arrive after the regime has changed.
Loosely put, the volatility of a variable captures how much its value moves through time, so a stock price with high volatility posts larger daily swings than one with low volatility. The definition used in risk management is sharper: volatility is the standard deviation of the return over one day, and the square of that quantity is the variance rate.
Write the return on day i as r subscript i. For an asset paying no income it is the price change over the day divided by the price at the start of the day.
Two simplifications to the textbook standard deviation
The sample standard deviation formula applied to the m most recent returns would divide the sum of squared deviations about the sample mean by m minus 1. Risk management trims that in two places. The divisor m minus 1 is replaced by m, and the sample mean is set to zero.
Each change has a defence. Dividing by m minus 1 delivers an unbiased estimate, while dividing by m delivers the maximum likelihood estimate, the one most likely given the data. Dropping the mean is reasonable because the quantity wanted is the expected mean return, not the historical one, and a high or low mean in the sample does not imply a high or low mean next period. An accurate estimate of a mean also needs far more data than an accurate standard deviation.
Averaging absolute returns instead
An alternative estimator averages absolute returns rather than squared ones. It is more robust when the distribution is not normal, and for fat-tailed distributions it often forecasts better. The squared-return formula above remains the one in widest use, and it is the basis for everything that follows.
Returns on ten successive days are given in the table below. Use Equation (3.1) to estimate the daily volatility.
| Day | Return | Squared return | Running total of squared returns |
|---|---|---|---|
| 1 | +1.1% | 0.000121 | 0.000121 |
| 2 | -0.6% | 0.000036 | 0.000157 |
| 3 | +1.5% | 0.000225 | 0.000382 |
| 4 | -2.0% | 0.000400 | 0.000782 |
| 5 | +0.3% | 0.000009 | 0.000791 |
| 6 | +0.4% | 0.000016 | 0.000807 |
| 7 | +1.8% | 0.000324 | 0.001131 |
| 8 | -0.3% | 0.000009 | 0.001140 |
| 9 | -0.3% | 0.000009 | 0.001149 |
| 10 | -0.4% | 0.000016 | 0.001165 |
Source: daily returns as given in the chapter. The last two columns are computed here.
Equation (3.1) leaves one decision open, and it is not a small one. How many days should m cover? Both answers are unattractive.
Set m too high and the estimate stops describing the present. With m equal to 250 the calculation runs across a calendar year, and since volatility is often cyclical, the first six months of that window can sit well above or below where volatility stands today. Set m low instead, at 10 or 20 days, and the data are relevant to current conditions, but so few observations enter that the estimate becomes unreliable.
Putting a number on the imprecision
A standard error measures the typical distance between an estimate and the true value, being the standard deviation of that gap. For a volatility estimate built from m observations it is approximately the estimate divided by the square root of 2 times the quantity m minus 1.
Apply this to the ten-day example, where the volatility estimate was 1.08% and m equals 10. The standard error is 1.08% divided by the square root of 18, which is 0.25%. Confidence intervals of two standard errors either side are the common way to show accuracy, and here it stretches from roughly 0.58% up to 1.58% per day, wide enough to make the estimate nearly useless. Raising the sample to 250 observations narrows it considerably, at the cost of staleness.
The ghost effect
A second problem afflicts the equally weighted formula. Suppose a large return, positive or negative, occurred 25 days ago and volatility is computed from 50 days of data. For the next 25 days that observation sits in the sample, carrying the same weight as every other day and holding the estimate up. On the twenty-sixth day it drops out of the window, and the estimate falls sharply in one step.
Nothing happened in the market on that day. The fall is an artefact of an observation leaving the window, and any fix has to abandon equal weighting.
Equation (3.1) hands every squared return the same weight of one divided by m. Exponential smoothing keeps the requirement that the weights add to one and drops the requirement that they be equal. The result is the exponentially weighted moving average, or EWMA, which RiskMetrics used to publish volatility estimates for a wide range of market variables in the early 1990s.
The rule is a single ratio. The weight on the squared return from k days ago equals lambda multiplied by the weight from k minus 1 days ago, with lambda a constant strictly between zero and one. Call the weight on the most recent squared return w subscript 0; the next weight back is lambda times w subscript 0, and so on.
Weights that shrink by a constant factor decay fast. At a decay factor of 0.94 the weight on the return from 250 business days ago is 0.94 raised to the power 249 times w subscript 0, which is 0.00000019 times w subscript 0. Data from that far back contribute almost nothing.
Fixing the first weight
Because old data carry so little weight, an implementation started from one or two years of history behaves almost like one reaching back forever, and the infinite case is easier to work with. Summing the geometric series gives w subscript 0 divided by one minus lambda, and setting that total equal to one fixes the first weight.
Applying the weights to past squared returns and then comparing the expression for day n with the expression for day n minus 1 collapses the whole series into one line, because everything except the newest term reappears inside the previous estimate.
This is sometimes called adaptive volatility estimation, since a prior belief about volatility is revised by one new observation. It is also cheap to run: only the most recent variance estimate has to be stored, and each return can be discarded once it has been used.
Volatility is estimated today at 2% per day, the return just observed is -1%, and the decay factor is 0.94.
Choosing the decay factor
One parameter governs the whole model, and its value decides how quickly the estimate reacts. The best choice is the one whose estimates carry the lowest error, and RiskMetrics settled on 0.94.
Consider the extremes. A decay factor of 0.995 puts 99.5% of the weight on yesterday’s variance rate and only 0.5% on the newest squared return, so even several turbulent days would move the estimate very little. A decay factor of 0.5 lets the newest observation claim half the weight, so the estimates chase every return and become too volatile to use.
RiskMetrics also wanted one decay factor for every market variable it covered, and separate fitting would probably have moved some of them away from 0.94. Two methods are used to fit it.
The first approach compares forecasts against outcomes. Compute the realized volatility for a given day from 20 or 30 days of subsequent returns, then search for the decay factor that minimises the gap between forecast and realized volatility.
The second is the maximum likelihood method, which picks the value maximising the probability of the observed data. Suppose a trial value forecasts a low volatility for some day and the return that day turns out to be large in absolute terms. Under that forecast such an observation is improbable, so the likelihood attached to the trial value falls, and the search moves away from it.
The same weighting idea transfers to historical simulation. Scenarios can be given weights that decline exponentially as the underlying observations recede into the past, determined much as EWMA determines them and required to sum to one. That turns historical simulation from a non-parametric procedure into a parametric one. The case for it is strongest when scenarios come from the immediately preceding year: when conditions turn stressed, as in the second half of 2007, heavier weights on recent scenarios let value at risk and expected shortfall register the change.
How the weights change the answer
Scenarios are ranked from the largest loss down, and the weights are accumulated down that ranking until the required tail probability is reached.
| Scenario number | Loss (USD millions) | Weight | Cumulative weight | Weight as a multiple of 0.002 |
|---|---|---|---|---|
| 490 | 7.8 | 0.0090 | 0.0090 | 4.50 |
| 492 | 6.5 | 0.0092 | 0.0182 | 4.60 |
| 2 | 4.6 | 0.0001 | 0.0183 | 0.05 |
| 23 | 4.3 | 0.0001 | 0.0184 | 0.05 |
| 48 | 3.9 | 0.0001 | 0.0185 | 0.05 |
| 367 | 3.7 | 0.0026 | 0.0211 | 1.30 |
| 235 | 3.5 | 0.0007 | 0.0218 | 0.35 |
Source: the weighted scenario table given in the chapter. The final column is computed here against the equal weight of 0.002.
Under equal weighting the 99% one-day value at risk was USD 3.9 million. With these weights it is USD 6.5 million, because the second scenario carries the cumulative weight past 1% on its own. Expected shortfall follows the same logic: inside the 1% tail the loss is USD 7.8 million with probability 0.9 and USD 6.5 million with probability 0.1, giving 6.5 X 0.1 + 7.8 X 0.9 = 7.67 million. The third, fourth and fifth scenarios barely register, being old data. These weights correspond to a decay factor of about 0.99, well above the values that work for squared returns.
Weighting by similarity instead of by age
Multivariate density estimation weights a historical day not by how recent it is but by how closely it resembles today. Interest rate volatility, measured by the standard deviation of daily percentage changes, tends to fall as rates rise, so days when rates stood near their current level are the informative ones, and the weight tapers off as the gap widens. Several conditioning variables can be combined, GDP growth alongside the level of rates being one example.
GARCH, developed by Robert Engle and Tim Bollerslev, stands for generalized autoregressive conditional heteroscedasticity, and the GARCH(1,1) model extends EWMA by adding a third input. Alongside the previous variance rate estimate and the latest squared return, weight is given to a long-run average variance rate.
The three weights must sum to one, so the third is determined by the other two, and both alpha and beta are positive.
EWMA is a special case, recovered by setting gamma to zero, alpha to one minus the decay factor and beta to the decay factor. Both models are first-order autoregressive, or AR(1), because the forecast depends on the immediately preceding value of the same variable. The pair of ones records that weight goes to one squared return and one variance rate estimate, both the most recent, while the general GARCH(p,q) uses p squared returns and q variance rate estimates. In practice GARCH(1,1) is used far more than any other version.
Substituting backwards through the series shows that GARCH(1,1) weights past squared returns exponentially, exactly as EWMA does. The extra term that never decays is the difference.
The three-parameter form and the long-run variance
Working with gamma and V subscript L together is awkward, so the product is usually collapsed into a single parameter, with omega defined as gamma times V subscript L. That leaves three parameters rather than four.
Recovering the long-run variance rate takes one line. Since the three weights sum to one, gamma equals one minus alpha minus beta, and omega equals gamma times V subscript L, so dividing omega by gamma returns V subscript L.
The sum of alpha and beta is the persistence of the model. The closer it sits to one, the smaller gamma becomes, and the weaker the pull of the long-run average variance rate on the next estimate.
A GARCH(1,1) model has been fitted with omega = 0.000003, alpha = 0.12 and beta = 0.87, so Equation (3.4) reads as 0.000003 plus 0.12 times the latest squared return plus 0.87 times the previous variance rate estimate.
The long-run term is the whole difference between the two updating models, and what it creates is a pull. When the current variance rate sits above the long-run average, GARCH(1,1) drags the forecast down toward it; when it sits below, the forecast is pulled up. EWMA has no such term and no pull. This tendency is mean reversion, and observed returns still scatter around the path it describes.
Mean reversion suits some market variables and not others. A tradable price should not revert predictably, because a predictable path in such a price is an inefficiency somebody would trade away. A volatility is not a traded price, so nothing prevents it from reverting, and unusually high or low volatility does not persist forever. The same argument covers an interest rate, which is not itself traded even though bond prices are.
From one day to T days
If volatility were constant, the variance rate over T days would be T times the one-day variance rate, so the volatility over T days would be the square root of T times the one-day volatility. That is the square root rule, matching the familiar idea that uncertainty grows with the square root of time, and it is how a one-day value at risk is scaled to a longer horizon.
Mean reversion improves on it. A currently high daily volatility is expected to fall, so scaling it up by the square root of T overstates value at risk, while a currently low one is expected to rise, so the rule understates it. Expected shortfall is distorted the same way. Under GARCH(1,1) the expected variance rate t days ahead is given below, and averaging it over the horizon gives the variance rate to use.
Suppose the current daily variance rate is 0.00010, a volatility of 1% per day, and mean reversion is expected to lift the average daily variance rate over the next 25 days to 0.00018. Basing the 25-day figure on that average rather than on today gives 6.7%, from the square root of 25 multiplied by the square root of 0.00018.
Current volatility is estimated at 3% per day, with a long-run average volatility of 2% per day, in a GARCH(1,1) model where alpha = 0.04 and beta = 0.94.
Every estimate so far has come from returns that have already happened. Implied volatility is the volatility implied by the market price of an option. Because the value of a call option or a put option rises as volatility rises, a traded option price can be inverted to recover the volatility consistent with it. That makes implied volatility forward looking, while EWMA, GARCH and every estimate built from history is backward looking. Participants pricing options are forming a view about volatility over the life of the option, and the evidence suggests they predict realized volatility better than estimates drawn from history do.
Quoting conventions and the term structure
Implied volatility is normally expressed per year. Converting to a daily figure means dividing by the square root of 252, the estimated number of trading days in a year, so 20% per year becomes 1.26% per day.
Maturity matters. A one-month option indicates the average volatility expected over the coming month and a three-month option what is expected over three months, and because volatility exhibits mean reversion these figures should not agree. A term structure sloping up from a low level, or down from a high one, is the option market pricing the pull that GARCH(1,1) models.
The VIX and the limits of the measure
The most closely watched implied volatility index is the VIX, built from the implied volatilities of options on the S&P 500 with 30 days to maturity. Typical values sit in the 10 to 20 range, which is a 30-day volatility for the index of 10% to 20% per year. In October 2008 the index twice reached 80, which implies 80% per year and about 5% per day. It spiked again in March 2020, when the coronavirus produced huge movements in the S&P 500.
Two limitations restrict the tool. Options are not actively traded on every asset, so for many exposures no reliable implied volatility exists at all. Neither limitation is a reason to ignore implied volatilities, which are best monitored alongside the estimates calculated from historical data.
Correlations move through time as volatilities do, and need monitoring for the same reason. The delta-normal model applied to a linear portfolio requires the correlations between daily asset returns as well as their volatilities, so a stale correlation matrix corrupts the portfolio figure even when every volatility is current. The updating rules carry over with one change of object: volatility rules work on variances, correlation rules on covariances. With mean daily returns assumed to be zero, the covariance between two returns is the expectation of their product.
The correlation is then recovered by dividing the covariance by the product of the two standard deviations. Both standard deviations should be updated by EWMA with the same decay factor as the covariance, since mixing decay factors can produce an inconsistent correlation estimate.
A covariance can equally be updated by GARCH(1,1). Doing so across a large matrix is another matter, since keeping many covariances mutually consistent that way becomes quite complex.
On day n minus 1 the volatility of X is 1% per day, the volatility of Y is 2% per day, and the coefficient of correlation between them is 0.2. On day n minus 1 both returns are 2%. The decay factor is 0.94.