QTA 12: Measuring Returns, Volatility, and Correlation
Every volatility estimate starts from a series of returns, so the definition picked at the outset shapes what follows. Buy at one close, sell at the next, and measure the gain against the price paid: the simple return.
Linking simple returns across periods calls for multiplication, since each period earns on whatever wealth the previous one left.
The alternative differences the natural logarithm of the price, producing the continuously compounded return, usually called the log return. Its advantage is addition across periods, and conversion runs through an exponential.
Over small moves the two barely differ: between -12% and 12% the approximation error stays under 0.85%. Stretch the range to plus and minus 50% and that comfort evaporates, since -40% simple pairs with -51% log and 40% simple pairs with only 33.6% log. The log return also ignores the floor at -100% that the simple return respects, falling past it once the simple return drops below -63%.
Each measure owns a domain. Log returns add through time, which is why volatility models on daily or weekly data use them. Simple returns add across assets, so a portfolio simple return is the weighted average of its holdings, a property the nonlinear logarithm destroys, leaving the geometric mean in its place.
In October 1987, on the session remembered as Black Monday, the Dow Jones Industrial Average closed at 1,738.74 against a previous close of 2,246.74.
Volatility is the standard deviation of returns, and its level decides how bad a routine bad day looks. With normal returns and annualized volatility of 10%, losses beyond 1% arrive on 5% of days and on 99.9% of days the loss stays under 2%. At 100% annualized volatility, 5% of days lose more than 10.2% and one day in four loses more than 4.2%.
Shocks are independent and identically distributed, and their distribution need not be specified, though a standard normal is the usual choice.
Why volatility scales with the square root of the horizon
Build a weekly return by adding five daily returns. The five means add to five times mu, and since the shocks are independent their variances add as well, so weekly variance is five times sigma squared and weekly volatility is sigma multiplied by the square root of five. Mean and variance grow in proportion to the holding period while volatility grows with its square root, a rule that leans on returns having no serial correlation.
Variance, also called the variance rate, is volatility squared and annualizes by multiplying rather than by taking a root. The standard estimator puts the sample average in place of the mean.
| JPY/USD | S&P 500 | Gold | |
|---|---|---|---|
| Daily | 8.5% | 16.0% | 15.0% |
| Weekly | 9.7% | 17.1% | 17.6% |
| Monthly | 9.8% | 14.5% | 17.1% |
Source: the London gold fix price, the S&P 500 index, and the Japanese yen against the US dollar, with weekly returns run Thursday to Thursday. Differences across frequencies are estimation error rather than different underlying volatilities.
Annualized volatility on an equity index is quoted at 20%, and its mean return is quoted at 9% per year.
Everything so far looks backwards, summarising dispersion that has already happened. Option prices read a market judgement about volatility still to come. A European call pays off with a kink at the strike.
Because that payout is convex in the asset price, its expected value responds to the variance of the return. The Black-Scholes-Merton model ties the call price to the riskless interest rate, the strike price, the price of the asset today, the years left until maturity, and the annual variance of the return.
Solving that relationship for the one unobserved quantity gives the implied volatility, the sigma reproducing the traded call price. It arrives already annualized, so none of the scaling above applies. An option quoted at 20% implied volatility gives 20% divided by the square root of 252, or 1.26%, as a daily figure.
The model leans on simplifications markets do not honour, above all a variance fixed through time. Data contradict that flatly, so implied volatilities refuse to agree across options on the same asset at different strike prices or maturities, and one asset generates a surface rather than a number.
The VIX Index takes a broader reading, reporting implied volatility on the S&P 500 for the coming 30 calendar days from options spanning a wide range of strike prices. The methodology has been carried to other underlyings, among them other key equity indices, individual stocks, gold, crude oil, and US Treasury bonds. It can only be computed for assets with a large and liquid derivatives market. Where it can be computed it is forward looking, unlike a backward-looking estimate built from historical returns.
A mean and a variance pin down a normal distribution completely, and a normal is symmetric and thin-tailed, with skewness of zero and kurtosis of exactly 3. Return series are skewed, often heavily, and fat-tailed, so the first two moments leave out most of what a risk manager cares about.
| Kurtosis | Skewness | p-value | JB statistic | |
|---|---|---|---|---|
| Gold, daily | 17.57 | 0.31 | 0.000 | 89571.3 |
| Gold, weekly | 13.08 | 0.93 | 0.000 | 9139.2 |
| Gold, monthly | 7.05 | 0.63 | 0.000 | 358.6 |
| Gold, quarterly | 6.65 | 1.03 | 0.000 | 116.7 |
| S&P 500, daily | 23.79 | -0.74 | 0.000 | 182548.7 |
| S&P 500, weekly | 11.68 | -0.68 | 0.000 | 6715.3 |
| S&P 500, monthly | 5.26 | -0.65 | 0.000 | 136.2 |
| S&P 500, quarterly | 3.99 | -0.62 | 0.000 | 16.6 |
| JPY/USD, daily | 6.95 | -0.38 | 0.000 | 6766.7 |
| JPY/USD, weekly | 7.13 | -0.61 | 0.000 | 1610.9 |
| JPY/USD, monthly | 4.09 | -0.21 | 0.000 | 27.4 |
| JPY/USD, quarterly | 3.14 | -0.20 | 0.557 | 1.2 |
Source: prices sampled daily, weekly on a Thursday to Thursday basis, monthly, and quarterly. The p-value tests zero skewness with kurtosis of 3.
All twelve series carry skewness away from zero and kurtosis above 3. Skewness barely moves as the interval lengthens from a day to a quarter while kurtosis falls sharply, so fat tails are far more a feature of high frequency data. The sign splits by asset: equities and the yen against the dollar are negatively skewed while gold is positively skewed, because equities sell off hard in a crisis and gold attracts money as a flight-to-safety asset.
Volatility clustering is what manufactures the fat tails
Heavy tails are largely the fingerprint of volatility that changes over time. Quiet weeks arrive in runs and violent weeks arrive in runs, a pattern called volatility clustering. Pool a calm regime with a turbulent one and the mixture carries more probability far from the centre than any single normal does.
Returns themselves show almost no autocorrelation, since yesterday’s direction says little about today’s, while squared returns display strong and slowly decaying autocorrelation, because the size of yesterday’s move is highly informative about today’s. That also explains why the normal approximation improves at longer horizons, as the falling kurtosis down each block shows.
A skewness of -0.65 in a table is not proof that the departure from normality is real rather than an accident of a finite sample. The Jarque-Bera test turns the two sample moments into a formal hypothesis test, with a null hypothesis of zero skewness and kurtosis of 3, against an alternative that either differs.
Under normality sample skewness is asymptotically normal with a variance of 6, so squared skewness divided by 6 is chi-squared with one degree of freedom. Sample kurtosis is asymptotically normal with a mean of 3 and a variance of 24, so its squared deviation from 3, divided by 24, is chi-squared with one degree of freedom too. The two are asymptotically independent, so the statistic is chi-squared with two degrees of freedom.
A 5% test then uses a critical value of 5.99 and a 1% test uses 9.21, and the p-value is one minus the chi-squared cumulative distribution function at the statistic. Applied to the twelve series above, the test rejects in eleven, the lone survivor being the quarterly yen against the dollar, whose statistic of 1.2 carries a p-value of 0.557. Equities are rejected even quarterly, though the quarterly 16.6 is a tiny fraction of the daily 182548.7.
An asset return series has skewness of -0.2 and kurtosis of 4.
Testing a whole distribution for normality is one route. Another ignores the middle and studies only the far tail, where losses that threaten a firm actually live. For a normal variable the probability of a return beyond k standard deviations collapses very quickly as k rises, while many other distributions let it fade far more slowly.
The tail index governs the speed of decay: a small alpha means extreme observations stay a live possibility far from the mean, which is what makes an asset heavy-tailed. The Student’s t is the standard example of a distribution with a power law tail.
Take the natural logarithm of the probability that a variable with mean zero and unit variance exceeds a level, and the shapes separate at once. The normal is quadratic in that level, so it plunges, while a Student’s t is close to linear and descends gently, which is what fat tails mean. At six standard deviations the log probability is around -20 for the normal and around -7 for a Student’s t with 4 degrees of freedom.
Assuming normality when returns are fat-tailed understates the probability of a large loss and never overstates it, because a thin tail assigns too little probability where the damage happens. A loss of four standard deviations carries a probability of 0.003% under a normal, which is 3 chances in 100,000, while under a standardized Student’s t with 4 degrees of freedom it carries 0.24%, roughly 80 times more likely. Any capital number, stress scenario, or value at risk estimate resting on normality inherits that optimism.
In a portfolio, the shape of the return distribution is governed less by any single holding than by how the holdings move together. Weak links mean large gains from diversification. Links that tighten in the tails mean the probability of a severe portfolio loss is far higher than the individual assets suggest.
The clean benchmark is independence, which holds when the joint density factors into the product of the marginal densities.
Any pair failing that condition is dependent, and financial assets fail it in linear and nonlinear ways at once. Covariance measures the average product of the two deviations from their means, and dividing by the two standard deviations rescales it into the linear correlation estimator, also known as Pearson’s correlation.
The same quantity surfaces again as the slope of a regression.
That slope is zero exactly when the correlation is zero, which is the precise sense in which correlation captures linear dependence alone. Nonlinear dependence comes in an unlimited number of shapes that no single statistic summarises. A vivid case is common heteroskedasticity, where the volatilities of several assets rise and fall together so that turbulent days are turbulent everywhere at once. Linear correlation is blind to it, so a correlation near zero is compatible with two assets that reliably blow up in the same week.
A correlation matrix reads as the covariance matrix of variables rescaled to unit variance. That imposes a constraint once three or more variables are involved, because any portfolio built from them must have a variance that is not negative, so correlations cannot be chosen freely one pair at a time.
Take a trivariate normal vector with mean zero whose three pairwise correlations are each set to -0.9. The variance of an equally weighted average is one ninth of three unit variances plus three doubled covariance terms, which comes to -0.267. A variance cannot be negative, so the matrix is invalid. That requirement, that every weighted average carry a non-negative variance, is positive definiteness.
Estimating a large matrix pair by pair from noisy data runs into this problem, so practitioners impose structure. One choice sets every pairwise correlation to a single value, a model known as equicorrelation. The other assumes correlations arise from shared exposure to a common factor, as the capital asset pricing model implies, and writes each pairwise correlation as a product of two loadings that each lie between -1 and 1.
What the one-factor model does to normally distributed variables
Several properties follow. Each variable is standard normal, and every pair is bivariate normal with correlation equal to the product of the two loadings, so the matrix is positive definite automatically and needs one parameter per asset instead of one per pair. When the loadings share a sign every pairwise correlation is positive, and none can exceed the smaller loading in absolute value. Conditioning on the factor makes the variables independent, which leaves a normal one-factor model with zero tail dependence: far into the tail, joint extreme losses behave as though the assets were independent. Real portfolios do not, which is why the structure is convenient rather than complete.
Two alternatives are used when linear correlation is suspected of missing the point: rank correlation, also called Spearman’s correlation, and Kendall’s tau. Each is scale invariant, bounded between -1 and 1, zero under independence, and signed with the direction of the relationship.
Rank correlation: the linear estimator applied to ranks
Replace every observation by its position within its own series, giving 1 to the smallest and n to the largest, then correlate those positions. A tied group takes the average of the positions it occupies, so 2, 5, 5, 5, 9 receives ranks 1, 3, 3, 3, 5. Without ties, a shortcut needs only the paired rank differences.
Discarding values and keeping order buys two things. Rank correlation is robust to outliers, since a wildly extreme observation is still merely the largest rank, and it is invariant to any monotonic increasing transformation rather than only a linear one, which suits it to comparing an asset with a derivative written on it. Linear correlation survives only transformations of the form a plus b times X with b positive. Rank correlation is less efficient, so it usually serves as a robustness check.
Kendall’s tau: counting concordant and discordant pairs
Kendall’s tau works from pairs of observations. Two are concordant when both coordinates of one exceed both coordinates of the other, discordant when they disagree, and neither when a tie appears in either coordinate.
Read as a probability statement, tau is the chance of concordance minus the chance of discordance, so all pairs concordant gives 1 and all discordant gives -1. Concordance turns on ordering alone, so monotonic increasing transformations leave tau untouched, and ties, which pull it toward the middle, rarely trouble return data. Both alternatives survive a monotonic transformation of either variable, where Pearson’s correlation can shift, because they read the ordering of the observations rather than their magnitudes.
| Linear DGP | Nonlinear DGP | |
|---|---|---|
| Kendall’s tau | 0.500 | 0.667 |
| Rank correlation | 0.690 | 0.837 |
| Linear correlation | 0.707 | 0.291 |
Source: population values where the conditional mean is linear in one process and exponential in the other.
For a bivariate normal the rank correlation tracks the linear correlation almost exactly, while tau rises along a curve and stays much closer to zero, so a tau smaller in magnitude than the sample correlation but of the same sign is ordinary rather than evidence of nonlinearity. Under the linear process, 0.707 and 0.690 are near neighbours and tau sits lower at 0.500. Under the nonlinear process the ordering inverts, with 0.837 and 0.667 both above the linear correlation of 0.291, and that inversion is what marks the dependence as nonlinear.
Ten paired observations are collected. The X series is 0.22, 1.41, -0.30, -0.59, -3.08, 1.08, -0.45, 0.40, -0.75, 0.24, and the matching Y series is 2.73, 6.63, -2.19, -6.51, -0.99, 2.63, -3.40, 5.10, -5.14, 1.14.
Five paired observations are collected. The X values are 3.12, -1.26, 2.08, -0.28, -1.96, and the matching Y values are 2.58, -0.05, -0.72, -0.52, -0.40.