QTA 8: Regression with Multiple Explanatory Variables
A regression built on a single explanatory variable teaches the mechanics of ordinary least squares cleanly, but it rarely matches the question an analyst wants answered. Real dependent variables respond to several things at once, and those things move together, so a slope fitted to one of them alone is contaminated by everything correlated with it that was left outside.
The k-variable model admits the other drivers explicitly. Its general form carries one constant and k slopes:
Two payoffs justify the extra structure. The first is isolation: with the other drivers in the equation, each coefficient measures the distinct contribution of its own variable rather than a blend of its own effect and those of its correlates. The second is the ability to judge something new, since a researcher proposing a fresh predictor of asset returns proves little by running it alone.
Models with several explanatory variables were historically called multiple regressions, but the distinction carries no mathematical content. The running application extends the capital asset pricing model with risk factors beyond the market, mapping a portfolio onto economically meaningful exposures so a manager’s performance can be judged once those exposures are priced out.
Moving from one regressor to many costs exactly one new assumption, with the other conditions carrying across in adjusted form. The new requirement is that the explanatory variables must not be perfectly linearly dependent. Put positively, every regressor has to contain some variation the remaining regressors cannot reproduce exactly. That surviving variation is the raw material the estimator works with, and a coefficient cannot be estimated without it.
When the requirement fails, the variables are perfectly collinear. Regressing a portfolio return on quarterly returns in local currency and on the same returns rescaled by a fixed exchange rate is the simplest illustration, since one column is a constant multiple of the other. Statistical packages return a warning or refuse to estimate such a model.
| Condition | Form in the k-variable model |
|---|---|
| Variation in the regressors | Every explanatory variable has a positive variance, for each of the k variables |
| Conditional mean of the error | The expected error is zero conditional on the full set of explanatory variables |
| Sampling | The vector of explanatory variables at each observation is independent and identically distributed |
| No outliers | The fourth moment of each explanatory variable is finite, applied to every variable |
| Constant error variance | The error variance conditional on the whole set of explanatory variables equals a constant |
Source: the assumption set as extended in the chapter, restated in words.
The conditional mean condition now demands much more, because the error must average zero given every explanatory variable simultaneously. A variable left out of the model and correlated with a variable inside it breaks that condition, which is the formal reason omitted variables damage estimates rather than merely reducing precision. Interpretation is otherwise unchanged, and only the no-perfect-collinearity requirement is genuinely new.
When the explanatory variables are distinct, meaning none is an exact function of the others, each regression coefficient reads as a partial effect: the change in the dependent variable produced by a small increase in that variable, holding all other included variables constant. Correlation among the regressors does not alter that reading. It affects precision, and how different estimators move together, but the quantity being estimated is a partial effect either way.
An industry factor model shows the interpretation in use. In a regression of industry portfolio returns on the market, size and value factor returns, the wholesale industry carries a coefficient of 0.26 on the size factor. Read as a partial effect, a 1 % return on the size factor raises the wholesale portfolio return by 0.26%, with the other two factor returns unchanged. Two stories fit: the firms may themselves be small, or their profitability may depend on the forces that drive small firms generally.
When holding other variables constant is impossible
Some models are built so that the variables are not separable. A quadratic specification is the standard case:
Here no experiment changes the first variable while the second stays put. The effect of a small change is not a single coefficient at all, and it depends on where in the range of the variable the change happens:
So the meaning of the coefficient becomes conditional on the level of the variable, and quoting a single number as its effect is wrong here. In the separable case each coefficient is estimated from the component of its own variable that no other included variable can explain, and that component is uncorrelated with the others by construction.
A two-variable regression can be taken apart into a sequence of simple regressions, which makes the partial interpretation concrete. Step one regresses the first explanatory variable on the second and keeps the residual:
Step two does the same to the dependent variable, and step three regresses one residual series on the other:
No constant appears in the last step because both residual series have mean zero by construction. The first two regressions split each variable into a piece perfectly correlated with the second explanatory variable, its fitted value, and a piece uncorrelated with it, its residual. Whatever the third regression finds cannot be an echo of the second variable, since that variable has been removed from both sides. Reversing the roles gives the other coefficient, and in a k-variable model the same route works with the other k minus 1 regressors projected out.
The chapter’s appendix holds annual data from 1999 until 2018 on the excess return of a shipping container industry portfolio, the excess return on the market, and the return on a value portfolio. Over those 20 years the three columns sum to 171.6, 120.5 and 38.9. The squared deviations of the market return around its mean sum to 6664.8, the cross-products of the industry and market deviations sum to 4638.9, and those of the value and market deviations sum to -1464.2. Fitting the full model gives an industry excess return of -4.06 plus 0.72 times the market excess return plus 0.101 times the value return, with an R-squared of 0.394 and standard errors of 4.16, 0.22 and 0.29.
Control variables are regressors included because they are known to relate to the dependent variable, not because anyone wants to read their coefficients. Their job is to occupy the ground they are entitled to, leaving the variable under investigation with only what belongs to it.
Since the market, size and value factor returns are the accepted description of how asset returns differ in the cross section, a study proposing a fourth measure includes those three as controls and asks whether it still earns a coefficient.
Leaving the controls out produces omitted variable bias. The residual procedure makes the mechanism visible: with a relevant variable missing, the regressor that remains keeps the part of itself that variable would have claimed, and its coefficient absorbs both effects at once.
An omission is therefore safe only if the missing variable has no effect on the dependent variable, or is uncorrelated with the variable of interest. Either condition alone kills the bias term, and in cross-sectional financial data neither is common.
An analyst tests a sentiment index as a predictor of monthly hedge fund returns. Regressing fund returns on the sentiment index alone produces a slope of 2.7. Adding the size factor return as a control changes the picture: the sentiment slope falls to 1.5 and the size factor earns a coefficient of 3.0. A separate regression of the size factor return on the sentiment index gives a slope of 0.4.
Controls are chosen because theory or evidence says they belong in the equation, not because adding variables improves the fit. A variable thrown in without that justification can itself distort the coefficient of interest.
Movements in the market explain a great deal of the variation in asset returns, and for a long time that was the whole model. It is not enough, because returns also differ systematically across firms and industries according to characteristics unrelated to market risk.
The Fama-French three-factor model is the leading example. It keeps the market excess return from the capital asset pricing model and adds two: a size factor, capturing the tendency of small-capitalisation stocks to outperform large-capitalisation stocks, and a value factor, capturing the extra return earned by value stocks over growth stocks. A value stock is one whose fundamentals are strong relative to its price, usually measured by a high ratio of book value of assets to market value. Firms at the other end of that ratio are called growth firms, since a high price relative to book value amounts to a forecast of rapid growth.
Both extra factors are returns on long-short portfolios, buying the stocks with the desirable characteristic and selling short those with the undesirable one. The two sides finance each other, so the risk-free rate cancels and neither factor needs an excess-return adjustment, while the market return does. Returns for all three factors, together with industry portfolio returns, come from Ken French’s data library.
Fitting the model industry by industry gives market coefficients that are all positive and close to one, with large t-statistics, which is unsurprising: industry portfolios move with the market almost by definition. The size and value coefficients are smaller, range from -0.6 to 0.8, and carry their information in the sign. A positive coefficient signals positive exposure, as with the wholesale size loading of 0.26. A large negative value coefficient, which the computer industry portfolio carries, says the typical firm there is a growth firm.
Why the factor coefficients shrink as controls are added
Estimating the exposure of the computer industry to the size factor on its own gives a much larger number than the three-factor model does. Adding the market return as a control roughly halves the size effect and reduces the value effect by 30%, and adding the other factor shrinks both again. The reason is correlation among the regressors.
| Market | Size | Value | |
|---|---|---|---|
| Market | 1 | 0.216 | -0.176 |
| Size | 0.216 | 1 | -0.256 |
| Value | -0.176 | -0.256 | 1 |
| Size and value, conditional on the market | -0.228 | ||
Source: estimated correlation matrix of the factor returns, the conditional figure computed from residuals after regressing each factor on the market.
Once the market is controlled for, the conditional correlation between the size and value factor returns is -22.8%. Because the factors overlap, a coefficient estimated without the others picks up shared movement.
Minimising the sum of squared residuals splits the variation in the dependent variable into two parts that do not overlap: what the model accounts for, and what it fails to account for and leaves in the residuals. Total variation is the total sum of squares, written TSS, the squared deviations of the dependent variable around its own sample mean:
Each observation splits into its fitted value and its residual. Subtracting the sample mean from both sides makes them mean zero, and squaring and summing produces the decomposition:
A cross-product term between the demeaned fitted values and the residuals should appear when the square is expanded, and it is absent because it equals zero. The first-order conditions force the residuals to sum to zero and to be uncorrelated with every explanatory variable, and fitted values are a linear function of those variables. Dividing the identity by the total sum of squares gives a share for each component, and the explained share is R-squared:
Because least squares makes the residual sum of squares as small as possible, it simultaneously maximises R-squared. In a single variable model R-squared is the squared correlation between the dependent variable and the regressor; with several regressors it is the squared correlation between the dependent variable and its own fitted values. The bounds follow from the decomposition: a model explaining nothing scores zero, a model fitting perfectly scores one, and every other model lands between.
A ten-observation sample from the chapter’s practice material pairs a dependent variable with two explanatory variables. The dependent values are -5.76, 0.03, -0.25, -2.72, -3.08, -7.1, -4.1, 0.14, -6.13 and 0.74. The first explanatory variable takes -3.48, -0.02, -0.5, -0.18, -0.82, -2.08, -1.06, 0.02, -1.66 and 0.68, and the second takes -1.37, -0.62, -1.07, -1.01, 0.39, 1.39, 0.75, -0.63, 1.31 and -0.15.
R-squared is easy to compute and easy to describe, and those virtues hide three problems. The first is that adding a variable can never lower it. The total sum of squares is untouched, because the dependent variable has not changed, while the residual sum of squares almost always falls. The one exception is a coefficient of exactly zero on the addition, so a rising R-squared is not evidence that a variable belongs in the model.
The second is that two models whose dependent variables differ cannot be ranked by it. A specification in levels and one in logarithms fail this test, and so do two models that are statistically identical: regressing the dependent variable on a regressor, and regressing the dependent variable minus that regressor on the same regressor, give equal residual sums of squares and identical predictions, yet different total sums of squares and so different R-squared values.
The third is that no value counts as good in the abstract. Forecasting tomorrow’s return on a liquid equity index futures contract from information available today, an R-squared of 5% would be implausibly high. Relating a well-diversified large-capitalisation portfolio to the contemporaneous market return, anything below 70% would be poor.
Adjusted R-squared addresses the first limitation, partly, by charging for the degrees of freedom that estimation consumes:
The adjustment factor g is the ratio of n minus 1 to n minus k minus 1, and it always exceeds one. Adding a variable pushes two quantities in opposite directions: the residual sum of squares falls, which raises the measure, while g rises, which lowers it. A variable earning too little reduction therefore sends the statistic down, and an exceptionally poor fit can drive it below zero. The protection is weak in large samples, where losing one degree of freedom barely moves g.
A model of quarterly portfolio returns uses 60 observations and three explanatory variables, reporting an R-squared of 0.520. A fourth explanatory variable is added and the R-squared rises to 0.524.
With the assumptions satisfied, the intercept estimator and every slope estimator obey a central limit theorem, which is what makes inference routine. Testing one coefficient in a model with many regressors works exactly as it does with one:
The formula for the standard error is messy once several correlated regressors are present, and it is never computed by hand: every statistical package reports it beside the coefficient. A confidence interval uses the same two quantities, bracketing the estimate by 1.96 standard errors on each side for a 95% interval, and the p-value for a two-sided test is twice the upper tail probability of the standard normal distribution beyond the absolute statistic.
The three-factor model fitted to the shipping container industry portfolio returns illustrates the calculation. Each t-statistic tests a null of zero, so dividing an estimate by its statistic recovers the standard error, and the interval and p-value follow.
| Coefficient | Estimate | t-statistic | Standard error | 95% interval | p-value |
|---|---|---|---|---|---|
| Market | 1.011 | 18.63 | 0.054 | 0.905 to 1.117 | 0.000 |
| Size | -0.012 | -0.16 | 0.075 | -0.159 to 0.135 | 0.128 |
| Value | 0.276 | 3.51 | 0.079 | 0.121 to 0.431 | 0.002 |
Source: three-factor regression results for the industry portfolio, with the remaining columns derived from the estimates and t-statistics.
Only the size interval contains zero, so the industry moves close to one for one with the market, tilts towards value firms, and shows no measurable size exposure once the other two factors are present.
One conversion is worth having straight, since factor models are often fitted to monthly data and quoted in annual terms. Multiplying a monthly intercept by 12 annualises it, so a beer and liquor industry intercept of 0.517 per month becomes 6.204% a year. The t-statistic does not change, because the standard error is scaled by the same factor, so a rescaling can never turn an insignificant result into a significant one.
The t-test handles one coefficient. It cannot be stretched to cover two or more, because the estimators are themselves correlated whenever the explanatory variables are, and stacking separate t-tests ignores that dependence. The F-test instead compares the fit when the null is imposed against the fit when the coefficients are free. Two models are estimated: the unrestricted full specification, and a restricted model that imposes the null and produces its own, necessarily larger, residual sum of squares.
An equivalent version runs on the two R-squared values:
If the restriction costs the model little, the two residual sums of squares are close and the statistic is small. If the unrestricted model fits far better, the gap is wide and the statistic large, the signal to reject.
Testing whether the extra factors matter
The natural application asks whether the capital asset pricing model would do just as well as a three-factor model. The null sets both the size and value coefficients to zero, so imposing it strikes those terms out and leaves a regression of the portfolio excess return on the market excess return alone. Two coefficients are restricted, so q is 2, and the statistic is compared against an F distribution with 2 and n minus 4 degrees of freedom.
Applied across industry portfolios, that test rejects the restriction at the 5% level for every portfolio except the electrical equipment and retail industries. A tighter null also fixes the market coefficient at one, so the portfolio has unit sensitivity to the market while both extra factors are absent. That version is rejected very strongly nearly everywhere, the exceptions being shipping containers, which needs a 10% significance level, and retail.
A three-factor model is fitted to 240 monthly returns on an industry portfolio. The unrestricted residual sum of squares is 1,800 and the total sum of squares is 3,000. Imposing the null that the size and value coefficients are both zero and refitting gives a restricted residual sum of squares of 1,980.
Software also reports an F-statistic for the regression as a whole, the joint counterpart of the t-statistic on a single coefficient. Its null sets every slope to zero at once, leaving the intercept untouched, so the restricted model is a regression on a constant alone whose residual sum of squares is the total sum of squares.
For the industry portfolio models this test rejects everywhere, which is no surprise given how closely industry returns track the market.
Confidence regions for two coefficients
A confidence interval for one parameter collects the values a test of a given size would fail to reject. Applying the same construction to an F-test produces the multi-parameter version, and the vocabulary shifts with it: interval is reserved for a single parameter, and the set defined over two or more jointly is a confidence region.
For two coefficients the region is an ellipse whose shape encodes the variance of each estimator and the correlation between them. When the regressors are uncorrelated the region is circular. Positive correlation between them tilts it downward, because the variables share a common component: if the first coefficient comes out above its true value, that shared movement has been over-credited to the first variable and the second must fall below its true value to compensate. Strong negative correlation reverses the tilt. Variances set size rather than tilt, since the extent of the region along an axis is inversely proportional to the variance of the corresponding explanatory variable, so halving that variance doubles the width in that direction.
Two patterns of disagreement
The first is an F-test that rejects while no t-test does, which points to multicollinearity. Take two regressors correlated at 98% or above. A fit with both coefficients equal to one is then almost indistinguishable from one with coefficients of 1.5 and -0.5, or any other pair summing to the same total. The joint test detects that the pair is doing real work, while neither individual test can attribute that work to one variable rather than the other.
The second is an individual test that rejects while the joint test does not, which says one variable contributes very little. For uncorrelated regressors the F-statistic is exactly the average of the two squared t-statistics.
If the second variable has no effect, the joint statistic collapses to half the square of the first. With 250 observations the joint test needs that square to exceed 6.06, requiring an absolute statistic above 2.46, while alone the coefficient is significant once its absolute statistic passes 1.96. Everything between 1.96 and 2.46 is a zone where the individual test rejects and the joint test does not.
Pulling the chapter together: the k-variable model changes interpretation rather than foundations, since each coefficient measures a partial effect built from the component of its variable the other regressors cannot explain, and the sums of squares that summarise fit are also the raw material for testing it.