VRM 15: The Black-Scholes-Merton Model
Two papers published in 1973 changed how the market thinks about options. Fischer Black and Myron Scholes wrote one, reaching the pricing relationship by way of the capital asset pricing model, which linked the return on a stock to the return on an option written on it. Robert Merton wrote the other, taking a no-arbitrage route of the kind that supports valuation on a binomial tree. Different starting points, identical formula.
The recognition came later and was incomplete. Scholes and Merton received the Nobel prize for economics in 1997 for developing the model. Black had died in 1995 and so could not be named alongside them.
What the formula covers and what it does not
In its basic form the pricing formula values European options on a stock that pays no dividends during the life of the option. The framework stretches further, to European options where the stock pays discrete dividends, and to European options written on stock indices, currencies and futures, each handled by a small adjustment to one input.
What it does not reach is the American option, exercisable at any time up to maturity. That right has no place in a formula built on a single terminal payoff, so American options go on a binomial tree instead. One exception matters: an American call on a stock paying no dividends should never be exercised early, so the European formula prices it exactly.
The model starts from one statement about a stock paying no dividends: over a short interval, the return on the stock is normally distributed. Write the mean return as mu and the volatility as sigma. Over an interval of length delta t, the return is normal with mean mu multiplied by delta t and standard deviation sigma multiplied by the square root of delta t.
Take a stock trading at USD 100, with mu set at 12% per year and sigma at 25% per year, and look one week ahead. The mean weekly return is 12% divided by 52, which is 0.23%. The standard deviation is 25% multiplied by the square root of one fifty-second, which is 3.4701%, or about 3.47%. A 95% confidence interval runs 1.96 standard deviations either side, so it covers 0.23% plus or minus 1.96 multiplied by 3.47%, roughly -6.6% to +7.0%. Those bounds put the stock at the end of the week between USD 93.4 and USD 107.
Normal returns over short periods give a lognormal price over long ones
Compounding a sequence of normally distributed short-period returns leaves the logarithm of the price normally distributed rather than the price itself, and a variable whose logarithm is normal has a lognormal distribution. A normal distribution is symmetric and can take any value from negative infinity to positive infinity. A lognormal distribution is skewed to the right and can take only positive values.
A share price cannot fall below zero, and the lognormal outcome respects that. The mechanism sits in the standard deviation of the price change itself, the price S multiplied by sigma and by the square root of delta t. That quantity shrinks as the price falls, so moves get smaller near zero and never carry the price through it.
Write the price today as S nought and the price at a future date T as S sub T. The expected future price grows at the mean return, continuously compounded.
The logarithm of the future price is normally distributed, with this mean and standard deviation.
Putting the first two side by side exposes a trap, because the mean of a logarithm and the logarithm of a mean are different numbers.
The gap is the half variance term, since taking logarithms of Equation 15.1 gives the logarithm of S nought plus mu multiplied by T. Averaging a variable and then transforming it is a different operation from transforming it and then averaging.
Confidence intervals over long horizons
Over a single week the price is close enough to normal for the shortcut used earlier. Over a year or two it is not, and the interval must be built on the logarithm of the price, then converted back by exponentiating each bound.
The stock from the previous section trades at USD 100. Its volatility is 25% per year and its mean return 12% per year.
The realized return over a period is the constant continuously compounded rate that would have carried the starting price to the ending price. Define it through the relation below, and invert.
Because R is a linear rescaling of the logarithm of the price ratio, its distribution follows from Equations 15.2 and 15.3: normal, with mean mu minus sigma squared over two and standard deviation sigma divided by the square root of T.
Why the expected return over a finite period is below mu
Two statements now sit next to each other and look contradictory. The expected return over an infinitesimally short period is mu, while the expected return over a finite period of length T is mu minus sigma squared over two, strictly smaller whenever the stock has any volatility. Both are correct: a return earned over T is not the arithmetic average of the returns earned over the short intervals inside it.
A three-month investment makes the point. Put USD 100 to work and suppose the monthly returns, with monthly compounding, come out at 4%, 10% and -8%. The value at the end is 100 multiplied by 1.04 multiplied by 1.1 multiplied by 0.92, which is 105.25. That same terminal value is 100 multiplied by 1.0172 cubed, so the investor earned 1.72% per month. The arithmetic average of the three returns is 2%, from 4% plus 10% minus 8%, divided by 3, and the realized 1.72% falls short of it.
Behind this sits the difference between a geometric average and an arithmetic average. Take the nth root of a product of n numbers and you have their geometric average, so 2, 3 and 4.5 give 3, against an arithmetic average of 3.1667, and unless every number is equal the geometric figure is the smaller. Compounding is a product, so realized returns follow the geometric average, and shrinking each period toward zero leaves the gap at exactly sigma squared over two.
Equation 15.5 supplies the definition the rest of the chapter uses: volatility is the annualized standard deviation of the continuously compounded return. A quick estimate favoured by risk managers takes the root mean square of recent daily returns as the daily volatility, which works well enough on daily data, but observations spaced further apart deserve a more careful calculation. Collect prices at intervals of tau years, so tau is one twelfth for monthly data and one fifty-second for weekly data, label the observations S sub zero through S sub n, and form the log price relative for each one.
Let s be the sample standard deviation of the u values, computed with n minus 1 in the denominator, and divide it by the square root of tau to annualize. The standard error deserves attention, because it is large at the sample sizes people use: doubling precision requires four times as many observations.
Ten weekly closing prices for a stock, with the price relatives and their logarithms. The running total is added here as a check.
| Week i | Stock price | Price relative | Log return u | Running total of u |
|---|---|---|---|---|
| 0 | 40.0 | |||
| 1 | 41.0 | 1.0250 | 0.0247 | 0.0247 |
| 2 | 43.0 | 1.0488 | 0.0476 | 0.0723 |
| 3 | 41.5 | 0.9651 | -0.0355 | 0.0368 |
| 4 | 39.0 | 0.9398 | -0.0621 | -0.0253 |
| 5 | 41.0 | 1.0513 | 0.0500 | 0.0247 |
| 6 | 42.5 | 1.0366 | 0.0359 | 0.0606 |
| 7 | 42.5 | 1.0000 | 0.0000 | 0.0606 |
| 8 | 43.0 | 1.0118 | 0.0117 | 0.0723 |
| 9 | 45.0 | 1.0465 | 0.0455 | 0.1178 |
| 10 | 46.0 | 1.0222 | 0.0220 | 0.1398 |
Source: prices, price relatives and log returns are the volatility calculation data of the source chapter. The running total column is added for this lesson.
Dividends need handling before the sum is taken. An investor holding the stock before its ex-dividend date collects the dividend while one buying after does not, so the price drops on that date, by an amount also shaped by how capital gains and investment income are taxed. That drop is not volatility, and the safest treatment is to strip price changes on ex-dividend dates out of the sample.
Nothing in the model says how the clock should be read. If variance accumulated evenly through calendar time, the variance of a return over three days would be three times the variance over one day, weekend included. Research on stock returns says otherwise, and the variance from the Friday close to the Monday close is much smaller than that. Work by Fama, French and Roll established the pattern. Most of what moves a price is generated by trading rather than by the passage of hours, and over a weekend information keeps arriving while positions cannot be adjusted.
The convention that follows
Derivatives markets settled on treating volatility as a trading time phenomenon rather than a calendar time phenomenon. Time is measured in trading days, days on which the market is open, both when volatility is estimated and when options are valued. The usual assumption is 252 trading days in a year.
Currencies trade on more days than most other assets, and 262 trading days per year is the figure normally used when currency options are valued. The pattern may soften as trading outside normal market hours spreads, but the convention holds for now.
The same clock has to be used for the life of the option. If an option has 56 trading days left, T in the pricing formula is 56 divided by 252, which is 0.222 years, not the calendar days divided by 365.
Seven assumptions carry the derivation. The first is the price process itself; the rest describe a market frictionless enough for the hedging argument to work.
| Assumption | What breaks without it |
|---|---|
| Stock prices behave as described earlier in this lesson, with mu and sigma held constant | The terminal distribution is no longer lognormal with known parameters |
| Trading is free of taxes and transaction costs, and every security is perfectly divisible | Rebalancing the hedge would cost more than the option is worth |
| The stock pays no dividend before the option matures | The stock drops on ex-dividend dates for reasons outside the process |
| No riskless arbitrage opportunity exists | The price is no longer pinned down by the hedge portfolio |
| Security trading is continuous | The hedge cannot be held between trades |
| Borrowing and lending are both available at one risk-free rate, held constant through time | The riskless portfolio has no single rate at which to grow |
| Early exercise is unavailable on the options considered | The terminal payoff is not the only condition to satisfy |
Source: assumption list from the source chapter. The second column is editorial.
Later research has loosened several of them. The risk-free rate r and the volatility sigma can be allowed to vary as functions of time, and results are available where dividends are anticipated, the extension developed in the final section.
Two are worth watching in practice. The market itself contradicts the constant volatility assumption, and the evidence appears when implied volatilities are compared across strike prices. The frictionless trading assumption is what makes the hedging argument tidy, since the derivation needs the hedge adjusted continuously and no real desk can do that. A desk rebalances at intervals and accepts the residual risk, a cost the model does not price.
Merton’s derivation is the one to hold on to, because the same idea drives the way derivatives desks hedge. Write the call price as c and the stock price as S. Delta is the sensitivity of the call price to the stock price: a small change in the option price divided by the change in the stock price that caused it.
Neither c nor delta is known at this stage. What is known is that one combination of the two carries no risk: sell one call option and buy delta shares of the stock, so that for a small move the gain on one leg cancels the loss on the other. Setting this riskless position up costs S multiplied by delta, less the premium c received. Since it bears no risk over the next instant, it must earn the risk-free rate, and any other return would be an arbitrage.
The binomial tree argument has the same structure, differing only in how long the hedge survives. On a tree it stays riskless for a full step. Here the call price is a continuous curve, so the gradient shifts as soon as the stock moves and the hedge holds only over an infinitesimally short interval. Merton turned that condition into a differential equation which the call price c must satisfy.
Boundary conditions turn one equation into many prices
The differential equation says nothing about the payoff or the maturity, so pinning it to a contract requires a boundary condition. For a European call with strike K and maturity T, the value at time T is the larger of S minus K and zero, and for a European put it is the larger of K minus S and zero. Other derivatives attach other boundary conditions, some considerably more complicated, and each selects a different solution from the same equation.
Solving the differential equation subject to those boundary conditions produces the two pricing formulas, for the European call price c and the European put price p.
The inputs are the current stock price S nought, the strike price K, the time to maturity T in years, the continuously compounded risk-free rate r for a maturity of T, and an estimate of the volatility sigma per year over the next T years. N is the cumulative normal distribution function, available from tables or from NORMSDIST in Excel. Four of the five inputs can be read off a screen. Volatility cannot, and that fact drives the next section but one.
A stock trades at USD 56. A European call and a European put both have a strike price of USD 60 and 18 months to maturity. Volatility runs at 30% per annum and the risk-free rate at 5% per annum.
Risk-neutral valuation appeared first alongside binomial trees, and it is a general principle rather than a trick for one model. Derivatives can be priced on the assumption that every investor is risk neutral, meaning nobody raises a required expected return to compensate for risk. The resulting price is not merely right in such a world; it is correct in the real world too. Three steps do the work: assume the expected return on the underlying asset is the risk-free rate, calculate the expected payoff on that assumption, and discount it back at the risk-free rate.
For a European call the payoff V is the larger of S sub T minus K and zero, and evaluating that expectation, which is where the calculus gets heavy, delivers Equation 15.6 exactly.
A cleaner illustration: the forward contract
The mechanics are easier to follow on an instrument whose expectation is trivial. Take a forward contract to buy a stock paying no dividends for a price K at time T. Its payoff is S sub T minus K, so its value today is the discounted expectation of that difference, which splits in two.
Set mu equal to r in Equation 15.1 and the expected terminal price under risk neutrality follows at once. Substituting it in, the two exponential factors cancel in the first term.
That is the standard forward valuation result, reached earlier by a pure arbitrage argument which never mentioned expectations or risk preferences. Two routes to one answer is the point: the shortcut is legitimate, and far easier than building a hedge portfolio for every new payoff.
Implied volatility is the value of sigma which, when substituted into the Black-Scholes-Merton formula, returns the option price observed in the market for that contract. It runs the formula backwards: price in, volatility out. No analytic expression exists, so it has to be found by iteration.
Successive bisection is the simplest method and it rests on one property of the formula: the price of an option is a continuous, increasing function of volatility. Raising the input far enough gives a price above the market price, a too high volatility, while a volatility of zero gives a price below it, the starting too low volatility. Each iteration averages the two and prices the option at that midpoint. An overshoot makes the midpoint the new too high volatility, an undershoot makes it the new too low volatility, and the bracket halves until it closes on the answer. Practitioners run faster methods for a non-linear equation of this kind, but bisection shows the logic without any machinery.
Take the option from Example 3 again: a European call with the stock at USD 56, a strike of USD 60, a risk-free rate of 5% and 1.5 years to maturity, quoted in the market at USD 7.00.
| Volatility tried | Model price | Model price less USD 7.00 | Bracket after the step |
|---|---|---|---|
| 30% | 8.3069 | +1.3069 | Too high volatility set at 30% |
| 20% | 5.6117 | -1.3883 | 20% to 30% |
| 25% | 6.9624 | -0.0376 | 25% to 30% |
| 27.5% | 7.6355 | +0.6355 | 25% to 27.5% |
Source: volatilities and model prices from the implied volatility illustration in the source chapter. The gap and bracket columns are derived for this lesson.
The same procedure works for American options, with the model price at each trial coming from a binomial tree.
Volatility smiles, skews and the volatility surface
Traders watch implied volatilities closely and often quote in them rather than in prices. Volatility indices track them at market level, the best known being the SPX VIX index, computed from 30-day options written on the S&P 500. Similar indices cover commodities, interest rates and currencies, and one even tracks the volatility of the VIX itself.
If the model assumptions held exactly, every option on one asset would carry the same implied volatility at all times. They do not. Implied volatility varies systematically with the strike price, direct evidence that quoted option prices are not consistent with the model. The pattern on currencies is roughly symmetric and is called a volatility smile. On equities it slopes down as the strike rises and is called a volatility skew.
Adding maturity gives the volatility surface, which maps implied volatility across strikes and across maturities together. To quote an option, a desk interpolates across the surface between known implied volatilities, substitutes the result into the formula, and reads off a price. That is how a model whose central assumption the market rejects continues to be used every day.
Everything so far assumed no dividends. Relaxing that is straightforward when the dividends paid over the life of the option are known or can be estimated well, which is fair for options running less than one or two years. Longer maturities are usually handled with a known dividend yield q, replacing S nought with S nought multiplied by the exponential of minus q multiplied by T and subtracting q from r inside d1 and d2.
Which dividends count, and at what size, needs care. The relevant ones are those whose ex-dividend date falls inside the life of the option, and the size that matters is the expected fall in the stock price on that date, which can differ from the declared dividend when capital gains and investment income are taxed differently.
The discrete dividend adjustment
Let D be the present value of the dividends paid over the life of the option, and split the current stock price in two: a component D that will be paid out, and a component S minus D that will still be there at maturity. Only the second is risky in the way the model assumes, so apply the volatility to it and replace S with S minus D.
A stock trades at USD 74 and a six-month European put option on it has a strike price of USD 70. The volatility parameter is 20%, and 5% per year, continuously compounded, is the risk-free rate. Two dividends of USD 1.50 each are expected, the first going ex-dividend after one month and the second after four months.
When dividends make early exercise worth it
A call option on a stock paying no dividends should never be exercised early, so Equation 15.6 prices the American call as well as the European one. Discrete dividends change that, but only at specific instants: exercise can be optimal immediately before an ex-dividend date and never between them, because holding preserves insurance value that exercising throws away. Let the ex-dividend dates be t sub 1 through t sub n, let the option mature at T, and let D sub i be the ith dividend. Exercise at an intermediate date is never optimal when the dividend there satisfies the first condition below, nor at the final ex-dividend date under the second.
When every dividend satisfies its condition, early exercise is never worthwhile and an American call on a dividend-paying stock can be valued as a European call. American puts are different. Exercising a put brings the strike forward, and cash received earlier is worth something, so early exercise can pay with or without dividends. American puts, and all American options on stock indices, currencies and futures, need a binomial tree.