FI 2 – The Arbitrage-Free Valuation Framework
Every valuation tool in this reading rests on one sentence: if a position requires no net outlay of your own money and carries no risk, its expected return must be zero. Prices in a well-functioning market are pushed around until that statement is true. Arbitrage-free valuation is simply the practice of assigning values that are consistent with that condition, so that no trader can assemble a package of securities that produces something out of nothing.
Valuing any financial asset, from a one-year discount bill to an interest rate swap, follows the same three steps.
- Step 1. Estimate the future cash flows.
- Step 2. Determine the discount rate, or set of discount rates, appropriate to those cash flows.
- Step 3. Compute the present value of the cash flows using the rates from Step 2.
Step 2 is where the difficulty lies. Imagine for a moment a world in which nothing can default and the benchmark yield curve is flat. One risk-free rate then applies to every cash flow whenever it arrives, and pricing a bond is a single discounting exercise. Leave that world and two things change at once: cash flows become uncertain, because borrowers can fail to pay, and the yield curve acquires a shape, so that money delivered in three years is not discounted at the same rate as money delivered in one. Throughout this reading the words yield, interest rate and discount rate are used interchangeably.
A bond is a portfolio of zero-coupon claims
The traditional method discounts every cash flow of a bond at a single yield-to-maturity, which implicitly treats the curve as flat. The arbitrage-free method treats the bond as what it actually is: a bundle of separate zero-coupon claims, each maturing on a different date and each entitled to its own discount rate. The term structure of those single-payment discount rates is the spot curve, and summing the present values of the individual claims produces a value that cannot be arbitraged.
The reason is straightforward. Ignore transaction costs and suppose the whole bond traded well below the combined value of its individual cash flows. A trader would buy the bond, sell claims on each of its cash flows separately, and keep the difference with no risk taken and no capital tied up. Because that trade is available to everyone, the gap closes. The same logic scales upwards: however complicated a security is, each of its components has to carry an arbitrage-free value. A bond with an embedded option, for instance, decomposes into an arbitrage-free option-free bond plus the arbitrage-free value of the option itself.
The law of one price
The principle behind all of this is the law of one price: two goods that are perfect substitutes must trade at the same price when trading is costless. Put two identical instruments side by side at different prices and anyone can buy the cheap one, sell the dear one, and bank the difference. Nothing stops that trade being repeated, so the prices converge. Arbitrage opportunities are, by definition, violations of this law.
Two kinds of arbitrage opportunity
An arbitrage opportunity is a transaction that requires no cash outlay and delivers a riskless profit. Opportunities come in two recognisable forms.
- Value additivity. The whole should be worth the sum of the parts. When a portfolio trades for less than the securities inside it, buy the portfolio and sell the components.
- Dominance. A risk-free future payoff must carry a positive price today, and two risk-free instruments must be discounted at the same rate. When one is priced generously relative to the other, sell the expensive one and buy the cheap one in the ratio that makes today net to zero.
Four risk-free instruments pay off one year from today. Prices today and payoffs are as follows.
| Asset | Price today | Payoff in one year |
|---|---|---|
| A | 0.952381 | 1 |
| B | 97 | 105 |
| C | 100 | 105 |
| D | 200 | 220 |
Sell 105 units of Asset A, raising 100, and use 97 of that to buy Asset B. Today the trade nets 100 − 97 = 3, received with certainty. In one year Asset B pays 105 and the 105 short units of Asset A require 105, so the future position nets to zero. The 3 is a riskless profit taken at inception, and the trade is repeated until the two prices meet.
Sell two units of Asset C for 2 × 100 = 200 and spend the whole 200 on one unit of Asset D. Today the construction costs nothing. In one year Asset D pays 220 while the two short units of Asset C demand 2 × 105 = 210, leaving 220 − 210 = 10 free of risk. A position built for nothing that pays 10 with certainty cannot survive: buyers of D and sellers of C will move the prices until it disappears.
Stripping and reconstitution
The decomposition is not merely an accounting device. In sovereign debt markets dealers can physically separate the individual cash flows of a bond and trade each as a zero-coupon security, a process called stripping, and they can buy back the right set of strips and reassemble the original coupon bond, a process called reconstitution. A five-year Treasury paying a 2% coupon semiannually is, in this sense, eleven instruments: ten coupon payments and one repayment of principal at maturity.
Because stripping and reconstitution are both available, value additivity is enforced by trading rather than by assumption. Arbitrage-free valuation is precisely the valuation method that leaves no profit in either direction. Two useful consequences follow. Bonds with the same maturity but different coupons are different bundles of zeros and must be valued as such, rather than by applying one bond yield to the other. And two cash flows carrying identical risk and arriving on the same date take the same discount rate even when they belong to different bonds.
Two 3% annual coupon, ten-year bonds are quoted in two locations each.
- Bond A. The yield in New York City is 2.5%. The same bond changes hands in Chicago at USD 104.376 per 100 of face value.
- Bond B. The yield in Hong Kong SAR is 3.2%. The same bond changes hands in Shanghai at RMB 97.220 per 100 of face value.
Bond A. 3/1.025 + 3/1.0252 + … + 103/1.02510 = 104.376. The Chicago price is 104.376. The two agree, so no arbitrage exists.
Bond B. 3/1.032 + 3/1.0322 + … + 103/1.03210 = 98.311. Shanghai is selling the identical bond at 97.220, which is below its arbitrage-free value.
Bond B is the answer. Buy in Shanghai at RMB 97.220 and sell in Hong Kong SAR at RMB 98.311, for a riskless RMB 1.091 per RMB 100 of bonds traded.
For a bond with no embedded options, the arbitrage-free value is obtained directly: discount each expected cash flow at the benchmark spot rate for its own date and add the results. Everything else in this reading exists because that simple recipe stops working once cash flows depend on the level of interest rates.
Which curve is the benchmark
Benchmark securities are the liquid, low-risk instruments whose yields anchor every other interest rate in a currency. Sovereign debt fills that role in most markets: on-the-run Treasuries in the United States, gilts in the United Kingdom and German bunds for euro-denominated issues. Where the government market is not deep enough to serve, the swap curve is the practical substitute. Throughout this reading the benchmark bonds are taken to be correctly priced, and the valuation model is built so that it reproduces their observed prices exactly.
Bootstrapping the spot curve from par rates
Quoted benchmark yields are par rates, the yields-to-maturity of bonds trading at 100. Spot rates are recovered from them one maturity at a time, an iterative procedure known as bootstrapping. Because a par bond is priced at 100 by definition, its cash flows discounted at the unknown spot rates must sum to 100.
The same relation can be written per unit of par, using the par rate itself in place of the coupon amount, which is often quicker to handle.
The first spot rate equals the first par rate, since a one-period par bond is already a zero. Substituting it into the two-year equation leaves one unknown, which is solved for directly. That result then feeds the three-year equation, and so on out along the curve.
Benchmark annual coupon bonds yield 2% at one year, 3% at two years and 4% at three years. A three-year, 5% annual coupon bond carrying the same risk and liquidity as the benchmarks is offered at 102.7751 today, priced to yield 4%.
For the two-year rate, write the two-year par bond as one unit of par:
1 = 0.03/(1 + 0.02) + (0.03 + 1)/(1 + z2)2.
The first term is 0.029412, leaving 0.970588 = 1.03/(1 + z2)2, so (1 + z2)2 = 1.061212 and z2 = 3.015%.
For the three-year rate:
1 = 0.04/(1 + 0.02) + 0.04/(1 + 0.03015)2 + (0.04 + 1)/(1 + z3)3.
The first two terms are 0.039216 and 0.037693, leaving 0.923091 = 1.04/(1 + z3)3, so z3 = 4.055%.
The spot curve is 2%, 3.015% and 4.055%.
P0 = 5/1.02 + 5/1.030152 + 105/1.040553 = 102.8102.
The offered price is too low. The bond is mispriced by 0.0351 per 100 of par value (102.8102 − 102.7751).
The reason is instructive. Pricing the bond at a single 4% yield discounts the early coupons at 4% when the market charges only 2% for one-year money and 3.015% for two-year money. With an upward sloping curve, yield-to-maturity discounting is too harsh on the near cash flows and the bond looks cheaper than it is. With a downward sloping curve the error runs the other way, and the near cash flows are discounted too lightly.
Where the spot-rate method runs out
Discounting at spot rates settles the matter for option-free bonds, and it will do so for the rest of this reading whenever a benchmark price is needed. It cannot handle a callable or putable bond, because the cash flows of such a bond are themselves a function of the interest rate path. If rates fall far enough, an issuer calls; if they do not, the bond runs to maturity. The size and timing of the payments are therefore not known in advance, and a single fixed schedule of cash flows discounted at spot rates does not describe the security.
What is needed is a structure in which interest rates are allowed to take different values in the future, in a way consistent with an assumed level of volatility. That structure is the interest rate tree, and it does two jobs at once: it generates the interest-rate-dependent cash flows, and it supplies the rates used to discount them. Because the picture resembles a lattice, these are commonly called lattice models. Models built on the evolution of a single short-term rate are one-factor models; those that let a second rate, such as a long rate, move independently are two-factor models.
A binomial interest rate tree is a map of the values the one-period rate might take at each future date. It is binomial because at every node the rate can move to one of exactly two values in the following period. Those two values are not chosen freely. They must satisfy three conditions simultaneously: the interest rate model that governs how rates move, the assumed level of interest rate volatility, and the current benchmark yield curve.
Par, spot and forward: three views of one curve
Building a tree starts with the benchmark par curve. From par rates, bootstrapping delivers spot rates, and from spot rates the one-year implied forward rates follow by no arbitrage. The three curves carry identical information, so any one of them generates the other two. They coincide only when the curve is flat.
| Maturity (years) | Par rate | Bond price | One-year spot rate | One-year implied forward rate |
|---|---|---|---|---|
| 1 | 1.00% | 100 | 1.0000% | 1.0000% (current one-year rate) |
| 2 | 1.20% | 100 | 1.2012% | 1.4028% (one year forward) |
| 3 | 1.25% | 100 | 1.2515% | 1.3521% (two years forward) |
| 4 | 1.40% | 100 | 1.4045% | 1.8647% (three years forward) |
| 5 | 1.80% | 100 | 1.8194% | 3.4965% (four years forward) |
Benchmark bonds are priced at par, so par rates and coupon rates coincide. The par curve slopes upward, so after Year 1 the spot rates sit above the par rates.
Because all three curves describe the same set of prices, valuing a benchmark bond gives 100 whichever route is taken: apply the par rate to every cash flow, discount each cash flow at its own spot rate, or roll each cash flow back one period at a time using the forward rates. The tree is built on the third of these routes, because a forward rate is precisely a single-period discount rate and a lattice moves one period at a time.
The lognormal random walk
The interest rate model used here is a lognormal random walk. Two features recommend it. Interest rates generated by a lognormal process cannot turn negative, because the absolute size of a rate change shrinks as the rate approaches zero. And volatility scales with the level of rates, so moves are large when rates are high and small when rates are low. For a lognormal distribution the standard deviation of the one-year rate is the product of the rate and the volatility parameter.
Under this model the two possible rates at any date are fixed multiples of one another. Writing i1,L for the lower one-year rate one year from now and i1,H for the higher one, the relationship is
The intuition is that the implied forward rate from the benchmark curve behaves as the centre of the distribution of possible one-year rates at that date. The lower branch sits one standard deviation below it and the higher branch one standard deviation above, so the two branches are two standard deviations apart and differ by the factor e raised to the power 2σ. Raise the volatility assumption and the multiplier grows, pulling the branches further apart while leaving them roughly centred on the same forward rate. For example, if i1,L is 1.194% and volatility is 15% a year, then i1,H = 1.194% × e2×0.15 = 1.612%.
Nodes, notation and recombination
Each point on the tree is a node, and the first node is the root, which is simply the current one-year rate at Time 0. The rate shown at a node is the discount rate that converts payments arriving at the end of that period back to the node itself. The gap between dates is the time step, and in these illustrations it is one year, matching the annual coupon frequency.
By Time 2 there are three possible one-year rates rather than four. The subscripts record the sequence of moves: i2,LL for two down moves, i2,HH for two up moves, and i2,HL for one of each, in either order. An up move followed by a down move lands on the same rate as a down move followed by an up move, which is why the tree is described as recombining. That property matters practically: the number of nodes at each date grows in a straight line rather than doubling, which keeps the computation manageable for long-dated instruments.
With i2,LL at 0.980% and volatility at 15%, the other two Time 2 rates are i2,HH = 0.980% × e4×0.15 = 1.786% and i2,HL = 0.980% × e2×0.15 = 1.323%. The middle rate sits close to the implied forward rate two years out, with the other two spread two standard deviations either side. At Time 3 there are four possible rates, and every one of them is the lowest Time 3 rate multiplied by e raised to 2σ, 4σ or 6σ as appropriate. The general shorthand is to write the rates at date t as the centring rate multiplied by e raised to an odd or even multiple of σ, positive above the centre and negative below it.
Estimating the volatility input
Two routes are used in practice to fix the volatility parameter. The first extrapolates from recent history, taking the observed variability of interest rates over some past window as indicative of the future. The second extracts implied volatility from the market prices of interest rate derivatives such as swaptions, caps and floors, which is forward looking but depends on the option pricing model used to invert the prices.
Once a tree of rates exists, valuing a bond on it is a mechanical exercise. The method is backward induction: begin at maturity, where the value of the bond is known with certainty barring default, fill in those terminal values, then work leftwards one date at a time until the root is reached.
The reason the calculation runs right to left is that a node cannot be valued until the two nodes it leads to have been valued. The value at any node depends on the coupon due at the end of the period plus the expected value of the bond one period on, and that expectation is a simple average of the two continuation values because the lognormal model assigns equal probability to the up and down branches.
Two details of the layout deserve care. The coupon due at the end of a period is recorded against the node at the start of that period, not against the node where it arrives. And the discount rate applied is the one-year rate observed at the node being valued, that is, the rate prevailing at the beginning of the period, not the rate at the destination.
A three-year, annual pay bond carries a 5% coupon. The one-year rate is 2.0% today. At Time 1 the rate is either 5.0% or 3.0%. At Time 2 it is one of 8.0%, 6.0% or 4.0%.
| Date | Possible one-year rates |
|---|---|
| Time 0 | 2.0% |
| Time 1 | 5.0% and 3.0% |
| Time 2 | 8.0%, 6.0% and 4.0% |
Time 2. There is no further uncertainty, so each node is simply 105 discounted one year at its own rate:
105/1.08 = 97.2222.
105/1.06 = 99.0566.
105/1.04 = 100.9615.
Time 1. Now the node formula is needed. At the 5.0% node the two successors are the 8.0% and 6.0% values:
[5 + (0.5 × 97.2222) + (0.5 × 99.0566)] ÷ 1.05 = 98.2280.
At the 3.0% node the successors are the 6.0% and 4.0% values:
[5 + (0.5 × 99.0566) + (0.5 × 100.9615)] ÷ 1.03 = 101.9506.
Time 0. Roll the two Time 1 values back at the current 2.0% rate:
[5 + (0.5 × 98.2280) + (0.5 × 101.9506)] ÷ 1.02 = 103.0287.
The bond is worth 103.0287 today. Notice that the middle Time 2 value of 99.0566 is used twice, once from each Time 1 node, which is the recombining property doing its work.
A tree of rates that satisfies the volatility condition is not yet useful. Any lowest rate at a date generates a legitimate-looking set of branches, but almost all of them price the benchmark bonds wrongly. Calibration is the process of selecting, at each date, the one set of rates that reproduces the observed benchmark prices. A tree that does so is arbitrage free by construction, because it is tied to the prices that actually trade.
The procedure is iterative and works outward one date at a time. Take the benchmark par curve of 1.00%, 1.20%, 1.25%, 1.40% and 1.80%, assume volatility of 15%, and build a four-year tree. The root is the current one-year rate of 1.0000%, which is given. The two-year benchmark bond, with a 1.20% coupon, then pins down the pair of rates at Time 1. The three-year bond pins down the three rates at Time 2, and so on.
Solving the Time 1 rates
Only one number has to be found, because the higher rate is tied to the lower one by the volatility multiplier. A sensible starting guess for the lower rate is something below the one-year implied forward rate of 1.4028%, since the lower branch sits beneath the centre of the distribution.
Calibrate the Time 1 rates to the two-year benchmark bond, which carries a 1.20% coupon and trades at 100. Volatility is 15% and the current one-year rate is 1.0000%.
At Time 2 the bond pays its maturity value of 100 plus the 1.20 coupon, so 101.20 arrives at both Time 2 nodes. Roll that back to each Time 1 node:
101.20/1.016873 = 99.5208.
101.20/1.012500 = 99.9506.
Then discount to Time 0 at the known root rate, remembering the 1.20 coupon due at Time 1:
[1.20 + (0.5 × 99.5208) + (0.5 × 99.9506)] ÷ 1.01 = 99.9363.
The trial produces 99.9363, below the target of 100.0000. The trial rates are too high: discounting is too heavy. They must come down.
Check the result. The Time 2 cash flow of 101.20 rolls back to:
VH = 101.20/1.016121 = 99.5944.
VL = 101.20/1.011943 = 100.0056.
And at Time 0:
[1.20 + (0.5 × 99.5944) + (0.5 × 100.0056)] ÷ 1.010000 = 100.0000.
The tree now reproduces the market price of the two-year benchmark exactly, so the Time 1 column is fixed.
Extending the tree to later dates
The Time 2 column is found the same way, using the three-year benchmark bond with its 1.25% coupon. Four things are now held fixed: the interest rate model, the 15% volatility, the current one-year rate of 1.0%, and the two Time 1 rates of 1.6121% and 1.1943% already solved. Only the Time 2 rates are unknown, and once again only one of them has to be searched over, since the upper and lower rates are multiples of the middle one.
A natural starting value for the middle rate is the two-year implied forward rate of 1.3521%, with the upper trial rate at 1.3521% × e2×0.15 and the lower at 1.3521% divided by the same factor. Adjusting the middle rate, and with it the other two, until the three-year bond values at exactly 100.0000 gives Time 2 rates of 1.7863%, 1.3233% and 0.9803%.
Working backward through the completed three-year tree confirms the result. The Time 2 nodes each receive 101.25 at Time 3, giving 99.4731 at 1.7863%, 99.9277 at 1.3233% and 100.2671 at 0.9803%. Rolling those back with the 1.25 coupon gives 99.3488 at the 1.6121% node and 100.1513 at the 1.1943% node, and one further step at 1.0000% returns 100.0000. Repeating the exercise with the four-year benchmark bond fills in the Time 3 column.
| Date | Rates from highest to lowest |
|---|---|
| Time 0 | 1.0000% |
| Time 1 | 1.6121%, 1.1943% |
| Time 2 | 1.7863%, 1.3233%, 0.9803% |
| Time 3 | 2.8338%, 2.0994%, 1.5552%, 1.1521% |
Each column is solved against the benchmark bond of the matching maturity, holding all earlier columns fixed.
What the volatility assumption does to the tree
Volatility controls the width of the fan and nothing else. Raise it and the branches spread further from the forward curve; lower it and they collapse onto the curve. At a volatility of 0.01% the Time 1 rates are 1.4029% and 1.4026%, on either side of the implied forward rate of 1.4028%, and the Time 2 rates cluster within a whisker of 1.3521%. In the limiting case of zero volatility the binomial tree is nothing more than the implied forward curve drawn as a straight line.
| Date | Volatility 20% | Volatility 15% | Volatility 0.01% |
|---|---|---|---|
| Time 0 | 1.0000% | 1.0000% | 1.0000% |
| Time 1 | 1.6806%, 1.1265% | 1.6121%, 1.1943% | 1.4029%, 1.4026% |
| Time 2 | 1.9415%, 1.3014%, 0.8724% | 1.7863%, 1.3233%, 0.9803% | 1.3523%, 1.3521%, 1.3518% |
| Time 3 | 3.2134%, 2.1540%, 1.4439%, 0.9678% | 2.8338%, 2.0994%, 1.5552%, 1.1521% | 1.8653%, 1.8649%, 1.8645%, 1.8641% |
All three trees price the same benchmark bonds at par. Volatility changes the dispersion of the rates, not the prices the tree is calibrated to reproduce.
Return to the par curve of 2.000%, 3.000% and 4.000% at one, two and three years. The bootstrapped spot rates are 2.000%, 3.015% and 4.055%, and the implied one-year forward rates are 2.000% for the current year, 4.040% one year forward and 6.166% two years forward. Volatility is 15% for every year.
The search settles on 4.646% and 3.442%. Confirm them by rolling the Time 2 cash flow of 103 back one year:
103/1.04646 = 98.427.
103/1.03442 = 99.573.
Then to Time 0, adding the 3 coupon due at Time 1:
[3 + (0.5 × 98.427) + (0.5 × 99.573)] ÷ 1.02 = 100.000.
The calibrated Time 2 rates are 8.167%, 6.050% and 4.482%. Verify them from the Time 3 cash flow of 104:
104/1.08167 = 96.148.
104/1.06050 = 98.067.
104/1.04482 = 99.538.
Roll back to Time 1, adding the 4 coupon:
[4 + (0.5 × 96.148) + (0.5 × 98.067)] ÷ 1.04646 = 96.618.
[4 + (0.5 × 98.067) + (0.5 × 99.539)] ÷ 1.03442 = 99.382.
And to Time 0:
[4 + (0.5 × 96.618) + (0.5 × 99.382)] ÷ 1.02000 = 100.000.
The tree now prices the one-year, two-year and three-year par bonds correctly, so it is calibrated and arbitrage free. It will also price the zero-coupon bonds behind the spot curve correctly, and provided the interest rate process and volatility have been chosen sensibly, it is ready to be used on bonds with embedded options and on their risk measures.
Two arbitrage-free methods have now been assembled: discounting each cash flow at its own spot rate, and rolling cash flows back through a calibrated lattice. Because both are arbitrage free and both are anchored to the same benchmark prices, they must agree on the value of an option-free bond. Demonstrating that agreement is the standard test that a tree has been built correctly.
The test case
Take an option-free bond with four years to maturity and a 2% annual coupon, valued against the benchmark curve of 1.00%, 1.20%, 1.25%, 1.40% and 1.80%. This is not itself a benchmark bond: its coupon is above the four-year par rate of 1.40%, so it must trade at a premium to par.
Priced off the spot curve, each cash flow is discounted at the zero-coupon rate for its own date:
Now price the same bond on the four-year tree calibrated at 15% volatility. Backward induction from the Time 4 cash flow of 102 produces the node values below.
| Date | One-year rate | Bond value at the node |
|---|---|---|
| Time 0 | 1.0000% | 102.3254 |
| Time 1 | 1.6121% | 100.6769 |
| Time 1 | 1.1943% | 102.0204 |
| Time 2 | 1.7863% | 99.7638 |
| Time 2 | 1.3233% | 100.8360 |
| Time 2 | 0.9803% | 101.6417 |
| Time 3 | 2.8338% | 99.1892 |
| Time 3 | 2.0994% | 99.9026 |
| Time 3 | 1.5552% | 100.4380 |
| Time 3 | 1.1521% | 100.8382 |
Each Time 3 value is 102 discounted one year at the node rate. Every earlier node applies the coupon of 2 and the average of its two successors, discounted at its own rate.
The lattice returns 102.3254, the identical figure produced by the spot curve. This is not a coincidence and it is not evidence that the volatility assumption was correct. It follows from the calibration: the tree was forced to reproduce the benchmark prices, the benchmark prices define the spot curve, and an option-free bond is a fixed bundle of the same dated claims however it is discounted. Change the volatility assumption to 20% or to 0.01% and the tree still returns 102.3254, because all three trees price the same benchmark bonds at par.
Use the tree calibrated in Example 6, with a root of 2.000%, Time 1 rates of 4.646% and 3.442%, and Time 2 rates of 8.167%, 6.050% and 4.482%. The bond is the same three-year, 5% annual coupon bond valued from the spot curve in Example 3.
105/1.08167 = 97.0721.
105/1.06050 = 99.0099.
105/1.04482 = 100.4958.
Roll back to Time 1, adding the 5 coupon due at Time 2:
[5 + (0.5 × 97.0721) + (0.5 × 99.0099)] ÷ 1.04646 = 98.4663.
[5 + (0.5 × 99.0099) + (0.5 × 100.4958)] ÷ 1.03442 = 101.2672.
And to Time 0 at the root rate:
[5 + (0.5 × 98.4663) + (0.5 × 101.2672)] ÷ 1.02 = 102.8105.
The spot-rate calculation in Example 3 gave 102.8102. The lattice gives 102.8105. The two agree, and the difference of 0.0003 is entirely a rounding artefact: the tree rates and the intermediate node values were carried to four decimal places rather than held at full precision. The economics are identical, because both methods discount the same three dated claims against the same benchmark curve.
The practical significance is that the lattice adds nothing for an option-free bond. Its value appears the moment cash flows stop being fixed. Because the tree makes the future rate at every node explicit, it can be used to decide at each node whether an issuer would call or a holder would put, alter the cash flow accordingly, and carry on with backward induction. The spot curve alone offers no way to make that decision.
Backward induction is not the only way to extract a price from a binomial tree. Pathwise valuation reaches the same answer by a different route: enumerate every possible journey through the tree, discount the bond cash flows along each journey separately, and average the results. The three steps are to list all the paths, value the bond along each one, and take the mean across paths.
Counting the paths
How many journeys are there? The counting problem is the same one a run of fair coin tosses poses, and the answer sits in Pascal’s Triangle. Build the triangle by starting with a single 1 at the apex, running 1 down both edges, and setting every interior entry equal to the sum of the two entries immediately above it. Writing U for a move to the higher rate and D for a move to the lower rate, the sequences after each period are as follows.
| Periods elapsed | Sequences of moves | Matching triangle row |
|---|---|---|
| 1 | U; D | 1, 1 |
| 2 | UU; UD, DU; DD | 1, 2, 1 |
| 3 | UUU; UUD, UDU, DUU; UDD, DUD, DDU; DDD | 1, 3, 3, 1 |
Each triangle entry counts the sequences that finish at one end state, and the row adds up to the number of sequences. After three periods there are 1 + 3 + 3 + 1 = 8 sequences, but a three-year bond only ever visits three rates, so it has four distinct paths.
For a three-year instrument the relevant paths are the four routes that traverse the Time 0, Time 1 and Time 2 columns: up then up, up then down, down then up, and down then down. Because the tree recombines, the second and third of those share the middle Time 2 rate but remain distinct paths, and each is weighted equally in the average.
Value a three-year zero-coupon bond with a face value of 100 on the four-year tree calibrated at 15% volatility, using pathwise valuation, then confirm the answer with backward induction.
For Path 1: 100/(1.01000 × 1.016121 × 1.017863) = 95.7291.
| Path | Forward rate Year 1 | Forward rate Year 2 | Forward rate Year 3 | Present value |
|---|---|---|---|---|
| 1 | 1.0000% | 1.6121% | 1.7863% | 95.7291 |
| 2 | 1.0000% | 1.6121% | 1.3233% | 96.1665 |
| 3 | 1.0000% | 1.1943% | 1.3233% | 96.5636 |
| 4 | 1.0000% | 1.1943% | 0.9803% | 96.8916 |
| Average | 96.3377 |
Time 2 at 1.7863%: 100/1.017863 = 98.2451.
Time 2 at 1.3233%: 100/1.013233 = 98.6940.
Time 2 at 0.9803%: 100/1.009803 = 99.0292.
Time 1 at 1.6121%: [(0.5 × 98.2451) + (0.5 × 98.6940)] ÷ 1.016121 = 96.9073.
Time 1 at 1.1943%: [(0.5 × 98.6940) + (0.5 × 99.0292)] ÷ 1.011943 = 97.6948.
Time 0 at 1.0000%: [(0.5 × 96.9073) + (0.5 × 97.6948)] ÷ 1.01 = 96.3377.
The two methods return the same 96.3377. Pathwise valuation and backward induction are two arrangements of the same arithmetic.
Use the tree calibrated in Example 6 to the par curve of 2%, 3% and 4%. Value the three-year, 5% annual coupon bond by pathwise valuation. Example 7 already priced this bond by backward induction at 102.8105.
Every path carries the same cash flows, since the bond has no options: 5 at Time 1, 5 at Time 2 and 105 at Time 3. Only the discount rates differ from path to path.
| Path | Time 0 | Time 1 | Time 2 |
|---|---|---|---|
| 1 | 2.000% | 4.646% | 8.167% |
| 2 | 2.000% | 4.646% | 6.050% |
| 3 | 2.000% | 3.442% | 6.050% |
| 4 | 2.000% | 3.442% | 4.482% |
Path 1. 5/1.02 + 5/(1.02 × 1.04646) + 105/(1.02 × 1.04646 × 1.08167) = 100.5298.
Path 2. Replacing the final rate with 6.050% gives 102.3452.
Path 3. 5/1.02 + 5/(1.02 × 1.03442) + 105/(1.02 × 1.03442 × 1.06050) = 103.4794.
Path 4. Replacing the final rate with 4.482% gives 104.8877.
The average across the four paths is:
(100.5298 + 102.3452 + 103.4794 + 104.8877) ÷ 4 = 102.8105.
This matches the backward induction result exactly. Note how wide the spread of individual path values is, from 100.5298 to 104.8877, even though the bond is option-free. No single path is the value of the bond; only the average is.
When pathwise valuation earns its keep
For an option-free bond, pathwise valuation is more work than backward induction for an identical answer. Its purpose is different. Backward induction collapses information at every node: once the two successor values have been averaged, the route taken to reach that node is forgotten. Some securities have cash flows that depend on that route, and for those the information must be preserved. Laying out the paths explicitly keeps it, which is why the pathwise arrangement is the conceptual bridge to simulation.
Enumerating every path stops being feasible very quickly. A thirty-year instrument with monthly steps has 360 periods, and the number of complete paths through such a lattice is astronomically large. The Monte Carlo method replaces exhaustive enumeration with sampling: generate a large number of interest rate paths at random, value the security along each, and average. It is an approximation to full pathwise valuation, and it is the standard tool wherever cash flows are path dependent.
Path dependency
Cash flows are path dependent when the amount received depends not only on where the interest rate is now but on the route by which it arrived. Mortgage-backed securities are the classic case, because their cash flows are dominated by prepayment behaviour. Borrowers refinance when rates fall, so a pool that has already experienced a period of low rates has lost many of the borrowers who would otherwise prepay later. The same current rate therefore implies different prepayment rates depending on history, and a backward induction lattice cannot represent that.
The procedure
Valuing a thirty-year bond with monthly coupons by simulation runs as follows.
- Simulate a large number of paths of the one-month interest rate, say 500, under an assumed volatility and probability distribution.
- Generate spot rates from the simulated future one-month rates.
- Determine the cash flow along each interest rate path, applying whatever path-dependent rules the security carries.
- Calculate the present value for each path.
- Average the present values across all paths.
The drift adjustment
A set of randomly generated paths will reproduce the market prices of the benchmark bonds only by accident. Left uncorrected, the model would neither fit the current spot curve nor be arbitrage free, which defeats the purpose. The fix is to add the same constant to every short rate on every path, chosen so that the average present value of each benchmark bond equals its observed market price. That constant is the drift term, and a model treated this way is described as drift adjusted. It is the simulation counterpart of calibrating a binomial tree.
Mean reversion
Modellers frequently add mean reversion to the simulation. The motivation is empirical rather than theoretical: history suggests interest rates seldom stay extremely high or extremely low indefinitely. What counts as too high or too low is left to the judgement of the modeller. In practice mean reversion is imposed by setting upper and lower bounds on the random process that generates future rates, and its effect is to pull simulated rates back towards the implied forward rates taken from the yield curve.
The four tree paths of Example 9 are replaced with eight randomly generated paths, calibrated to the same par and spot curves. The bond is unchanged: three years, 5% annual coupon, cash flows of 5, 5 and 105.
| Path | Time 0 | Time 1 | Time 2 | Present value at Time 0 |
|---|---|---|---|---|
| 1 | 2.000% | 2.500% | 4.548% | 105.7459 |
| 2 | 2.000% | 3.600% | 6.116% | 103.2708 |
| 3 | 2.000% | 4.600% | 7.766% | 100.9104 |
| 4 | 2.000% | 5.500% | 3.466% | 103.8543 |
| 5 | 2.000% | 3.100% | 8.233% | 101.9075 |
| 6 | 2.000% | 4.500% | 6.116% | 102.4236 |
| 7 | 2.000% | 3.800% | 5.866% | 103.3020 |
| 8 | 2.000% | 4.000% | 8.233% | 101.0680 |
| Average | 102.8103 |
5/1.02 + 5/1.04550 + 105/1.09305 = 105.7459.
Averaging all eight paths gives 102.8103, which reproduces the value obtained from the binomial tree in Examples 7 and 9. That agreement is the diagnostic: it shows the simulation has been calibrated correctly, since a correctly drift-adjusted model must return the arbitrage-free price for a security whose cash flows are not path dependent.
The simulated paths are far more varied than the four tree paths, with Time 1 rates ranging from 2.500% to 5.500%. That variety is the point. It is what allows path-dependent instruments such as mortgage-backed securities to be analysed, and Monte Carlo methods also make it easier to let parameters change over the life of the security than a binomial lattice does.
Everything so far has taken the shape of the tree for granted. Something had to decide how far apart the top and bottom nodes sit and how the centre of the distribution moves through time, and in this reading that job was done by an assumed lognormal random walk with a constant volatility. Term structure models are the general answer to that question. They describe how interest rates evolve, precisely enough to be fitted to a lattice and used for pricing and hedging.
All of them simplify. No model captures interest rate dynamics completely, and the modeller trades simplicity against accuracy in choosing one. What matters for an analyst is to know the families, the assumptions that define them, and how those assumptions show up in prices and hedges.
How many factors
Valuation and hedging require information about the whole term structure, so a model has to say something about every maturity. The simplest class uses one factor, the short rate, and derives everything else from it. That does force all rates to move in the same direction over any short interval, but it does not force them to move by the same amount, so a one-factor model can still produce changes in slope. The restriction is defensible empirically, since parallel shifts commonly account for more than 90% of the variation in yields. Multi-factor models add drivers such as the slope of the curve and can represent curvature more faithfully, at the cost of complexity in estimation and use.
How rates move
Interest rate dynamics are written as a stochastic process with two parts. The general one-factor form for the short rate r is
The drift term describes the expected path of the rate if all randomness were switched off. It can be a constant, or it can pull the rate back towards a long-run level, which is mean reversion. The second term supplies the randomness, and it is what makes option features and interest rate derivatives priceable at all. The symbol Z denotes a Wiener process, which is normally distributed. Because the normal distribution is symmetric and unbounded below, models built directly on it can and often do generate negative interest rates. The processes are written in continuous time, but they are fitted to binomial or trinomial lattices in discrete form for computation.
Equilibrium models against no-arbitrage models
The sharpest division between models is whether they are forced to agree with today’s market prices.
No-arbitrage models take the current term structure as given. They assume that the observed bond prices, and the curve bootstrapped from them, are correct, and they are parameterised so that the model reproduces those prices. The word parameterised is doing real work: these models let their parameters vary deterministically through time, which creates enough free quantities to match many points on the curve rather than a handful. That is why they can fit the market yield curve closely enough to value derivatives and bonds with embedded options, and why practitioners often prefer them. The cost is a larger number of parameters to estimate.
Equilibrium models start from fundamental economic variables assumed to drive interest rates and impose restrictions that allow equilibrium prices for bonds and interest rate options to be derived. Their parameters are not forced to match current market prices. Some market participants regard that as a serious weakness for pricing and hedging today, since an equilibrium model will frequently produce a term structure inconsistent with the one actually observed. Others regard it as a strength, because such a model describes not only today but the range of futures the economy might reach, which suits scenario work better. A further technical contrast is that equilibrium models work with real probabilities while arbitrage-free models use risk-neutral probabilities.
The Cox–Ingersoll–Ross model
The Cox–Ingersoll–Ross model, usually shortened to CIR, is an equilibrium model in which rates revert to a long-run mean and the variance of rate changes depends on the level of rates.
Read the drift term in three pieces. The quantity θ is the long-run mean rate, rt is the level of the rate at time t, and their difference measures how far the rate currently sits from its mean. When the rate is exactly at its long-run mean the difference is zero and the drift disappears. The parameter k governs how quickly the rate is pulled back, so it is the speed of mean reversion. The distinctive feature of CIR is the square root in the random term: volatility rises with the level of rates, and as rates approach zero the random term shrinks, which keeps rates from turning negative.
The Vasicek model
The Vasicek model was not derived from a general equilibrium of individuals optimising consumption and investment, as CIR was, but it is classified as an equilibrium term structure model. It shares the CIR drift term exactly, so it too reverts to a long-run mean at speed k.
The difference is in the random term. Vasicek assumes constant volatility over the period of analysis: the stochastic term draws from a normal distribution with mean zero and standard deviation one, scaled by a fixed σ that does not depend on the level of rates. There is still only one stochastic driver. The consequence worth remembering is that because volatility does not shrink as rates fall, the Vasicek model can generate negative interest rates.
The Ho–Lee model
Ho and Lee produced the first arbitrage-free term structure model in 1986. Like all models in that family, it begins with the current term structure extracted from the market prices of a reference set of instruments, taking those prices to be correct. It is calibrated to market data and uses a binomial lattice to generate the distribution of future rates, so it sits very close to the machinery built earlier in this reading. In the Ho–Lee model the short rate follows a normal process.
The subscript on the drift term is the whole point. There is a separate value of θ at every time step, and it is exactly that freedom which lets the model reproduce observed market prices at many maturities. Volatility, by contrast, is constant, and as with Vasicek the symmetry of the normal distribution combined with constant volatility means negative rates are possible.
The Kalotay–Williams–Fabozzi model
The Kalotay–Williams–Fabozzi model, abbreviated KWF, is structured like Ho–Lee: no mean reversion and constant volatility, with a time-dependent drift. The difference is what the stochastic differential equation describes.
Because the logarithm of the short rate is normally distributed, the short rate itself is lognormally distributed. The obvious implication is that rates cannot go negative. The less obvious implication matters more for valuation: modelling the log changes the shape of the tails of the rate distribution, and the value of an interest rate option depends heavily on those tails. Two models that agree about the centre of the distribution can disagree materially about option values.
| Model | Type | Short rate modelled | Drift term | Volatility |
|---|---|---|---|---|
| CIR | Equilibrium | drt | Mean reversion at speed k | Varies with the square root of rt |
| Vasicek | Equilibrium | drt | Mean reversion at speed k | Constant |
| Ho–Lee | Arbitrage free | drt | Time dependent | Constant |
| KWF | Arbitrage free | d ln(rt) | Time dependent | Constant |
The two equilibrium models share a drift term; the two arbitrage-free models share a time-dependent drift. Only CIR ties volatility to the level of rates, and only KWF works on the logarithm of the rate.
Modern models
These one-factor models are building blocks rather than the state of the art. Current practice extends them with additional factors, or combines observed forward curves with volatilities backed out of interest rate option prices. The Gauss+ model is a widely used multi-factor example. It carries short, medium and long-term rate factors. The long-term factor is mean reverting and picks up macroeconomic trends; the medium-term rate reverts towards the long-run rate; and the short-term rate carries no random component at all, which reflects the reality that a central bank controls the very short end of the curve. The combination produces a hump-shaped volatility profile across tenors, with medium-term rates the most volatile. Knowing the assumptions behind the classic models is what makes models of this kind intelligible.
An analyst is choosing a term structure model for two separate tasks: pricing a callable bond against today’s market, and running a long-horizon scenario study of how the curve might evolve.