VRM 4: External and Internal Credit Ratings
A credit rating is a compressed answer to one question: how likely is it that a borrower fails to pay what it owes? Everything else about the letter grade follows from that. Three names dominate the credit rating agency business: Moody’s, Fitch, and Standard and Poor’s (S&P). All are headquartered in the United States though all keep offices elsewhere, and smaller agencies such as DBRS work around the world.
What these firms sell is an independent opinion formed against published criteria, and they try to make a grade mean the same thing whatever the region, industry or year, with mixed success. Rating bonds and money market instruments issued by corporations and governments has been the core business for more than a hundred years, and on that work the record has been good.
The rating attaches to the instrument, not to the firm
What carries the grade is normally an instrument the entity has issued rather than the entity itself, and collateral, the term of the paper and the position of the claim can all move it. In practice an agency often assigns the same grade to everything a borrower has issued, which is why people say firm X is rated BBB when the grade belongs to particular bonds firm X has sold. Issuer ratings are published alongside these issue-specific ones.
The agencies do not all claim to measure the same quantity. S&P and Fitch describe the target as probability of default, while Moody’s describes its ratings as measuring expected loss, which is probability of default multiplied by loss given default. That has a testable consequence: where a default would cause little or no loss, Moody’s ought in theory to award a higher grade than S&P for the same borrower.
Why risk managers cannot rely on external ratings alone
Agencies normally look only at firms with publicly traded bonds or money market instruments. A company funded by bank borrowing that issues no debt securities is often left unrated, so a lender facing such borrowers has no external opinion to lean on. Banks therefore built internal rating systems structured along the same lines as the agency scales.
Regulators leaned on ratings early. Banks were barred from investing in poorly rated firms as far back as the 1930s, and the Basel Committee uses credit ratings in setting credit risk capital. Since the crisis of 2007-2008 the United States no longer wishes to, so some capital calculations come in two versions, one for jurisdictions willing to use external ratings and one for those that are not.
Agencies run two scales because they rate two kinds of paper. Long-term ratings apply to bonds, which pay periodic coupons. Short-term ratings apply to money market instruments, which run for a year or less and deliver the whole return in one final payment.
The long-term scale
At the top of the Moody’s ladder sits Aaa, for bonds thought to have almost no chance of defaulting. Falling from there come Aa, then A, then Baa, then Ba and B, and finally Caa, Ca and C. The matching S&P categories are AAA, then AA, then A, then BBB, then BB and B, and finally CCC, CC and C, with Fitch very close to S&P. A grade of D marks issuers already in default.
Finer distinctions come from modifiers. Moody’s splits Aa into Aa1, Aa2 and Aa3 and A into A1, A2 and A3, continuing that way further down, while S&P splits AA into AA+, AA and AA- and A into A+, A and A-. Neither subdivides its top category or, usually, its two lowest. Grades from different agencies are read as equivalent, so a BBB+ from S&P says what a Baa1 from Moody’s says.
One line matters more than any other. Instruments graded BBB- or Baa3 and above are investment grade, and anything below is non-investment grade, also called speculative grade or a junk bond. Mandates, collateral rules and regulatory tests cluster at that boundary, which is why crossing it does far more damage than a downgrade within a category.
| Moody’s | S&P and Fitch | Grade band | Modifier form used by S&P |
|---|---|---|---|
| Aaa | AAA | Investment grade | Not subdivided |
| Aa | AA | Investment grade | AA+, AA, AA- |
| A | A | Investment grade | A+, A, A- |
| Baa | BBB | Investment grade down to BBB- and Baa3 | BBB+, BBB, BBB- |
| Ba | BB | Speculative grade | BB+, BB, BB- |
| B | B | Speculative grade | B+, B, B- |
| Caa | CCC | Speculative grade | CCC+, CCC, CCC- |
| Ca | CC | Speculative grade | Not usually subdivided |
| C | C | Speculative grade | Not usually subdivided |
| No equivalent | D | Already in default | Not applicable |
Source: assembled from the rating categories described in the chapter. The grade band and modifier columns are a study summary.
The short-term scale
Moody’s uses three prime categories for money market paper: P-1 where the capacity to repay short-dated obligations is superior, P-2 where it is strong and P-3 where it is acceptable, with NP marking non-prime. S&P splits the P-1 equivalent into A-1+ and A-1, uses A-2 and A-3 below it, and holds B, C and D lower still. Fitch follows the same shape with F1+, F1, F2, F3, B, C and D.
What the published default tables show
S&P publishes cumulative average default rates for 1981 to 2018, and Moody’s and Fitch produce comparable tables. Each entry gives the fraction of issuers starting at a rating that had defaulted by the end of year one, year two and so on to fifteen years. An issuer rated A has a 0.06% chance of defaulting within one year, six in ten thousand, a 0.14% chance within two years and a 0.23% chance within three. Read down any column and the figures climb as the initial grade falls, which is the basic evidence that grades carry information.
Two probabilities, two different questions
Two different probabilities attach to the same future year, and confusing them is an easy way to misprice a loan. The unconditional probability for year n is measured today and asks what share of issuers now holding the rating will fail during that year. The conditional probability starts the clock later: given the borrower is alive at the end of year n, what is the chance it fails during year n plus one?
Getting from cumulative figures to both quantities
Subtracting neighbouring entries in a cumulative row gives the unconditional series directly, since the extra defaults between one column and the next are those that happened during the intervening year. Dividing that difference by the probability of surviving that far gives the conditional figure.
A bond was rated B when it was issued. The S&P cumulative average default rates put the chance that such an issuer has defaulted within four years at 14.95%, and within five years at 17.33%.
Neither number is more correct. They are the same evidence seen from two vantage points, as a mortality table holds both the chance a newborn dies at ninety-nine and the much larger chance a ninety-eight year old does.
Why the shapes differ across the rating scale
Run the calculation across every rating and year and a pattern appears. For investment grade issuers the annual rate rises over the first four years. Such a firm was judged sound when rated, so unless the agency erred it is unlikely to fail immediately; what grows with time is the chance its finances have since deteriorated.
Issuers rated CCC/C run the other way, the annual rate falling over the first few years. That firm was already in trouble when the grade was assigned, so surviving a year is evidence that its position improved and the hazard it faces declines as that evidence accumulates.
| Years elapsed | AAA | AA | A | BBB | BB | B | CCC and C |
|---|---|---|---|---|---|---|---|
| 1 | 0.00 | 0.02 | 0.06 | 0.17 | 0.65 | 3.44 | 26.89 |
| 2 | 0.03 | 0.04 | 0.08 | 0.29 | 1.37 | 4.66 | 12.83 |
| 3 | 0.10 | 0.06 | 0.09 | 0.34 | 1.65 | 4.26 | 7.63 |
| 4 | 0.11 | 0.10 | 0.12 | 0.42 | 1.68 | 3.51 | 4.77 |
| 5 | 0.11 | 0.10 | 0.14 | 0.43 | 1.61 | 2.80 | 3.78 |
| 6 | 0.10 | 0.10 | 0.14 | 0.42 | 1.49 | 2.33 | 1.72 |
| Derived: peak rate and its year | 0.11 in yr 4 | 0.10 in yr 4 | 0.18 in yr 7 | 0.43 in yr 5 | 1.68 in yr 4 | 4.66 in yr 2 | 26.89 in yr 1 |
Source: S&P 2018 global corporate default and rating transition study, transposed here so years run down the table. The last row is derived from the full fifteen-year series.
Shrink the window over which a conditional default probability is measured until it is very short, of length delta t, and what survives is a rate rather than a probability. The hazard rate h at time t is defined so that h multiplied by delta t gives the chance of failure inside that short window, conditional on no earlier default. Analysts also call it the default intensity, and some models treat it as stochastic.
From a hazard rate to a default probability
The appeal of the hazard rate is that unconditional default probabilities follow from it in closed form. Write h with a bar over it for the average rate applying from today out to time t. The chance of defaulting before that date and the chance of surviving to it are one minus the other.
A credit asset carries a hazard rate that stays constant at 1% per year.
The relationship runs in both directions. Published studies give the cumulative figure, so the practical move is to invert the formula and recover the average hazard rate that number implies. Taking logarithms of both sides of the survival expression does it in one step.
The result averages the whole period, so it says nothing about whether the true hazard was front loaded or back loaded. Where the rate changes at a set date, averaging is done on the rates themselves, weighted by the time each applies.
Part A uses published default statistics. Part B uses a hazard rate that is known to step up partway through.
Default probability answers only half of a lender’s question. A bankruptcy sets creditors filing claims; sometimes assets are sold off so that claimants take partial payment, and sometimes the parties agree a reorganisation in which claims are adjusted.
The recovery rate for a bond is normally defined as the value of that bond shortly after the default, expressed as a percentage of face value. Loss given default carries exactly the same information from the other side, being the percentage recovery rate subtracted from 100.
What moves the recovery rate
Seniority and security drive most of the variation: seniority is where the claim ranks once the estate is divided, security is whether collateral has been posted. Agency statistics put the average recovery rate near 25% for junior bonds, which rank behind others, and near 50% for senior secured bonds. A lender holding subordinated paper therefore faces roughly double the loss given default of one holding secured paper from the same borrower.
The correlation that makes credit portfolios awkward
Recovery rates are negatively correlated with default rates. In a recession, default rates on bonds run high and recovery rates run low, while in a strong economy defaults are rare and recoveries generous. Both move with the value of a defaulting firm’s assets, which sinks in bad conditions and holds up in good ones.
For a bond portfolio manager high default rates are therefore doubly bad: the years with an unusual number of failures are the years in which each failure returns least. Multiplying average default by average recovery understates losses in the worst years, and a model treating the two as independent gives a loss distribution with too thin a tail.
An instrument is normally rated when it is issued and reviewed at least every twelve months thereafter. An agency rates only where the information available is good enough for it to form an opinion, and that opinion rests on a mixture of analysis and judgement rather than on a formula.
The inputs S&P describes cover financial information, both historical and forecast, data on the industry and the wider economy, comparisons against peers, and the financings the borrower plans to undertake. Qualitative material sits alongside that core, covering matters such as the institutional or governance framework, and a meeting with management is commonly part of the exercise.
Who pays, and why that is contentious
The fee is paid by the firm being rated. It comes to a handful of basis points on the notional amount of the bond, and S&P has quoted its own at 6.75 basis points. Where an issuer declines to pay, the agency may decline to publish a rating.
The awkwardness is plain. The issuer pays, while the user of the product is the purchaser of the bond. Having the purchaser pay would align incentives better, but the free rider problem makes it organisationally very difficult: once one investor has bought the rating, others obtain it from that investor at no cost.
Critics argue that an agency paid by the issuer drifts toward whatever grade the issuer believes it deserves. The counterargument is that an agency trades on its reputation, and such a business cannot afford to withhold a low grade when one is warranted. The structured products episode tested whether that discipline suffices.
Outlooks and watchlists
Agencies also publish outlooks, indicating the most likely direction of travel over the medium term. A positive outlook signals that the rating may be raised, a negative one that it may be lowered, and a stable one that no change is expected. A developing or evolving outlook says the grade may move but the agency cannot yet call which way.
A watchlist carries a shorter horizon, indicating a change anticipated within about three months. A positive watchlist flags a review for a possible upgrade and a negative one a possible downgrade. These signals often matter more to a market participant than the grade itself, because they arrive first.
Stability is an objective in its own right. Bond traders are heavy users of grades and many operate under rules about what they may hold, so a grade that jumped about would force them to trade and pay away transaction costs. Ratings also sit inside financial contracts and, in some countries, inside regulatory rules: a bond rated A- might qualify as collateral where one rated BBB+ does not, and a grade oscillating between them would make that contract unworkable. Agencies move a rating only when the change looks long lasting.
Two ways to answer the same question
Economies swing between fast growth, slow growth and outright contraction, and a firm’s probability of default swings with them. That leaves an agency with a choice about what its letter grade means.
A through-the-cycle rating aims at average creditworthiness across several years and is deliberately insulated from the wider economy. A point-in-time rating aims at the best current estimate of future default probability, so it moves whenever conditions move. In theory the through-the-cycle estimate understates default probability in the downward phase of the cycle and overstates it in the upward phase, which is the cost of the stability it buys.
Consistent with their stability objective, the agencies produce through-the-cycle estimates. So it is not always the case that grades worsen when the economy weakens or improve when it strengthens, and a user expecting them to track conditions will be disappointed.
Converting one into the other
Users wanting a point-in-time view sometimes adjust agency grades themselves. That requires an index for the health of the economy, applied so that ratings rise when the economy is strong and fall when it is weak. The adjustment must be calibrated against empirical default data from different stages of past cycles, which is what makes it difficult.
The agencies apply one scale to every borrower, wherever it operates and whatever it does. Whether a BBB+ awarded to a Californian firm in one industry means what a BBB+ awarded to a German firm in another means is an empirical question, and the evidence is uneven.
Part of the difficulty is the length of the record. Moody’s, S&P and Fitch are all based in the United States and much of what they report rests on United States data, while histories elsewhere are shorter. S&P publishes default statistics separately for the United States, Europe and emerging markets, but the samples differ in depth: the United States data covers fifteen years, the European data seven and the emerging markets data just five.
| Initial rating | European firms | U.S. firms | Emerging markets firms | Derived: Europe as a multiple of the U.S. figure |
|---|---|---|---|---|
| AAA | 0.00 | 0.41 | 0.00 | 0.00 |
| AA | 0.20 | 0.43 | 0.00 | 0.47 |
| A | 0.26 | 0.69 | 0.04 | 0.38 |
| BBB | 0.56 | 1.92 | 2.24 | 0.29 |
| BB | 3.71 | 7.89 | 5.26 | 0.47 |
| B | 12.43 | 18.70 | 12.10 | 0.66 |
| CCC/C | 43.37 | 51.42 | 26.32 | 0.84 |
| Investment grade | 0.34 | 1.12 | 1.45 | 0.30 |
| Speculative grade | 10.14 | 16.47 | 9.35 | 0.62 |
| All rated | 2.62 | 7.47 | 5.87 | 0.35 |
Source: S&P 2018 global corporate default and rating transition study. The final column divides the European figure by the United States figure.
A European grade has historically been worth more than the same grade in the United States. The derived column sizes the gap: for investment grade issuers as a group the European default rate has been around one-third of the United States rate, and the advantage narrows as grades fall, until at CCC/C the two regions are close together.
Emerging markets follow neither pattern. Firms rated AAA, AA and A there recorded very low five-year default rates, but firms rated BBB fared slightly worse than their United States counterparts and much worse than European ones.
Industry differences and the warning attached to all of this
Less evidence exists across industries. Historically, banks carrying a given grade have defaulted more often than non-financial corporations with the same grade, and the agencies have agreed with each other less about banks than about other borrowers.
Every comparison here carries the same caveat. The agencies strive continuously for geographic consistency and for sector consistency, so past differences may not persist, and extrapolating from a short record is dangerous.
Agency grades are not the only default estimates on sale. Organisations such as KMV, now part of Moody’s, and Kamakura run models that produce default probabilities and sell the output for a fee. The inputs are market observables rather than analyst judgement: debt in the capital structure, the market value of the equity and its volatility.
These vendors supply point-in-time estimates and carry none of the stability objective that constrains an agency. Their key input, the equity price, changes continuously, whereas a grade is revisited only periodically. The output therefore responds far faster, an advantage for a trading desk and a nuisance for anyone writing a collateral contract.
The structural model underneath
The framework comes from Merton. In its simplest form default can happen at only one future date, and it happens when the assets of the firm are worth less on that date than the debt repayment then falling due. Writing V for the asset value and D for the face value of that debt, the firm defaults when V is less than D, and the equity is worth the following.
That is exactly what a call option written over the firm assets pays, with a strike price set at the face value of the debt, and default is the state in which the option finishes out of the money. Posed that way, the probability of default follows from standard option pricing theory, which is why an equity price and an equity volatility are enough to produce a number.
Banks and other financial institutions run internal rating systems built on their own assessment of borrowers. The usual construction takes several factors, among them financial ratios, cash flow projections and a judgement on management quality, scores each one, and combines the scores into a weighted average that sets the final grade.
Three reasons make the effort worthwhile. External grades are not always available, since most bank borrowers never issue public debt. Regulatory credit risk capital depends on probabilities of default. And IFRS 9, with its FASB counterpart, makes default probability an input to the balance sheet valuation of a loan.
Point-in-time or through-the-cycle, and why regulators care
Internal grades can be built either way, with a tendency toward point-in-time, though through-the-cycle grades may suit long-term lending commitments better. Regulators have reason to prefer the stable measure, and the probabilities used to determine regulatory capital are through-the-cycle. IFRS 9 pulls the other way, requiring point-in-time estimates when loans are valued, so a bank maintains both views of one borrower.
The reason is procyclicality. Point-in-time probabilities rise in bad conditions and banks become less inclined to lend, so firms struggle to fund working capital and fixed assets, pushing conditions down further. In good conditions the mechanism reverses and easier credit amplifies the expansion.
Validating an internal rating system
Banks must back-test the procedures behind their internal grades, and validation is the hard part. The test builds the internal equivalent of a cumulative default rate table and checks that borrowers given better grades did default less often, work that normally needs a decade or more of history. Few internal portfolios hold enough defaults in the higher grades for that to mean much, ten years spans one cycle at best, and a change of lending policy partway through breaks the sample.
Scoring models and machine learning
Statistical scoring predates all of this. Altman proposed the Z-score in 1968, applying discriminant analysis to five ratios. Three of them divide by total assets: working capital, then retained earnings, then earnings before interest and taxes. The fourth sets the market value of equity against the book value of total liabilities, and the fifth is sales over total assets. For publicly traded manufacturing firms the score took this form, weighting the five in turn by 1.2, 1.4, 3.3, 0.6 and 0.999.
A Z-score above 3 indicated a firm unlikely to default, and default probability rose as the score fell, to the point where a firm below 1.8 was very likely to fail. Some banks now automate lending with machine learning, feeding an algorithm data on firms and on whether they defaulted and letting it derive the separating rule. Modern versions use far more inputs and far more data than Altman had, and the function need not be linear.
A rating transition matrix records the probability that an issuer migrates from one rating category to another over one year. The version below covers 1981 to 2018 and is laid out with the starting rating across the columns, so each column is a distribution over where the issuer stands twelve months later. D denotes default and NR that the issuer was no longer rated.
| Rating one year later | From AAA | From AA | From A | From BBB | From BB | From B | From CCC/C |
|---|---|---|---|---|---|---|---|
| AAA | 86.99 | 0.50 | 0.03 | 0.01 | 0.01 | 0.00 | 0.00 |
| AA | 9.12 | 87.06 | 1.69 | 0.09 | 0.03 | 0.02 | 0.00 |
| A | 0.53 | 7.85 | 88.17 | 3.42 | 0.11 | 0.08 | 0.11 |
| BBB | 0.05 | 0.49 | 5.16 | 86.04 | 4.83 | 0.17 | 0.20 |
| BB | 0.08 | 0.05 | 0.29 | 3.62 | 77.50 | 4.93 | 0.59 |
| B | 0.03 | 0.06 | 0.12 | 0.46 | 6.65 | 74.53 | 13.21 |
| CCC/C | 0.05 | 0.02 | 0.02 | 0.11 | 0.55 | 4.42 | 43.51 |
| D | 0.00 | 0.02 | 0.06 | 0.17 | 0.65 | 3.44 | 26.89 |
| NR | 3.15 | 3.94 | 4.48 | 6.10 | 9.67 | 12.41 | 15.50 |
| Derived: downgrade or default | 9.86 | 8.49 | 5.65 | 4.36 | 7.85 | 7.86 | 26.89 |
Source: S&P 2018 global corporate default and rating transition study, transposed here so columns are the starting rating. The final row is derived by adding every downgrade probability to the default probability. Column totals differ from 100 by small rounding amounts.
Take the column headed A. An issuer just assigned that grade has an 88.17% probability of still being A-rated a year later, a 1.69% chance of moving to AA and a 0.03% chance of reaching AAA, against a 5.16% chance of falling to BBB and a 0.06% chance of defaulting. Investment grade ratings hold better: a AAA issuer stays put with probability 86.99% and a BB issuer 77.50%, while a CCC/C issuer manages only 43.51%.
The derived bottom row makes a less obvious point: deterioration risk is not monotonic in the grade. It falls from 9.86% at AAA to a minimum of 4.36% at BBB, because the highest grades have nowhere to go but down, then climbs again through the speculative categories. NR entries also grow as quality falls, so the weakest issuers are likeliest to leave the sample. For analysis NR is often allocated proportionally across the remaining outcomes.
An investor holds a bond from an issuer currently rated BBB and wants cumulative default probabilities out to three years, built from the one-year matrix above.
Uses and the limits of the independence assumption
Extending a one-year matrix to n years by matrix multiplication assumes each year is independent of the one before it. Published multi-year matrices differ from calculated ones, and the culprit is ratings momentum: an issuer downgraded once is more likely to be downgraded again, and the same holds for upgrades. Successive years are therefore not statistically independent.
Transition matrices are built for internal ratings as well, and one test of a rating system is whether its transitions stay roughly stable from year to year. Research has found differences across sectors, especially for investment grade issuers, and the matrices depend on the cycle, with downgrades rising significantly during recessions despite the through-the-cycle intent behind the grades.
If a downgrade carries genuine news, stock and bond prices should fall on the announcement and credit default swap spreads should widen. The alternative is that the market already knew, making the agency a follower rather than a leader. Researchers have produced mixed results, and the mix is informative.
Most agree that stock and bond market reactions to downgrades are significant, and strongest when the downgrade carries an issuer from investment grade into non-investment grade. Reactions to upgrades are much less pronounced. Part of that asymmetry is not about information: crossing the boundary changes who may hold the bond, and issuers may have contracts containing rating triggers.
Work by Hull, Predescu and White in 2004 examined credit default swap spreads, taking in outlooks and watchlists alongside the rating changes. Reviews for downgrade were found to contain significant information, while downgrades themselves and negative outlooks did not, and positive rating events mattered less again. Broadly, credit default swap spreads anticipate rating changes rather than respond to them.
The relationship also runs the other way, so credit spread changes help estimate the probability of a negative rating action to come. Of the downgrades in their sample, 42.6% came from the quartile of issuers with the largest credit default swap spread increases, along with 39.8% of reviews for downgrade and 50.9% of negative outlooks.
Where credit ratings failed
Rating a structured product differs in kind from rating a bond, because the grade depends almost entirely on a model. Agencies were open about those models: S&P and Fitch rated on the probability that the product would give a loss, Moody’s on expected loss as a percentage of principal. The inputs, above all the assumed correlations between defaults on different mortgages, proved far too optimistic, and grades on products built from other structured products fared worse still.
Publishing the models had a second consequence. Once the creators of these products understood them, a deal could be designed to hit whatever grade was wanted: present the plan, obtain an advance ruling, adjust the structure until the desired rating appeared. The work was very profitable, and the agencies were perhaps not as independent as they should have been.
Many structured products created from subprime mortgages defaulted during the 2007-2008 crisis, and agency reputations have not fully recovered. The Dodd-Frank Act now obliges agencies to disclose the assumptions and methodologies behind their ratings and has widened their legal liability, the Office of Credit Ratings was set up inside the Securities and Exchange Commission to supervise them, and in the United States they are no longer used by bank supervisors to determine regulatory capital. The lesson is narrow: the historical record supports grades on corporate bonds and money market instruments, and does not extend to instruments whose rating is the output of a model.