VRM 7: Operational Risk
Financial institutions have had to take operational risk far more seriously over the past twenty years, and part of what makes the subject awkward is that the term carries several meanings. Read broadly, operational risk is a residual: whatever remains once market risk and credit risk have been carved out. Read narrowly, it covers only mistakes made in running the business, so a transaction processed incorrectly would count while fraud, a cyberattack or the loss of a building would not.
Regulation settled between those two extremes. The Basel Committee defines operational risk as “the risk of loss resulting from inadequate or failed internal processes, people, and systems or from external events.” The International Association of Insurance Supervisors uses a parallel definition for insurers, pinning the loss on controls, procedures, personnel or internal systems that prove inadequate or fail, and on events outside the firm.
Either reading is wide. It picks up losses from computer hacking, fines imposed by regulatory agencies, litigation, rogue traders, terrorism and systems failures. Strategic risk and reputational risk stay outside the perimeter, even though a large operational loss usually damages a reputation as well.
Why operational risk is harder to quantify than market risk or credit risk
Market risk measurement rests on the volatilities of risk factors, and those can be estimated from long price histories, so a measure such as value at risk can be produced. Credit risk measurement rests on default and recovery experience, published by rating agencies and accumulated inside a bank’s own records, which yields workable estimates of expected loss and unexpected loss. Operational risk offers much less. How likely is it that a cyberattack destroys a bank’s records, and how large would a rogue trader loss be? Events of this kind are rare and each tends to be novel, so the sample a firm can learn from stays small.
Working from its sound practices guidance, the Basel Committee sorted operational risk events into seven categories, and that classification is still the common language banks use when they record losses. They are arranged by where the failure sits: intent inside the firm, intent outside it, obligations owed to staff, obligations owed to clients, the physical estate, the technology stack, and the daily grind of processing.
Intent separates the first two: internal fraud requires at least one insider, while external fraud covers the same acts by a third party, and discrimination matters go into the employment category rather than either. The client category turns on negligent failure to meet a professional obligation, or on selling something unsuitable, rather than on any intent to defraud.
| Category | Example events | Usual loss profile |
|---|---|---|
| Internal fraud | Employee theft, intentional misreporting of positions, insider trading on an employee’s own account | Rare, severe |
| External fraud | Forgery, robbery, check kiting, losses from computer hacking | Frequent, mostly small |
| Employment practices and workplace safety | Discrimination claims, organized labor activities, workers compensation claims, a customer who falls at a branch office | Frequent, small |
| Clients, products, and business practices | Money laundering, fiduciary breaches, sale of unauthorized products, misuse of confidential customer information | Rare, very severe |
| Damage to physical assets | Fires, floods, earthquakes, vandalism, terrorism | Rare, moderate to severe |
| Business disruption and system failures | Utility outages, telecommunication problems, failures of hardware or software | Moderate on both counts |
| Execution, delivery, and process management | Collateral management failures, data entry errors, incomplete legal documentation, vendor disputes | Very frequent, small |
Source: category names and example events from Basel Committee on Banking Supervision, “Sound Practices for the Management and Supervision of Operational Risk”. The loss profile column is editorial.
The categories matter beyond reporting. Crossed with the eight regulatory business lines, they generate the grid of loss cells that the advanced measurement approach required a bank to model one by one.
Cyber risk is the largest operational exposure most financial institutions carry. Mobile wallets, online banking, electronic funds transfer and the credit and debit card have all been good for banks and their customers, and every one widened the surface an attacker can reach. Threats arrive from nation states, organized crime, individual hackers and insiders, and the standard defenses are firewalls, cryptography, intruder detection software and user account controls.
The damage runs from money stolen and data destroyed to intellectual property copied. A 2013 cyberattack at Yahoo produced a data breach covering three billion user accounts, exposing passwords, dates of birth and the answers to security questions. Equifax, a large consumer credit reporting agency, disclosed a 2017 cyberattack that affected 143 million people in the United States. Forbes has estimated the cost of cyber crime to all businesses at as much as USD 6 trillion per year.
Customers are attacked through phishing, an email inviting the recipient to confirm account details. The graver threat runs against the institution itself, since an attacker inside a bank’s systems can read client information, delete records, enter false transactions and move money out. In March 2016 the Central Bank of Bangladesh was hacked through multiple entry points, and the attackers planned to embezzle over USD 1 billion through a series of international transactions. A data entry mistake on their side held the theft to USD 80 million, an embarrassingly large sum even so. Any firm should assume it will be breached eventually and hold plans for attacks of different severities, up to a response as extreme as refusing new transactions for a few days.
Compliance risk and the size of the fines
Compliance risk is the risk of fines or other penalties for failing, knowingly or not, to act in line with laws and regulations, internal policies, or accepted best practice. Money laundering, terrorism financing and assisting tax evasion all attract heavy penalties. Volkswagen was fined about USD 2.8 billion after cheating during emissions testing in breach of United States standards. HSBC paid a USD 1.9 billion penalty in 2012, having failed to run anti-money laundering programs at its Mexican branches, which let Mexican drug traffickers deposit large sums in cash, and it settled with the United States Department of Justice under a deferred prosecution agreement. In 2014 BNP Paribas agreed to pay USD 8.9 billion, roughly one year’s profit, for routing dollar transactions through the United States banking system for Sudanese, Iranian and Cuban parties despite economic sanctions, and was barred from certain United States transactions for a year.
Rogue trader risk arises when an employee takes unauthorized actions that end in a large loss. The best known case involved Nick Leeson at Barings Bank, whose job in the firm’s Singapore office was to run relatively low-risk trades. Flaws in the Barings systems let him take large positions instead and bury the losses in a secret account. Trying to trade his way back, he lost more, and the losses exceeded USD 1 billion. He fled Singapore, leaving behind a note of apology, and was later returned, prosecuted and imprisoned. Barings Bank, in existence for 200 years, was forced into bankruptcy.
Societe Generale lost more. Jerome Kerviel was ostensibly hunting arbitrage in equity index products such as the Euro Stoxx 50, the French CAC 40 and the German DAX, the kind of opportunity that appears when an index futures contract trades at different prices on two exchanges. He found instead a way of speculating while appearing to arbitrage, taking large positions and booking fictitious trades that made him look hedged. The unauthorized trading came to light in January 2008, and closing the positions cost SocGen EUR 4.9 billion. UBS lost USD 2.3 billion in 2011 and Allied Irish Bank lost USD 700 million in 2002.
One theme runs through all of them. A single trader took very large risks without the firm knowing or authorizing them, and concealed the exposure with fictitious offsetting trades or some equivalent device, hoping that continued speculation would recover the losses and that the episode would then be forgiven.
Two controls that address it
The structural control is independence between the front office, which trades, and the back office, which keeps records and verifies transactions. When one reporting line spans both, concealment becomes straightforward.
The cultural control is harder. Unauthorized trading that ends in a loss brings consequences, while unauthorized trading that ends in a profit is tempting to overlook, and overlooking it teaches everyone that risk limits are negotiable.
The Basel Committee on Banking Supervision writes global rules that national supervisors then implement. In 1999 it issued an early draft of what became Basel II, largely a revision of the credit risk capital rules, carrying one unexpected element: banks would have to hold capital for operational risk on top of what market risk and credit risk already required.
Many risk managers thought this unworkable, on the ground that operational risk cannot be quantified with any precision. The Committee pressed ahead, because a large share of the biggest losses banks had suffered were operational rather than market or credit losses, and because a capital charge would push banks to devote real resources to managing them. Insurance regulation moved the same way: Solvency II, the European Union framework issued in 2016, also requires capital for operational risk, using a definition close to the one in Basel II.
Basel II offered three approaches, taken in sequence as a bank qualified for each: the basic indicator approach, the standardized approach, and the advanced measurement approach. Banks typically began with the first and met further criteria to move up. The first two are simple; the third is not.
The basic indicator approach
Capital was set at 15% of average annual gross income over three years, with gross income defined as interest earned less interest paid, plus non-interest income.
The standardized approach
The standardized approach ran the same calculation business line by business line, with a coefficient that varies by line. Corporate finance, trading and sales, and payment and settlement attracted the highest coefficient, retail banking, asset management and retail brokerage the lowest.
| Business line | Capital (% of gross income) | Capital per USD 1 billion of business line gross income (USD million) |
|---|---|---|
| Corporate finance | 18% | 180 |
| Trading and sales | 18% | 180 |
| Payment and settlement | 18% | 180 |
| Commercial banking | 15% | 150 |
| Agency services | 15% | 150 |
| Retail banking | 12% | 120 |
| Asset management | 12% | 120 |
| Retail brokerage | 12% | 120 |
Source: Basel II standardized approach coefficients. The final column is derived, and the rows are ordered by coefficient rather than by the order used in the source.
The advanced measurement approach, usually shortened to AMA, asked banks to treat operational risk the way they treat credit risk. Regulatory capital became the 99.9 percentile of the one-year operational loss distribution less the expected operational loss, on the reasoning that expected losses are a cost of doing business and belong in pricing rather than in capital.
The work was heavy. A bank had to take every combination of the eight business lines and the seven event categories, which gives 56 cells, estimate in each cell the loss sitting at the 99.9 percentile over a single year, then aggregate the 56 estimates into one capital requirement.
Regulators eventually judged the approach unsatisfactory, and the reason was comparability rather than principle. Two banks given the same data could implement AMA in different ways and arrive at very different capital requirements. In March 2016 the Committee announced that all three Basel II approaches would be replaced by a single new one, the standardized measurement approach.
AMA has not disappeared from banking, only from the regulatory calculation. Estimating a loss distribution and reading a high percentile from it is what economic capital requires, so many banks still run the model internally, usually at a percentile above 99.9%.
The standardized measurement approach, the SMA, was designed to remove the discretion that made AMA results incomparable. Two ingredients go into it: a measure of how large the bank is, and what the bank has actually lost to operational events over the previous ten years.
The size measure is the business indicator, written as BI. It follows the lines of gross income but is adjusted to reflect bank size more faithfully: items such as trading losses and operating expenses reduce gross income, and under the business indicator they add to it instead. The BI converts into the BI Component through a piecewise linear relationship, a function assembled from straight line pieces.
The loss element is the loss component, built from ten years of the bank’s own losses.
The three terms deliberately overlap. A single loss of EUR 150 million enters X, Y and Z, so it carries a weight of 7 plus 7 plus 5, while a loss of EUR 2 million enters X alone. Severe losses therefore drive the loss component far harder than a long tail of routine ones. The calibration sets the loss component and the BI Component equal for an average bank, and the Committee supplies the formula that combines the two into required capital.
Over the last ten years a bank has recorded operational risk losses, in millions of euros, of 280, 140, 70, 50, 12, 6 and 4.
Economic capital needs a full distribution of the one-year loss for each category and for the categories combined. Two inputs generate it: average loss frequency, the number of times a year that large losses of that type occur on average, and loss severity, the distribution of the size of an individual loss.
Loss frequency
Frequency is usually modelled with a Poisson distribution, which counts events arriving at a steady rate and independently of one another. Writing lambda for the average number of losses a year brings, the probability of exactly n of them is given by the Poisson formula.
| No. of losses | Lambda = 2 | Lambda = 4 | Cumulative, lambda = 4 | Lambda = 6 | Cumulative, lambda = 6 |
|---|---|---|---|---|---|
| 0 | 0.135 | 0.018 | 0.018 | 0.002 | 0.002 |
| 1 | 0.271 | 0.073 | 0.091 | 0.015 | 0.017 |
| 2 | 0.271 | 0.147 | 0.238 | 0.045 | 0.062 |
| 3 | 0.180 | 0.195 | 0.433 | 0.089 | 0.151 |
| 4 | 0.090 | 0.195 | 0.628 | 0.134 | 0.285 |
| 5 | 0.036 | 0.156 | 0.784 | 0.161 | 0.446 |
| 6 | 0.012 | 0.104 | 0.888 | 0.161 | 0.607 |
| 7 | 0.003 | 0.060 | 0.948 | 0.138 | 0.745 |
| 8 | 0.001 | 0.030 | 0.978 | 0.103 | 0.848 |
| 9 | 0.000 | 0.013 | 0.991 | 0.069 | 0.917 |
| 10 | 0.000 | 0.005 | 0.996 | 0.041 | 0.958 |
| 11 | 0.000 | 0.002 | 0.998 | 0.023 | 0.981 |
| 12 | 0.000 | 0.001 | 0.999 | 0.011 | 0.992 |
Source: Poisson probabilities for average loss frequency of 2, 4 and 6. The two cumulative columns are derived.
Loss severity
Severity is usually fitted to a lognormal distribution, in which the natural logarithm of the loss is normally distributed. The shape suits operational losses, being bounded below at zero and skewed to the right. Given estimates mu and sigma for the mean and the spread of an individual loss, the mean and variance of the logarithm follow directly.
Take a loss size with a mean of 80 and a standard deviation of 40, in USD million. Then w is 0.5 squared, or 0.25, and the logarithm of the loss size has a mean of 4.27, found by dividing 80 by the square root of 1.25 and taking the logarithm, and a variance equal to ln(1.25), which is 0.223.
Neither frequency nor severity on its own gives the annual loss. What is needed is the distribution of the sum of a random number of random amounts, and there is no convenient closed form for it, so Monte Carlo simulation supplies one by brute force. Once lambda, mu and sigma are estimated, the procedure runs in four steps.
Step one draws from the Poisson distribution to fix the number of loss events, n, in the simulated year. Step two draws n times from the lognormal severity distribution. Step three adds those n amounts to give the total loss for that trial. Step four repeats the cycle many times, and the collection of annual totals is the loss distribution from which any percentile can be read.
Sampling the frequency means drawing a random number between zero and one and reading it as a percentile of the Poisson distribution.
A loss category has an average loss frequency of 4 per year. Loss severity has a mean of 80 and a standard deviation of 40, in USD million, fitted to a lognormal distribution as above. The random number drawn for frequency is 0.31.
Estimating loss frequency and loss severity mixes data with judgement. Frequency should come either from the institution’s own history or from the subjective estimates of operational risk professionals who have examined the controls in place. Severity is harder, since a firm that has never suffered a given event has nothing internal to fit.
External data fills that gap. Banks have built mechanisms for sharing loss data with one another, and data vendors, Factiva and Lexis-Nexis among them, assemble publicly reported losses at other institutions. Both carry problems that have to be handled first.
Reporting bias in vendor data
Publicly reported losses are the large ones, so a severity distribution fitted directly to vendor data is pulled towards the upper end and overstates what a typical event costs. The way around it is to use vendor data for relative severity only, anchoring the level on the firm’s own experience. If vendor data shows loss type A is on average twice as severe as loss type B, and the bank holds data on loss type B but none on loss type A, it can set the mean for loss type A at twice its own mean for loss type B. If the vendor data puts the standard deviation for loss type A 50% above that for loss type B, the bank scales its own figure by the same amount.
Scale bias between banks of different sizes
A loss observed elsewhere also needs adjusting for the difference in size between the two firms, since a larger bank has more transactions, more staff and more clients. Scaling in proportion to revenue overcorrects, because the relationship is far from one for one. Shih, Samad-Khan and Medapa fitted a model in which the observed loss is multiplied by the ratio of the two revenues raised to a power, and found that 0.23 for that exponent gives a good fit. Revenue here means gross income.
Bank B has revenues of USD 20 billion and has suffered a loss of USD 300 million. Bank A, with revenues of USD 10 billion, wants to use that event to estimate a similar loss of its own.
Some of the events that matter most have never happened to the firm. No amount of data cleaning produces a frequency estimate for those, and scenario analysis takes over. It is aimed at losses of low frequency and high severity, the events that shape the far right tail and therefore the capital number.
The work starts with a list. Some scenarios come from the institution’s own history, some from losses known to have hit other banks, and some are hypothetical situations built by risk professionals, occasionally with consultants brought in to help. Each scenario then receives an estimate of loss frequency and one of loss severity, usually from a committee of operational risk experts, and the frequency estimate should reflect the controls in place and the business the firm is doing.
Estimating the frequency of an event that has never happened
Asking a committee for a probability of 0.017 invites false precision. A more workable method offers a small set of buckets and asks the experts to place each scenario in one. A scenario expected once every 1,000 years on average carries a lambda of 0.001; once every 100 years gives 0.01; once every 50 years gives 0.02; once every ten years gives 0.1; and once every five years gives 0.2.
Fitting loss severity to a percentile range
Severity can be elicited the same way. Rather than ask for a mean and a standard deviation, which few people hold intuitions about, ask for the 1 percentile and the 99 percentile of the loss and fit a lognormal distribution to that range. Suppose the answers are 20 and 200. Their logarithms, ln(20) which is 2.996 and ln(200) which is 5.298, are the 1 and 99 percentiles of the logarithm of the loss. A normal distribution has its mean midway between those two points, so averaging 2.996 and 5.298 gives 4.147, and the 99 percentile lies 2.326 standard deviations above the mean, so dividing 5.298 minus 4.147 by 2.326 gives 0.49.
Those parameters feed the same Monte Carlo machinery used for categories that do have data. The number is not the whole return on the exercise. Building the scenarios forces an active discussion of how such losses could occur, which tends to produce a plan for responding and sometimes a change that makes the event less likely.
Economic capital is read at very high confidence levels, where the simulated distribution rests on few observations. The power law offers a way of extrapolating into that region. If v denotes a random variable and x sits high in the range v can take, the probability of exceeding x follows a particular form.
As alpha falls, the tail grows fatter and extreme values become more likely. The law describes the right tail alone rather than the whole distribution, which is why it holds only for high values of x, in practice the top 5% or so.
The result traces to the Polish mathematician B.V. Gnedenko, who showed that the tails of a very wide range of distributions share this property, strictly in the limiting case where x grows without bound. Empirically it fits the incomes of individuals, the magnitude of earthquakes, the sizes of cities measured by population, the sizes of corporations, the daily hit count of a website, how often a given word turns up in a text, and the trading volume of a stock. The common thread is a variable produced by many independent random effects, often multiplicative rather than additive, since adding independent variables tends to give a normal distribution instead. Work by de Fontnouvelle and co-authors indicates that the law also holds for operational risk losses.
For a category of operational risk losses measured in USD millions, K and alpha have been estimated as 10,000 and 3 respectively, using maximum likelihood methods applied to the observed large losses.
Economic capital is pushed down to business units so that each unit’s return on capital can be measured. The mechanics resemble those used for credit risk capital: the total is calculated for the firm, and each unit is charged the share attributable to it, allowing for the diversification benefit that arises because the units do not all suffer their bad years together.
What makes the allocation worth doing is the incentive it creates. A manager who can show a reduction in loss frequency or loss severity is allocated less capital, so the unit’s return on capital improves and the manager’s bonus prospects improve with it. The same logic runs at the level of the whole bank under the standardized measurement approach, where the loss component depends on the firm’s own ten-year record. Even where the effect on the allocated number is modest, the process makes a manager aware that operational losses arising in the unit carry a cost.
Reducing operational risk is not always the right answer
Some operational risk is inherent in running a business unit, and none of it can be driven to zero. Any decision to reduce it raises operating costs, so it needs a cost-benefit analysis behind it rather than an appeal to principle. A study might show that transaction processing errors could be cut by 5% by building a new computer system and giving employees extra training. Where that cost greatly exceeds the present value of the losses avoided, the right decision is to leave the risk where it is.
Measurement is only half of the discipline. Operational risk units are also expected to cut the chance of large losses and limit their severity, and firms learn from each other: when a large loss occurs at one institution, risk managers elsewhere study what happened and ask whether it could happen to them.
Part of that work is identifying causal relationships, since losses often trace back to something manageable. It may be possible to show that a category of losses falls when employees receive more training or when a role carries a higher educational requirement, or that a run of losses comes from an outdated computer system. The fix can then be costed against the losses avoided.
Risk and control self assessment
Risk and control self assessment, known as RCSA, asks line managers and their staff, rather than operational risk professionals, to identify the exposures in their own area. The word self is what distinguishes it. It covers losses that have already occurred and losses that could occur, and it is repeated periodically, typically once a year, with the frequency and severity of loss events quantified each time.
Firms run it in many ways: interviews with line managers and their staff; risk questionnaires; a review of the history of risk incidents; third-party reports from auditors, regulators and consultants; a review of what has happened to comparable managers elsewhere; suggestion boxes and intranet reporting portals; a whistle blowing process; and brainstorming in a workshop. Some loss events prove unavoidable, while others yield, with frequency, severity or both coming down.
Key risk indicators
A key risk indicator, or KRI, is a data point that signals a raised chance of operational losses in an area, ideally early enough for remedial action. Simple examples are unfilled positions, positions filled by temporary staff, failed transactions and staff turnover. Movement matters more than level, so indicators are tracked through time and unusual behaviour is investigated. Some are subtler: an employee who will not take a vacation may be concealing unauthorized trading or embezzlement, since concealment needs daily attention.
Education and risk culture
Teaching employees which business practices are unacceptable matters, and building a culture in which they are seen as unacceptable matters more. Goldman Sachs drew adverse publicity in 2007 over a product called ABACUS, which arguably served one client at another’s expense. The firm then worked on its risk culture, and chief executive Lloyd Blankfein made a 23-country tour of its regional offices to talk about ethics and about not selling products to clients who do not fully understand the possible outcomes. The legal department has a related job, reminding staff that emails and recorded phone calls generally have to be produced in a legal dispute.
Many operational risks can be insured, and the operational risk manager has to judge whether the premium is worth paying. Insurance now does two jobs. It reduces the severity of a loss when one occurs, and because the standardized measurement approach is driven by the frequency and magnitude of losses incurred over the previous ten years, an insured loss also feeds through into a lower capital requirement later.
Moral hazard
Moral hazard is the risk that having the insurance changes the behaviour of the insured party in a way that makes a claim more likely. Rogue trader cover is the sharp example. An insurer writing it has reason to worry that traders will take larger unauthorized positions, since a gain would be kept by the bank while a loss would fall on the insurer. That such cover can be bought at all is somewhat surprising. Insurers manage the problem by specifying how trading limits are implemented and monitored, by negotiating the policy with risk managers and commonly requiring that its existence is not revealed to traders, and by investigating claims carefully, with the payout at risk where the bank has failed to meet the conditions.
The general devices are the same across operational risk cover. A deductible leaves the firm carrying the first part of any loss. A co-insurance provision pays only a percentage of the loss rather than the whole amount. A policy limit caps the total payout. Premiums are also raised after a loss has been incurred. Every one of these leaves the insured party exposed to part of the outcome, which is the point.
Adverse selection
Adverse selection is the insurer’s difficulty in telling low-risk situations from high-risk ones. Charge every financial institution the same premium for a given cover, and those that know their own risk is high will find the price attractive while those with strong internal controls will consider it expensive, leaving the insurer with the worst risks. Cyber insurance behaves the same way: the institutions that have invested least in their defenses find the cover most attractive.
Insurers respond by learning about the customer before quoting, as motor insurers do when the initial quote reflects past accidents and speeding tickets and the premium is adjusted as new information arrives. For rogue trader cover and cyber cover the equivalent is an underwriting process in which the institution must satisfy the insurer that its risk controls are sound before cover is offered at all.