Counterparty Insight Bank Risk Rating Methodology
1. Overview
The Counterparty Insight Bank Risk Rating model reports each FDIC-insured depository institution's condition as a 1–100 strength score, where 100 is strongest. The model is based on the CAMELS framework used by U.S. federal banking regulators, adapted for quantitative analysis using publicly available regulatory financial data.
The model evaluates six components across 17 financial ratios, applies component-level weights, and produces a single composite rating. Each bank is scored against its asset-size peer group. The FDIC failed-bank record informing the model's distress signals and its back-testing spans 2000 to 2026 (573 bank failures), including the 2008–2010 financial crisis and the 2023 rate-shock failures.
2. Rating Scale
Counterparty Insight reports a continuous 1–100 strength score (100 = strongest). The underlying composite is itself continuous (a weighted average of seventeen ratios across six CAMELS components, smoothed across quarters), and the strength score simply rescales it onto the 1–100 range. A stronger composite always maps to a higher score, and because the score moves continuously, two otherwise similar institutions still receive distinct scores.
The scale is anchored so the median bank scores near 65, which spreads the sound majority across the upper-middle of the scale and leaves resolution to separate healthy banks from one another rather than compressing them near the top.
For reference, the score maps to six descriptive tiers:
| Tier | Strength score |
|---|---|
| Strongest | 80–100 |
| Strong | 60–80 |
| Moderate | 45–60 |
| Elevated | 30–45 |
| Weak | 20–30 |
| High Risk | 1–20 |
3. CAMELS Components
The model scores six components drawn from the CAMELS framework:
- Capital adequacy (C): the capital cushion available to absorb losses, including the unrealized erosion in held-to-maturity securities.
- Asset quality (A): the credit quality of the loan portfolio and the adequacy of loss reserves; the most heavily weighted component, since credit risk drives most bank failures.
- Management (M): quantitative proxies for management effectiveness (operating efficiency, growth discipline, revenue mix, reserve prudence), since governance quality itself is not observable in public data.
- Earnings (E): the capacity to build capital organically, measured by trailing-twelve-month return on average assets.
- Liquidity (L): the reliability of the bank's funding and its ability to meet obligations under normal and stressed conditions.
- Sensitivity to market risk (S): exposure to interest-rate movements and asset concentrations, chiefly commercial real estate and securities holdings.
Each component carries a weight in the composite. The starting weights are grounded in the CAMELS supervisory framework and established credit-analysis principles, and are then refined by a machine-learning model trained on historical bank failures; the two are blended evenly, with a floor so no component drops out. Section 4.3, Composite Rating, shows both the starting weights and the weights currently deployed.
3.1 Capital Adequacy
Capital adequacy measures the institution's ability to absorb losses and support growth while maintaining sufficient capital buffers above regulatory minimums. The component captures both the reported regulatory capital cushion and the hidden capital erosion in the held-to-maturity (HTM) securities portfolio, which is carried at amortized cost on the balance sheet and not reflected in book equity.
| # | Ratio | Description |
|---|---|---|
| 1 | Leverage (Core Capital) Ratio | Tier 1 capital divided by average total assets. The primary measure of capital strength under the prompt corrective action (PCA) framework. Well-capitalized threshold: 5%. Higher is better. |
| 2 | HTM Unrealized Loss / Tier 1 | Unrealized mark-to-market loss on the held-to-maturity securities portfolio, expressed as a percentage of Tier 1 capital. HTM securities are carried at amortized cost on the balance sheet, so rate-driven losses do not appear in book equity but would materialize as realized losses if the portfolio had to be liquidated under funding stress (the SVB pattern). Contribution floor: this ratio is scored for every bank but only contributes to the Capital component average when the loss exceeds 5% of Tier 1. Below 5% the row is displayed but excluded from the average, so well-managed banks with no HTM losses are not credited with a perfect-zero score that would artificially boost Capital. The 5% threshold sits near the 92nd percentile of the universe-wide distribution, so contribution effectively begins around a strength score in the 50–61 band and worsens fast above it. Lower is better. |
3.2 Asset Quality
Asset quality is the most heavily weighted component because credit risk is the dominant risk for most banks. These ratios measure the quality of the loan portfolio and the adequacy of loss reserves.
| # | Ratio | Description |
|---|---|---|
| 3 | Noncurrent Loans to Total Loans | Loans 90+ days past due or on nonaccrual status, as a percentage of total loans. Directly measures the volume of problem credits. Lower is better. |
| 4 | Texas Ratio | Problem assets (loans 90+ days past due, nonaccrual loans, and other real estate owned), net of the U.S.-government-guaranteed portion of those non-current loans, divided by loss-absorbing capacity (tangible equity plus the allowance for credit losses). Below 30% is healthy; above 60% signals trouble. Replaces NPA/Assets for a more meaningful distress signal. Lower is better. |
| 5 | Net Charge-Offs to Loans | Annualized net loan losses written off during the period. Measures actual credit losses realized, as opposed to reserves or delinquencies. Lower is better. |
| 6 | Loss Allowance to Loans | Allowance for credit losses as a percentage of total loans. While adequate reserves are positive, a high ratio may indicate a troubled portfolio requiring large provisions. Scored as lower-is-better because elevated allowances generally correspond to elevated risk. |
Government-guaranteed non-current loans. The Texas Ratio numerator subtracts the guaranteed portion of the bank's own non-current loans. A loan carrying an SBA, USDA, FHA, or VA guarantee is reported as past due or nonaccrual like any other, but the guaranteeing agency, not the bank, absorbs the loss on the guaranteed share. Counting those balances as problem assets overstates the credit risk the bank actually bears, and it does so most severely at exactly the banks that specialize in guaranteed lending, where a large book of government-backed paper can push the ratio past the distress floors while the bank's own exposure is modest.
The guaranteed balances are taken from Call Report Schedule RC-N, Memoranda items 1.a and 1.b, which report the guaranteed portion of non-current loans (MDRM K040/K041) and the separately reported rebooked GNMA balances (K043/K044). Item 1.a excludes rebooked GNMA loans by definition, so the two items are additive and do not double count. The subtraction is floored at zero, and where a bank does not report the memoranda items the numerator is the unadjusted figure, so the adjustment can only reduce a bank's Texas Ratio, never raise it. The same netted ratio is used both for the Asset Quality ratio above and for the Texas Ratio hard floor in Section 7.
3.3 Management
The regulatory CAMELS Management component is inherently qualitative (board governance, compliance, audit, strategic planning). Since these factors are not available from public data, this component uses quantitative proxy metrics that reflect management effectiveness through operational efficiency, growth discipline, revenue strategy, and reserve prudence.
| # | Ratio | Description |
|---|---|---|
| 7 | Efficiency Ratio | Non-interest expense as a percentage of net revenue (net interest income + non-interest income). Measures how much it costs the institution to generate each dollar of revenue. Lower is better. |
| 8 | Asset Growth Rate (YoY) | Year-over-year percentage change in total assets, scored as deviation from a 5% optimal growth midpoint. Both extremes are penalized: rapid growth can signal aggressive risk-taking, loosened underwriting standards, or reliance on volatile funding, while sustained asset shrinkage signals franchise erosion, deposit flight, or business contraction. Moderate growth (~5%) indicates a healthy, organically expanding institution. A bank growing at 5% receives the best score; banks at −5% and +15% are penalized equally. U-shaped scoring around the 5% midpoint. |
| 9 | Revenue Diversification | Non-interest income as a percentage of total revenue (net interest income + non-interest income). Higher diversification indicates less reliance on the interest rate cycle and a broader business model. Higher is better. |
| 10 | Provision Coverage | Provision for credit losses divided by net charge-offs. Values above 100% indicate the institution is building reserves faster than it is realizing losses, a sign of conservative, forward-looking management. Higher is better. |
3.4 Earnings
Earnings quality and sustainability determine the institution's ability to build capital organically, absorb losses, and fund growth without excessive leverage.
| # | Ratio | Description |
|---|---|---|
| 11 | Return on Average Assets (ROAA) | Net income divided by average total assets, computed on a Trailing Twelve Month (TTM) basis when sufficient cross-quarter history is available (falling back to the single-quarter annualized value otherwise). The primary profitability measure for banks; the industry benchmark is approximately 1.0%. TTM ROAA is the sole Earnings metric because it captures comprehensive bottom-line profitability after accounting for all revenue sources, expenses, credit losses, and taxes; other earnings metrics (ROE, NIM, Pre-tax ROA) are highly correlated with ROAA and introduce redundancy without adding independent information. Higher is better. |
Why TTM ROAA. Reported quarterly net income is annualized and noisy: seasonal revenue patterns, one-time gains/losses, and provisioning lumpiness can move single-quarter ROAA materially even when the underlying franchise is unchanged. The TTM computation smooths across four consecutive quarters of profitability so the Earnings score reflects sustained capital generation rather than quarter-to-quarter noise. Banks lacking the cross-quarter history needed for TTM (newly chartered institutions, recent reporting-format conversions) fall back to single-quarter annualized ROAA.
3.5 Liquidity
Liquidity ratios assess whether the institution can meet its obligations under both normal and stressed conditions without incurring unacceptable losses.
| # | Ratio | Description |
|---|---|---|
| 12 | Net Loans & Leases to Deposits | Measures the proportion of deposits deployed into illiquid loans. Higher ratios indicate greater asset-liability mismatch and potential funding stress in a deposit run scenario. Lower is better. |
| 13 | Earning Assets to Total Assets | The proportion of total assets that generate interest or investment income. Non-earning assets (fixed assets, intangibles, cash in vault) do not contribute to income generation. Higher is better. |
| 14 | Uninsured Deposits to Total Deposits | SVB-style flight-risk indicator. Deposits above the $250K FDIC insurance cap have skin in the game and tend to run at the first sign of trouble (SVB and Signature were both ~85-90% uninsured at the quarter before failure). Uses the bank's RC-O memo disclosure when available, with a cap-based residual estimate as fallback. Lower is better. |
3.6 Sensitivity to Market Risk
Sensitivity measures the institution's exposure to adverse movements in interest rates, foreign exchange, and equity prices. For community banks, interest rate risk and CRE concentration are the dominant concerns.
| # | Ratio | Description |
|---|---|---|
| 15 | CRE Concentration (CRE / Capital) | Income-producing commercial real estate loans (construction, non-owner-occupied non-residential, and multifamily) divided by Tier 1 Capital plus the allowance for credit losses, per FFIEC SR 07-1. Owner-occupied non-residential CRE is excluded because the borrower repays from business cash flow rather than property income. Regulatory guidance flags CRE concentrations above 300% of capital as warranting heightened risk management. Lower is better. |
| 16 | Securities to Assets | Investment securities as a percentage of total assets. Large securities portfolios carry interest rate risk (unrealized losses in rising rate environments) and can indicate limited lending opportunities. Lower is better. |
4. Scoring Algorithm
4.1 Step 1: Ratio-Level Scoring
Each ratio is scored on a continuous 1–16 scale by interpolating the ratio's value within its 15 frozen percentile threshold boundaries (see Section 5, Threshold Calibration). Rather than rounding the value to one of 16 integer buckets, the model locates the two adjacent thresholds that bracket the value and returns a fractional score reflecting where the value sits between them. Two banks whose ratios differ only slightly receive slightly different scores, even within the same percentile band, so no resolution is lost to rounding.
For "higher is better" ratios (e.g., ROAA, Equity/Assets), thresholds are in descending order:
Threshold[0] > Threshold[1] > ... > Threshold[14]
value >= Threshold[0] → score 1 (top performers)
Threshold[i] > value >= Threshold[i+1] → score interpolated continuously
between (i+1) and (i+2), by the
value's position within the band
value < Threshold[14] → score 16 (worst performers)
For "lower is better" ratios (e.g., Noncurrent Loans, Efficiency Ratio), thresholds are in ascending order:
Threshold[0] < Threshold[1] < ... < Threshold[14]
value < Threshold[0] → score 1 (top performers)
Threshold[i] <= value < Threshold[i+1] → score interpolated continuously
between (i+1) and (i+2), by the
value's position within the band
value >= Threshold[14] → score 16 (worst performers)
A value falling exactly on a threshold scores the integer at that boundary; any value between two thresholds takes the continuously interpolated fraction.
For "U-shaped" ratios (e.g., Asset Growth Rate), the raw value is first transformed to a deviation from the optimal midpoint before scoring with the standard lower-is-better logic. For Asset Growth, the optimal midpoint is 5%, so the scored value is |rawGrowth - 5|. A bank growing at exactly 5% has a deviation of 0 (best possible score), while banks at −5% and +15% both have a deviation of 10 (equally penalized). The deviation-based thresholds are computed from the same peer-group percentile distribution as other lower-is-better ratios.
If a ratio value is null, undefined, or NaN (e.g., the institution does not report that field), the ratio is excluded from the component calculation.
4.2 Step 2: Component-Level Scoring
Each CAMELS component score is the simple average of its constituent ratio scores, excluding any ratios that returned null and any ratios whose raw value is below their contribution floor:
Component Score = Sum(eligible ratio scores) / Count(eligible ratio scores)
A ratio is eligible when:
- Its scored value is non-null, AND
- If it defines a minimum-value floor, the raw value meets or exceeds it
The contribution floor is used selectively for ratios where a benign reading (typically zero) would mechanically pull the component average toward the best possible score for every bank, an unintended side effect of simple averaging. The floor is currently applied to HTM Unrealized Loss / Tier 1 in the Capital component (5% floor), so banks with no HTM losses score Capital purely from the Leverage Ratio, while banks with material HTM losses average the two. The ratio is still shown in the per-bank breakdown when below the floor, with an annotation that it is not counted toward the average, so the rationale is transparent in both the UI and exported credit reviews.
If all ratios in a component are null or below their floors, the component score is null and excluded from the composite.
4.3 Step 3: Composite Rating
The composite score is a weighted average of the six component scores, with weights renormalized if any component is null:
Composite = Sum(Component Score_i x Weight_i) / Sum(Weight_i)
The six component weights start from the CAMELS supervisory framework and standard credit-analysis principles, and are refined by machine learning; the weights currently deployed, and how they are derived, are shown in the subsection below. The same component weights apply to every bank; large banks are not given a different weight mix. Instead, G-SIBs and public-market banks receive fixed composite adjustments (−2 / −1; see Section 7, Qualitative Adjustments).
Weight rationale: The starting framework places the highest weight on Asset Quality because credit risk (the probability and severity of borrower default) is the primary driver of bank failures. The refinement then shifts weight toward the components most predictive of distress in the recent training window, currently trimming Asset Quality and raising Liquidity and Sensitivity (the 2023 rate-shock cohort), with a 12% floor so no CAMELS dimension is neglected. Capital is the loss-absorption pillar; Earnings reflects internal capital generation; and Management uses quantitative proxy metrics for governance quality.
Machine-Learned Weight Refinement
The starting weights were refined by a machine-learning model trained on four years (16 quarters) of historical bank distress. (This 16-quarter weight-training window is distinct from the 20-quarter pool used for threshold calibration in Section 5: because the distress label looks four quarters ahead, the four most recent quarters cannot yet be labeled and are held out, which leaves 16.) The model fits a logistic regression on the six CAMELS component scores to predict whether a bank becomes distressed within the next four quarters, where "distressed" means any of: actual failure (FDIC failed-bank list), Texas Ratio crossing 50% in a future quarter, the composite score crossing into the distress range, or a sharp quarter-over-quarter composite deterioration.
Why keep the framework at all, rather than let the data choose the weights outright? Two reasons matter most:
- It keeps the weights from chasing the last crisis. The training window is only four years (16 quarters), so an unconstrained fit reflects whichever kind of failure dominated it: the 2008–2014 credit crisis pushes weight onto Asset Quality, while the 2021–2025 rate-shock era pushes it onto Liquidity and Sensitivity. Anchoring to the framework gives a through-the-cycle view, so the weights do not whipsaw with each retrain. (It also avoids a subtler trap: one of the distress signals, a rising Texas Ratio, is itself an Asset Quality measure, so an unconstrained fit would over-weight that component by partly predicting itself.)
- It is defensible for model-risk review. "The CAMELS supervisory framework is the starting point and the model refines it from data" is a structural, reviewable rationale; "the optimizer put this percentage on one component" is not. Anchoring to the framework also steadies the weights where distress events are sparse, so a single noisy training window cannot dominate.
How the two are combined. The deployed weight for each component is an even blend of the machine-learned weight and the starting weight, so the data re-ranks the components without overriding the framework:
Deployed weight = 0.5 × machine-learned weight + 0.5 × starting weight
The machine-learned half re-ranks toward what the recent record favors; the starting half retains the through-the-cycle framework. Both sets of weights sum to 100%, so the blend does too.
Minimum-weight floor. A 12% floor is then applied so every CAMELS dimension carries weight that materially affects the composite. Components below the floor are raised to 12% and the deficit is redistributed proportionally from above-floor components, keeping the total at 100%. The floor exists because machine learning can drive a component close to zero during cycles where its predictive lift temporarily collapses; the floor ensures the framework remains diversified and that no CAMELS dimension can be silently neglected.
Deployed weights (model version 3, most recently retrained 2026-08-30 on 16 quarters / 74,300 records / 3,504 distress events):
| Component | Starting Weight | Machine-Learned | Deployed (after blend + floor) |
|---|---|---|---|
| C -- Capital Adequacy | 15% | 5.6% | 12.0% (floored) |
| A -- Asset Quality | 30% | 20.3% | 24.7% |
| M -- Management | 10% | 15.1% | 12.3% |
| E -- Earnings | 20% | 12.6% | 16.0% |
| L -- Liquidity | 15% | 24.6% | 19.4% |
| S -- Sensitivity | 10% | 21.8% | 15.6% |
Capital currently sits at the 12% floor (its pre-floor blend value is roughly 10.3%). Liquidity and Sensitivity carry above-framework weight because the recent training window includes the 2023 rate-shock cohort. Every deployed version is stored with its calibration parameters, so the weights in force on any date are reproducible and auditable. The weights update as new quarterly data is released, but the even blend with the framework and the 12% floor hold them within a narrow band rather than letting them swing with each refresh.
5. Threshold Calibration
Frozen Pooled Percentile Thresholds
The primary scoring method uses frozen peer group percentile thresholds computed from 20 quarters of pooled historical data (~94,000 institution-quarter records). Each bank is scored against institutions of similar asset size, eliminating the systematic bias that occurs when applying a single set of absolute thresholds across banks of vastly different size and business model.
Why pooled (frozen) thresholds: If percentile thresholds were recomputed each quarter from that quarter's population, scores would be unstable in the trend view: a bank's composite could shift from quarter to quarter even when its fundamentals were unchanged, purely because the peer distribution moved (for example, industry-wide NIM compression would move the NIM threshold and penalize all banks equally). Pooling 20 quarters of data into a single, fixed threshold set keeps the boundaries stable, so quarter-to-quarter score changes reflect actual changes in a bank's financial condition.
Thresholds persist across scoring runs. The pooled threshold set is computed once, stored, and reused by every subsequent run. It is deliberately not re-derived when a new quarter is added.
This matters because pooling alone does not deliver stability. A pooled set that is recomputed on each run still moves: the pool is a rolling 20-quarter window, so each new quarter shifts the percentile boundaries slightly and recalculates every previously published score against them. Computing the set once and holding it fixed is what prevents that.
A published rating is a record. It does not change because an unrelated later quarter arrived. Only the bank's own reported figures move its score.
Recalibration is a versioned event. Re-deriving the thresholds is performed deliberately, not as part of routine processing, and restates historical scores by construction. It is therefore treated as a model-version change: dated, accompanied by a before-and-after impact analysis, and recorded in the model's change history. Expected cadence is roughly annual. The tradeoff accepted here is that a fixed threshold set gradually diverges from the industry distribution over time; periodic recalibration is what corrects that.
Why peer scoring is necessary: Large GSIBs (e.g., JPMorgan, Bank of America) naturally operate with lower capital ratios (~7% leverage vs ~10%+ for community banks), lower net interest margins (~2.5% vs ~3.5%), and higher efficiency costs due to their global operations and regulatory complexity. Absolute thresholds would penalize these structural differences as weaknesses, even though they are normal for the peer group.
Peer Groups
| Peer Group | Label | Asset Range |
|---|---|---|
| 1 | >$100B | Total assets >= $100 billion |
| 2 | $10B--$100B | $10 billion <= assets < $100 billion |
| 3 | $3B--$10B | $3 billion <= assets < $10 billion |
| 4 | $1B--$3B | $1 billion <= assets < $3 billion |
| 5 | $300M--$1B | $300 million <= assets < $1 billion |
| 6 | <$300M | Assets < $300 million |
Each bank is assigned to its peer group based on total assets from its most recent quarterly filing.
Bell-Curve Score Distribution
The percentile thresholds use a right-skewed bell-curve distribution, so the median bank in each peer group scores near the center of the distribution:
| Strength Score | Width | Cumulative % |
|---|---|---|
| 99–100 | 1% | 1% |
| 97–99 | 2% | 3% |
| 95–97 | 4% | 7% |
| 93–95 | 7% | 14% |
| 91–93 | 11% | 25% |
| 77–91 | 14% | 39% |
| 62–77 | 14% | 53% |
| 52–62 | 12% | 65% |
| 43–52 | 10% | 75% |
| 38–43 | 7% | 82% |
| 32–38 | 5% | 87% |
| 26–32 | 4% | 91% |
| 19–26 | 3% | 94% |
| 12–19 | 2.5% | 96.5% |
| 5–12 | 2% | 98.5% |
| 1–5 | 1.5% | 100% |
Each row is one grade of the internal 1–16 composite, and its score range is the deployed anchor mapping evaluated at that grade's boundaries (grade g spans composite g−0.5 to g+0.5), so the table is reproducible from the mapping rather than maintained by hand. The width of each band (its share of the population) is set by the composite percentile spacing and is unchanged by the choice of anchors; the score ranges are wider through the middle of the scale, where the anchors deliberately give the sound majority more room to separate, and narrow at the very top. The shape reflects the empirical reality that most FDIC-insured institutions are fundamentally sound, while a smaller tail exhibits material weaknesses. The median bank scores near 65, with the healthy majority spread across the upper-middle of the scale and a long thin tail toward the weakest scores. This spacing gives more resolution among sound banks, where most counterparties sit, and reduces reliance on large qualitative adjustments by letting the quantitative scorecard itself separate strong institutions from one another.
For each ratio direction (lower-is-better or higher-is-better), a set of 15 percentile points is computed from the pooled peer-group distribution and used as the boundaries between adjacent grades. Score 1 corresponds to the best ~1% of the peer-group distribution; Score 16 corresponds to the worst ~1.5%. The percentile spacing is denser near the median (where most banks cluster and small ratio movements should not shift grades) and sparser in the tails (where any extreme reading is meaningful).
Recalibration Cadence
Routine processing runs on an automated schedule and re-scores institutions as new Call Report data is released, but always against the frozen threshold set; the thresholds themselves are not recomputed as part of that routine. Re-deriving the thresholds (recalibration) is instead a deliberate, dated model-version event on a roughly annual cadence, as described under "Recalibration is a versioned event" above. At a recalibration, the latest 20 quarters are pooled, the percentile boundaries are recomputed for each of the 17 ratios across each peer group plus the all-institutions population, and the historical scores are restated against the new set so the full history stays comparable.
If the pre-computed frozen threshold store is unavailable, the system falls back to single-quarter live computation (see Frozen vs. Per-Quarter Thresholds below).
Frozen vs. Per-Quarter Thresholds
| Aspect | Frozen | Per-Quarter Recompute |
|---|---|---|
| Data pool | 20 quarters (~94K records) | 1 quarter (~4,500 records) |
| Stability | Thresholds are constant across all scored quarters | Thresholds shift with each quarter's population |
| Score changes | Reflect actual financial changes only | Mix of financial changes and threshold drift |
| Peer Group 1 sample | ~600 institution-quarter records | ~30 records |
6. Data Sources
All input data is sourced from public regulatory filings made by FDIC-insured depository institutions.
The model uses the 20 most recent quarters of available financial data for threshold calibration and trend scoring:
- All 20 quarters are pooled to compute the frozen percentile thresholds.
- Each of the 20 quarters is scored individually against those frozen thresholds, so a bank's full history is comparable to its current quarter.
7. Qualitative Adjustments
After computing the raw quantitative CAMELS composite score, the model applies qualitative adjustments that shift the composite score by points. These reflect structural advantages and disadvantages that peer-group-relative scoring cannot capture.
These adjustments and floors operate on the internal 1–16 composite (the intermediate quantity from Sections 2 and 4), not directly on the 1–100 strength score, and the two run in opposite directions: on the composite, 1 is strongest and 16 is weakest. So an adjustment that lowers the composite (the −2 and −1 below) raises the strength score, while a floor holds the composite no better than a given weak value, which caps the strength score near the bottom of the range. Every floor therefore caps a bank's strength score in the Elevated tier or below; the approximate strength score for each floor value is shown alongside it in the summary below. (A floor holds the composite no better than the boundary of the named grade, so the capped strength score is the score at that boundary, not the score at the grade's centre.)
| Adjustment | Composite Adjustment | Applies To | Rationale |
|---|---|---|---|
| G-SIB Government Support | −2 | 8 U.S. G-SIBs (JPMorgan, BofA, Citi, Wells Fargo, Goldman, Morgan Stanley, BNY, State Street) | Implicit too-big-to-fail guarantee reduces default risk; G-SIBs benefit from extraordinary government support frameworks and resolution regimes. The −2 adjustment lowers the internal composite (raising the strength score) to reflect that reduced default risk, an uplift that peer-group-relative scoring of the raw ratios would otherwise miss. |
| Public Market Access | −1 | Category III-IV banks plus Northern Trust (Category II) | Ability to issue preferred stock, subordinated debt, and other capital instruments in public markets provides additional loss-absorbing capacity and funding flexibility not available to smaller institutions. |
| Loans-to-Deposits Floor | Tiered floor | Institutions with Net Loans & Leases to Deposits > 115% | A loans-to-deposits ratio meaningfully above 100% means the institution has deployed more into illiquid loans than it holds in deposits and is relying on wholesale or brokered funding to bridge the gap. The 115% first tier accommodates online and wholesale-funded business models. The floor escalates with severity: >115% → floor at 11, >130% → floor at 12, >150% → floor at 14. |
| Texas Ratio Floor | Tiered floor | Institutions with Texas Ratio > 40% | Texas Ratio above 40% signals material credit distress that the simple-average Asset Quality component score under-penalizes (only two of four A ratios pin at 16 even in deep stress). The floor escalates with severity: >40% → floor at 12, >60% → floor at 14, >100% → floor at 16. The ratio tested here is the guarantee-netted ratio defined in Section 3.2, so a bank is not floored on non-current balances a government agency will make whole. |
| Sustained-Loss ROAA Floor | Tiered floor | Institutions with TTM ROAA below −1% | A bank generating sustained losses is eroding capital organically, not just earning less. Because Earnings carries only ~11% of composite weight, a max-distress E score alone cannot move the composite to the distress band. The floor escalates with severity: <−1% → floor at 13, <−3% → floor at 15. |
These are hard floors, not additive adjustments: each one overrides the weighted-average composite when applicable, and the highest applicable floor wins.
The final composite score is calculated as:
Adjusted Score = clamp(Raw Quantitative Score + Qualitative Adjustments, 1, 16)
Hard floors (highest applicable floor wins):
Loans/Deposits > 115% → floor at 11 (strength ≈ 43)
Loans/Deposits > 130% → floor at 12 (strength ≈ 35)
Loans/Deposits > 150% → floor at 14 (strength ≈ 19)
Texas Ratio > 40% → floor at 12 (strength ≈ 35)
Texas Ratio > 60% → floor at 14 (strength ≈ 19)
Texas Ratio > 100% → floor at 16 (strength ≈ 5)
TTM ROAA < −1% → floor at 13 (strength ≈ 26)
TTM ROAA < −3% → floor at 15 (strength ≈ 12)
Final Score = max(Adjusted Score, applicable floor)
where the score is bounded to the [1, 16] rating scale. The hard floors are applied after all additive adjustments, ensuring that no combination of favorable adjustments (e.g., G-SIB support) can override an asset-quality, earnings, or liquidity constraint that's already in distress territory.
Floor stickiness across blending: Floors are evaluated against the current quarter's ratios, but the published composite is a 3-quarter weighted moving average (50% current, 30% prior, 20% two-quarters-ago). To preserve the "highest applicable floor wins" rule across the blending step, the floor is re-applied after blending: if the bank's current-quarter floor is more severe than the blended composite, the composite is set to the floor's boundary value. This prevents floor signals from being diluted by prior quarters where the floor wasn't triggered (e.g., a bank whose Loans/Deposits ratio jumped from 95% to 105% this quarter immediately reflects the loans-to-deposits floor rather than blending back toward its two prior in-band quarters).
Why these specific floors: The CAMELS composite is a weighted average across six components, and each component is itself a simple average across multiple ratios. That double-averaging dampens single-ratio extremes: a 62% Texas Ratio only drives 2 of the 4 Asset Quality ratios to their weakest, and even a maxed-out Earnings component (one ratio at weight ~11%) can only shift the composite by a point or so. For most banks this is the right behavior because credit risk is multi-dimensional, but a bank already in distress on a single dimension can still land in the Moderate tier or higher on the strength score, understating the credit risk a creditor would actually face. The hard floors recover the right behavior at the tail without distorting the middle of the distribution.
8. Model Validation
The rating is backtested against the historical record of bank failures to confirm that lower grades correspond to genuinely higher failure risk. Two questions are tested separately: does the rating rank-order failures (discrimination), and do the failure rates implied by each grade match what actually happened (calibration). A third test checks that ratings are stable quarter to quarter rather than jumping around.
8.1 Discrimination
Discrimination is measured with the Accuracy Ratio (Gini) derived from the cumulative accuracy profile: banks are ranked from worst grade to best, and the curve plots the share of eventual failures captured against the share of the population reviewed. A Gini of 0 is random ordering; 1.0 is perfect ordering. Two tests are run:
Out-of-sample (2008 crisis). The strictest test. The score thresholds are calibrated using only pre-crisis quarters (2005–2007), then frozen and used to score the 2008–2010 crisis window (~89,000 institution-quarters, 363 failing banks). This isolates hindsight in the threshold fitting: the percentile boundaries were set without any crisis-period data. It does not claim the model was designed in ignorance of crisis dynamics; the choice of ratios, weights, and hard floors reflects general knowledge of what stresses banks.
Full history. The same scoring run measured across the full available panel (2007–2026; ~457,000 institution-quarters, 495 in-window failures) on the deployed frozen thresholds. Because these thresholds are pooled from the same history they are tested on, this is an in-sample-threshold measurement, reported as a consistency check on the out-of-sample result rather than a stricter test.
| Test | Lead horizon | Gini (Accuracy Ratio) |
|---|---|---|
| Out-of-sample (2005–07 → 2008–10) | 4 quarters ahead | 0.89 |
| Out-of-sample (2005–07 → 2008–10) | 8 quarters ahead | 0.80 |
| Full history | 4 quarters ahead | 0.92 |
| Full history | 8 quarters ahead | 0.87 |
These two figures are not independent: the 2008–2010 cluster supplies 363 of the 495 full-history failures, so the full-history run is largely the same crisis event measured over a longer, mostly-quiet panel. The load-bearing evidence is the out-of-sample test itself, pre-crisis-fit thresholds still rank-ordering the crisis failures. A bank-level bootstrap (resampling institutions, not rows) puts the 95% confidence interval on the full-history Gini at 0.90 to 0.94 at the 4-quarter lead. The companion Back-testing document reports these results in full, with a year-by-year walk-forward, single-ratio baselines, a pre-failure trajectory, and a candid 2023 case study.
8.2 Calibration
For calibration the empirical failure rate within four quarters is computed for each strength-score band across the full history. The five upper bands are the product display tiers; the lowest tier (High Risk, 1–20) is split into finer bands to show how sharply failure risk rises through the bottom of the scale. A well-calibrated rating shows failure rates that rise monotonically as the score falls, and here the rise is monotonic across every band:
| Strength Score | Population | Failures (4Q) | Failure rate |
|---|---|---|---|
| 80 and above | 57,961 | 4 | 0.01% |
| 60–80 | 185,766 | 31 | 0.02% |
| 45–60 | 145,308 | 69 | 0.05% |
| 30–45 | 40,062 | 129 | 0.32% |
| 20–30 | 15,219 | 260 | 1.71% |
| 10–20 | 3,881 | 119 | 3.07% |
| 5–10 | 6,672 | 731 | 10.96% |
| 1–5 | 2,341 | 611 | 26.10% |
Failure risk climbs steeply through the bottom of the scale: from 1.7% at 20–30, to 3.1% at 10–20, to 11.0% at 5–10, to 26.1% below 5. A bank scoring in the lowest band (1–5) failed within a year in about one case in four, versus 0.01% for banks scoring 80 or above (the Strongest tier): more than two thousand times the realized failure rate. (The Strongest tier saw only 4 failures across roughly 58,000 institution-quarters, so the top of the scale is best read as "failure is vanishingly rare" rather than as a precise multiple.)
8.3 Stability
A useful rating should change because a bank's condition changed, not because of measurement noise. Across ~82,000 consecutive-quarter observations (the most recent five years):
- 69.8% of banks' strength score moved by 2 points or less quarter over quarter.
- 91.7% moved by 5 points or less; the median quarter-over-quarter move was 1 point.
This stability comes from two design choices: the frozen pooled thresholds remove the quarter-to-quarter threshold drift that would otherwise move a score even when a bank's fundamentals were unchanged, and the three-quarter weighted moving average dampens single-quarter noise before it reaches the reported score.
8.4 Caveats
- Bank failures are sparse, especially in the post-2015 window, so the confidence intervals around the Gini figures are wide. Treat them as directional evidence of strong rank-ordering, not as exact point estimates.
- The out-of-sample test scores a concentrated stress event (the 2008–2010 failure cluster); its grade-level failure rates run higher across the board than a normal period and are not directly comparable to the full-history calibration above.
- The backtest measures the quantitative scorecard. The qualitative hard floors and adjustments described above are applied on top of the score and are not separately captured in these figures.
9. Limitations and Disclaimers
Not a credit rating. Counterparty Insight is not a credit rating agency and is not registered as a Nationally Recognized Statistical Rating Organization (NRSRO) under Section 15E of the Securities Exchange Act. The 1–100 strength scores this model produces are independent, quantitative analytical assessments for informational use only. They are not "credit ratings" as defined under the Credit Rating Agency Reform Act, are not a recommendation to buy, sell, or hold any security or to extend or deny credit, and must not be used as the sole basis for any investment, lending, or counterparty decision. The numeric scale is a proprietary risk indicator used for interpretability; it is not affiliated with, endorsed by, or derived from any rating agency.
Not a regulatory rating. This model is an independent analytical tool. Actual CAMELS ratings assigned by bank examiners are confidential and incorporate qualitative factors (management interviews, compliance reviews, IT audits) that are not available from public data.
Management component is proxy-based. The real CAMELS "M" component evaluates board governance, strategic planning, risk management culture, internal controls, and compliance. The proxy metrics used here (efficiency, growth discipline, diversification, reserve coverage) are correlated with management quality but are not a substitute.
Lagging indicators. Financial ratios are backward-looking, reflecting the institution's condition as of the most recent quarterly report date. Rapidly deteriorating conditions may not be reflected until the next filing.
Small peer groups. Peer Group 1 (>$100B) contains ~30 institutions per quarter. Pooling 20 quarters mitigates this (~600 institution-quarter records), but the group remains the smallest. Ratios with fewer than 10 valid data points fall back to all-institutions thresholds.
Data availability. Some ratios may be null for certain institution types (e.g., thrifts may report different capital metrics). The model handles nulls by excluding them from the averaging, which may result in component scores based on fewer ratios.
Threshold staleness. The frozen thresholds are held fixed between recalibrations rather than recomputed as new quarters arrive, which keeps published scores stable (see Section 5). The tradeoff is that when macroeconomic conditions shift significantly (e.g., rapid rate hikes), the fixed thresholds gradually diverge from the current industry distribution until the next recalibration, a deliberate, dated model-version event on a roughly annual cadence. Between recalibrations the thresholds reflect the 20-quarter pool as of the last recalibration, not the latest quarter.
Data quality and lineage. Inputs are sourced from public regulatory filings and are subject to the filers' own reporting accuracy, restatements, and framework differences (call report field availability varies by charter type and reporting framework). The pipeline applies field-mapping and reconciliation controls, but residual data-quality risk is inherent to any model built on third-party-reported financials.
Not investment advice. This model is for informational and analytical purposes only. It should not be used as the sole basis for investment, lending, or counterparty decisions.
Appendix A. Model Change Log
Dated model-version events. Each entry records what changed, why, and what evidence was produced. Routine quarterly data refreshes are not model changes and are not listed.
2026-09-09 — Texas Ratio: net of government-guaranteed non-current loans
Change. The Texas Ratio numerator now subtracts the U.S.-government-guaranteed portion of the bank's non-current loans, sourced from Call Report Schedule RC-N Memoranda 1.a and 1.b (MDRM K040/K041 for the guaranteed portion excluding rebooked GNMA, K043/K044 for rebooked GNMA). The subtraction is floored at zero and applies to both the Asset Quality ratio (Section 3.2) and the Texas Ratio hard floor (Section 7).
Why. The unadjusted ratio treated agency-guaranteed balances as problem assets, overstating the credit risk actually borne by banks that specialize in SBA, USDA, or FHA/VA lending. Because the Texas Ratio feeds both an Asset Quality ratio and a hard floor, the overstatement was counted twice for the affected banks.
Evidence. The adjustment is self-targeting: of the 49 banks carrying a Texas Ratio floor before the change, 11 cleared it (guaranteed shares of 25% to 97% of non-current balances) and 38 were unchanged (guaranteed shares of 0% to 7%). Verified against a worked case in which 96.8% of a $4.71B non-current balance was government-guaranteed. Full-history discrimination was unaffected (Gini 0.924 at 4Q, unchanged), which is expected: the change corrects the level of a small number of banks rather than the ordering of the population. See docs/risk-rating/texas-ratio-guaranteed-loan-adjustment.md.
Backward compatibility. MDRM codes are mapped at read time, so the change took effect on the existing Call Report store without a data rebuild. Banks that do not report the memoranda items are scored on the unadjusted numerator, so the adjustment can only lower a Texas Ratio, never raise one.
2026-09-09 — Score anchor re-calibration (lower half of the scale)
Change. The anchors mapping the internal 1–16 composite onto the 1–100 strength score were re-set through the lower half of the scale. The composite anchor points are unchanged; the score values attached to them moved from [100, 91, 84, 74, 65, 54, 42, 30, 18, 1] to [100, 91, 84, 74, 65, 60, 53, 44, 30, 1].
Why. Sound banks with no distress signal were landing in the 40s, low enough to read as a warning when the underlying composite did not warrant one. The re-anchoring lifts the lower-middle of the scale so that score bands correspond more closely to the risk they are meant to convey.
What did and did not change. The mapping remains strictly monotonic in the composite, so rank ordering, percentile position, peer comparisons, and every discrimination statistic are unchanged by construction (full-history Gini 0.924 at 4Q and 0.866 at 8Q, identical before and after). The median bank still scores 66. What moved is the score attached to a given composite below roughly the 60th percentile, and therefore the population of each display tier: Elevated fell from 9.8% to 4.0% of banks and the two weakest tiers from 6.5% to 2.8% combined.
Consequences to read carefully. Because the tier boundaries are fixed points on the score scale and the scale moved beneath them, the tiers now sit at different composites. Two published figures move as a direct result and should not be read as a change in model performance: the four-quarter failure rate within each of the lower bands rises (fewer, weaker banks now occupy them), and the share of eventual failures already scoring below 30 four quarters out falls from 92% to 79%. Rank ordering is unchanged; the line labelled "30" simply sits deeper in the distribution than it did. Quarter-over-quarter stability improved (69.8% of scores move by 2 points or less, against 58.7% before). Section 8 reports all of these on the re-anchored scale.
Tier naming. The 20–30 tier was renamed from "Vulnerable" to "Weak" in the same release. This is a label change only; no boundary or score moved.