← Main site

Decision brief · Market risk

When a 95% VaR model fails its own backtest

A historical 95% VaR forecast produced too many breaches, and the breaches clustered instead of arriving independently.

Depth 1 · Answer

The historical VaR forecast failed on both breach frequency and breach timing.

Across 2,140 one-step forecasts, the rolling 500-observation historical 95% VaR model recorded 135 breaches versus 107 expected: a 6.31% breach rate. Kupiec coverage and Christoffersen independence and conditional-coverage tests all rejected calibration at the 5% level. The misses were too frequent and clustered.

Published 8 August 2026 · Underlying investigation last updated 3 August 2026

Observed breaches
135
Expected breaches
107
Breach rate
6.31%
Forecast days
2,140

Why it matters

Risk owners, model validators, and control functions

Anyone setting limits, capital buffers, or escalation thresholds from historical VaR should care whether a model's exceptions match its stated confidence level and arrive independently.

Decision affected

Whether VaR is fit to govern the portfolio by itself

This evidence says no for this portfolio, model, and sample. VaR can remain one lens, but it needs recalibration or redesign and must sit beside expected shortfall, drawdown, scenario analysis, and correlation stress.

Depth 2 · Evidence

Three tests separate frequency error from clustered failure

The study used a look-ahead-free, monthly rebalanced strategic ETF portfolio with 10 basis points of proportional turnover cost. After warm-up, it produced 2,640 net daily returns and 2,140 forecasts. Each VaR estimate used only the preceding 500 observations.

The unconditional-coverage test rejected the stated 5% breach probability (Kupiec p = 0.0075). The independence test rejected non-clustered arrivals (Christoffersen p = 0.00025). The combined conditional-coverage test also rejected (p = 0.000035). Reporting all three matters: the model failed both in how often and in how consecutively losses crossed the threshold.

The study's full-sample historical VaR95 point estimate was 1.079%, with a 95% moving-block bootstrap interval of 0.930%–1.219%. ES95 was 1.718%, with an interval of 1.421%–2.126%. Those intervals quantify estimator uncertainty; they do not rescue a rejected forecast process.

Depth 3 · Application

Backtest the forecast process, not just the latest risk number

  • Test unconditional coverage and exception independence separately; a correct average rate can still hide clustering.
  • Treat an exception cluster as information about regime adaptation, not as a sequence of unrelated surprises.
  • Report expected shortfall beside VaR so the severity beyond the threshold remains visible.
  • Use date-fixed scenarios and correlation stresses for risks a rolling empirical quantile may adapt to slowly.
  • Preserve costs, timing, and no-look-ahead controls in the backtest that generates the evidence.

Depth 4 · Limits

A rejection is specific evidence, not a universal verdict on VaR

  • The bootstrap conditions on one realized history, and the 21-day block length is judgmental.
  • Historical quantiles can adapt slowly after regime changes; other windows and weighting schemes may behave differently.
  • The portfolio uses fixed strategic weights with monthly rebalancing; dynamic de-risking was not modeled.
  • Close-to-close ETF prices omit intraday liquidity and execution stress.
  • Historical scenarios replay known crises and do not bound future loss paths.

Depth 5 · Method and code

Inspect the forecasts, tests, data version, and code

Related research