Empirical model validation · Market risk
Historical VaR backtesting under coverage, independence, and stress-window analysis
The lab tests a multi-asset risk forecast against later observations. Sampling uncertainty and three dated stress periods are reported separately.
Revised · research record
Decision summary
A plausible risk number failed when its forecasts met the future
Interpretation
The rolling 95% historical VaR model was rejected. It produced 135 breaches against 107 expected. The breaches clustered rather than arriving independently.
Decision context
Backtest and monitor historical quantiles as forecasts. Do not accept them as fixed capital or limit inputs. Slow adaptation to regime changes is part of the model risk.
Intended analytical use
Risk managers, portfolio teams, model validators, audit and governance functions, and limit owners use it to test coverage.
Principal limitation
This rejection applies to one fixed multi-asset portfolio, dated daily data, and a 500-observation window. It does not establish that every alternative risk model is better.
Data and design
A leak-free, version-pinned backtest
- Backtest sample
- 2016-01-04 – 2026-07-06
- Price vintage
- 2026-07-06
- Snapshot hash
sha256:5dc8433e1c99bf8e…- Pinned source
ad24c4999583- Daily observations
- 2,640
- Total modeled costs
- 0.183% of NAV
Target weights: 30% SPY, 10% QQQ, 5% IWM, 10% EFA, 20% AGG, 10% TLT, 7.5% GLD, and 7.5% VNQ. Rebalancing is monthly after a 252-day warm-up. Turnover is charged 10bp on one side. No future observation enters a rebalance decision.
Observed path
Portfolio risk results
- Ann. return
- 10.14%
- Ann. volatility
- 11.50%
- Sharpe (rf 3%)
- 0.64
- Max drawdown
- -24.3%
- VaR 95 (1d)
- 1.08%
- ES 95 (1d)
- 1.72%
- Skewness
- -0.54
- Excess kurtosis
- 12.98
Annualized where applicable. VaR and ES are one-day historical estimates at 95%. Uncertainty intervals appear in the next section.
View key chart data
| Measure | Value |
|---|---|
| Ending value of $1 | $2.750 |
| Maximum drawdown | -24.33% |
| VaR 95, one day | 1.08% |
| Expected Shortfall 95, one day | 1.72% |
| Backtest observations | 2,640 |
Sampling uncertainty
Bootstrap intervals for VaR and Expected Shortfall
The point estimates are recomputed across 2,000 moving-block bootstrap
samples, using 21-day blocks and seed 20260803. Blocks preserve some local dependence.
The percentile interval still describes uncertainty conditional on this observed history. It does
not manufacture crises missing from that history.
| Measure | Point estimate | Bootstrap 95% interval |
|---|---|---|
| Historical VaR 95 | 1.079% | 0.930% – 1.219% |
| Historical ES 95 | 1.718% | 1.421% – 2.126% |
Falsification test
The 95% historical VaR forecast is rejected
Each forecast uses only the preceding 500 observations. In the holdout sequence, breaches occur too often and cluster. Both Kupiec unconditional coverage and Christoffersen conditional coverage reject the model at 5%. A red result is useful. It marks the exact boundary of a method that looks plausible in a static histogram.12
| Test | Statistic | p-value | Decision |
|---|---|---|---|
| Observed breaches | 135 / 2140 | 6.31% rate | Expected 107 |
| Kupiec coverage | 7.148 | 0.00751 | Reject at 5% |
| Christoffersen independence | 13.392 | 0.00025 | Reject at 5% |
| Conditional coverage | 20.540 | 0.00003 | Reject at 5% |
Historical stress
Three dated windows inside the sample
Stress windows are descriptive slices fixed by event dates, not optimized after inspecting portfolio minima. They show how the same allocation behaved through the Q4 2018 volatility shock, the COVID-19 selloff, and the 2022 inflation and rate shock.
| Window | Dates | Cumulative return | Max drawdown | Worst day | VaR breaches |
|---|---|---|---|---|---|
| Volatility shock of Q4 2018 | 2018-09-20 – 2018-12-24 | -10.2% | -10.8% | -1.86% | 16 |
| COVID-19 selloff | 2020-02-19 – 2020-03-23 | -21.3% | -21.7% | -7.08% | 9 |
| 2022 inflation and rate shock | 2022-01-03 – 2022-10-14 | -24.2% | -24.0% | -3.37% | 26 |
Interpretation
Implications of coverage and independence rejection
VaR marks a quantile. Expected Shortfall averages the losses beyond it. Neither becomes a law of nature because it has three decimal places. The rejected coverage tests say that a rolling 500-day empirical distribution adapts too slowly to regime changes in this portfolio. They do not say the portfolio is uninvestable.
The next research iteration should compare filtered historical simulation and volatility- scaled forecasts on the same locked dates. The present model stays as the baseline. That comparison belongs in the ledger before any method is declared superior.
Reproduce
Code, commit, and references
Inspect the risk notebook at the source commit and the backtest engine used to construct the return stream.
- Kupiec, P. H. (1995), “Techniques for Verifying the Accuracy of Risk Measurement Models,” Journal of Derivatives 3(2), 73–84. doi:10.3905/jod.1995.407942 ↩
- Christoffersen, P. F. (1998), “Evaluating Interval Forecasts,” International Economic Review 39(4), 841–862. doi:10.2307/2527341 ↩
- Artzner, P., F. Delbaen, J.-M. Eber, and D. Heath (1999), “Coherent Measures of Risk,” Mathematical Finance 9(3), 203–228. doi:10.1111/1467-9965.00068