Kyle Wisniewski

Every result here is reproducible: open-source code and 229 automated checks, from theory-anchored quantitative tests to publishing checks. Browse the source on GitHub →

Empirical model validation · Market risk

Historical VaR backtesting under coverage, independence, and stress-window analysis

The lab tests a multi-asset risk forecast against later observations. Sampling uncertainty and three dated stress periods are reported separately.

Revised · research record

Decision summary

A plausible risk number failed when its forecasts met the future

Interpretation

The rolling 95% historical VaR model was rejected. It produced 135 breaches against 107 expected. The breaches clustered rather than arriving independently.

Decision context

Backtest and monitor historical quantiles as forecasts. Do not accept them as fixed capital or limit inputs. Slow adaptation to regime changes is part of the model risk.

Intended analytical use

Risk managers, portfolio teams, model validators, audit and governance functions, and limit owners use it to test coverage.

Principal limitation

This rejection applies to one fixed multi-asset portfolio, dated daily data, and a 500-observation window. It does not establish that every alternative risk model is better.

Data and design

A leak-free, version-pinned backtest

Backtest sample
2016-01-04 – 2026-07-06
Price vintage
2026-07-06
Snapshot hash
sha256:5dc8433e1c99bf8e…
Pinned source
ad24c4999583
Daily observations
2,640
Total modeled costs
0.183% of NAV

Target weights: 30% SPY, 10% QQQ, 5% IWM, 10% EFA, 20% AGG, 10% TLT, 7.5% GLD, and 7.5% VNQ. Rebalancing is monthly after a 252-day warm-up. Turnover is charged 10bp on one side. No future observation enters a rebalance decision.

Observed path

Portfolio risk results

Ann. return
10.14%
Ann. volatility
11.50%
Sharpe (rf 3%)
0.64
Max drawdown
-24.3%
VaR 95 (1d)
1.08%
ES 95 (1d)
1.72%
Skewness
-0.54
Excess kurtosis
12.98

Annualized where applicable. VaR and ES are one-day historical estimates at 95%. Uncertainty intervals appear in the next section.

Growth of $1 after modeled transaction costs. The horizontal axis uses calendar dates from the frozen price snapshot.
Drawdown from the portfolio’s running peak. Depth is encoded by position and a shaded area.
Rolling 63-observation annualized volatility.
Pre-binned daily returns with empirical VaR and ES thresholds. A matched Gaussian is overlaid only as a diagnostic reference.
View key chart data
Key values represented in the empirical risk charts
MeasureValue
Ending value of $1$2.750
Maximum drawdown-24.33%
VaR 95, one day1.08%
Expected Shortfall 95, one day1.72%
Backtest observations2,640

Sampling uncertainty

Bootstrap intervals for VaR and Expected Shortfall

The point estimates are recomputed across 2,000 moving-block bootstrap samples, using 21-day blocks and seed 20260803. Blocks preserve some local dependence. The percentile interval still describes uncertainty conditional on this observed history. It does not manufacture crises missing from that history.

Moving-block bootstrap uncertainty; 2,000 replications, 21-day blocks
MeasurePoint estimateBootstrap 95% interval
Historical VaR 951.079%0.930% – 1.219%
Historical ES 951.718%1.421% – 2.126%

Falsification test

The 95% historical VaR forecast is rejected

Each forecast uses only the preceding 500 observations. In the holdout sequence, breaches occur too often and cluster. Both Kupiec unconditional coverage and Christoffersen conditional coverage reject the model at 5%. A red result is useful. It marks the exact boundary of a method that looks plausible in a static histogram.12

Rolling 500-observation historical VaR forecast tests
TestStatisticp-valueDecision
Observed breaches135 / 21406.31% rateExpected 107
Kupiec coverage7.1480.00751Reject at 5%
Christoffersen independence13.3920.00025Reject at 5%
Conditional coverage20.5400.00003Reject at 5%

Historical stress

Three dated windows inside the sample

Stress windows are descriptive slices fixed by event dates, not optimized after inspecting portfolio minima. They show how the same allocation behaved through the Q4 2018 volatility shock, the COVID-19 selloff, and the 2022 inflation and rate shock.

Fixed historical stress windows within the frozen sample
WindowDatesCumulative returnMax drawdownWorst dayVaR breaches
Volatility shock of Q4 20182018-09-20 – 2018-12-24-10.2%-10.8%-1.86%16
COVID-19 selloff2020-02-19 – 2020-03-23-21.3%-21.7%-7.08%9
2022 inflation and rate shock2022-01-03 – 2022-10-14-24.2%-24.0%-3.37%26

Interpretation

Implications of coverage and independence rejection

VaR marks a quantile. Expected Shortfall averages the losses beyond it. Neither becomes a law of nature because it has three decimal places. The rejected coverage tests say that a rolling 500-day empirical distribution adapts too slowly to regime changes in this portfolio. They do not say the portfolio is uninvestable.

The next research iteration should compare filtered historical simulation and volatility- scaled forecasts on the same locked dates. The present model stays as the baseline. That comparison belongs in the ledger before any method is declared superior.

Reproduce

Code, commit, and references

Inspect the risk notebook at the source commit and the backtest engine used to construct the return stream.

  1. Kupiec, P. H. (1995), “Techniques for Verifying the Accuracy of Risk Measurement Models,” Journal of Derivatives 3(2), 73–84. doi:10.3905/jod.1995.407942
  2. Christoffersen, P. F. (1998), “Evaluating Interval Forecasts,” International Economic Review 39(4), 841–862. doi:10.2307/2527341
  3. Artzner, P., F. Delbaen, J.-M. Eber, and D. Heath (1999), “Coherent Measures of Risk,” Mathematical Finance 9(3), 203–228. doi:10.1111/1467-9965.00068