Skip to content
← Investigations

Market risk · Forecast calibrationEmpirical finding

Historical VaR under coverage and independence tests

Breaches arrived too often and in clusters; the forecast failed its calibration tests.

A historical 95% VaR forecast produced too many breaches. The breaches also clustered instead of arriving independently.

Research by Quantitative Markets & Institutions LabRevised 3 sources8 figuresLedger entry

Commit 4806df9Evidence ad24c49Data · 2,892 rowsUniverse 15 instrumentsTests 239 / 239 passed

  • 135breaches at 95% VaR
  • 107expected breaches at 5%
  • 6.31%observed breach rate
  • 2,140one-step forecasts
  • 2,640net daily returns after warm-up
  • 0.0075Kupiec unconditional-coverage p-value

The historical VaR forecast failed on both breach frequency and breach timing.

The rolling 500-observation historical 95% VaR model made 2,140 one-step forecasts. It recorded 135 breaches against 107 expected, a 6.31% breach rate. Kupiec coverage, Christoffersen independence, and conditional-coverage tests all rejected calibration at the 5% level. The misses came too often, and they came together.

Growth of $1 after modeled transaction costs

$ · daily · 2016-01-04 to 2026-07-06 · notebook 06

Drawdown from the running peak

% · daily · notebook 06

Rolling 63-day annualised volatility

% · daily, 63-observation window · notebook 06

Daily return distribution with a matched Gaussian

count · daily return (%) · 56 bins · notebook 06

Histogram of daily portfolio returns in percent with the expected counts of a Gaussian of the same mean and variance overlaid, and vertical rules at the historical VaR 95 and ES 95 thresholds.
daily return (%)observedGaussian, same mean and variance
-6.9619111.80868e-18
-6.7253204.02217e-17
-6.4887308.04016e-16
-6.2521411.44469e-14
-6.0155502.33342e-13
-5.7789603.3878e-12
-5.5423704.42129e-11
-5.3057805.18664e-10
-5.0691905.46928e-9
-4.832615.18418e-8
-4.5960114.4171e-7
-4.3594200.00000338299
-4.1228300.0000232901
-3.8862410.000144128
-3.6496520.000801737
-3.4130610.00400887
-3.1764700.0180185
-2.9398820.0727984
-2.7032920.264382
-2.466740.863074
-2.2301152.53262
-1.9935276.68037
-1.756931515.8393
-1.520342633.7582
-1.283754464.6738
-1.0471658111.374
-0.810565122172.404
-0.573974157239.892
-0.337384296300.048
-0.100793442337.344
0.135797560340.927
0.372387371309.711
0.608978214252.906
0.845568136185.638
1.0821680122.485
1.318753972.6447
1.555341738.7286
1.791931618.5595
2.0285277.99481
2.2651153.09568
2.501711.07748
2.7382900.33711
2.9748800.0948067
3.2114700.023967
3.4480610.00544621
3.6846500.00111245
3.9212400.000204256
4.1578300.0000337113
4.3944210.0000050013
4.6310116.66955e-7
4.867617.99497e-8
5.104208.61477e-9
5.3407918.34405e-10
5.5773807.26469e-11
5.8139705.68543e-12
6.0505613.9996e-13

Exception record · one mark per day beyond forecast VaR

1 = breach · 2,140 forecast days · notebook 06

Timeline strip of the 2,140 one-step forecasts marking each day the realised return fell below the rolling 500-observation historical VaR forecast; clusters of consecutive marks are visible in 2018, 2020 and 2022.
DateReturn (%)VaR 95 forecast (%)
2018-01-29-0.6149370.570531
2018-01-30-0.7074870.57505
2018-02-02-1.467770.570531
2018-02-05-2.277990.57505
2018-02-07-0.5981150.594871
2018-02-08-2.174560.597847
2018-02-27-1.026790.598137
2018-03-01-0.7049890.599059
2018-03-19-0.8536990.608902
2018-03-22-1.228120.615557
2018-03-23-1.152970.627784
2018-03-27-0.8461280.636859
2018-04-02-1.210730.642594
2018-04-06-0.9595610.644716
2018-04-20-0.7329580.664067
2018-04-24-0.699470.677527
2018-05-15-0.8336660.677527
2018-06-25-0.8985440.645417
2018-10-04-0.799580.636859
2018-10-10-1.863480.636859
2018-10-11-0.8367880.642594
2018-10-18-0.9282770.645417
2018-10-24-1.568350.678491
2018-10-26-0.9277470.699746
2018-11-12-1.074520.705114
2018-11-19-0.9758720.70876
2018-11-20-1.076420.733338
2018-12-04-1.68040.74157
2018-12-07-1.209550.761585
2018-12-14-0.9923470.761585
2018-12-17-1.108710.777232
2018-12-19-0.8267570.800103
2018-12-21-1.312150.810881
2018-12-24-1.274790.827103
2019-01-03-0.9085350.833822
2019-03-22-0.9577130.837255
2019-05-07-0.9791540.846507
2019-05-13-1.292990.855941
2019-08-05-1.406280.899044
2019-08-14-1.317690.909495
2019-08-23-1.020560.927774
2020-02-24-1.769020.899044
2020-02-25-1.768870.899044
2020-02-27-2.560510.909495
2020-03-05-1.454860.927774
2020-03-09-4.538570.929749
2020-03-11-3.817380.957806
2020-03-12-7.080210.960377
2020-03-16-6.332090.976036
2020-03-18-4.876130.979814
2020-03-27-1.278020.974693
2020-03-31-1.189670.976036
2020-04-01-2.912680.979814
2020-04-15-1.312140.993758
2020-04-20-1.058011.02326
2020-04-21-1.689381.05884
2020-04-30-1.07881.07461
2020-05-01-1.614311.07654
2020-05-12-1.276541.0803
2020-06-11-3.605571.11276
2020-06-24-1.549911.19066
2020-06-26-1.218731.21281
2020-09-03-2.04311.22153
2020-09-08-1.522741.27487
2020-09-23-1.619821.27661
2020-10-28-2.198651.27487
2021-01-27-1.511611.19112
2021-01-29-1.277741.22162
2021-02-25-2.080471.2766
2021-03-18-1.296561.27775
2021-05-12-1.648681.27775
2021-09-28-1.582531.2766
2022-01-05-1.49861.27775
2022-01-18-1.426641.27894
2022-02-03-1.618621.29734
2022-02-10-1.519831.31786
2022-03-07-1.899721.26015
2022-04-05-1.36741.11278
2022-04-11-1.171631.11278
2022-04-21-1.134771.11278
2022-04-22-1.665831.117
2022-04-26-1.619681.13588
2022-04-29-2.404461.13588
2022-05-05-2.819811.1577
2022-05-09-2.123971.1577
2022-05-18-2.043471.17399
2022-06-09-1.487581.17815
2022-06-10-1.868631.22075
2022-06-13-3.365321.26021
2022-06-16-1.783291.27868
2022-08-19-1.284521.27868
2022-08-22-1.510761.28512
2022-08-26-2.124181.3001
2022-09-13-2.917351.28512
2022-09-23-1.350081.28512
2022-09-29-1.509111.30093
2022-10-07-2.007441.35094
2022-10-14-1.74871.37036
2022-11-02-1.737161.37036
2022-12-05-1.536921.42968
2022-12-15-1.578671.48813
2023-02-21-1.664971.48813
2023-09-21-1.621711.42968
2024-02-13-1.548571.31861
2024-04-10-1.48661.29934
2024-04-30-1.337841.28495
2024-07-24-1.577161.08616
2024-08-05-1.786951.09381
2024-10-31-1.238260.944043
2024-12-18-2.455110.929984
2025-01-10-1.165360.927022
2025-02-27-1.118650.886173
2025-03-06-1.266680.886173
2025-03-10-1.476010.886173
2025-04-03-2.703550.888084
2025-04-04-3.576150.912106
2025-04-07-1.26230.915113
2025-04-08-1.253780.931241
2025-04-10-2.561930.956226
2025-04-21-1.284730.97381
2025-05-21-1.302940.986209
2025-10-10-1.271370.888084
2025-11-13-1.239660.855362
2025-11-20-0.8799540.876526
2026-01-20-1.084070.880298
2026-01-30-1.352690.888084
2026-02-12-0.8937090.888084
2026-03-03-1.186550.894623
2026-03-12-1.185290.912888
2026-03-18-1.290760.9345
2026-03-20-1.827291.02322
2026-03-26-1.538261.07502
2026-05-15-1.360961.02322
2026-06-05-2.059671.07502
2026-06-10-1.184561.0858

Rolling 500-observation historical VaR 95 forecast

% loss · one-day · notebook 06

The same portfolio, six risk numbers per confidence level

% of NAV · one day · notebook 06

Grouped columns of historical, parametric, Cornish–Fisher and Monte Carlo VaR plus historical and parametric ES at the 95 and 99 percent levels.
estimatorα = 95%α = 99%
historical VaR1.079061.91245
parametric (normal) VaR1.150991.64484
Cornish-Fisher VaR1.06944.04734
Monte Carlo VaR1.165611.65666
historical ES1.717843.02151
parametric ES1.45381.8904

Diversification decay under correlation stress

% annualised · λ from 0 (observed) to 1 (all correlations → 1) · notebook 06

Columns of annualised portfolio volatility as the correlation matrix is blended toward perfect correlation.
λportfolio volatility
011.6441
0.0511.8839
0.112.119
0.1512.3496
0.212.576
0.2512.7984
0.313.017
0.3513.232
0.413.4435
0.4513.6518
0.513.8569
0.5514.0591
0.614.2584
0.6514.4549
0.714.6488
0.7514.8402
0.815.0291
0.8515.2157
0.915.4
0.9515.5821
115.7622

Coverage and independence tests · rolling 500-observation historical VaR 95

TestStatisticp-valueVerdict
Kupiec unconditional coverage7.147820.0075055Reject at 5%
Christoffersen independence13.39170.00025274Reject at 5%
Conditional coverage20.53950.00003467Reject at 5%

135 breaches over 2,140 forecasts against 107 expected; likelihood-ratio statistics.

Breach transition counts

n00n01n10n11Expected n11 under independence
1,889115115208.51636

Exception run lengths

Run length (days)Runs
198
215
31
41

VaR and ES by estimator (% of NAV, one day)

Estimatorα = 95%α = 99%
historical VaR1.079061.91245
parametric (normal) VaR1.150991.64484
Cornish-Fisher VaR1.06944.04734
Monte Carlo VaR1.165611.65666
historical ES1.717843.02151
parametric ES1.45381.8904

Moving-block bootstrap uncertainty · 2,000 replications, 21-day blocks, seed 20260803

MeasurePoint estimate (%)Bootstrap 95% lower (%)Bootstrap 95% upper (%)
Historical VaR 951.079060.9301961.21882
Historical ES 951.717841.421212.12557

Fixed historical stress windows within the frozen sample

WindowStartEndCumulative returnMax drawdownWorst dayVaR breaches
Volatility shock of Q4 20182018-09-202018-12-24-0.102264-0.108447-0.018634816
COVID-19 selloff2020-02-192020-03-23-0.212531-0.217114-0.07080219
2022 inflation and rate shock2022-01-032022-10-14-0.241591-0.240258-0.033653226

Five deepest drawdown episodes

DepthStartTroughRecoveryDuration (days)
0.2433052021-12-272022-10-142024-03-21561
0.2171142020-02-192020-03-182020-07-1099
0.1126882018-08-292018-12-242019-03-15135
0.1095042025-02-192025-04-082025-05-1661
0.07201422026-02-252026-03-272026-04-1736

Replaying four historical episodes against today's weights

ScenarioPortfolio P&LAssets shocked
GFC-2008-0.3058
Covid-2020-0.2068
RateShock-2022-0.2478
DotCom-2000-0.239258

Target weights

ETFWeight
SPY0.3
QQQ0.1
IWM0.05
EFA0.1
AGG0.2
TLT0.1
GLD0.075
VNQ0.075

Summary card (rf 2%)

MetricValue
annualized_return0.101367
annualized_vol0.115035
sharpe_ratio0.723273
sortino_ratio1.00797
calmar_ratio0.416624
hit_rate0.568561
max_drawdown0.243305
skew-0.542907
kurtosis12.9753
VaR950.0107906
ES950.0171784
Sharpe (rf 3%)0.636342
Total cost (fraction of NAV)0.00182581

Audience and decision

Research significance

Risk owners, model validators, and control functions

Risk managers, treasury teams, model validators, and governance committees setting limits, capital buffers, or escalation thresholds require evidence that exceptions match the model's stated confidence level and arrive independently.

Decision context

Whether VaR is fit to govern the portfolio by itself

This evidence says no for this portfolio, model, and sample. VaR can remain one lens, but it needs recalibration or redesign and must sit beside expected shortfall, drawdown, scenario analysis, and correlation stress.

Three tests separate frequency error from clustered failure

The evidence comes from one forecast record. The portfolio was a strategic ETF book, rebalanced monthly and built without look-ahead. It paid 10 basis points of proportional turnover cost. After warm-up it produced 2,640 net daily returns and 2,140 forecasts. Each VaR estimate used only the preceding 500 observations.

The unconditional-coverage test rejected the stated 5% breach probability (Kupiec p = 0.0075). The independence test rejected non-clustered arrivals (Christoffersen p = 0.00025). The combined conditional-coverage test also rejected (p = 0.000035). All three belong in the report. The model missed too often, and its misses ran back to back.

The full-sample historical VaR95 point estimate was 1.079%. Its 95% moving-block bootstrap interval ran 0.930%–1.219%. ES95 was 1.718%, with an interval of 1.421%–2.126%. Those intervals measure estimator uncertainty. They do not rescue a rejected forecast process.

Backtest the forecast process, not just the latest risk number

  • Test unconditional coverage and exception independence separately. A correct average rate can hide clustering.
  • Treat an exception cluster as information about regime adaptation, not as unrelated surprises.
  • Report expected shortfall beside VaR so the severity beyond the threshold stays visible.
  • Add date-fixed scenarios and correlation stresses for risks a rolling empirical quantile may adapt to slowly.
  • Keep costs, timing, and no-look-ahead controls inside the backtest that produces the evidence.

Inspect the forecasts, tests, data version, and code

Clustered breaches establish calibration failure, not its cause

Lab measurement The rolling historical 95% VaR model produced 135 breaches against 107 expected across 2,140 forecasts. Coverage, independence, and conditional-coverage tests all rejected calibration. Its 500-observation window changes only as new returns enter it.

Institutional record The sample spans a period when inflation, policy rates, Treasury yields, equity valuations, and cross-asset correlation all changed together.

Interpretive synthesis An exception cluster should prompt a review of the return distribution, dependence structure, and volatility process. It should also prompt scenarios drawn from outside the window. Regime change is one possible explanation, not a state the backtest measures.

Causal boundary: the backtests identify frequency error and dependence among exceptions. They do not identify inflation, monetary policy, correlation reversal, or another macro factor as the cause.

Examine the calibration regime

Numbers

StatementValueAs statedNote
breaches at 95% VaR135135
expected breaches at 5%107107
observed breach rate6.31%6.31%
one-step forecasts2,1402,140
net daily returns after warm-up2,6402,640
Kupiec unconditional-coverage p-value0.00750.0075
Christoffersen independence p-value0.000250.00025
Conditional-coverage p-value0.0000350.000035
Historical VaR 95 (one day)1.079%1.079%
VaR 95 bootstrap interval · lower0.930%0.930%
VaR 95 bootstrap interval · upper1.219%1.219%
Historical ES 95 (one day)1.718%1.718%
ES 95 bootstrap interval · lower1.421%1.421%
ES 95 bootstrap interval · upper2.126%2.126%
breach-follows-breach days observed2020
breach-follows-breach days expected under independence≈8.5≈8.5
isolated single-day exception runs (from the breach dates)9899 unverifiedThe page's run structure was derived from the transition counts alone; the recorded breach dates give 98 run(s) of 1 day(s), 15 run(s) of 2 day(s), 1 run(s) of 3 day(s), 1 run(s) of 4 day(s).
two-day exception runs (from the breach dates)1512 unverifiedThe page's run structure was derived from the transition counts alone; the recorded breach dates give 98 run(s) of 1 day(s), 15 run(s) of 2 day(s), 1 run(s) of 3 day(s), 1 run(s) of 4 day(s).
three-day exception runs (from the breach dates)14 unverifiedThe page's run structure was derived from the transition counts alone; the recorded breach dates give 98 run(s) of 1 day(s), 15 run(s) of 2 day(s), 1 run(s) of 3 day(s), 1 run(s) of 4 day(s).
four-day exception runs (from the breach dates)1
Annualised return10.14%10.14%
Annualised volatility11.50%11.50%
Sharpe (rf 3%)0.640.64
Sharpe (rf 2%)0.72≈0.72
Maximum drawdown−24.3%-24.3%
Skewness of daily returns−0.54-0.54
Excess kurtosis of daily returns12.9812.98
Ending value of $1$2.750$2.750
Total modeled costs (% of NAV)0.183%0.183%
Hit rate (days with positive return)57%57%
Monthly rebalances127127
First net return2016-01-042016-01-04
Last net return2026-07-062026-07-06

3 statements the exporter could not reproduce to printed precision; each note says why.

Notes

Fat tails and the failure of normal VaR

5 min · Prerequisites: quantiles, moments, and basic risk metrics

Parametric-normal VaR takes a mean and standard deviation and reads the quantile off the Gaussian: $ \mathrm{VaR}\alpha = -( \mu + z{1-\alpha},\sigma )$. The procedure is exact when returns are normal and can materially underestimate tail risk when they are not, because the Gaussian density dies like ex2/2e^{-x^2/2} while empirical return distributions die like a power law, P[r>x]xα\mathbb{P}[|r| > x] \sim x^{-\alpha} with tail index α\alpha around 3 to 4 for daily equity returns.1 Every moment of the comparison fails in the same direction:

The counting argument. Daily equity moves of five standard deviations should occur, under normality, about once per 14,000 years. The realized record produces them every few years; October 19, 1987 was, on a Gaussian yardstick, roughly a 20σ event — a probability so small it has no physical interpretation. The model is not slightly wrong in the tail; it is wrong by factors of 10310^{3} to 105010^{50}, depending on the tail threshold.

The moment argument. Excess kurtosis of daily index returns is far above the Gaussian's zero. Since sample variance is dominated by the very observations the Gaussian deems nearly impossible, σ^\hat\sigma is inflated by past crises while the normal quantile formula simultaneously understates how much worse than zσ^z\hat\sigma the next crisis will be. The errors do not cancel; at high confidence levels the understatement wins.

The structural argument. Volatility clustering means returns are a mixture of distributions — calm-regime and stress-regime — and mixtures of normals with different variances are themselves fat-tailed. So even if each day were conditionally Gaussian, unconditional normal VaR would still be miscalibrated. Worse, the stress regime arrives with correlations lurching toward one, so the portfolio-level tail is fatter than any asset-level analysis suggests.

Available responses include historical-simulation VaR, as used on the risk dashboard; Expected Shortfall, which averages the tail beyond a quantile; and extreme-value methods that fit an asymptotic tail family such as the GPD or GEV. Estimates at very high confidence levels remain extrapolations beyond limited observed data.

Footnotes

  1. McNeil, A. J., R. Frey, and P. Embrechts (2015), Quantitative Risk Management: Concepts, Techniques and Tools , revised 2nd ed., Princeton University Press, chs. 2 and 5, ISBN 978-0-691-16627-8. publisher catalog

Limitations

  • The bootstrap conditions on one realized history, and the 21-day block length is judgmental.
  • Historical quantiles can adapt slowly after regime changes. Other windows and weightings may behave differently.
  • The portfolio holds fixed strategic weights and rebalances monthly. Dynamic de-risking was not modeled.
  • Close-to-close ETF prices omit intraday liquidity and execution stress.
  • Historical scenarios replay known crises and do not bound future loss paths.

Sources

  1. 1Kupiec, P. H. (1995), “Techniques for Verifying the Accuracy of Risk Measurement Models,” Journal of Derivatives 3(2), 73–84. doi:10.3905/jod.1995.407942
  2. 2Christoffersen, P. F. (1998), “Evaluating Interval Forecasts,” International Economic Review 39(4), 841–862. doi:10.2307/2527341
  3. 3Artzner, P., F. Delbaen, J.-M. Eber, and D. Heath (1999), “Coherent Measures of Risk,” Mathematical Finance 9(3), 203–228. doi:10.1111/1467-9965.00068

Cite this

Wisniewski, K. (2026, August 8). Historical VaR under coverage and independence tests. Quantitative Markets & Institutions Lab. https://www.kylewisniewski.com/lab/var-backtest

@misc{wisniewski2026var,
  author = {Wisniewski, Kyle},
  title = {Historical VaR under coverage and independence tests},
  year = {2026},
  month = {aug},
  howpublished = {\url{https://www.kylewisniewski.com/lab/var-backtest}},
  note = {Empirical finding · Quantitative Markets & Institutions Lab · commit 4806df9}
}