Module 03 · Empirical risk analytics
A risk model that can fail in public
A monthly rebalanced multi-asset portfolio, built from frozen adjusted-close snapshots with transaction costs. Historical VaR and Expected Shortfall are shown with bootstrap uncertainty, rolling forecast tests, and named market stress windows.
Revised · research record
Data and design
A leak-free, version-pinned backtest
- Status
- Loading frozen research artifact…
Target weights: 30% SPY, 10% QQQ, 5% IWM, 10% EFA, 20% AGG, 10% TLT, 7.5% GLD, and 7.5% VNQ. Rebalanced monthly after a 252-day warm-up; 10bp is charged on one-sided turnover. No future observation enters a rebalance decision.
Observed path
Portfolio risk results
Annualized where applicable. VaR and ES are one-day historical estimates at 95%; uncertainty intervals appear in the next section.
View key chart data
Sampling uncertainty
Tail estimates are intervals, not talismans
The point estimates are recomputed across 2,000 moving-block bootstrap
samples, using 21-day blocks and seed 20260803. Blocks preserve some local dependence;
the percentile interval still describes uncertainty conditional on this observed history and
does not manufacture crises missing from it.
Falsification test
The 95% historical VaR forecast is rejected
Each forecast uses only the preceding 500 observations. In the holdout sequence, breaches occur too often and cluster. Both Kupiec unconditional coverage and Christoffersen conditional coverage reject the model at 5%. A red result is useful: it marks the exact boundary of a method that looks plausible in a static histogram.12
Historical stress
Three dated windows inside the sample
Stress windows are descriptive slices fixed by event dates, not optimized after inspecting portfolio minima. They show how the same allocation behaved through the Q4 2018 volatility shock, the COVID-19 selloff, and the 2022 inflation and rate shock.
Interpretation
What the rejection means
VaR marks a quantile; Expected Shortfall averages the losses beyond it. Neither becomes a law of nature because it has three decimal places. The rejected coverage tests say that a rolling 500-day empirical distribution does not adapt fast enough to regime changes in this portfolio. They do not say that the portfolio is uninvestable or that every alternative model is better.
The next research iteration should compare filtered historical simulation and volatility- scaled forecasts on the same locked dates, preserving the present model as a baseline. That comparison belongs in the ledger before any method is declared superior.
Reproduce
Code, commit, and references
Inspect the risk notebook at the source commit and the backtest engine used to construct the return stream.
- Kupiec, P. H. (1995), “Techniques for Verifying the Accuracy of Risk Measurement Models,” Journal of Derivatives 3(2), 73–84. doi:10.3905/jod.1995.407942 ↩
- Christoffersen, P. F. (1998), “Evaluating Interval Forecasts,” International Economic Review 39(4), 841–862. doi:10.2307/2527341 ↩
- Artzner, P., F. Delbaen, J.-M. Eber, and D. Heath (1999), “Coherent Measures of Risk,” Mathematical Finance 9(3), 203–228. doi:10.1111/1467-9965.00068