Module 03 · Empirical risk analytics

A risk model that can fail in public

A monthly rebalanced multi-asset portfolio, built from frozen adjusted-close snapshots with transaction costs. Historical VaR and Expected Shortfall are shown with bootstrap uncertainty, rolling forecast tests, and named market stress windows.

Revised · research record

Data and design

A leak-free, version-pinned backtest

Status
Loading frozen research artifact…

Target weights: 30% SPY, 10% QQQ, 5% IWM, 10% EFA, 20% AGG, 10% TLT, 7.5% GLD, and 7.5% VNQ. Rebalanced monthly after a 252-day warm-up; 10bp is charged on one-sided turnover. No future observation enters a rebalance decision.

Observed path

Portfolio risk results

Annualized where applicable. VaR and ES are one-day historical estimates at 95%; uncertainty intervals appear in the next section.

Growth of $1 after modeled transaction costs. The horizontal axis uses calendar dates from the frozen price snapshot.
Drawdown from the portfolio’s running peak. Depth is encoded by position and a shaded area.
Rolling 63-observation annualized volatility.
Pre-binned daily returns with empirical VaR and ES thresholds; a matched Gaussian is overlaid only as a diagnostic reference.
View key chart data

Sampling uncertainty

Tail estimates are intervals, not talismans

The point estimates are recomputed across 2,000 moving-block bootstrap samples, using 21-day blocks and seed 20260803. Blocks preserve some local dependence; the percentile interval still describes uncertainty conditional on this observed history and does not manufacture crises missing from it.

Falsification test

The 95% historical VaR forecast is rejected

Each forecast uses only the preceding 500 observations. In the holdout sequence, breaches occur too often and cluster. Both Kupiec unconditional coverage and Christoffersen conditional coverage reject the model at 5%. A red result is useful: it marks the exact boundary of a method that looks plausible in a static histogram.12

Historical stress

Three dated windows inside the sample

Stress windows are descriptive slices fixed by event dates, not optimized after inspecting portfolio minima. They show how the same allocation behaved through the Q4 2018 volatility shock, the COVID-19 selloff, and the 2022 inflation and rate shock.

Interpretation

What the rejection means

VaR marks a quantile; Expected Shortfall averages the losses beyond it. Neither becomes a law of nature because it has three decimal places. The rejected coverage tests say that a rolling 500-day empirical distribution does not adapt fast enough to regime changes in this portfolio. They do not say that the portfolio is uninvestable or that every alternative model is better.

The next research iteration should compare filtered historical simulation and volatility- scaled forecasts on the same locked dates, preserving the present model as a baseline. That comparison belongs in the ledger before any method is declared superior.

Reproduce

Code, commit, and references

Inspect the risk notebook at the source commit and the backtest engine used to construct the return stream.

  1. Kupiec, P. H. (1995), “Techniques for Verifying the Accuracy of Risk Measurement Models,” Journal of Derivatives 3(2), 73–84. doi:10.3905/jod.1995.407942
  2. Christoffersen, P. F. (1998), “Evaluating Interval Forecasts,” International Economic Review 39(4), 841–862. doi:10.2307/2527341
  3. Artzner, P., F. Delbaen, J.-M. Eber, and D. Heath (1999), “Coherent Measures of Risk,” Mathematical Finance 9(3), 203–228. doi:10.1111/1467-9965.00068