Questions before conclusions. Assumptions beside results. Failures preserved
rather than edited away. This ledger records what the lab has tested, what the evidence says,
how uncertain the answer remains, and the exact artifact from which it can be reproduced.
The rule
A result that cannot be reproduced from a clean clone is an anecdote.
Proposed question registered; test not begunRunning evidence being producedReplicated result survived an independent or theoretical checkRejected stated claim failed its criterionRevised evidence required a narrower claim
Investigation 01 · Numerical validation
Four roads to one price
Replicated
Research question
Do four independent implementations of European option pricing agree within their distinct, pre-declared numerical error budgets?
Testable claim
Black–Scholes, a CRR tree, a Crank–Nicolson PDE solver, and Monte Carlo will converge to the same arbitrage-free value; put–call parity will hold for every method.
Assumptions
GBM under the risk-neutral measure; constant rate and volatility; frictionless markets; European exercise; identical contract inputs across methods.
Dataset / version / seed
Synthetic benchmark contracts with S₀ = 100, r = 5%, σ = 20%, T = 1 year, K ∈ {90, 100, 110}; 200,000 antithetic paths; PCG64 seed 42.
Method
Cross-method price comparison, put–call parity, CRR convergence against a 1/n guide, PDE error checks, Monte Carlo standardized error, and implied-volatility round trips.
Falsification criterion
A deterministic method exceeds its discretization tolerance; Monte Carlo differs materially beyond its reported standard error; parity fails; or implied volatility does not round-trip to solver tolerance.
Result and uncertainty
The tree and PDE land within roughly 10⁻³ of closed form; Monte Carlo prices are statistically consistent with it; implied volatility returns to about 10⁻⁸. Monte Carlo uncertainty is reported as estimator standard error, while tree and PDE errors remain discretization-dependent.
Failure modes
Agreement validates numerics, not GBM. Exact terminal sampling flatters Monte Carlo. Barrier, digital, and American contracts stress different weaknesses. Model Greeks do not measure misspecified-dynamics hedging error.
How does portfolio construction change when classical return optimization gives way to uncertainty-aware risk allocation?
Testable claim
Methods that reduce or remove dependence on estimated returns will produce more stable out-of-sample risk profiles than unconstrained maximum-Sharpe optimization.
Assumptions
Long-only, fully invested allocations; liquid ETF universe selected in 2026; historical moments stand in for future moments; identical raw covariance inputs; frozen weights during evaluation.
Dataset / period / version / seed
Frozen adjusted-close snapshot for 15 ETFs, 2,891 daily observations from 2015-01-05 to 2026-07-06. Fit: 2015–2021. Sealed evaluation: 2022–2026. Seed 42.
Method
Compare max-Sharpe MVO, box-uncertainty robust MVO, equal-risk-contribution risk parity, and hierarchical risk parity in sample and on one untouched out-of-sample window.
Falsification criterion
The claim fails if return-optimized weights retain their in-sample advantage without greater concentration or degradation, or if risk-based portfolios fail to reduce realized volatility and drawdown out of sample.
Result and uncertainty
MVO Sharpe fell from about 1.12 in sample to 0.44 out of sample. Risk-based portfolios ran at roughly half the volatility and two-thirds the drawdown of optimized books, but did not reliably win on return. Sharpe differences of ±0.2 are well inside an estimated sampling error near 0.5 for this 4.5-year window. The broad claim was narrowed: risk-based methods delivered risk shape, not superior return.
Failure modes
One evaluation window; survivorship-tilted universe; no rolling-origin test; frozen weights; raw rather than shrunk covariance; robust MVO preserved a concentrated ranking and benefited from a specific mega-cap regime.
Are the constant-volatility and Gaussian-return fingerprints required by Black–Scholes compatible with observed equity-index behavior?
Testable claim
If constant-volatility GBM is an adequate description, realized volatility should fluctuate around a stable level, return magnitudes should not remain autocorrelated, standardized returns should be approximately Gaussian, and one implied volatility should price every strike.
Assumptions
SPY is used as an equity-index proxy; adjusted daily closes represent the observable return process; Heston parameters are illustrative rather than calibrated to a dated option chain.
Dataset / period / version / seed
Frozen SPY adjusted-close snapshot, approximately 2,900 daily observations from 2015-01 to 2026-07. Seed 42 for simulated Heston paths.
Method
Measure rolling realized volatility, skew, kurtosis, extreme standardized moves, and return/absolute-return autocorrelation; then compare flat Black–Scholes volatility with Heston-generated strike and maturity skews.
Falsification criterion
Materially unstable realized volatility, persistent magnitude autocorrelation, fat or asymmetric tails, or a strike-dependent implied-volatility curve rejects the stated GBM fingerprints.
Result and uncertainty
The claim was rejected. Realized volatility ranged from roughly 3% to 93%; excess kurtosis was about 14; the worst day was near 10σ under a Gaussian yardstick; lag-1 absolute-return autocorrelation was about 0.35 and persisted. An equity-like Heston specification generated a short-dated skew near 25% to 16%. The empirical rejection is sample-specific; the Heston surface demonstrates mechanism, not a live market calibration.
Failure modes
One underlying and one historical window; close-to-close returns omit intraday structure; parameters treated as constants still drift in practice; Heston omits jumps and rough volatility; simulation discretization carries bias.
Do the simulation library's errors behave as theory requires, and when does Monte Carlo earn its computational cost?
Testable claim
Exact schemes are free of time-step bias at observation points; Euler bias shrinks with Δt; independent-replication RMSE decays as N⁻¹ᐟ²; and antithetic sampling lowers estimator variance for monotone payoffs.
Assumptions
Specified stochastic processes are the data-generating truth; PCG64 pseudorandomness is adequate; payoff and scheme comparisons use common random numbers where required.
Benchmark exact and Euler GBM against closed form, estimate antithetic variance reduction across independent replications, test OU/CIR moments and boundaries, fit the Monte Carlo log–log convergence slope, and price a discretely monitored barrier.
Falsification criterion
Persistent bias from an exact transition, an Euler error that does not shrink with step size, a fitted RMSE slope materially different from −0.5, or confidence intervals that systematically miss known benchmarks.
Result and uncertainty
The N⁻¹ᐟ² law was recovered within replication noise; exact GBM remained unbiased across step counts; Euler bias contracted with Δt; antithetic pairing cut empirical standard error by about 30% at the same path budget. Every Monte Carlo estimate retains its reported sampling error.
Failure modes
CIR and Heston discretization bias was not fully budgeted by step-halving; regime parameters were chosen rather than estimated; antithetics are weak for some payoff shapes; all results share one pseudorandom generator family.
Do widely held style ETFs load on their advertised factors, and does statistically credible alpha remain after those exposures are controlled?
Testable claim
Each branded ETF should have a positive, economically meaningful loading on its named factor; most apparent excess return should disappear after six-factor attribution.
Assumptions
Linear exposures over each estimation window; the Fama–French five factors plus momentum span the relevant systematic effects; adjusted-close ETF returns approximate investor experience; Newey–West HAC(5) is sufficient for short-range residual dependence.
Six-factor time-series OLS for MTUM, VLUE, QUAL, IWM, USMV, and QQQ with Newey–West HAC(5) intervals; 252-observation rolling QQQ betas; chronological 50/50 frozen-coefficient holdout; Holm and Benjamini–Hochberg correction across six alpha tests.
Falsification criterion
An advertised factor loading is absent, incorrectly signed, or economically negligible; or alpha remains broad, stable, and statistically credible after exposure and multiple-testing considerations.
Result and uncertainty
The advertised exposures were recovered. QQQ's +3.29% annualized alpha has a raw HAC p-value of 0.024, but its Benjamini–Hochberg q-value is 0.147; no alpha survives the 5% false-discovery threshold. Frozen-coefficient holdout R² decays by 0.4 to 29.6 percentage points across ETFs, with USMV showing the largest decay. Confidence intervals and adjusted p-values are published on the factor page.
Failure modes
The factor menu defines the alpha left behind; factor portfolios are not costless tradables; eleven years is one regime; the 50/50 split is only one holdout path; HAC lag choice and vendor price adjustments can move inference.
Does a rolling 500-observation historical 95% VaR forecast achieve correct unconditional coverage and independent breaches for the cost-aware multi-asset portfolio?
Testable claim
If the model is calibrated, the breach rate should be statistically compatible with 5% and breach arrivals should not cluster.
Assumptions
Fixed strategic weights; monthly rebalancing; 10bp proportional costs; frozen liquid-ETF universe; historical shocks are informative but not exhaustive; static correlation stress preserves asset volatilities.
Dataset / period / version / seed
Frozen adjusted-close snapshot through 2026-07-06, SHA-256 5dc8433e1c99…; 2,640 net backtest returns after warm-up and 2,140 one-step VaR forecasts. Moving-block bootstrap seed 20260803.
Method
Look-ahead-free monthly backtest with 10bp turnover costs; one-day historical VaR/ES; 2,000-replication 21-day moving-block bootstrap; rolling 500-observation VaR; Kupiec coverage and Christoffersen independence/conditional-coverage tests; three date-fixed stress windows.
Falsification criterion
Reject calibration if Kupiec or Christoffersen conditional-coverage p-value is below 0.05; report independence separately to distinguish frequency error from clustered failures.
Result and uncertainty
Rejected. The model records 135 breaches versus 107 expected (6.31%); Kupiec p = 0.0075, independence p = 0.00025, and conditional-coverage p = 0.000035. VaR95 is 1.079% with a block-bootstrap 95% interval of 0.930%–1.219%; ES95 is 1.718% with interval 1.421%–2.126%. The failure is both excess breach frequency and clustering.
Failure modes
The bootstrap conditions on one realized history; block length is judgmental; empirical quantiles adapt slowly after regime changes; historical scenarios replay known crises; the fixed portfolio and close-to-close prices omit intraday liquidity and dynamic de-risking.
Every entry above is pinned to the commit that produced its executed notebook. Clone that
revision, install the project, run the tests, and execute the selected notebook against the
committed snapshots. Refreshing market data is deliberately a separate, opt-in operation
because a revised dataset is a new experiment.