Decision brief · Portfolio construction
Why “optimal” portfolios fail out of sample
In this frozen-window study, the maximum-Sharpe portfolio’s performance declined when fixed weights were evaluated on later data.
Depth 1 · Answer
Risk-based methods delivered a more stable risk shape—not a reliable return advantage.
Maximum-Sharpe mean–variance optimization fell from a Sharpe ratio of about 1.12 in the 2015–2021 fit period to 0.44 in the sealed 2022–2026 evaluation. Risk parity and hierarchical risk parity ran at roughly half the volatility and two-thirds the drawdown of the optimized books, but they did not reliably win on return. That narrower claim is what the evidence supports.
Published 8 August 2026 · Underlying investigation last updated 6 July 2026
- MVO Sharpe · fit
- 1.12
- MVO Sharpe · evaluation
- 0.44
- Fit window
- 2015–2021
- Sealed evaluation
- 2022–2026
Why it matters
Allocators, risk committees, and model reviewers
The result matters to anyone using an optimizer to convert noisy expected returns and covariances into precise weights—and to anyone evaluating whether an analytical system distinguishes numerical sophistication from decision reliability.
Decision affected
How much authority to give an optimized allocation
Do not interpret an in-sample efficient frontier or Sharpe ranking as a forecast. Decide first whether the mandate is return maximization, concentration control, or a stable risk budget, then evaluate the construction method against that stated objective.
Depth 2 · Evidence
The ranking changed when the estimates left the fit window
The experiment used 2,891 daily observations for 15 liquid ETFs from a frozen adjusted-close snapshot. Maximum-Sharpe MVO, box-uncertainty robust MVO, equal-risk-contribution risk parity, and hierarchical risk parity were fit on 2015–2021 data. Their weights were then frozen and applied to one untouched 2022–2026 window.
The robust portfolio posted the best out-of-sample Sharpe, about 0.6, but it remained concentrated in QQQ and benefited from the 2023–2025 mega-cap rally after drawing down more than 30% in 2022. The box uncertainty set reduced estimated return levels without changing their ranking, so it did not solve concentration. One 4.5-year window cannot distinguish a robust method from a fortunate regime.
Differences of ±0.2 in Sharpe are well inside an estimated sampling error near 0.5 for this evaluation window. The defensible conclusion is about realized risk shape, not which method will earn the highest future return.
Depth 3 · Application
Treat optimization as an assumption-sensitive proposal
- Separate the return objective from the risk-budget objective before comparing methods.
- Report concentration, risk contribution, drawdown, and out-of-sample behavior beside the in-sample optimum.
- Make the uncertainty treatment explicit: shrinking return levels is not the same as changing uncertain rankings.
- Use multiple rolling-origin evaluations before making comparative performance claims; this study did not perform them.
- State each method's regime dependency. Risk-based allocations can embed a bond and stock–bond-correlation bet.
Depth 4 · Limits
What this result does not establish
- There is one sealed evaluation window, not a distribution of rolling-origin outcomes.
- The 15-ETF universe was selected in 2026 and is therefore survivorship-tilted.
- Weights were frozen; rebalancing, transaction costs, and drift were outside this experiment.
- All methods received the same raw covariance estimate. Covariance shrinkage was available but deliberately not used.
- The study does not prove risk parity or HRP will outperform, nor that MVO will always underperform.
Depth 5 · Method and code
Move from the implication to the implementation
Related research