Skip to content
← Investigations

Portfolio construction · Estimation riskEmpirical finding

Portfolio optimization under estimation error and out-of-sample evaluation

The fitted max-Sharpe portfolio weakened materially on later data.

One maximum-Sharpe portfolio was fitted, frozen, and then run on later data. Its measured performance fell.

Research by Quantitative Markets & Institutions LabRevised 5 figuresLedger entry

Commit 4806df9Evidence ad24c49Data · 2,892 rowsUniverse 15 instrumentsTests 239 / 239 passed

  • 1.12MVO Sharpe · 2015–2021 fit
  • 0.44MVO Sharpe · sealed 2022–2026 evaluation
  • 0.68Gap between the two published Sharpe ratios
  • ≈0.5Estimated sampling error of a Sharpe ratio over this window
  • ≈0.6Robust MVO Sharpe · evaluation (best of the four)
  • 34.3%Robust MVO maximum drawdown · evaluation window

Risk-based methods delivered a more stable risk shape—not a reliable return advantage.

Maximum-Sharpe mean–variance optimization scored a Sharpe ratio near 1.12 in the 2015–2021 fit period. The same frozen weights scored 0.44 in the sealed 2022–2026 evaluation. Risk parity and hierarchical risk parity ran at roughly half the volatility of the optimized books. Their drawdowns were about two-thirds as deep. Neither reliably won on return. The evidence supports that narrower claim.

In-sample efficient frontier, 2015–2021 estimates

% annualised · in-sample moments · notebook 02

Scatter of annualised volatility against expected return: the constrained efficient frontier, the fifteen ETFs, and the four construction methods placed by their in-sample moments.
Pointannualised volatility (%)annualised expected return (%)
frontier4.133183.30723
frontier4.188383.96457
frontier4.335734.62191
frontier4.567485.27925
frontier4.871475.9366
frontier5.243896.59394
frontier5.664087.25128
frontier6.105087.90862
frontier6.562538.56596
frontier7.033229.2233
frontier7.514669.88064
frontier8.0049110.538
frontier8.5024511.1953
frontier9.0060711.8527
frontier9.514812.51
frontier10.027913.1674
frontier10.544613.8247
frontier11.098214.482
frontier11.71515.1394
frontier12.385515.7967
frontier13.101516.4541
frontier13.855917.1114
frontier14.645117.7687
frontier15.471418.4261
frontier16.330119.0834
frontier17.216219.7408
frontier18.125820.3981
frontier19.055521.0555
frontier20.002521.7128
frontier20.964422.3701
SPY17.674415.4237
QQQ20.964422.3701
IWM22.182612.7561
EFA17.48048.17667
EEM21.5067.73937
AGG4.337162.94934
TLT14.08825.3667
LQD7.991394.88487
GLD13.92156.75544
DBC17.08653.74975
VNQ20.90911.1479
USMV15.233113.007
MTUM20.136917.2725
VLUE20.162711.8096
QUAL17.759315.4057
MVO max Sharpe11.076414.4575
Robust MVO18.940720.975
Risk parity7.539477.69801
HRP5.330265.19272

Sharpe ratio by construction method · fit window and sealed evaluation

Sharpe (rf 2%) · 2015–2021 vs 2022–2026 · notebook 02

Dots comparing each method's in-sample Sharpe ratio with its out-of-sample Sharpe ratio; every method degrades and the ranking reshuffles.
construction methodin-sample 2015–2021out-of-sample 2022–2026
MVO max Sharpe1.124690.435431
Robust MVO1.001810.586694
Risk parity0.7557570.355055
HRP0.598980.120618

Out-of-sample equity curves, frozen 2021 weights

growth of $1 · 2022-01 = 1 · daily · notebook 02

Weights fitted on 2015–2021

weight (%) · long-only, fully invested · notebook 02

Grouped columns of the portfolio weight each method assigns to each of the fifteen ETFs.
ETFMVO max SharpeRobust MVORisk parityHRP
SPY01.53083e-143.680741.93356
QQQ52.516391.79473.312970.833922
IWM003.220990.803273
EFA1.0623e-1503.932581.29356
EEM02.54178e-143.300991.20391
AGG2.45467e-143.14652e-1521.680953.6186
TLT35.87118.205317.36316.41027
LQD9.27459e-1509.923615.7936
GLD11.61264.0381e-149.022876.56466
DBC02.14586e-146.328472.99082
VNQ7.09841e-154.16991e-153.34281.27363
USMV2.09683e-147.32338e-154.251272.8401
MTUM2.00786e-141.09009e-133.344850.903865
VLUE01.04656e-143.576171.62112
QUAL3.65742e-145.13112e-143.717671.91513

Fractional risk contributions by construction method

fraction of portfolio variance (%) · in-sample covariance · notebook 02

Grouped columns of each asset's fractional contribution to portfolio variance under the four methods, against the one-over-N line.
ETFMVO max SharpeRobust MVORisk parityHRP
SPY01.30349e-146.666673.10052
QQQ85.8728101.4296.666671.48136
IWM006.666671.44996
EFA1.0174e-1506.666672.00767
EEM02.15787e-146.666672.24447
AGG3.91655e-153.87382e-176.6666735.8623
TLT9.6943-1.429116.666678.47121
LQD3.21556e-1506.6666719.8684
GLD4.432891.07574e-156.666678.35406
DBC06.50499e-156.666672.89091
VNQ7.20969e-152.74168e-156.666672.71972
USMV1.98284e-144.75443e-156.666674.30471
MTUM2.9375e-141.07247e-136.666671.64508
VLUE08.53807e-156.666672.54444
QUAL4.25797e-144.29262e-146.666673.05526

Performance by method · in-sample 2015–2021 and out-of-sample 2022–2026

WindowMethodAnnualised returnAnnualised volatilitySharpe (rf 2%)SortinoMax drawdownHit rateSkewExcess kurtosisVaR 95 (1d)ES 95 (1d)
in-sample 2015–2021MVO max Sharpe0.148430.1107641.124691.591420.140470.575482-0.4715225.767090.01101110.0168343
in-sample 2015–2021Robust MVO0.2112780.1894071.001811.405160.2553420.573212-0.5174368.993310.01888160.029439
in-sample 2015–2021Risk parity0.07693040.07539470.7557571.019170.1670890.566969-1.5466723.58240.005918840.0109146
in-sample 2015–2021HRP0.05179140.05330260.598980.7982160.1303880.553348-2.3581243.73990.00418460.00744273
out-of-sample 2022–2026MVO max Sharpe0.07577480.1463840.4354310.6277390.3038080.5172720.2399924.349380.01560760.0203249
out-of-sample 2022–2026Robust MVO0.1315910.2165270.5866940.8436570.3433730.5456160.2026694.965440.02220370.0307724
out-of-sample 2022–2026Risk parity0.05025760.09429310.3550550.5051060.2024050.5376440.125193.562730.009640460.0131716
out-of-sample 2022–2026HRP0.02649660.07319570.1206180.1698610.1797640.5270150.06767242.348740.007492750.0101249

Returns, volatility, drawdown and VaR/ES are decimals (0.146 = 14.6%).

Portfolio weights fitted on 2015–2021

ETFMVO max SharpeRobust MVORisk parityHRP
SPY01.53083e-160.03680740.0193356
QQQ0.5251630.9179470.03312970.00833922
IWM000.03220990.00803273
EFA1.0623e-1700.03932580.0129356
EEM02.54178e-160.03300990.0120391
AGG2.45467e-163.14652e-170.2168090.536186
TLT0.3587110.0820530.1736310.0641027
LQD9.27459e-1700.0992360.157936
GLD0.1161264.0381e-160.09022870.0656466
DBC02.14586e-160.06328470.0299082
VNQ7.09841e-174.16991e-170.0334280.0127363
USMV2.09683e-167.32338e-170.04251270.028401
MTUM2.00786e-161.09009e-150.03344850.00903865
VLUE01.04656e-160.03576170.0162112
QUAL3.65742e-165.13112e-160.03717670.0191513

Effective number of holdings (1 / Σ w²)

MethodEffective N
MVO max Sharpe2.3926
Robust MVO1.17736
Risk parity8.93028
HRP3.08402

Fractional risk contributions

ETFMVO max SharpeRobust MVORisk parityHRP
SPY01.30349e-160.06666670.0310052
QQQ0.8587281.014290.06666670.0148136
IWM000.06666670.0144996
EFA1.0174e-1700.06666670.0200767
EEM02.15787e-160.06666670.0224447
AGG3.91655e-173.87382e-190.06666670.358623
TLT0.096943-0.01429110.06666670.0847121
LQD3.21556e-1700.06666670.198684
GLD0.04432891.07574e-170.06666670.0835406
DBC06.50499e-170.06666670.0289091
VNQ7.20969e-172.74168e-170.06666670.0271972
USMV1.98284e-164.75443e-170.06666670.0430471
MTUM2.9375e-161.07247e-150.06666670.0164508
VLUE08.53807e-170.06666670.0254444
QUAL4.25797e-164.29262e-160.06666670.0305526

In-sample moments and the standard error of the annualised mean

ETFAnnualised meanAnnualised volSE(mean)mean / SE
QQQ0.2237010.2096440.07928292.82156
MTUM0.1727250.2013690.07615362.26812
SPY0.1542370.1767440.0668412.30753
QUAL0.1540570.1775930.06716182.29382
USMV0.130070.1523310.05760852.25783
IWM0.1275610.2218260.08389011.52058
VLUE0.1180960.2016270.07625111.54878
VNQ0.1114790.209090.07907361.40981
EFA0.08176670.1748040.06610721.23688
EEM0.07739370.215060.0813310.951589
GLD0.06755440.1392150.05264821.28313
TLT0.0536670.1408820.05327841.00729
LQD0.04884870.07991390.03022181.61634
DBC0.03749750.1708650.06461750.580299
AGG0.02949340.04337160.01640221.79814

SE(mean) = σ / √Y with Y = 6.99 years of in-sample data.

Efficient frontier points (in-sample)

Expected returnVolatilitySharpeSPYQQQIWMEFAEEMAGGTLTLQDGLDDBCVNQUSMVMTUMVLUEQUAL
0.03307230.04133180.31627707.62127e-1801.38194e-1700.92067905.38686e-173.23357e-170.04279453.3169e-181.68322e-1700.03652633.11706e-17
0.03964570.04188380.4690531.99685e-170.0342551.99204e-171.03127e-165.31432e-180.899798.24102e-1703.06954e-170.036120900.0068190900.009514460.0135009
0.04621910.04335730.6047221.79929e-170.076229804.42227e-184.80605e-180.876337000.005809980.02685648.09248e-190.014766509.37688e-181.7372e-18
0.05279250.04567480.7179571.15967e-170.11259101.62894e-1700.84695406.76723e-170.02301010.01293345.26401e-180.00451165000
0.0593660.04871470.8080913.11439e-170.145965.79581e-1807.78612e-180.8139472.29964e-1700.040093502.22739e-1701.27597e-1701.26165e-17
0.06593940.05243890.8760544.89644e-190.1766988.85812e-193.82443e-188.65296e-190.7673454.89644e-1900.05595713.743e-1806.96303e-181.12603e-181.37278e-192.75464e-18
0.07251280.05664080.9271191.82879e-170.205078001.58887e-170.6999960.0303500.064576102.19012e-180003.57231e-18
0.07908620.06105080.967826.22297e-180.2331561.26022e-1708.34505e-180.6299850.06459224.72905e-170.07226618.1614e-1803.95984e-183.14423e-1704.70799e-18
0.08565960.06562531.0005200.2612345.16615e-181.31906e-1800.5599750.09883431.74271e-170.07995603.07915e-1801.5013e-184.82165e-180
0.0922330.07033221.027033.43798e-170.289312001.27796e-180.4899650.1330771.32892e-170.08764591.27555e-1703.17287e-17001.33442e-17
0.09880640.07514661.048700.317395.12567e-181.16016e-177.40095e-180.4199550.1673192.64453e-170.09533622.62586e-1700000
0.105380.08004911.0665900.3454683.81848e-1801.00199e-170.3499450.20156100.103026002.56967e-1703.01285e-180
0.1119530.08502451.081499.25257e-180.3735472.89161e-184.2051e-1800.2799350.2358031.00558e-170.110716001.13037e-1707.05339e-190
0.1185270.09006071.0941.74559e-170.4016253.86594e-182.9926e-171.40128e-170.2099250.27004500.1184063.18203e-175.25165e-183.0292e-1701.55199e-179.50887e-18
0.12510.0951481.10469.8004e-180.4297034.28833e-182.02625e-173.5074e-170.1399140.3042873.56359e-170.1260968.0533e-1801.4644e-1703.5556e-191.91112e-17
0.1316740.1002791.1136300.45778107.03971e-186.86282e-180.06990420.3385300.1337861.98883e-172.01509e-184.23491e-1705.70509e-180
0.1382470.1054461.121391.42237e-190.4858781.52971e-267.19843e-182.32138e-251.5545e-170.3726988.47923e-260.1414247.20028e-254.49001e-265.68295e-244.59557e-236.16784e-193.89918e-17
0.144820.1109821.1246900.5266838.15763e-182.75761e-172.82521e-173.00254e-170.3581700.1151472.78027e-1700009.90611e-18
0.1513940.117151.1215900.5674892.85192e-170000.3436429.28242e-190.08886952.52519e-1801.4407e-177.05009e-1707.35824e-18
0.1579670.1238551.1139400.6082942.36201e-172.72726e-176.29895e-181.23943e-170.3291131.76964e-170.062592305.17784e-188.81319e-1806.26743e-183.52052e-19
0.1645410.1310151.103242.66839e-170.6491002.72442e-181.71e-170.31458500.03631506.79449e-180008.14622e-17
0.1711140.1385591.090616.45656e-170.68990500000.30005700.0100377001.60916e-17006.03269e-17
0.1776870.1464511.076739.90126e-170.72938503.55386e-172.08907e-174.97382e-170.2706155.53112e-1702.50667e-171.71943e-1707.64517e-173.57625e-172.17339e-17
0.1842610.1547141.06172.64942e-170.7680445.59602e-170000.23195607.90255e-178.4421e-181.71387e-170002.43435e-17
0.1908340.1633011.046135.37406e-170.8067030001.0414e-170.19329700000000
0.1974080.1721621.030471.43328e-160.84536306.84843e-185.5133e-194.29623e-170.1546374.14493e-173.68285e-171.09106e-172.30044e-174.27591e-1701.055e-166.93113e-17
0.2039810.1812581.0150200.88402202.99622e-1703.02091e-170.1159783.08411e-172.75685e-175.43616e-186.24216e-183.71163e-175.16901e-184.11801e-180
0.2105550.1905550.9999983.00367e-170.9226816.23547e-180000.07731873.80169e-1701.18233e-173.88217e-188.21793e-1801.19718e-170
0.2171280.2000250.9855187.33051e-170.96134100000.038659303.77579e-17000000
0.2237010.2096440.971655012.98063e-172.34613e-174.51358e-173.14083e-179.67571e-161.4783e-172.05579e-181.68394e-172.66241e-175.12837e-1801.97442e-182.68099e-17

Audience and decision

Research significance

Allocators, risk committees, and model reviewers

Portfolio managers, asset allocators, investment committees, and model validators use optimizers to convert noisy expected returns and covariances into precise weights. The evidence distinguishes numerical sophistication from decision reliability.

Decision context

Authority assigned to an optimized allocation

Do not interpret an in-sample efficient frontier or Sharpe ranking as a forecast. Decide first whether the mandate is return maximization, concentration control, or a stable risk budget, then evaluate the construction method against that stated objective.

The ranking changed when the estimates left the fit window

Four construction methods sat the same test. The experiment drew 2,891 daily observations for 15 liquid ETFs from a frozen adjusted-close snapshot. Maximum-Sharpe MVO, box-uncertainty robust MVO, equal-risk-contribution risk parity, and hierarchical risk parity were each fit on 2015–2021 data. Their weights were then frozen and applied to one untouched 2022–2026 window.

The robust portfolio posted the best out-of-sample Sharpe, about 0.6. It also stayed concentrated in QQQ and drew down more than 30% in 2022, then rode the 2023–2025 mega-cap rally. The box uncertainty set lowered estimated return levels without changing their ranking, so concentration survived. One 4.5-year window cannot separate a robust method from a fortunate regime.

Differences of ±0.2 in Sharpe sit well inside an estimated sampling error near 0.5 for this evaluation window. So the defensible conclusion describes realized risk shape. It does not name the method that will earn the highest future return.

Treat optimization as an assumption-sensitive proposal

  • Separate the return objective from the risk-budget objective before comparing methods.
  • Report concentration, risk contribution, drawdown, and out-of-sample behavior beside the in-sample optimum.
  • State the uncertainty treatment plainly: shrinking return levels does not change an uncertain ranking.
  • Rank methods only after multiple rolling-origin evaluations. This study ran one window.
  • State each method's regime dependency. Risk-based allocations can embed a bond and stock–bond-correlation bet.

Implementation and reproduction path

Fitted return and covariance relationships entered a different observed environment

Lab measurement Maximum-Sharpe optimization scored about 1.12 in the 2015–2021 fit period. It scored 0.44 in the sealed 2022–2026 evaluation. Risk-based methods held a steadier risk shape without establishing a return advantage.

Institutional record The two windows straddle a documented change in stock–bond correlation and monetary conditions. That change brought inflation pressure, rapid tightening, and weakness in equities and nominal bonds at once.

Interpretive synthesis An optimized portfolio assumes its fitted return and covariance relationships still bear on the decision. Precise weights carry relationships measured in one observed environment into a different one.

Causal boundary: the experiment does not identify regime change as the cause of the performance ranking. Correlation, estimation error, concentration, asset selection, and the realized equity path can each move the result.

Examine the allocation regime

Numbers

StatementValueAs statedNote
MVO Sharpe · 2015–2021 fit1.121.12
MVO Sharpe · sealed 2022–2026 evaluation0.440.44
Gap between the two published Sharpe ratios0.680.68Difference of the two published (rounded) Sharpe ratios; unrounded gap 0.689.
Estimated sampling error of a Sharpe ratio over this window≈0.5near 0.5sqrt((1 + SR²/2) / Y) with Y the evaluation window in years.
Robust MVO Sharpe · evaluation (best of the four)≈0.6about 0.6
Robust MVO maximum drawdown · evaluation window34.3%more than 30% in 2022
Daily observations · 15 ETFs2,8912,891
Liquid ETFs in the universe1515
Fit window2015–2021
Sealed evaluation window2022–2026
First daily observation2015-01-052015-01-05
Last daily observation2026-07-062026-07-06

Limitations

  • The result rests on one sealed evaluation window, not a distribution of rolling-origin outcomes.
  • The 15-ETF universe was selected in 2026, so it tilts toward survivors.
  • Weights stayed frozen. Rebalancing, transaction costs, and drift sat outside this experiment.
  • Every method received the same raw covariance estimate. Shrinkage was available and deliberately unused.
  • The study does not prove risk parity or HRP outperform, or that MVO always underperforms.

Cite this

Wisniewski, K. (2026, August 8). Portfolio optimization under estimation error and out-of-sample evaluation. Quantitative Markets & Institutions Lab. https://www.kylewisniewski.com/lab/portfolio-estimation

@misc{wisniewski2026portfolio,
  author = {Wisniewski, Kyle},
  title = {Portfolio optimization under estimation error and out-of-sample evaluation},
  year = {2026},
  month = {aug},
  howpublished = {\url{https://www.kylewisniewski.com/lab/portfolio-estimation}},
  note = {Empirical finding · Quantitative Markets & Institutions Lab · commit 4806df9}
}