Close-to-close covariance wins this allocation contest. Across all four forecasting setups at the reported configuration, it delivers a higher Sharpe than Yang-Zhang OHLC covariance, alongside lower VaR and CVaR at both the 1% and 5% levels. The authors reach the same verdict: adding intraday volatility during optimisation yields no systematic improvement in risk-adjusted performance. They argue that the value of refined risk estimates depends on the portfolio's correlation structure and the forecasting architecture. With no significance test behind that qualification, its survival remains doubtful.

We ran no replication. Only part of the paper's 28-token universe overlaps the roughly 50 cryptocurrencies for which we have prices. Any result from us would therefore use a smaller universe and would be incomparable with theirs. Available data are sufficient to implement rolling ML forecasts and OHLC-based covariance estimators. We can establish no more than that.

The test: 28 tokens and ten out-of-sample periods

Toscano, Roca and Jareño study 28 fungible NFT-ecosystem tokens, including SAND, MANA, AXS, GALA, THETA, RARE and UFO. Their daily OHLC data come from Yahoo Finance and run from 1 January 2021 to 31 March 2024. These exchange-traded assets are platform utility and governance tokens. Collectible NFTs form a separate category, described by the paper as unique, illiquid and heterogeneous by design.

The study has two stages. First, random forest, gradient-boosted regression trees, NNET and BPNN forecast next-day returns. BPNN is a back-propagation feedforward network. Each model receives the same 15 OHLC-based log-price predictors plus four technical indicators: true range, ATR, momentum and RSI. Grid search with cross-validation sets the hyperparameters.

Allocation changes in the second stage while forecasts remain fixed. A correlation filter removes assets whose pairwise correlation exceeds 0.8, after which the top seven by forecast return remain. A Monte Carlo sweep evaluates 50,000 candidate portfolios in each rolling window and selects the highest Sharpe. Three rules use the identical forecasts: classical mean-variance with close-to-close covariance, modified mean-variance, or MMV, with Yang-Zhang OHLC covariance, and equal weight. The main windows contain nine months of training and three months of testing across ten out-of-sample periods. The paper also reports 4+1 and 6+2 as robustness checks. One permille is deducted from daily portfolio return for costs.

Intraday movement is plainly present in these assets. Across 32,976 NFT token-days, the pooled mean intraday range reaches 10.72%, compared with 6.26% for Bitcoin, Ethereum, BNB, XRP and SOL. A range above 30% occurs on 5.51% of NFT sample days and 0.58% of days for the established coins. Across the 28 tokens, Yang-Zhang annualised volatility averages 130.5%, versus 117.4% close-to-close, giving a ratio of 1.14. The corresponding mature-coin ratio is 1.01. Closing prices omit information here. That extra information then fails to improve allocation systematically.

Close-to-close wins

Before costs, GBRT plus MV records a Sharpe of 4.414, against 3.957 for GBRT plus MMV. Random forest produces 4.020 against 3.567. BPNN gives 4.020 against 3.953, while NNET gives 3.687 against 3.524.

Four for four at the reported configuration, which uses the 0.8 filter.

MMV generally earns the higher raw cumulative return. GBRT+MMV reaches 5.099 versus 4.978 for GBRT+MV, while RF+MMV reaches 4.944 versus 4.806 for RF+MV. The price appears in volatility and tail losses. RF annual SD rises to 1.386 under MMV from 1.195 under MV. At 1%, RF CVaR is 0.062 under MMV and 0.051 under MV. For GBRT, the same comparison is 0.047 against 0.041.

The paper's second hypothesis proposed that intraday-informed covariance would reduce downside tail risk. Its own results reject that hypothesis, and the authors state the conclusion directly: "The marginal gains from intraday information are consistently offset by higher estimation risk, making daily MV optimization more robust for NFT-ecosystem tokens." Chen, Zhang and Jia (2022) find the opposite sign in equities. This paper extends their framework.

The reported Sharpe ratios of 3.5 to 4.4 for a seven-token crypto basket lack credibility as stated. We could not locate the annualisation convention used for Table 11. Nor could we find whether a "cumulative return" between 3.0 to 5.1 denotes a multiple or a percentage. An annual standard deviation of roughly 1.0 to 1.4 also fits awkwardly beside token-level Yang-Zhang volatilities averaging 130.5% annualised across the 28 tokens. These ambiguities leave the ordering intact. The levels should not be quoted.

When does OHLC help?

Averaging over the 50, 75 and 100-day covariance windows produces a few positive MMV gaps. For BPNN, the Sharpe difference is positive by 0.37 at the tightest correlation filter, rho = 0.40. GBRT gains 0.15 at rho = 0.80 and 0.16 at rho = 1.00. The same averaging gives RF a negative difference of 0.34 at rho = 0.80 and NNET a negative difference of 0.25 at rho = 0.40. Intraday information changes sign with both the correlation regime and the signal model. This conditional result is the paper's actual claim.

The supporting evidence is thin. There are no Sharpe-difference tests or bootstrap results, and each cell rests on ten out-of-sample periods. The 50, 75 or 100 days of look-back and the rho threshold are selected from the out-of-sample performance tables. The best reported configuration is therefore selected. A 0.15 Sharpe gap drawn from a grid of 16 rho-by-model cells offers little basis for a trade.

The missing forecast benchmark

The abstract claims that machine learning forecasts create substantial economic value. Its evidence comes from comparisons with equal weight: RF+MMV returns 4.944 against 3.516 for RF+1/N, GBRT+MMV returns 5.099 against 3.491, and NNET+MMV returns 3.956 against 3.050. Yet each 1/N portfolio owns the same seven tokens, selected through the same forecasts after the same 0.8 filter. Those comparisons isolate weighting rather than forecast skill. We did not find a portfolio using historical-mean expected returns, random selection or a zero signal, any of which could attribute the gap more clearly.

Average directional accuracy across the ten evaluation periods is 53.53% for RF, 52.78% for NNET, 48.34% for GBRT and 47.45% for BPNN. Two of the four models fall below a coin flip. BPNN still generates a 4.020 Sharpe under MV and returns 4.747, compared with 3.380 for BPNN+1/N. A sub-50% hit ratio financing that spread leaves return prediction as an uncertain mechanism.

The authors also identify an inversion between forecast errors and directional accuracy. Models with the lowest errors tend to call direction less successfully. BPNN has the lowest average RMSE at 0.0079 and the worst hit ratio, whereas RF combines the best hit ratio with higher errors.

Trading frictions and capacity

The paper deducts one permille from portfolio return every day, regardless of whether any trading occurred, and never reports turnover. A flat daily deduction can overstate costs for a low-turnover portfolio or understate them for a high-turnover one. The authors acknowledge the first possibility. They also state that network fees and slippage are excluded. After the flat charge, GBRT+MV still leads with a Sharpe of 3.622 and return of 4.085. MV remains ahead of MMV throughout, leaving the central conclusion insensitive to this cost treatment.

Liquidity presents the larger obstacle. RARE has a daily log-return standard deviation of 0.423 and a sample range of 5.781. Daily rebalancing into RARE at scale is implausible, yet the paper provides no capacity or market-impact analysis. We could not find a stated inclusion rule for the 28 tokens. If full-window Yahoo Finance coverage determines the screen, the sample becomes a survivor set spanning the 2021 boom and the 2022-2023 crypto winter, periods the authors themselves name.

Diversification is limited as well. Across the 378 pairs, mean correlation is 0.53 and median correlation is 0.60. Correlation exceeds 0.50 for 59% of pairs, while only 6% are negative. WAXE and WAXP reach 0.99. We have raised a related concern about connectedness within a crypto sub-sector (/articles/the-minimum-variance-optimizer-keeps-rejecting-defi).

Keep it out of the covariance matrix

DeMiguel's estimation-error argument holds in a setting deliberately chosen to favour the richer estimator, which makes this a fair test. Yang-Zhang volatility runs 14% above close-to-close for these tokens, compared with 1% for Bitcoin, Ethereum, BNB, XRP and SOL. Even so, the more accurate estimator produces worse allocations because the optimiser magnifies the noise carried by the additional information.

A bootstrapped Sharpe-difference test, applied to a universe broad enough that seven names do not consume the available parameter space, would resolve the conditional claim. Intraday range may still belong in the signal. It has not earned a place in this covariance matrix.