The next NIFTY 50 close looks too predictable
A 22% RMSE cut against close persistence is too good to accept without checking the target. Sain and Singh report that edge for plain linear regression in their comparison of twelve regressors. Their models forecast the next NIFTY 50 open and close from one daily OHLCV bar. For a liquid index, the close result deserves more suspicion than the paper gives it.
The data are daily ^NSEI bars from Yahoo Finance. NIFTY 50 is the National Stock Exchange of India's index of 50 large companies. The five, ten and twenty-year windows all end on 1 June 2026. Each model uses today's open, high, low, close and volume to predict tomorrow's open or close as a price level. A second feature set tests whether indicators help: it adds 10- and 20-day SMAs, daily return, 20-day rolling volatility and a 14-day Relative Strength Index (RSI). Each window has one 80:20 chronological split. Naïve Persistence, the benchmark, uses today's price as tomorrow's forecast. Only Random Forest and XGBoost receive tuning, through RandomizedSearchCV within a 5-fold TimeSeriesSplit.
This is a forecasting benchmark with no trading rule or claimed source of profit. The authors conclude that "only the simple linear models consistently did better than the basic naïve benchmark across all timeframes." Indicators change little. With them, linear regression's 10-year close RMSE drops from 215.56 to 201.08; XGBoost's 5-year close RMSE climbs from 259.52 to 374.46.
For the open, today's close supplies an obvious advantage. As we read the benchmark definition, persistence uses today's open to forecast tomorrow's open. Regression can use today's close instead. The close-to-next-open gap contains the overnight move; the open-to-next-open gap also contains today's session. On the 20-year raw-OHLCV test, linear regression records MAE 72.74 against 142.97 for persistence. The authors highlight that halving. They acknowledge that R² above 0.99 "mostly show this natural day-to-day market stability", then take the lower MAE as additional predictive value. Much of it is probably today's close, which a desk would already use for a naive open forecast.
Can the close result survive an alignment check?
The shortcut for the open does not explain the close result: persistence already forecasts from today's close. Yet on the 20-year test, linear regression has RMSE 182.35 against 233.82 for persistence, with MAE 130.60 against 172.75. The 5-year RMSE comparison is 195.77 against 254.89. Squaring the 20-year RMSE ratio gives about 39% less mean squared error than persistence. Under zero drift, persistence error is the daily move itself. From a single bar, the regression is thus claiming to forecast roughly two-fifths of the variance of NIFTY's next close-to-close change.
Target alignment needs checking first. The paper defines next-day open and next-day close targets, and its account of missing rows at the sample's end is consistent with a forward (t+1) shift. We cannot inspect the implementation: code is available only on request. The paper also leaves open whether StandardScaler was fit solely on training rows. Scaler leakage cannot account for this gap, since OLS with an intercept makes identical predictions under any per-feature affine rescaling.
Trees at the training-price ceiling
The 10- and 20-year tree ensembles nearly memorize their training samples and then fail out of sample. For the 10-year open, Random Forest and XGBoost each attain train R² 0.9999. Test R² falls to -2.3956 (RMSE 1832.49) for Random Forest and -3.0999 (RMSE 2013.59) for XGBoost. Random Forest's 20-year open test MAE reaches 3836.93. The authors give the reason: tree models "can never predict a number higher than the highest price they saw during training." New highs in the test set leave their forecasts stuck at a flat ceiling.
Tuning cannot raise that ceiling.
Tuned Random Forest scores -2.5985 (RMSE 1886.44) on the 10-year open, worse than its untuned -2.3956 (RMSE 1832.49). Trees fare better in the 5-year window: open RMSE is 183.76 for Random Forest and 212.33 for XGBoost, against 236.80 for persistence. Linear regression still leads at 159.62. Predicting price levels creates the collapse. The authors list returns or directional movements as possible future targets, without connecting that suggestion to the tree results.
Different histories, different test markets
Because all windows end on 1 June 2026, the 80:20 splits test roughly the last one, two and four years respectively. The 20-year model faces a long stretch of new highs; the 5-year model faces a single recent year. The authors acknowledge the comparison problem and call for aligned out-of-sample periods, so window length and market regime can be separated.
Their close-price summary still describes linear performance as improving with more history. In the raw OHLCV results, test R² rises from 0.9479 to 0.9960 as RMSE falls from 195.77 to 182.35. The wider price range in a four-year trending test period accounts for much of the R² rise: the same error produces a higher R² there. The 7% RMSE reduction also compares different periods. The paper reports no significance test on forecast errors, another item the authors leave for future work.
No trade to test
The reported outcomes are forecast errors, MAE and RMSE in index points, plus R². We found no strategy returns, transaction costs, hit rates or Sharpe ratios. We did not rerun the study because we had no NIFTY 50 daily history back to 2006.
The close-price result remains the reason to read it. If the 22% RMSE edge over persistence survives a return target, an aligned walk-forward test and a verified target shift, it would merit testing as a trading signal on India's benchmark index. Until those checks are run, we regard it as a likely bookkeeping artifact.