A 72% hit rate on the direction of individual stocks' 12-day returns would get a trader's attention. So would an R2 of 0.18 at that horizon, especially when the model sees only the previous 36 daily returns. The validation split gives us more reason to doubt the result than to trade it.

Balakrishnan P, A. Anny Leema and two co-authors at VIT Vellore assemble a decision-support tool for Indian retail investors from four components. First comes a health score. Annualized return is clamped to [-20%, +25%], and Sharpe, using a 6% risk-free rate, to [-1, 2]. The authors map each linearly to 0 to 100 and blend them 0.6/0.4. Scores of 70 or above receive an Overperforming label; 40 or below, Underperforming.

Next, a 2-layer LSTM with 54,787 parameters takes the last 36 standardized daily returns and predicts simple returns 12, 36 and 60 trading days ahead. The authors explicitly decline to treat it as a price oracle. A Monte Carlo simulation then draws daily returns from a normal distribution using the trailing 252-day mean and volatility. It treats 21 trading days as a simulated month, feeds each month back into the LSTM's 36-day window, and repeats the process 150 times per asset. Those paths produce distributions of CAGR, volatility, Sharpe, Sortino and drawdown. Retail users see the simulation alongside the health score.

The LSTM affects the simulated drift. Its first output is clipped to ±0.03, divided by trailing 252-day volatility and passed through tanh. A factor k scales the resulting tilt, ranging from 0.12 for stocks down to 0.02 for debt funds. The fourth component adds Integrated Gradients, which attributes output to inputs in the window, and natural-language event narratives. The paper claims consistent evaluation, stable long-term forecasts and interpretable output. Any edge depends on the LSTM's direction call.

Training covers 159 stocks and 152,426 daily rows from 2017-01-02 to 2021-01-01. Across the three horizons, the authors report mean directional accuracy of 72%, 67% and 62%, with R2 of 0.18, 0.09 and 0.04. Diebold-Mariano statistics reach 3.42 against persistence and 2.15 against a rolling mean. In an ablation, Monte Carlo without the LSTM signal records 58% average accuracy and 0.073 RMSE, versus 67% and 0.061 for the full pipeline. The abstract describes an evaluation on Indian equity and mutual fund data. We did not find a stated mutual fund sample; the described training set contains stocks only.

What does the holdout tell us?

The 137,162 samples are rolling windows advanced one day at a time. The authors randomly split them 80/20 using seed 42. Samples a day apart share 35 of their 36 inputs and 59 of the 60 days in their +60 target. After shuffling, almost every validation row has a near-twin in training. The 27,433-row validation set measures interpolation within a 2017 to 2020 panel. It tells us little about the next month.

The authors acknowledge the choice: "In this work, we report results under the implemented holdout protocol and treat walk-forward validation as a rigorous extension for future evaluation." They also say time-series practice recommends time-respecting evaluation "to avoid optimistic estimates under temporal autocorrelation." Efficiency is their stated reason for the split.

Their defense is that the system is "still good for a prototype that focuses on publishing because it puts reproducibility, stability, interpretability... first." Accuracy figures presented as findings need more than a prototype smoke test. The significance section calls the gains "not attributable to chance variation," while the abstract promises "stable long-term forecasts." The DM tests use those same contaminated validation rows. They establish a win over persistence and a rolling mean on rows shuffled from the training panel.

Other results leave questions. A paired t of 2.89 on roughly 27,000 observations implies a two-sided p near 0.004; the paper prints 0.013. Removing the tanh and k bounds increases RMSE to 0.079, attributed in the paper to "overconfident out-of-distribution predictions." Those bounds operate inside the simulation. We could not work out how they alter point forecast error.

The forecast's 4%-to-16% floor and ceiling

The final simulated CAGR blends historical and simulated CAGR at 0.4/0.6. The authors then clip that blend to [4%, 16%] for equities and [6%, 14%] for funds. They rescale the displayed path to reach the clipped terminal value, subject to hard caps of 3.0x the current price for horizons of five years or less and 2.5x beyond. They call the bands "consistent with long-run return expectations in Indian equity and debt markets."

The floor decides what the headline can say.

An equity forecast CAGR cannot fall below 4% a year. Because the displayed path is rescaled to match it, the headline forecast never shows a loss, even for a stock whose history is a halving. Intra-path drawdowns remain: rescaling preserves the path's shape, so the drawdown statistic still conveys some information.

The LSTM bound allows a larger move than its appearance suggests. The paper gives a maximum tilt of ±0.00144 a day for a typical equity. Multiply that by 252 and it is about ±36% a year before compounding; the CAGR clip binds well before tanh. The scenario health score weights return 0.35, Sharpe 0.25, and volatility and drawdown 0.20 each. Labels remain consistent in over 89% of cases under alternative weights. That 89% speaks to sensitivity to weighting. Nothing in the paper tests whether a score of 70 or more precedes higher returns.

A US substitution

We did not find a portfolio return, cost assumption or out-of-sample Sharpe for the health score or forecasts. The paper makes no claim of trading profits. Our proposed rule has no results yet: hold names with scores of 70 or above and a positive +12-day forecast, retrain the LSTM walk-forward, and purge training rows within 60 days of each test block to match the longest target.

We cannot trade the paper's market. We would substitute US equities for Indian equities and US ETFs for Indian mutual funds. Historical returns alone drive the scores and LSTM, but the authors calibrated the CAGR bands and volatility bounds for Indian markets. Applied to US names, those constraints are arbitrary, and the reported results cannot be assumed to transfer.

Table 7 raises a separate trust question. Its narratives mention the RBI's 2022 to 23 hike cycle and a 2022 tech selloff, both later than the training sample's January 2021 end. The authors' account of narrative generation did not let us determine where the events came from. Their own design principle says "unverifiable event-based statements (e.g., 'the move was caused by earnings news') should not be used unless the external event dataset is very clearly included." Table 7 contains such statements, and no event dataset is described.

Walk-forward accuracy would change our view. If its average across the three horizons stays above the 58% reported for simulation without the LSTM, and anywhere near 67%, the signal is real.