The number from this paper that belongs on a trading desk is 82.0%.

After an Extreme Fear day, the next day remains Extreme Fear 82.0% of the time. Greed persists at 80.0%, Extreme Greed at 77.2% and Fear at 76.4%; Neutral lasts only 61.2% of the time. The corresponding implied spells are 5.6 days for Extreme Fear, 5.0 for Greed, 4.4 for Extreme Greed, 4.2 for Fear and 2.6 for Neutral. Direct transitions from Extreme Fear into Greed or Extreme Greed are essentially absent from the sample. Those five labels on retail dashboards, refreshed each morning, act like a slow ordered regime variable whose spells last three to six days.

The paper's test

Szeberényi and Kovács study Bitcoin alone, using 2,945 daily observations from 1 February 2018 to 27 February 2026. CoinGecko supplies prices. Alternative.me supplies the index score and its five labels, with both series retrieved 27 February 2026. Each runs on a seven-day calendar, which removes any need for non-trading-day alignment. Forward returns are simple returns, with log returns excluded, over 1, 7 and 30 calendar days.

Public market commentary often treats Extreme Fear as an automatic contrarian entry and Extreme Greed as automatic evidence of overheating. The authors divide that story into separate empirical questions. Does the index describe a state? Does it alter the conditional distribution of forward returns? Does it retain value in real-time forecasting? Their originality claim rests on separating those three questions. By the end, persistence carries the contribution.

The analysis starts with state frequencies and a one-day transition matrix. OLS mean regressions then relate forward returns to the standardised index, controlling for lagged FGI change, lagged return, lagged absolute return, seven-day momentum and seven-day volatility. Newey-West errors use horizon-matched lags of 1, 7 and 30. The authors next estimate categorical state dummies with Neutral omitted, an asymmetric model that separates greed intensity from fear intensity, and a quadratic specification. Quantile regressions cover tau 0.10 to 0.90. Logits estimate the probability of a negative forward return, while expanding-window ridge forecasts face a historical-mean benchmark.

Dependence from overlapping returns gets its own treatment. The paper uses a circular moving-block bootstrap with 999 replications: 30-day blocks for the mean regressions, then horizon-matched 7-day and 30-day blocks for the quantiles. Adjacent 7-day and 30-day forward returns reuse much of the same price path.

The bootstrap is the paper's most candid exercise. Most of the reported effects disappear inside it.

Which coefficient makes it through?

With HAC errors, a one-standard-deviation increase in the index corresponds to +0.14 percentage points of one-day forward return (p = 0.046, N = 2,936). At seven days, the estimate is +0.88 pp (p = 0.039, N = 2,930). The 30-day coefficient reaches +3.23 pp, though its p = 0.118 leaves it imprecise.

The paper calculates 30-day block intervals at all three horizons. The one-day interval is [0.00002, 0.00284], barely excluding zero. Seven days gives [-0.00092, 0.01817], while thirty days gives [-0.00686, 0.07634]. Both longer-horizon intervals cross zero. The abstract acknowledges the seven-day result; the results section also reports the thirty-day interval.

Changing the controls changes the ranking. Once thirty-day momentum and volatility replace the seven-day measures, the FGI coefficient is 0.0014 at one day with HAC p = 0.151. It becomes 0.0105 at seven days with p = 0.061 and 0.0267 at thirty days with p = 0.263. Under this control set, the one-day result loses 5% significance as the seven-day estimate approaches it. Which coefficient survives depends on the momentum and volatility lookback.

Quantile estimates weaken further. At seven days, the point estimates climb toward the upper tail: 0.0054, 0.0033, 0.0037, 0.0054, 0.0142 from tau 0.10 to 0.90. At thirty days they form a U-shape: 0.0469, 0.0201, 0.0073, 0.0273, 0.0540. Every one of the ten 95% moving-block intervals in the appendix contains zero. Only the 30-day tau = 0.10 coefficient survives at the 10% level, where its 90% interval is [0.0016, 0.0649]. Conventional quantile standard errors would have suggested a tail effect. Respecting the overlapping windows removes it. The abstract's distributional claim consequently rests on the pattern of point estimates, as the authors acknowledge when they advise reading them "as evidence about the shape of the point-estimate pattern rather than as uniformly precise tail effects."

Fit remains slight. Adjusted R-squared is 0.002 at one day, 0.008 at seven and 0.021 at thirty. Relative to a fully sentiment-free benchmark, adding the index raises fit from 0.0007 to 0.0022 at one day, from 0.0014 to 0.0083 at seven days and from 0.0018 to 0.0208 at thirty days. Joint HAC tests for the level and lagged change produce p = 0.061, 0.117, 0.279. At all three horizons, the diagnostics reject homoskedasticity, independence and no-ARCH.

Greed carries on

Extreme Greed has the highest average forward return at every horizon: +0.36% (1d), +3.80% (7d), +13.03% (30d). Its seven-day downside frequency, 42.0%, is also the lowest. Plain Fear performs worst. Its seven-day return is -0.31%, compared with +0.91% after Extreme Fear, and its seven-day downside frequency is 53.2% against 45.0%. Over thirty days, the comparison is 50.4% against 42.5%.

The contrarian interpretation fails by its own standard. At the greedy end, the index appears to capture short-run continuation, and the controlled estimates tell the same story. Seven-day greed intensity is 0.0513 (p = 0.017), implying that a full move from neutral to maximum greed corresponds to roughly +5.13 pp over a week. Seven-day fear intensity is 0.0080 (p = 0.649). The squared standardised index reaches 0.0086 at seven days (p = 0.048), giving the relation an upward bend at the extremes. Both 5% findings sit near the cutoff. Meanwhile, the paper estimates mean, categorical, intensity, quadratic, fifteen quantile specifications and three logits without a multiple-testing adjustment.

A trader should pause at the entry-day results. Extreme Greed begins 64 times, after which returns average +1.42% over seven days and +2.67% over thirty. Extreme Fear begins 116 times and is followed by -0.82% and -1.15%. Across all days classified as Extreme Greed, the seven-day average was +3.80%. The tradeable event is therefore roughly a third of the state statistic. Capturing the full state average requires a position established before the label changes. The authors reach the sensible conclusion: state persistence matters more than the entry day.

Forecasts lose to the mean

The expanding-window ridge is compared with a historical-mean benchmark recalculated from the training window in each block. Out-of-sample R-squared is -0.010 at one day, -0.074 at seven and -0.130 at thirty. RMSE is 0.0247 against 0.0246, 0.0658 against 0.0635, and 0.1573 against 0.1480. Directional accuracy comes in at 49.8%, 51.5%, 48.1%. The authors make this their own headline: the abstract says plainly that expanding-window forecasts do not outperform the historical mean benchmark.

For the downside logits, AUCs are 0.509 against a 0.494 base rate at one day, 0.498 against 0.454 at seven, and 0.428 against 0.405 at thirty. The paper offers a restrained reading. The logit slightly improves one-day AUC, while its seven-day and thirty-day improvements lack economic force. In absolute terms, a 0.428 AUC falls below a coin flip. Brier scores are 0.250, 0.249, 0.250.

Transaction costs appear nowhere. The paper has no strategy P&L, no Sharpe and no drawdown because it never constructs a rule. Given its thesis, omitting costs is defensible, and the authors write that the index "should not be used as independent trading rules." The economic question remains open. With no transaction-cost assumption reported anywhere, the +5.13 pp weekly greed-intensity slope remains a gross in-sample association. The coefficients also receive no test based on splitting the 2018-2026 period in half. We did not find a sub-period stability analysis, a material omission when one pool contains both the 2022 deleveraging and the 2025-2026 high-price regime.

The interpretation problem

The index is composite. In defining the market-state variable, the paper says it combines recent price behaviour, volatility, market momentum, attention and sentiment-related information. The authors list their inability to decompose it as a limitation. An association between this index and forward returns therefore sits close to an association between momentum, realised volatility and forward returns passed through a wrapper. Reverse causality also remains unresolved. The authors write that sentiment "often looks like an interpretation of recent price movement rather than an independent signal arriving clearly before the price."

Controls for lagged return, lagged absolute return, seven-day momentum and seven-day volatility help with that problem. They cannot identify the channel. And with adjusted R-squared of 0.8% at the weekly horizon, little remains available to attribute.

Decomposing the index into its attention, volatility and momentum components would change my reading if the non-price part retained the greed-intensity coefficient. Until such evidence appears, persistence is the finding. It rests on the sample's 2,944 next-day transitions rather than any regression.

We could not test this ourselves. The explanatory variable is the published Alternative.me index with its five labels, which we do not hold. Rebuilding it would require the attention, volatility, social-media activity and momentum inputs named in the paper but never decomposed, together with the historical weighting. A news-derived sentiment score would substitute a different signal.

This pattern has appeared before: an effect clears conventional inference, then collapses when the estimator matches the horizon of the measured quantity. In a rotation-premium paper, |t| fell under 1.2 after horizon matching (/articles/the-correlation-rotation-premium-that-no-traded-instrument-spans). Szeberényi and Kovács performed that check on their own work and published what happened.