PPO-HRAP's best SPY result could owe as much to a seven-row lookup table as to PPO. The paper does not report what that table earns on its own (ρ = 1). We did not train PPO; our backtest below combines the rule with a substitute policy and cannot test the paper's claim.
The rule's share of every trade
Duong and Huynh trade one ETF against cash each day, allowing weights from -1 to +1. Their method, PPO-HRAP, uses PPO with a "hybrid regime-aware policy" built around a hand-written rule. The rule takes direction from the 20-day return and the spread between the 5-day and 20-day moving averages. Two positive signs mean bull; two negative signs mean bear. A VIX z-score measured over 252 days sets the stress level, with cutoffs at 0.8 for stress and 1.8 for crisis. The resulting targets are 1.0 for a calm bull, 0.5 for a stressed bull, 0.3 for neutral, 0.1 for weakly negative, 0.0 for a calm bear, -0.1 for a stressed bear and -0.8 for a bear crisis.
The PPO actor takes 13 inputs: nine market features and four portfolio features, namely cash ratio, prior weight, unrealized PnL and drawdown. It proposes a weight, but the trade uses 0.4 times that proposal plus 0.6 times the rule's target. The reward starts with log return and also steers the actor toward the rule. It deducts a penalty for increasing drawdown, whose base coefficient of 0.05 rises with the clipped VIX z-score. Further deductions are 0.02 times the squared distance from the rule's target and 0.0015 times the absolute change in weight. The intended source of gains is lower exposure when trend breaks or VIX jumps, alongside a long position during calm advances.
Training spans 2010 to 2017; validation covers 2018 to 2019. A grid from 0.3 to 0.7 selected the 0.6 blend weight on validation, where its reported Calmar was 0.2299. The test covers 2020 to 2022, or 756 SPY days. Every method faces a 0.1% proportional cost, and PPO gets 30,000 training timesteps. The authors describe the attribution plainly: "the reported gains should be interpreted as the joint effect of the prior and the learned correction."
Did the risk control pay on SPY?
For the headline seed, 42, it did. PPO-HRAP returned 27.62% in total versus 17.60% for buy and hold. Its Sharpe was 0.6447 versus 0.3420, with maximum drawdown of 18.47% versus 34.10%. Average exposure was 0.4337 in high-VIX periods and 0.8839 in low-VIX periods, the shift the design seeks.
The seed-42 comparison gives an incomplete impression. Profit-only PPO and SAC both match buy and hold to four decimals there: total return 0.1760 and Sharpe 0.3420. Across five seeds, SAC profit has a mean Sharpe of 0.8644, Calmar of 1.4733 and drawdown of 4.69%, though its mean return is lower at 0.1978. PPO-HRAP's five-seed means are 0.6219 Sharpe, 0.4508 Calmar, 18.76% drawdown and 0.2725 return.
The authors acknowledge those results and narrow the claim to "a return-leading hybrid policy with improved drawdown control relative to passive long exposure, not as a universal winner on every risk metric." They deserve credit for saying it. The abstract and conclusion do not carry that concession through: the abstract says PPO-HRAP "remains stable" across seeds, while the conclusion claims improvements in "total return and risk-adjusted metrics over both financial and RL baselines" for the main SPY run. Their stated defence is that PPO-HRAP "has the best mean total return among the RL methods and a small return standard deviation", 0.0109. Table 6 offers another defence they leave unstated. SAC's Sharpe standard deviation is 0.2417, against PPO-HRAP's 0.0565. SAC Profit (Sharpe std 0.2417), PPO Static MDD (0.3882) and PPO Variance (0.2117) all vary sharply across seeds. PPO Static MDD records the best seed-42 drawdown at 9.51%, yet its five-seed mean reaches 35.54%, worse than buy and hold.
I would put the most weight on PPO-HRAP's tight return dispersion, a standard deviation of 0.0109. Even that does not distinguish the rule-anchored policy: profit-only PPO, which collapsed onto buy and hold, is tighter at 0.0059. With 60% of the position fixed by a deterministic rule and a reward penalty for departing from it, limited variation across seeds is close to built in.
Where is the rule-only result?
The paper defines the missing endpoint. At ρ = 1, "the learned actor is ignored and the system becomes a deterministic regime rule." The validation sweep includes 0.3, 0.4, 0.5, 0.6 and 0.7. We found no result for 1.0 in validation or test.
Calmar rises from 0.1792 at ρ = 0.3 to 0.2299 at 0.6, then reaches 0.2291 at 0.7. Drawdown is shallower at 0.7, 16.50% against 17.13%, while validation return falls from 0.0800 to 0.0767 between those points. The paper warns that "an overly strong prior can suppress learned correction," but its top two settings are nearly tied. A Calmar gap of 0.0008 cannot establish whether 0.6 is a peak or the start of a plateau extending toward 1.0.
Both actor and rule are capped at 1.0. Their blend can therefore reach full long only when each reaches 1.0. The paper says the strategy spent 54.10% of the test period at full long exposure, without defining the term. If it means a weight of exactly 1.0, the actor was pinned to the same boundary as the rule on those days. For more than half the test, it contributed no different position. The concession about joint effects is candid, yet the abstract recommends blending "learned actions with a volatility-aware regime prior," subject to caveats about turnover and single-run cross-asset evidence. A desk considering a neural network needs to see whether its 40% share pays for itself. Running the rule alone at ρ = 1 over 2020 to 2022 would require no training.
We also could not find how the 0.8 and 1.8 VIX cutoffs or the seven target levels were selected. Only ρ appears in the validation sweep. And the test includes the 2020 crash and 2022 bear market, the episodes most likely to favour VIX-driven de-risking.
The cost of moving
PPO-HRAP's SPY turnover is 51.22. The paper gives 1.0000 for buy and hold on QQQ and DIA, though the SPY table omits buy-and-hold turnover. Its per-step penalty uses |w_t - w_{t-1}|, without explaining how the reported 51.22 is aggregated. Buy and hold's 1.0 looks consistent with a sum of absolute weight changes that counts the initial allocation. Under that reading, 0.1% per unit amounts to about 5.1 points of cumulative return across three years, already included in the 27.62%. The assumption deserves scrutiny: the cost is a flat 10 bps with no slippage, and we found no borrow charge for shorts generated by the -0.8 crisis target. Doubling the cost would take roughly another five points on the same arithmetic. The 10-point advantage over SPY buy and hold would remain, with little margin to spare.
The authors call turnover "far above passive or reward-only baselines" and acknowledge that QQQ and DIA are single runs. On QQQ, PPO-HRAP earns 0.4149 with Sharpe 0.7811 versus buy and hold's 0.2306 and 0.3826, at turnover 39.42. SAC profit has the higher Calmar there, 0.5739 versus 0.4799, with a 6.09% drawdown. DIA returns are 0.2811 and 0.6623 versus 0.1468 and 0.3088, at turnover 55.59. We did not find whether ρ was selected again for those assets or transferred from SPY. The paper says its features use information available before trading, but leaves the fill price unnamed. It uses the close C_t in its features and NAV_{t+1}/NAV_t in the reward. A daily strategy trading a close-derived signal needs that timing settled.
Our substitute run cannot answer the PPO question
We could not reproduce the paper's experiment. Its claim depends on PPO training for each asset and seed, and we lack the compute to run it as specified. Instead, we ran a long-only, fully funded allocation among cash, SPY, QQQ and DIA, capping each ETF at one-third. We clipped regime targets at zero, turning the -0.1 and -0.8 bear targets into flat positions, then blended in a policy action at ρ = 0.6. This differs from the paper's single-ETF long-short strategy. It has no ρ = 1 arm either, so it says nothing about whether a learner improves on the rule.
The figures in the strip above are ours, from our backtest covering 2020-01-03 to 2024-07-01. They include commissions of $0.004 per share, subject to a $1 minimum per order. Under the closing-auction fill convention, slippage is zero. Our run made 1,270 trades over four and a half years. Its Sharpe of 0.89 exceeds the paper's 0.6447, and its maximum drawdown of 15.61% is shallower than the paper's 18.47%. Those figures measure different books and periods: our long-only, three-ETF allocation from 2020 to mid-2024 and the paper's single-ETF long-short policy from 2020 to 2022. None of our result is evidence about PPO-HRAP. Our allocation cannot short during a crisis, holds at most a third in any ETF, and includes 2023 and early 2024. Two choices in our run remain unresolved. Signal-to-fill timing was specified as next open but recorded under a close convention; a separate 0.1% notional charge may also have been applied on top of per-share commissions. Whatever their effect on the figures, those questions concern our implementation first.
A rule-only SPY Sharpe well below 0.6447 at ρ = 1 over 2020 to 2022 would make the case for the actor's 40% share. Until that result is reported, the paper has not shown that the network earns it.
Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.