SOTA (Stock Options Trading Agents) has an interface worth copying. Its agent chooses among nine option-strategy families instead of thousands of listed contracts. The 18.32% return comes from one seed, midpoint fills and a six-month test. The reported policy is one of several RL checkpoints the authors scored on that test window.

What does the agent choose?

Xie and Liu give the agent a family decision and let code select the contracts. An action specifies an underlying, a strategy family, an orientation, a tenor bucket and delta-grid coordinates. The nine families are outright, debit vertical, credit vertical, defined-risk reversal, long straddle, long strangle, iron butterfly, iron condor and butterfly. They cover directional, volatility, skew and curvature exposure.

For each underlying, the model offers up to six candidates and sees their prices before committing. Deterministic resolvers handle expiry, strikes, size and hedging. Expiry is the listed date nearest an anchor of 5, 14, 45 or 120 days; strikes sit nearest the requested delta on a 0.05 grid. Each package is sized to equalize second-order scenario risk at 0.005 of NAV. The five volatility families use a Whalley-Wilmott no-trade band for their delta hedge.

The proposed edge lies in reading the option-implied return distribution. Each underlying supplies 10 state features, including the IV minus 21-day RV wedge, the 90-minus-30-day ATM slope, 25-delta skew, the 25-delta butterfly, Lee-Ready signed delta flow and the change in open interest. News and filing summaries are available too.

Training has two stages. A frontier teacher first produced 1,000 trajectories over September to November 2024, with ticker, price level and calendar date masked and news included. The 144 trajectories whose annualized Sharpe exceeded 0.75 supplied supervised fine-tuning data for Qwen3.8-27B. The hyperparameter table, however, lists 414 SFT episodes; the paper does not reconcile those counts. GRPO then trained on December 2024 to February 2025, with reward tied to the log change in NAV. Testing runs from 2025-03-03 to 2025-08-29, starts at $1,000,000 and covers SPY plus AAPL, AMZN, GOOGL, META, MSFT, MU, NVDA, PLTR and TSLA.

Five losing baselines

SOTA returns 18.32% in the paper's test, with a Sharpe of 1.60, max drawdown of 8.96% and annualized volatility of 21.52%. Under the same environment, resolvers and costs, every baseline loses money. Equal-weighted loses 5.22%; the threshold rule loses 28.46%, GARCH 49.24%, GBDT 49.78% and logistic 55.45%. Table 2 supports the paper's claim that SOTA alone has positive total and risk-adjusted returns among the feasible policies. It sets a low bar for the comparison: three of five baselines lose about half the book in six months. Those losses show how harshly the environment treats naive family selection. They give little evidence of how SOTA would fare against a competent selector. We did not find a test-window run of either the base Qwen model or the frontier teacher.

The payoff distribution is uneven. Fewer than three trades in ten win (29.17%), while the average winner is 4.68 times the average loser. On the supervised window, the teacher's attribution shows where some of the drag arose. Across the 144 trajectories retained by the Sharpe filter (total +7.99%), delta hedging cost 2.83% and long straddles lost 0.57%; iron condors added 2.21% and outrights 3.16%. A hindsight oracle choosing from the same candidates returns 608.31%. The available strategy space is large. Whether this policy found a durable part of it remains open.

Four-leg trades at the midpoint

The execution model fills at the bid-ask midpoint, then charges $0.65 per contract per leg, $5 per assignment, $0.005 per share and 0.5% annual borrow. Iron condors and iron butterflies each have four legs, with zero spread charged on every leg at a midpoint fill. The appendix also specifies a half-spread multiplier of 1.0, the "full quoted half-spread", for marking and forced exit. We could not determine from the text whether that charge enters exit P&L or only marks. Its treatment matters if the 60% stop-loss uses marks, a point the paper does not specify.

We could not rerun the result. Our options data contain end-of-day prices and Greeks but no bid or ask quotes, leaving midpoint execution without an input. We also lack option trade prints and open-interest history, so two of the 10 state features cannot be built. The policy requires frontier-teacher trajectories and a 27B fine-tune we cannot replicate. A conventional ML selector on our data would address the GBDT row, a different claim from the agent's. Nothing in this review tests SOTA.

A checkpoint scored on the test window

The note beneath the training-hyperparameter table says: "Checkpoints are written every 10 optimizer steps, and each was evaluated on the test window." The paper reports the step-20 checkpoint. With GRPO running 40 optimizer steps, at least four checkpoints were scored on March to August 2025. We did not find a rule fixed before testing for choosing 20. If the authors chose the best of those four, the 18.32% is a maximum over test draws. The run has one seed, 0, and an RL batch size of 3. Twenty GRPO steps at batch size 3 leave little optimization behind the headline.

When news disappears during RL

Dropping news during RL moves the test return 21.04 points, from -2.72% to 18.32%. Keeping news yields a 31.69% drawdown; dropping it yields 8.96%. The authors' careful conclusion is: "news is useful for constructing frontier-model supervision, but retaining it during policy optimization does not improve out-of-sample performance." Their abstract makes the stronger numerical claim that retaining news during RL "reduces out-of-sample return from 18.3% to -2.7%." Each arm has one RL run.

The design leaves another question unresolved. Both RL arms begin from the same supervised checkpoint, which was trained with news. SOTA is therefore a news-trained model evaluated without news. The SFT-only row retains news and returns -10.19%, with a Sharpe of -1.25. With those inputs unchanged, RL lifts the return to -2.72% and Sharpe to -0.16. Max drawdown roughly doubles from 15.56% to 31.69%, as the paper notes. This is the paper's clean RL comparison. We did not find an evaluation of the SFT checkpoint without news. That row would help distinguish the effect of RL from the removal of news in the 28.51-point move between -10.19% and 18.32%.

My view would change if the checkpoint were fixed on the December to February window, results held across seeds, the quoted half-spread were shown to reach exit P&L, and the same pipeline stayed positive in a second six-month test.

Take the nine-family interface. Leave the return figure pending that evidence.