AQAI QuantAI research lab for systematic strategies

Automated analysis

This analysis was drafted by our research engine and has not been checked by a human editor. It may contain errors. It separates the paper’s own results from our tests, and any figures called ours come from our own backtest.

Our automated analysisOur backtest

Dirac-3 gains 0.045 Sharpe over 33 months

Each peak came from 16 configurations on one window, without a reported significance test

2026-09-08 · 8 min read · US equity factor portfolios constructed from US stocks

Reviewing: Photonic Quantum Computing vs. Classical Solvers in Constrained Factor Portfolio Optimization · Nirvik Sahoo, Chyng Wen Tee and Paul Robert Griffin · Read it on arxiv

Our backtest of this idea

Our automated quick test, not the paper's

Classical Two-Stage QUBO Multi-Factor Sleeve Allocation

Backtest period 2020-01-01 to 2024-07-01 · hypothetical, net of modelled costs

Why these figures are not the paper's (3)

Run on a different market than the paper

The paper allocates across precomputed US equity anomaly-factor portfolios from the Jensen-Kelly-Pedersen library, which is not available in the platform. We would construct comparable long-only US equity factor sleeves from daily prices, point-in-time fundamentals, and market-cap data, then allocate across those sleeves monthly. The constrained factor-allocation mechanism survives this substitution, but the paper's reported performance and exact 13-factor exposures do not transfer to the reconstructed sleeves.

The paper's own figures describe its universe and do not carry over to ours.

This is not a replication of the paper (3)

  • No Dirac-3 photonic quantum annealer or quantum-sampling access is available; only a classical mixed-integer optimizer or heuristic QUBO solver can be run, so the backtest cannot test the paper's claimed photonic-hardware advantage.
  • The exact Jensen-Kelly-Pedersen 13-factor return library is unavailable. Reconstructed factor sleeves from platform equity, price, and fundamental data test an analogous allocation process rather than the authors' exact inputs.
  • A full Soft Actor-Critic comparison may be computationally feasible only as a separately trained classical model, but it will not reproduce the paper's training environment, tuning protocol, or seed-level comparison exactly.

The figures below measure what we could run, not the paper's own method, so they are not evidence for or against its claim.

Our own audit found this run does not follow the paper faithfully (15)

  • Equation 1 portfolio objective (invalidates: All paper-reported beta-sweep performance optima and solver rankings as direct expectations for this strategy.)
  • Equation 2 QUBO objective (invalidates: Paper energy minima and all reported solver-comparison results based on its unconstrained binary domain.)
  • Standalone SAC softmax weights (invalidates: All standalone SAC performance, concentration-collapse, seed-instability, and efficient-frontier results.)
  • Uniform-allocation HHI reference (invalidates: Paper HHI levels, HHI ranges, and cross-solver concentration comparisons as direct targets for this six-sleeve strategy.)

11 further finding(s) are described in the note.

These are our findings about our own implementation, not criticisms of the paper. Read the figures below as a description of what we ran.

Jan 2020Total 11.4%Jul 2024
Sharpe
0.25
Total Return
11.4%
Max Drawdown
-18.7%
CAGR
2.4%
Volatility
9.9%
Beta vs SPY
0.28
Trades
4,873

What the paper reports for its own strategy

  • Dirac-3, best overall configuration (beta1=0, beta2=1): Sharpe 0.760, Sortino 0.841, Calmar 0.567, MDD -3.47%, CVaR5% -1.278%, annual return 1.70%, annual volatility 2.23% (seed mean of 3 seeds; 33-month out-of-sample test window from the 164-month panel; JKP returns net of imputed costs plus 0.002 proportional cost in the objective)
  • Dirac-3 (beta1=0, beta2=0.5): Sharpe 0.721, Calmar 0.538, MDD -2.63% (lowest in both sweeps), CVaR5% -1.107%, annual return 1.39%, annual volatility 1.96%
  • Dirac-3 beta1 sweep (beta2=0): peak Sharpe 0.715 at beta1=2 (vol 2.05%, skewness +0.265); best Calmar 0.376 at beta1=1 with MDD -3.37%, Sharpe 0.697, CVaR5% -1.015%
  • Gurobi beta1 sweep: peak Sharpe 0.718 at beta1=5, Sortino 0.764, annual volatility 1.81%, MDD -4.33%, CVaR5% -1.049%, HHI 0.157, CR-5 0.816
  • Gurobi joint sweep: best CVaR5% -0.863% at (0,1) (Sharpe 0.565, Calmar 0.190, MDD -5.37%); peak Sharpe 0.715 at (0,50) with CVaR5% -0.911%, MDD -4.95%, Calmar 0.257, annual return 1.09%
  • Dirac-3 (beta1=1, beta2=10): CVaR5% -0.895%, Sharpe 0.655, MDD -3.41%, Calmar 0.381

A 0.045 Sharpe gap over 33 months cannot rank a photonic annealer against branch-and-bound. Each portfolio was selected as the best of 16 joint configurations run for its own pipeline on the same monthly-return window. The paper's tables then reverse the ranking when the skewness term is removed.

Sahoo, Tee and Griffin acknowledge the narrow result in the abstract. Photonic hardware "can locate superior risk-return topologies within a narrow operating range", while mixed-integer programming "remains superior for risk-constrained mandates requiring tight tail-risk control and cross-seed stability". Yet the conclusion says Dirac-3 "demonstrates clear statistical advantages in out-of-sample risk-adjusted returns", and Exhibit 11 calls Dirac-3 (0,1) the "Best overall risk-adjusted configuration". That claim deserves scrutiny. Quantum supremacy is beside the point.

The paper also says "No single pipeline dominates" and describes its mandate recommendations as "starting points for further validation". No t-statistic or standard error for any Sharpe difference appears anywhere. The authors answer the short window with operational controls: composition alerts, HHI thresholding above 0.30, and post-optimization weight capping. But their proposed HHI < 0.25 ceiling would disqualify the recommended configuration because (0,1) has HHI 0.340. We did not find a backtest of the capped version.

One QUBO, three pipelines

The investable universe contains 13 value-weighted, capped US equity anomaly factors from the Jensen-Kelly-Pedersen library: value, momentum, short-term reversal, quality, profitability, investment, low-risk, low-leverage, seasonality, accruals, debt issuance, idiosyncratic volatility and size. The dataset has 164 months of monthly returns and no missing observations. Returns are already net of costs imputed at the JKP level.

Allocation is long-only across the 13 series. Weights lie within [0.00, 0.60], sum to one, and are smoothed with eta = 0.10. The objective combines a rolling volatility penalty beta1, a rolling Fisher skewness reward beta2, and a 0.002 proportional charge on L1 turnover. Any return comes from shifting the portfolio each month toward the anomaly premia favored by the optimizer.

Dirac-3, an entropy-based photonic annealer, and Gurobi share nearly the entire pipeline. A policy network produces an expected-return vector alpha, alongside a 24-month Ledoit-Wolf covariance estimate. Those inputs are assembled into a QUBO, the quadratic unconstrained binary optimization form consumed by both solvers. The solution selects a binary subset of factors. A restricted mean-variance solve then assigns weights to that subset before clipping and renormalization.

SAC, the Soft Actor-Critic reinforcement learning agent, bypasses those stages. It applies a softmax directly to the raw policy action and uses a 6-month rolling window for the penalty terms. Dirac-3 against Gurobi therefore isolates the solver while holding the pipeline fixed. SAC represents a separate investment process.

The chronological 80/20 split leaves 131 months for training and 33 for testing. According to the paper, "all reported performance metrics are computed exclusively over the test window." The abstract instead describes evaluation "across 164 months test window". Those statements differ, and every headline result rests on the 33-month sample.

Can 0.045 Sharpe identify the better solver?

Dirac-3 reaches its global optimum at beta1 = 0, beta2 = 1. Averaged over three seeds, the portfolio records Sharpe 0.760, Sortino 0.841, Calmar 0.567, MDD -3.47%, CVaR5% -1.278%, annual return 1.70% and annual volatility 2.23%. Gurobi's best-Sharpe joint configuration is (0, 50). Relative to that peak, Dirac-3 delivers Sharpe +0.045, Sortino +0.100, Calmar +0.310 and MDD 1.48 percentage points shallower, while CVaR5% is 0.367 points worse.

The other sweep changes the order. With beta2 = 0, Gurobi peaks at Sharpe 0.718 when beta1 = 5, ahead of Dirac-3's 0.715 at beta1 = 2. Gurobi also posts the best CVaR5% in that sweep, -0.983%. At beta1 = 5, the setting where Gurobi peaks, Dirac-3 produces 0.551.

Dirac-3's 0.551 is the toughest figure for the narrow-window argument. The paper's peak-metrics exhibit prints the result plainly, awarding Gurobi Sharpe and CVaR while Dirac-3 wins on Calmar and drawdown. Dirac-3 clears Gurobi at the peak only after the skewness term enters.

We found no standard error for a Sharpe difference, no t-statistic, and no adjustment for the 16 joint configurations searched per pipeline. The full search covers 24 per pipeline across both sweeps and 48 joint runs in total across the three. Reported cross-seed dispersion addresses another question: solver reproducibility. It does not estimate sampling error from 33 monthly returns. A 0.045 gap on that sample lies inside sampling noise. So does 0.003. Both figures were chosen as maxima.

The repeated configurations raise another problem. Identical setting (0, 0) appears in both sweeps with different outcomes: Dirac-3 Sharpe 0.630 against 0.646, Gurobi 0.563 against 0.603, and SAC 0.492 against 0.536. Gurobi's 0.040 discrepancy on the repeated setting nearly matches the headline advantage attributed to the solver.

An accruals bet at the sweet spot

At the Dirac-3 optimum, accruals carries 25.9% and investment 19.6%. Together they make up 45.5% of the portfolio. HHI reaches 0.340, versus an equal-weight reference of 0.077, while CR-5 is 0.895. No Dirac-3 configuration across either sweep has CR-5 below 0.825. Its mean HHI in the beta1 sweep is 0.228, compared with 0.152 for Gurobi.

Monthly L1 turnover over the tabulated beta1 grid ranges from 0.017 to 0.028 for all three pipelines. The paper flags one exception: Dirac-3 at beta1 = 0.5 reaches 0.044. Similar turnover across the three leaves composition as the source of the Sharpe and CVaR spread rather than trading intensity. The tables show it directly.

The authors recognize the concentration. Their monitoring protocol calls for post-optimization weight capping whenever HHI exceeds 0.30, naming (0, 0) and (0, 1) as examples. Their return-seeking recommendation is (0, 1). Capping would materially alter the trade because the 45.5% assigned to accruals and investment generates the Sharpe.

Both of Dirac-3's best configurations in the joint sweep occur at beta1 = 0, with no explicit volatility aversion in the objective. Its primary-sweep window appears at beta1 = 1 and 2.

Gurobi's value appears in tail risk

Gurobi's best CVaR5% in either sweep is -0.863% at (0, 1). Its Sharpe there is only 0.565 and its Calmar 0.190. Across all 16 joint configurations, HHI remains between 0.135 and 0.228, with no factor exceeding 17.7%.

At beta1 = 20, Gurobi's cross-seed Sharpe standard deviation stays strictly below 0.05. Dirac-3 remains below 0.08 throughout its beta1 grid. The classical pipelines received five seeds, while Dirac-3 received three seeds because of per-run photonic cost, as the authors state openly. That imbalance weakens the paper's stability comparison.

The paper's reproducibility recommendation gives Gurobi a return standard deviation of 2.10pp at beta1 = 20. Text discussing the same configuration reports 0.10 percentage points.

These results come from annualized returns of roughly 0.84% to 1.85% with volatility between 1.5% and 3.7%. The paper's best configuration earns 1.70% annually at 2.23% volatility. It reports no capacity or leverage analysis.

Where SAC broke

At beta1 = 0, beta2 = 20, SAC drops to Sharpe 0.088. MDD reaches -14.64%, CVaR5% -2.237%, HHI 0.593 and CR-5 0.985. Under the same setting, Gurobi retains Sharpe 0.617, CVaR5% -1.230% and HHI 0.140. SAC's CR-5 rises to 0.996 at beta2 = 50.

Another collapse occurs at (1, 20), where low-risk alone takes 45.5%, the study's largest single-factor weight. The skewness reward was meant to reduce tail risk. Instead, this configuration produces the deepest drawdown reported anywhere in the paper, -14.64%.

Reward design and constraint enforcement account for the failure. The policy network supplying alpha to the QUBO pipelines belongs to the same class of object. Dirac-3 and Gurobi place a selection stage and bounded mean-variance projection between that network and the portfolio. SAC lacks those layers.

Even when SAC functions, it trails. Its best joint Sharpe is 0.602 at (1, 0.5), and its best beta1-sweep Sharpe is 0.572. The highest nominal return, 1.85% at beta1 = 2, arrives with MDD -12.23% and CVaR5% -2.388%. Gurobi at the same setting returns 1.34% with CVaR5% -1.334%.

Our adaptation, with limits

We had no access to Dirac-3 and performed no quantum sampling of any kind. None of our results tests the photonic claim. We also lacked the JKP factor return library and did not train an SAC comparison. Our adaptation used a classical branch-and-bound solver with the same QUBO structure.

We built six long-only large-cap US equity factor sleeves: value, quality, profitability, investment, momentum and low volatility. The universe comprised the 500 largest non-ADR names. Each month, a cardinality-constrained QUBO selected 2 to 4 sleeves. A bound-constrained mean-variance solve assigned weights with a 60% sleeve cap and 10% stock cap. We applied eta = 0.10 smoothing and rebalanced monthly at the next day's close. Without a trained policy network, alpha was each sleeve's trailing 12-month mean. Costs were 0.002 on stock-level L1 turnover plus $0.004 per share.

Our run covers 2020-01 to 2024-07. It produced Sharpe 0.25, Sortino 0.27, volatility 9.94%, max drawdown -18.67%, Calmar 0.13, total return 11.36%, beta 0.28 to SPY and 4,873 trades. The paper's Gurobi peak in the beta1 sweep has Sharpe 0.718, volatility 1.81% and drawdown -4.33%. Our Sharpe is roughly a third of theirs, with five times the volatility.

The objects differ. Their allocation spans 13 factor return series with annualized volatility from 1.5% to 3.7%. Ours holds long-only equity sleeves. The substitution mechanically produces our 9.94% volatility and -18.67% drawdown, compared with the reported Gurobi MDD of -4.33%.

The paper supplies no calendar start or end date for its 33-month test window, leaving regime overlap with our 2020-01 to 2024-07 run unknown. Two choices in our adaptation push in the same direction. A trailing 12-month sleeve mean is probably much weaker than a trained signal and likely follows leadership into reversals. Selecting 2 to 4 sleeves from six is also much coarser than binary selection across 13 factors, costing us breadth.

Gurobi's HHI ranges from 0.136 to 0.170 in the beta1 sweep, with mean 0.152. Across all 16 joint configurations, it spans 0.135 to 0.228. Our figures measure our implementation. They are one automated pass rather than a verdict on the authors' work.

A repeat of the same grid on a second factor library or a held-out block of months would change my mind, provided Dirac-3 and Gurobi peaks were compared with standard errors accounting for the 16 configurations searched per pipeline. As presented, the paper uses one factor library, follows one historical path, and gives no standard error for any Sharpe difference. Its lasting results are the SAC failure map and Gurobi's HHI band of 0.135 to 0.228. Both are useful. Neither establishes an advantage for quantum hardware.

Our backtest stops at 2024-07-01, and everything after that date is deliberately left untouched so the same strategy can be checked out of sample later.

How our backtest worked

The steps the code we ran actually executed, from its strategy card. Ours, not the paper's — it is one automated implementation of the idea, not the authors' own.

At each annual universe refresh:
  Select the 500 largest non-ADR US stocks with at least 252 price observations.

At each calendar month-end:
  1. Build six point-in-time factor scores using prices and filings available by the signal date.
  2. Winsorize each component at the 1st/99th percentiles and combine available components,
     requiring at least 50% component coverage.
  3. For each sleeve, select its top score quintile and construct capped stock weights.
  4. Reconstruct causal monthly sleeve returns; require a complete 24-month panel.
  5. Set alpha to each sleeve's trailing 12-month arithmetic mean return.
  6. Estimate a 24-month Ledoit-Wolf covariance matrix and stabilize it with epsilon I.
  7. Form QUBO terms:
       Q_ii = -alpha_i + 0.01 + Sigma_ii + delta_i
       Q_ij = 2 * Sigma_ij
     where delta_i is +0.10 for a prospective entry and -0.10 for retention.
  8. Use deterministic branch-and-bound to select 2-4 of the six sleeves.
  9. Over the selected subset, solve a long-only mean-variance allocation with
     sum(weights) = 1 and each sleeve weight &lt;= 0.60.
 10. Smooth targets: w_t = 0.90 * w_previous + 0.10 * w_target.
 11. Aggregate sleeve holdings to stocks, enforce the 10% stock cap, and retain cash
     if full investment is infeasible or an observed execution price is unavailable.
 12. Execute at the next trading day's close; skip orders lacking an observed close.
 13. Deduct stock-level transaction costs from full-L1 changes in actual holdings.

If covariance data, the binary solution, or the restricted allocation is invalid:
  retain the preceding allocation and place no rebalance orders.