At alpha 60% and 50%, the best price forecast loses the battery trade. Lipiecki and Weron trace the full chain from day-ahead forecasting to bidding across Germany, Poland and Spain over 1,826 test days. TabPFN-3 records the lowest CRPS in every market. DDNN-JSU leads on none of the error measures, yet earns more once limit prices tighten.

The signal comes first. The trade then follows easily.

What reaches the model?

Before each day-ahead auction closes, the trader already has a substantial information set. ENTSO-E publishes tomorrow's day-ahead load forecasts, along with day-ahead onshore wind and solar generation forecasts. Germany also gets offshore wind. The latest closing prices for nearest-to-delivery EUA carbon, TTF gas, Brent and API2 coal are available as well. Lipiecki and Weron lag those prices by two days, matching the information a bidder could actually use.

TabPFN-2, TabPFN-3, Mitra+CP and the two benchmarks receive price history directly. Prices enter at lags of 1, 2, 3 and 7 days. Load and renewable generation forecasts use lags 0, 1 and 7. A day-of-week column completes the table. Each model then produces a predictive distribution for tomorrow's 24 hourly prices.

The economic logic is familiar. The merit order sets prices, with residual load, demand after wind and solar, and marginal fuel cost doing most of the work. Tomorrow's residual load is public before bids are due.

Two established electricity price forecasting models provide the comparison. LEAR+CP is a Lasso regression with 250 parameters, averaged across calibration windows of 56, 84, 728 and 1,092 days. Its quantiles come from conformal prediction using a one-year rolling window of out-of-sample errors.

DDNN-JSU uses two hidden layers. For each hour, its output layer supplies the four parameters of a Johnson's SU distribution. The model has 0.4M parameters and is retrained every day on the latest 1,449 observations. Annual hyperparameter tuning runs through four Optuna studies of 500 trials each.

The challengers are nine foundation model variants drawn from five families: three Chronos-2 variants, Moirai-2, TimesFM-2.5, three TabPFN variants, and Mitra. Zero-shot means the weights receive no task-specific updates, although observations from the target market still appear as in-context data. TabPFN-2 and TabPFN-3 each receive 1,449 observations, while Mitra gets 1,085. Mitra's conformal stage also uses a year of out-of-sample errors.

All three markets are hourly: Germany (BZN|DE-LU), Poland (BZN|PL) and Spain (BZN|ES). The test runs from 1 January 2021 to 31 December 2025, covering 1,826 days. Earlier work, the authors argue, suffered because "the evidence is based predominantly on short test periods". Most of those studies used a single test year. This paper uses five.

A 1 MWh battery, with 25 EUR charged per trade

Charge efficiency is 0.95, as is discharge efficiency, putting the round trip at about 0.9. The battery begins empty. Each completed trade incurs 25 EUR of round-trip operating cost, covering investment expenditure, fixed and variable O&M, and degradation. The strategy allows one cycle per day.

For each day, the trader chooses the buy and sell hours that maximise predicted profit under the median forecast. The forecast profit already includes the 25 EUR charge. An order is submitted only when that profit is positive, which requires the predicted spread to cover the round-trip cost.

Execution uses a loop bid, EPEX's coupled two-leg order. Both legs fill together or neither fills, removing partial-execution risk by construction. The paper studies two policies. Unlimited bids become market orders whenever predicted profit is positive. Quantile-based bids impose limit prices.

For the buy hour, the limit equals the (1+alpha)/2 quantile. For the sell hour, it equals the (1-alpha)/2 quantile. Alpha takes values of 90, 80, 70, 60 and 50%. At alpha 90%, the strategy bids to buy at the 95% quantile and sell at the 5% quantile, making execution relatively easy. Tightening alpha reduces fills and increases selectivity. The paper's main economic result appears as that risk-appetite dial moves.

TabPFN wins the forecast contest

TabPFN wins.

TabPFN-3 delivers the lowest CRPS in every market: 9.15 for Germany, 9.61 for Poland and 8.15 for Spain. DDNN-JSU records 10.70, 10.96 and 9.30. Those differences correspond to reductions of roughly 14.5%, 12.3% and 12.4%.

Every TabPFN variant beats DDNN-JSU at the 1% level on daily MAE, RMSE and CRPS in all three markets. The authors use the one-sided multivariate Giacomini-White test of conditional predictive ability. They say the result also holds against LEAR+CP, although they do not tabulate it.

The general-purpose time series models fare much worse. Chronos-2 and Chronos-2-synth beat DDNN-JSU on all three measures in Poland, with no corresponding win elsewhere. In Poland, Chronos-2-synth posts MAE 14.61, RMSE 24.23 and CRPS 10.62. DDNN-JSU comes in at 14.97, 24.70 and 10.96.

Moirai-2 and TimesFM-2.5 never beat it. Their German CRPS values are 16.20 and 14.26, compared with DDNN-JSU's 10.70. TimesFM-2.5 is the study's largest model at 200M parameters. TabPFN-2 has 11M and still beats all five general-purpose time series models, meaning the three Chronos-2 variants, Moirai-2 and TimesFM-2.5.

The paper locates the difference in the covariate channel. TabPFN consumes the full table directly. Moirai-2 accepts neither exogenous variables nor native multivariate forecasts, so the authors treat prices as a univariate hourly series. It receives a 364-day context (8,736 hours) and produces a 24-step horizon. TimesFM-2.5 can use the load and wind forecasts only through XReg, an external linear regression fitted to residuals. With tomorrow's residual load already public, the authors call a univariate design "a serious limitation in EPF".

Synthetic pretraining matters here. TabPFN is trained entirely on synthetic data. Chronos-2-synth is synthetic-only as well, and matches or beats the default Chronos-2 on MAE and CRPS in Germany and Poland.

Release timing complicates the comparison. Moirai-2 arrived in August 2025, TimesFM-2.5 in September 2025, Chronos-2 in October 2025 and TabPFN-3 in May 2026. The contamination caveat therefore covers the mixed-pretraining models: Chronos-2 default, Chronos-2-small, Moirai-2 and TimesFM-2.5. According to the authors, it "does not apply to models pretrained exclusively on synthetic data, such as the TabPFN variants and Chronos-2-synth".

Within that synthetic-only group, TabPFN beats DDNN-JSU across Germany, Poland and Spain. Chronos-2-synth manages the same result only in Poland. In Spain it trails DDNN-JSU, with MAE 13.23 against 12.66.

Profit changes hands at alpha 60%

The abstract concedes the economic reversal. "This statistical dominance does not translate directly into economic dominance," the authors write. TabPFN "performs best under unlimited bids and riskier quantile-based strategies, whereas the Distributional Deep Neural Network benchmark is more profitable when risk tolerance is lower." Their conclusion adds that foundation models "cannot universally replace market-specific models, and their value depends on both model architecture and the decision problem".

That reversal belongs to the paper. Our narrower concern is the degenerate quantile spreads visible in one model's profit column, together with the missing test of profit differences.

With unlimited bids, a TabPFN variant leads all three markets, though the margins are small. In Germany, TabPFN-TS-3, the family's third variant, earns 132,716 EUR over five years. DDNN-JSU earns 130,434, while Oracle perfect foresight reaches 143,474. TabPFN-TS-3 has German CRPS of 9.33 against 10.70, about a 12.8% reduction. The corresponding profit advantage is roughly 1.8%.

Poland's unlimited-bid ranking starts with TabPFN-3 at 106,685 EUR. Chronos-2-synth follows at 106,119, then TabPFN-TS-3 at 106,060. Across Spain's 1,826 days, the three TabPFN variants finish within 350 EUR of one another. The authors report no test of those profit differences.

Limit prices alter the order. TabPFN-TS-3 leads Germany and Poland at alpha 90, 80 and 70%. DDNN-JSU takes over at 60% and 50%. In Germany, DDNN-JSU earns 95,842 and 86,123 EUR, while TabPFN-TS-3 makes 94,907 and 85,416.

Spain turns earlier. DDNN-JSU wins at every setting except alpha 90%, and at alpha 50% earns 46,714 EUR against 42,150. Across all three markets, DDNN-JSU never has the lowest error on any statistical measure. Under the conservative policies, it nevertheless earns the most everywhere.

The split also appears inside the TabPFN family. TabPFN-TS-3 earns more than TabPFN-3 at every alpha examined, while TabPFN-3 retains the lowest CRPS in all three markets.

The authors identify the limitation themselves: quantile-based trading profits fall outside proper scoring functions. The resulting economic ranking applies to one battery, one cost assumption and one bid type. The paper attributes this point to the recent review by Hirsch and Ziel that it cites.

Evidence that would shift the verdict

I accept the statistical result. It spans five years and three markets with distinctly different fuel stacks. In 2025, Poland generated 33.0% from hard coal and another 19.4% from brown coal. Spain generated 20.8% from nuclear and 21.0% from solar. Across every market, all three TabPFN variants beat DDNN-JSU at the 1% level. The earlier studies reviewed by the authors usually covered a single test year; this study covers five.

I treat the profit ranking with more caution because of a pattern in the table. TimesFM-2.5 reports the same German profit at alpha 90% and 80%: 98,639 EUR in both cases. Another repeat occurs at 70% and 60%, with 77,542 both times. The paper leaves those repetitions unexplained. My reading is that the model's quantile spread ceases to move continuously as alpha tightens, turning the risk setting into a step function. Other models change winner between alpha 70% and 60%. That switch needs an error bar, and the paper supplies none.

The simulation includes no exchange fees, bid-ask spread, market impact or intraday and balancing settlement. Its only charge is the 25 EUR round-trip cost. Loop bids remove partial fills by assumption. The profit column therefore compares forecast policies inside an idealised clearing mechanism.

The benchmark implementations also differ from the published originals. Changes include shorter LEAR windows, 500 instead of 1,000 tuning trials, no variable selection, AdamW replacing Adam, a 512-neuron cap, and no batch normalization. The authors describe their DDNN-JSU implementation as "computationally faster while retaining comparable predictive" accuracy.

Their LEAR may be stronger in one respect. It chooses lambda through seven-fold cross-validation instead of the original's "faster but less accurate" LARS Lasso. The judgment that these changes produce a slightly weaker benchmark is mine. The authors make no such claim.

Moirai-1 and Moirai-MoE were removed after initial tests. The authors selected Moirai-2's univariate multi-horizon configuration because it "performed better in our preliminary tests". They do not identify the sample used for those tests. Since this family remains among the weakest models on MAE, RMSE and CRPS in all three markets, the choice raises a small concern rather than a large one.

We did not run this. We do not hold day-ahead wholesale power prices for Germany, Poland or Spain. The arbitrage leg requires auction settlement mechanics and battery constraints that an equity or futures substitute cannot reproduce. Testing a supported asset class would amount to a different strategy.

In Germany, TabPFN-3 improves CRPS over DDNN-JSU by 14.5%. Beside it sits an unlimited-bid profit advantage below 2% for TabPFN-TS-3. Tight limit prices reverse even that sign. Score the policy.