For a battery paying 25 EUR per cycle, the forecast target matters more than adding temporal blocks. The operator earns the gap between one hour's price and a later hour's price, after round-trip losses and cycling cost. Two forecast sets can have identical RMSE yet select entirely different charge and discharge hours. One may lose money on a day when the other earns. Lipiecki, Kourentzes and Weron therefore change the forecasting target rather than the loss function.

From blocks to 276 spreads

The hierarchy in the authors' earlier work was temporal. It combined 24 hourly prices with non-overlapping blocks of 2 to 24 hours, producing 36 block series in all. They call the method Block THieF. Its gains reached 5.5% for hourly prices and 13.4% for baseload prices.

Spread THieF replaces those aggregation levels with price differences. Every one of the C(24,2) = 276 intraday price spreads is a linear combination of two hourly prices. Stack the 276 spreads with the 24 hourly prices and the result is a 300x24 summing matrix. Each spread series receives its own direct forecast, built from lagged spreads and the corresponding differences in day-ahead load and renewable-generation forecasts.

MinT reconciliation pulls the 300 base forecasts back onto the 24-dimensional coherent subspace. The weighting comes from a 300x300 error covariance with Schäfer-Strimmer diagonal-target shrinkage, re-estimated daily. Directly fitted spread models therefore influence the hourly price path and contribute information about the day's relative shape that separately fitted hourly models would miss.

Three base architectures cover different levels of model capacity. ARX uses OLS on 20 regressors after an asinh transformation. NARX has one hidden layer, five tanh neurons and a committee of ten nets. The final model is zero-shot TabPFN-2. The markets are Germany (EPEX-DE) and Spain (OMIE), using ENTSO-E load and renewable forecasts alongside TTF gas and API2 coal futures as daily inputs. The training window contains 1085 usable observations and rolls forward one day at a time. The out-of-sample period runs from 01.01.2021 to 31.12.2025, with T = 1826 days.

Reconciliation only helps when the targets matter

The actionable result comes from comparing the two hierarchies on the strongest base model.

Block THieF damages TabPFN. In Germany the changes are -1.1% RMSE, -1.9% MAE, -2.0% daily RMSE and -1.9% daily MAE. Spain records -0.2% RMSE and -0.8% MAE. These are small losses, though the direction is wrong throughout all four German measures. Apply Spread THieF to the same model and German RMSE falls from 27.30 to 25.69, a 5.9% gain. Daily RMSE improves 7.3% and daily MAE 5.4%. Three of four measures land between 5.4% and 7.3%. Only hourly MAE is weak in the Germany TabPFN row of Table 1, improving 1.8% from 15.49 to 15.22.

Extra temporal aggregates added noise to an already capable forecaster. Differences supplied signal.

The structure can also substitute for model capacity near the lower end of the range. In Spain, ARX with spread reconciliation posts RMSE 21.09 and MAE 14.56, beating an unreconciled NARX at 22.25 and 15.53 on all four accuracy measures. Across markets and error measures, skill scores against Base range from 1.8% to 19.7%. The German ARX produces the upper end. CPA tests reject equal predictive ability in favour of Spread THieF at the 1% level for every model-market pair, against Base and Block THieF, under absolute and squared loss.

A large error reduction buys a small profit gain

The economic test uses a 1 MWh battery with one charge and one discharge each day, a 25 EUR round-trip cost and an initial state of charge of zero. Trades clear at the day-ahead auction. The tested efficiencies are 1, 0.95 and 0.9.

For Germany at eta = 1, Spread THieF earns 153,274 / 154,265 / 153,485 EUR across ARX / NARX / TabPFN. Oracle perfect foresight earns 166,424, while a seasonal naive forecast earns 132,684. Spain produces 81,704 / 81,859 / 81,000 against Oracle at 95,336. At eta = 0.9, the best German Spread THieF result is NARX at 108,681, compared with Oracle at 120,827.

Holding the baseline fixed shows how forecast accuracy converts into money. German ARX at eta = 1 cuts daily MAE by 19.7% versus Base. Profit moves from 150,281 to 153,274 EUR, an increase of 2,993 EUR, or 2.0% over five years for a 1 MWh battery. A fifth off the error measure buys two percent more money. Compared with Block THieF on the same base model, profit rises by roughly 1.1% to 9.7%, while relative opportunity cost declines by 0.9 to 6.9 percentage points. The abstract places the 19.7% accuracy figure beside profit gains of up to 10.4%, both measured against unreconciled hourly forecasts.

TabPFN records the lowest errors in both markets. Under Spread THieF, its German RMSE is 25.69 versus NARX's 26.45. Yet NARX with Spread THieF earns the most in all six market-efficiency combinations. In Germany at eta = 1, NARX makes 154,265 against TabPFN's 153,485. The authors acknowledge the distinction in their abstract, writing that exploiting coherent relationships between economically relevant forecasting targets "can improve both predictive accuracy and decision value", while "better statistical forecasts do not necessarily imply better economic decisions". Both statements survive the results. Accuracy and profit move together within each architecture, while the ordering reverses across architectures.

The paper also includes the Serafin and Weron approach as a fourth comparison, an unreconciled spread-only benchmark. At eta = 1 in Germany, it fails to beat hourly forecasts consistently. NARX benefits, earning 149,662 against 148,743. TabPFN loses 3,535 EUR, falling to 147,815 from 151,350. Reconciliation and blending generate the payoff.

Spain's relative opportunity cost under Spread THieF ranges from 14.1% to 22.7% of Oracle profit, compared with 7.3% to 11.1% in Germany. Yet Spanish forecast errors are lower in level terms. TabPFN RMSE is 19.12 there, against Germany's 25.69. Our reading is that flatter Spanish price profiles make the within-day ordering harder to identify. The paper limits its own explanation to the observation that error consequences depend on whether they alter "the economically relevant ordering of prices within the day".

Five-year totals without dispersion

The profit table gives five-year EUR totals, relative opportunity cost and a rolling 180-day ROC curve. It provides no daily P&L distribution, no volatility and no test of whether profit differences of 1.1% to 9.7% are distinguishable from zero. The authors say the elevated rolling ROC for TabPFN in Germany during the second half of 2023 is "largely driven by a single trading decision". On 2 July 2023, the German daily average was -53.9 EUR/MWh and hour 15 printed -500 EUR/MWh. Base TabPFN selected hour 15 for the sale and lost over 300 EUR on a 1 MWh battery. A five-year total needs an error bar when one day can move a model's rolling cost by that amount. At eta = 1, only 780 EUR separates the best and third-best Spread THieF runs in Germany.

The authors describe the storage model as stylized. It assumes price-taking, a single daily charge-discharge cycle and a fixed round-trip cost. State of charge begins at zero and does not carry over, a consequence of the setup rather than part of their concession. The model has no intraday or balancing exposure and no market impact. Its 25 EUR charge is a single lumped estimate for capex, O&M and degradation. Battery efficiency enters at the decision stage only, never inside the hierarchy. Once eta falls below 1, the unreconciled spread benchmark disappears because an efficiency-adjusted margin cannot be recovered from an unadjusted spread. One of four benchmarks is therefore absent at eta = 0.95 and 0.9.

The estimation design raises a separate attribution question. Both the 300x300 covariance and the shrinkage intensity come from in-sample forecast errors over the same rolling window used for the base models. MinT requires that construction, and its weights are set before the forecast day, so there is no look-ahead. The authors compare the shrinkage target with Ledoit-Wolf constant-correlation shrinkage and obtain qualitatively similar results. They provide no run that separates gains from the weights from gains due to the spread structure. Nor do they evaluate the 276 individual spread forecasts on standalone accuracy. The paper's reason is that trading relevance differs across spread pairs.

We could not run this ourselves. Reproduction requires hourly EPEX-DE and OMIE auction prices together with point-in-time ENTSO-E load and renewable forecasts, neither of which is in our data. A commodity ETF substitution would replace within-day power spreads with a different problem entirely.

When the decision depends on a difference, that difference belongs in the hierarchy. The choice of reconciliation targets determines whether a strong model benefits. For TabPFN in Germany, block aggregates cost between 1.1% and 2.0% across the four measures, while spreads gain between 1.8% and 7.3%. A distributional test of daily P&L would change my view of the economics. Until then, I would not size against a five-year total drawn from a market containing the July 2023 episode.