For a German or Austrian imbalance-price model, I would start with VWAP for the target quarter-hour and the products delivering after it. That is the finding in this paper by Yu, Cremer, Pinson, Kazempour, Semmelmann, Matsumoto and Bunn most likely to travel. The other results depend more on the country, and the evaluation ends at pinball loss rather than money earned.

The test behind the feature

Traders close positions in the continuous intraday market as delivery nears. The authors conjecture that positions left open, together with physical forecast errors arriving after gate closure, shape the imbalance price. On that account, intraday trades should offer an early read on settlement.

They ask how much of the information behind imbalance-price formation is already reflected in the EPEX continuous intraday orderbook. The test uses EPEX Spot trades and ENTSO-E data for Germany and Austria from January 2022 through January 2025. Products deliver in 15-minute intervals; German traded volume is roughly ten times Austrian volume. If the orderbook has absorbed the relevant information, it should forecast the imbalance price on its own. Any gap indicates information the intraday market has yet to price.

The models forecast the 0.1, 0.5 and 0.9 imbalance-price quantiles at 15, 60 and 180 minutes ahead. The learners are a feedforward artificial neural network (ANN) and linear quantile regression (LQR). Executed trades go into 15-minute buckets, then into one of three summaries:

Each summary uses either the target product alone (Self), the next 1, 4 or 12 products as well (Neighbor), or the matching delivery slot in the other country (Country). Validation determines the Neighbor count. Testing covers 2024 in three four-month rolling folds, with three seeds per run and paired Diebold-Mariano tests. The main measure is AQL, or pinball loss averaged across the three quantiles, in EUR/MWh.

VWAP-Neighbor lands in the statistically strongest group in 11 of 12 metric-horizon cells for Germany and 10 of 12 for Austria. At 15 minutes, its German AQL is 25.88, against 30.03 for OHLCV-Self and 27.04 for VWAP-Self. ANN beats LQR in 17 of 18 orderbook cases; German OHLCV-Self is the tie.

Why look beyond the target product?

VWAP already has a history as an intraday feature, and the paper cites Hirsch and Ziel on cross-product effects in intraday price forecasting. Here the forecast target is the settlement price. The authors' proposed mechanism is that traders spread large orders across adjacent quarter-hours, leaving traces of a position being closed in a neighbor's trades. The losses fit that account, as the authors acknowledge, without identifying it.

The gain from neighbors varies by market. In Austria, adding neighbors reduces loss for every VWAP and LMP setting, reaching 14.4% for LMP with 12 products. Liquidity offers another possible explanation: an Austrian product has roughly ten times fewer trades than a German one, so nearby products may help wash out noise. German OHLCV goes the other way. Adding 4 or 12 products raises its loss by up to 40%. The choice of neighbor count therefore needs validation for each representation, as the authors do in each fold.

One sentence in the paper goes further than its table. It says VWAP-Neighbor has the lowest RMSE at every horizon in both countries. At 15 minutes in Austria, LMP-Neighbor records 331.60 to VWAP-Neighbor's 332.64. The difference is 1.04 EUR/MWh; VWAP-Neighbor still leads the Count column at 10 of 12.

The lead at longer horizons

Mostly, yes.

At 180 minutes, VWAP-Neighbor still beats VWAP-Self. German AQL moves from 25.88 through 28.43 to 29.68 across the three horizons, about 15% worse by 180 minutes. Yet the advantage over VWAP-Self remains 4.3% at 15 minutes (27.04 versus 25.88) and 3.6% at 180 (30.78 versus 29.68). Austria shows 6.8% (28.81 versus 26.85) and 7.2% (31.72 versus 29.43), respectively. Across every orderbook case, coverage error for the 80% interval lies between 0.018 and 0.108. The authors interpret the differences among representations as accuracy differences rather than calibration differences. A 0.108 error on an 80% interval amounts to about 11 points of coverage.

Fundamentals split the markets

The authors add four nested fundamental-data sets, beginning with renewables and extending to cross-border positions and system imbalance. Their stated result cuts in opposite directions: combining orderbook and fundamentals reduces testing loss in Germany and increases it in Austria. All four combined German models beat VWAP alone; all four Austrian additions raise test loss. Both directions have triple-star DM marks. The bar values are confined to the figures, with neither result given as a number in the text.

For the paper's question about information flow, the losses suggest that Austria's orderbook already reflects what the fundamentals contribute, while Germany's leaves some physical information unpriced. The authors attribute the Austrian pattern to positions still open after intraday trading, which the orderbook can reflect directly. Germany's large wind and solar share, in their account, brings uncertainty after gate closure that trades cannot fully anticipate. They say the losses are consistent with this explanation, without identifying it.

Austria's thinner book makes its result surprising. The German tail ablation complicates the picture too: among the 690 positive and 144 negative German sample-horizon pairs beyond ±1000 EUR/MWh, VWAP alone has the lowest loss. The combined model's German edge comes from ordinary quarter-hours.

Austrian tails are heavier. Austria has 666 negative extreme sample-horizon pairs, compared with Germany's 144. At 15 minutes, Austrian VWAP-Neighbor RMSE is 332.64 against Germany's 198.01, even though Austrian MAE is lower (69.12 versus 77.51). Excluding 2022 produces the lowest VWAP loss on extremes in both countries; training on all data wins overall. The authors caution against reading this as a general reason to discard crisis data.

There is a live-use consideration as well. The paper says it uses only information available at the forecast origin, while its appendix notes that irregular publication delays make some ENTSO-E variables unreliable at every origin. A desk using the German combined model would depend on that feed. The trade feed has no corresponding problem.

Two details matter if you copy the construction. The "orderbook" features come from executed trades and contain no depth; the authors themselves list order-level depth features as future work. Empty 15-minute buckets receive zero masks. Since an Austrian product has roughly ten times fewer trades than a German one, Austria should have more empty buckets.

Forecast loss and a trading decision

The authors say the economic value for trading and storage bidding "remains to be quantified." Their motivating example is imbalance speculation under a single price in Belgium and the UK, neither of which they test. Their aim is to explain price formation through information flow between markets, and they acknowledge the PnL gap as a scope limit. At 180 minutes, German VWAP-Neighbor has a 1.10 EUR/MWh AQL edge over VWAP-Self (29.68 versus 30.78). Against German RMSE near 250, that does not establish how many positions would flip.

We could not test this ourselves. We lack German and Austrian intraday trade prints and imbalance-settlement price series. Commodity futures bars cannot stand in for a quarter-hourly continuous market settling against a TSO price.

For a desk already forecasting imbalance prices with EPEX data, VWAP from the next few products is a cheap addition. In the pooled 2024 results, VWAP-Neighbor reaches the top group in 11 of 12 German cells and 10 of 12 Austrian cells, with its neighbor count selected on validation. Turning that feature preference into a position would take a storage-bidding test showing that the 3.6 to 7.2% AQL gap over VWAP-Self appears in euros.