The trading case for JSU(4) rests on Spanish day-ahead price tails: 1%, 5% and 99%. At the 95% quantile it ties the Normal (3.41 each, p = 0.901), and on MAE it ties SHASH(2) (p = 0.372).

What Ciarreta, Muniain and Zarraga built

The authors forecast the distribution of prices in the Spanish day-ahead auction. This sample runs from a minimum of -2.00 to a maximum of 700.00 EUR/MWh, a range that gives the tails practical weight. Its mean is 92.66, with excess kurtosis of 2.65.

Spanish power prices changed shape as well as level from 2020 to 2024. The annual mean rose from 33.96 EUR/MWh in 2020 to 167.54 in 2022, then fell to 63.03 in 2024. Excess kurtosis reached 9.14 at hour 20 in 2022. Moving only the mean and variance risks putting the tails in the wrong place.

GAMLSS (generalised additive models for location, scale and shape) lets covariates move each parameter of a chosen distribution. In this study, skewness and tail weight can respond to wind, solar and load forecasts, along with weekday and month dummies. Regime dummies mark four events: the COVID lockdown, the July 2021 price-limit change, the Ukraine war and the Iberian gas exception. The inputs also include lagged daily min and max price and seven days of own-price lags.

Each of the 24 hourly series gets a separate fit, using lasso for variable selection and a rolling two-year (730-observation) window. Forecasts are one-day-ahead over the last three years. The five specifications are a Normal, Johnson's SU with constant shape (JSU(2)), JSU with covariate skewness (JSU(3)), JSU with covariate skewness and kurtosis (JSU(4)), and sinh-arcsinh with constant shape (SHASH(2)). A naive copy of yesterday's or last week's price serves as the benchmark. The authors score MAE and pinball loss at 1%, 5%, 95% and 99%, then use Diebold-Mariano tests. Their overall choice is JSU(4).

Where does JSU(4) win?

The point forecasts are close. Average MAE is 19.34 for SHASH(2), 19.39 for JSU(4) and 19.44 for the Normal; the DM test comparing JSU(4) with SHASH(2) gives p = 0.372. The authors describe JSU and SHASH point forecasts as "broadly comparable". They also say JSU(4) and SHASH(2) "significantly outperform" the Normal, while the conclusion calls the advantage "substantially outperform". For JSU(4), that claim reaches only the 10% level (p = 0.094). SHASH(2) reaches 5% (p = 0.017). The difference is 0.05 EUR/MWh against an average error of about 19.4.

Tail scores make the stronger case. At 99%, JSU(4) records 1.29, versus 1.42 for SHASH(2) and 1.55 for the Normal. At 1%, the corresponding scores are 1.13, 1.24 and 1.87. JSU(4) scores 3.53 at 5%, compared with 3.67 for JSU(3), 3.86 for SHASH(2) and 3.92 for the Normal, all three at p = 0.000. Its 1% and 99% advantages over SHASH(2) also register p = 0.000.

The 95% result runs the other way. JSU(4) and the plain Normal each score 3.41 (p = 0.901). The Normal beats the other three shape-flexible fits: JSU(2) scores 3.85, JSU(3) 3.99 and SHASH(2) 4.01.

Two hourly failures inside the averages

The authors dropped some candidate distributions after "non-convergence", yet two surviving fits still go badly wrong. At hour 3, JSU(2) has an MAE of 71.04 and a 1% pinball loss of 57.16. The other GAMLSS models are around 12.6 to 12.8 and 0.76 to 0.98 at that hour. At hour 8, JSU(3) records a 99% pinball loss of 7.75, against 1.24 to 1.42 for the other GAMLSS specifications.

That hour-3 result lifts JSU(2)'s averages: its 21.89 average MAE and 3.53 average 1% pinball loss both include the failed fit. It also bears on the lower-tail comparison. JSU(2) is never significantly different from any model there. The paper reports that finding while saying the averages "consistently favour JSU(4)"; those averages include the same bad hour.

Hourly variation in point-forecast performance is acknowledged in the paper. For pinball loss, it says JSU(4) has the lowest losses "across most hours and quantiles". The evening peak tests that description. At hours 19 and 20, the naive model has the lowest 99% pinball loss of any model, 1.93 at both hours; at hour 20, the GAMLSS scores range from 2.41 to 5.79. At 95%, the naive scores are 5.04 and 4.96, against 5.43 to 8.45. This is where 2022 kurtosis was worst. A tail model driven by covariates ought to have an advantage there, yet yesterday's price wins.

Across all hours, the naive model also beats the Normal at 1% (1.66 vs 1.87, p = 0.000).

Pooled tests and dependent hours

The DM tests pool all 24 hours. For a one-day horizon, the paper argues that "the long-run variance estimator simplifies to the sample variance" and makes no autocorrelation correction. Dependence across hours remains a problem for that argument. The 24 deliveries on a given day share a gas market and weather system, while daily loss differentials through 2022 are unlikely to be independent either.

Correlation across hours would reduce the effective sample. Borderline findings would be most exposed: p = 0.094 for MAE and p = 0.026 for the Normal over JSU(3) at 95%. JSU(4)'s tail gaps over SHASH(2) are larger: 1.13 vs 1.24 at 1% and 1.29 vs 1.42 at 99%, both at p = 0.000. We expect gaps of that size to hold. We would put little weight on the claim that JSU(4) significantly outperforms the Normal on point forecasts, given its 0.05 EUR/MWh difference and p = 0.094.

Quantile accuracy, then a hedging claim

The conclusion says JSU(4)'s tail accuracy provides "more accurate estimates of downside and upside risk, leading to improved hedging design decisions". We found no bidding, hedging or trading exercise in the paper. Its evidence consists of MAE and pinball-loss scores.

The comparison is limited to the naive model and a Normal GAMLSS. The paper cites the open-access benchmark of Lago et al., though we did not find it run here. Gas price is also absent from the covariates, despite the sample's worst months falling in a gas crisis.

We could not run our own check. We hold no Spanish day-ahead price series or the ENTSO-E load and renewable forecasts required by the model. Nothing in our futures set reproduces the hourly Iberian auction.

A result that would change our view is a JSU(4) 99% quantile beating the naive copy at hours 19 and 20 when used to set an actual bid or size a cap hedge. For now, JSU(4) has the lowest average 1% and 99% pinball loss among the five GAMLSS fits over 2022 to 2024. Its value to a hedger remains unmeasured.