A return model trained and selected on MSE has a strong incentive to predict zero. Pfeffer, Kruijssen, Stecker and Longmore show how costly that incentive can be. On intraday BTC, their CZAR (Composite Zero-Agnostic Return) loss reduces forecast shrinkage from about an order of magnitude to a factor of 2.6 to 6.6. Whether that improvement makes money remains open.
CZAR is a convex loss designed to treat forecast errors differently according to their direction. The authors use it as a LightGBM objective on BTC 15-minute and 1-hour bars, comparing it with L1 and L2 over 2,000-candle test windows. Its treatment of overshoots is the attraction. Its trading evidence is much thinner.
Why zero wins
The paper starts with a toy forecaster. True returns follow a standard normal distribution; the forecast is rho times the true return plus Gaussian noise with scale sigma_N. Predict zero and the scores are MAE 0.795 and MSE 1.007. Set rho = 0.1 and sigma_N = 0.5, and directional accuracy reaches 58%. Yet that forecaster scores 0.826 on MAE and 1.079 on MSE. Both losses prefer zero.
The authors derive the accuracy needed to tie a zero forecast. Every symmetric, strictly increasing loss has the same breakeven curve, whether evaluation averages losses or their logs. MAE, MSE and Huber therefore share the problem. When sigma_N exceeds the scale of returns, a Gaussian linear forecaster cannot beat zero at any hit rate under a symmetric loss. Changing the symmetric loss leaves that incentive intact.
What CZAR charges for
CZAR operates on z-scored returns. For the BTC tests, the scale comes from a trailing 100-candle standard deviation, with the mean set to zero. An undershoot or wrong-direction forecast incurs a near-linear penalty and a quadratic one. A correct-direction overshoot pays a discounted quadratic penalty. The discount, 1/(1+beta|z|), makes overshooting a large move almost free; when the true return is zero, the loss is symmetric.
A separate floor raises the loss on small true returns. Because that term depends only on the truth, it contributes nothing to the gradient. Its effect comes during tuning and early stopping, when the hyperparameters and tree count are chosen. The defaults are alpha = 1, beta about 31.2, floor level about 2.19 and hinge width 0.5. Monte Carlo fits the asymmetry rate beta and the floor level to alpha on the Gaussian toy.
Under log-averaged evaluation, CZAR needs accuracy within roughly 5pp of 50% to break even as sigma_N runs from 0 to 1.5. The symmetric-loss curve climbs steeply toward sigma_N = 1. CZAR also keeps its margin over MSE with Student-t truths at nu = 5 and nu = 3, although both breakeven curves rise.
There is a LightGBM detail worth watching. For an overshoot, the Newton step is exactly y_hat minus y, regardless of alpha and beta. Discounting that error leaves the per-sample step unchanged. The outward pressure comes from undershoots, where the step acquires a constant offset of sigma(1 minus beta_eff)/alpha. Smaller alpha consequently spreads the forecasts: on 15-minute bars, the log10 aspect ratio shifts from about -0.82 at alpha = 1 to about -0.46 at alpha = 0.005.
The forecasts spread out
The real-data experiment begins with one-minute Tiingo BTC/USD candles aggregated into 15-minute and 1-hour bars. Each horizon has one split of 20,000 training candles, a one-candle gap and 2,000 test candles. The features are deliberately basic: lagged returns, candle geometry, volume and calendar encodings.
Shrinkage is where CZAR earns attention. L1 and L2 forecasts have log10 aspect ratios between about -0.9 and -1.2, leaving predictions roughly an order of magnitude too small. CZAR variants range from -0.41 to -0.82. On moves larger than one sigma, L1 and L2 record hit rates of 0.47 to 0.50; CZAR records 0.50 to 0.56.
Across the full sample, the directional result is slight. CZAR ranges from 0.512 to 0.531, versus 0.510 to 0.518 for the baselines, and the binomial standard error is about 1.1pp. The authors acknowledge the limited separation. Their more consequential test changes the selection rule: retrain the 15-minute CZAR models, then select and early-stop on symmetric log-L1. The aspect ratio falls back to -1.0 to -1.4, and accuracy on large moves drops to about 50%. As they put it, "Training with CZAR is thus necessary but not sufficient." Keeping an old validation metric while replacing the training objective can quietly erase the gain.
Can the sign trade pay?
The trading rule holds one unit long for a positive forecast and one unit short for a negative one, rebalancing every candle. Magnitude never enters the position, so less aspect-ratio shrinkage cannot help this rule directly. Accuracy on large moves can help: each candle's payoff reflects the size of the realised return. The tuned 15-minute model (alpha about 0.054) hits 0.561 on one-sigma moves, against 0.499 for L2. Its naive Sharpe is 7.09, against 0.96.
That 7.09 excludes costs and slippage. The authors say it is incomparable with live trading and describe square-root annualization as a reporting convention. Even so, the abstract claims improved long-short performance without the cost-excluded qualifier retained in the conclusion. They defend the exercise as a controlled comparison across losses, where relative results matter more than the level. That works for the forecast metrics. The Sharpe comparison is gross of costs too, and different forecasts can change sign at different rates. We did not find a turnover figure in the paper.
Nor does the table provide standard errors for the Sharpe ratios. A rough calculation under independent candles puts them high. The 15-minute annualization factor is about 187, making 7.09 roughly 0.038 per candle. With 2,000 observations, one standard error for a per-period Sharpe is about 0.022, or about 4.2 annualized. The headline lies near 1.7 standard errors. L1 and L2 are both close to chance on direction; their Sharpes are -4.97 and 0.96.
At 1 hour, the highlighted alpha = 0.1 row posts 5.96, the highest of seven CZAR rows at that horizon. The tuned 1-hour model, which corresponds to the 15-minute headline, posts 3.00. The same rough arithmetic gives each 1-hour figure a standard error near 2.1 and puts the tuned row at about 1.4 standard errors. Against L2's 1.29, the 5.96 row has a gap of 4.67. If the strategies were uncorrelated, that would be about 1.6 standard errors (SE about 3.0); correlated signals would reduce the SE.
The alpha grid is uneven as well. At 15 minutes, alpha = 0.005 earns 5.62, alpha = 0.01 earns -0.95, and alpha = 0.05 earns 3.95. The authors can point to fixed-alpha CZAR beating both L1 and L2 on naive Sharpe in 6 of 6 settings at 1 hour and 5 of 6 at 15 minutes. Those settings reuse the same features and data. Their wins are positively correlated, so they carry less weight than independent trials. Each horizon has seven CZAR rows, six fixed alphas and one tuned, facing two baselines; we found no multiple-comparison adjustment.
The loss is a real contribution. The selection result alone makes the paper worth reading. Tradability needs a cost-inclusive walk-forward across many chronological windows, in place of one 2,000-candle block (about 21 days at 15 minutes, 83 days at 1 hour), with L1, L2 and CZAR selected on the same validation criterion. A BTC walk-forward in which CZAR still beats L2 by more than two standard errors would change my view.