Ten times tighter on the policy costs a hundred times tighter on the value gap. Huh proves the exponent is sharp when only value information is available. Every practical question about the paper starts there.
Consider the problem Huh has in mind. You train a neural policy for constrained dynamic allocation, perhaps with no-short weights and a borrowing cap. The network returns feasible weights at every date. The optimum is unavailable for comparison, which was the reason to use the network in the first place. Classical portfolio duality, from He and Pearson and Cvitanić and Karatzas through the simulation method of Haugh, Kogan and Wang, supplies a bracket. Simulating the policy gives a primal value, while an artificial market gives a dual upper value. Their gap bounds utility loss. It does not tell you the distance between your weights and the optimal weights. Nor can it distinguish an economically excluded asset from a zero produced when the output layer clips at the boundary.
Huh extracts both answers from the same bracket. Everything is certified on the declared discrete-time grid that also handles simulation and policy evaluation. With polyhedral weight sets, an exact conditional budget identity expresses the primal-dual residual as the expectation of a pathwise nonnegative quantity. The quantity combines a terminal Fenchel defect with a date-by-date, row-by-row ledger of complementary slackness. Doob compensation strips out the budget martingale, whose variance would otherwise bury the small positive residual.
A computable lower bound on Bellman action curvature then turns the residual upper bound into a confidence region around the unknown optimal policy. The region is weighted according to the states the deployed policy visits. Constraint identification uses separate machinery. One constraint row is relaxed by a preset epsilon, and the paired gain is evaluated on common paths. This gives a lower bound for the optimal multiplier. A positive lower bound certifies that the row binds at the optimum.
There is no market data.
The paper studies four simulated models. They are a solved five-asset Merton benchmark (T=1, K=20, r=0.03, gamma=3, volatilities 0.16 to 0.24), a one-asset predictable-return model with an OU predictor and K=30, a two-asset simplex model, and a no-short Merton stress test at d=5, 20 and 50. The path budgets are substantial: 262,144 residual paths, 96 restart states with 16,384 conditional paths each, and 32,768 paired paths for each face test. Reference dynamic programs are consulted only after each certificate has been fixed.
We did not test the method on our own data because there is nothing here to test. These guarantees apply to a declared simulator. Running the method on history would assess a fitted implementation instead of the certificate. A practical implementation could use liquid US ETFs or single names as the risky assets. That would provide a universe and no bound.
Compensation changes everything
The variance reduction works without ambiguity. At the final Merton checkpoint, the compensated residual estimate is 6.526e-6 and its standard error is 5.92e-9. Separate primal and dual evaluation produces 8.500e-5. Squaring the comparison gives an effective path gain of 2.06e8, versus 1.21e4 early in training. Meanwhile, the mean certified policy radius declines from 0.10385 to 0.01409. It still covers all five coordinate RMS errors, which range from 8.4e-5 to 5.52e-3.
The failed shortcut is more informative. Common paths without compensation produce a standard error of 1.764e-4, worse than the 8.500e-5 obtained from separate primal and dual evaluation. They also produce a residual mean of -3.704e-4 when the true residual is 6.526e-6. Weak duality makes the population residual nonnegative, leaving sampling noise as the source of the negative sample mean. Huh states the point exactly: "A negative CRN realization is therefore a variance event, not a duality violation, and must not be clipped into a false certificate."
The 2.06e8 gain comes from an analytically calibrated benchmark with an exactly solved dual. In the short-budget learned-dual audit, the terminal Fenchel residual is 3.42e-5 and the dynamic complementarity residual is 4.10e-5. They are comparable in size. Huh explicitly warns against carrying the extreme path gains over to this setting.
Can the test certify zero?
A symmetric radius cannot establish that a holding is exactly zero. The face test handles that task separately. In the solved five-asset benchmark, it identifies exactly the two analytically excluded assets and makes no false exclusion. Certified multiplier lower bounds are 9.16e-4 and 1.04e-2, compared with true multipliers of 2.73e-3 and 1.35e-2. Its leakage bounds are 7.63e-3 and 6.70e-4. Those bounds cover actual weighted allocations of 4.50e-4 and 8.34e-5, placing them roughly 17x and 8x above the observed leakage.
Recall is the weakness, especially for the constraint a risk manager would watch most closely. The two-asset study certifies 59/140 asset-1 no-short pairs and 46/120 asset-2 pairs. For borrowing-cap pairs, it certifies only 6/114, or 5.3%. Yet the same cap accounts for 24.6% of the positive residual in the KKT ledger, second to the terminal Fenchel term at 40.4%.
Structural certification and loss attribution therefore order the constraints differently, as the paper acknowledges. Huh also explains that the row-level allocation depends on a minimum-norm tie-break selected for reproducibility. The percentages are a convention rather than an economic decomposition. Failure to certify remains inconclusive by design.
The width problem
Coverage is clean throughout the experiments: 10/10 in the locked one-asset audit, 5/5 in two assets and 15/15 in the high-dimensional replication, with zero false face declarations anywhere. Width decides whether that coverage is useful.
For one asset, the mean radius is 0.05016 and the mean fine-grid error is 0.01741. Their ratio is 3.76, ranging from 1.63 to 6.11. The two-asset covariance-norm ratio is 3.95. Across the five reported directions, averaged over five seeds, the ratios are 3.95, 7.03, 4.86, 6.26 and 5.53. Even in the favorable cases, an acceptance tolerance must be set three to seven times looser than the error one would actually accept. At tau=0.05, 6 of 10 policies certify. At 0.10, all ten do.
Dimension causes no visible deterioration when the dual wrapper is an exact convex program. Mean radius/error stays between 1.005 and 1.007 from d=5 through d=50, and the maximum is 1.023. One of fifteen cells misses its residual bound by 4.39e-9. Bonferroni recalibration covers all fifteen, increasing the maximum ratio only to 1.025. At d=50, median audit time is 0.091 seconds.
Once the wrapper becomes a learned state-dependent dual, the ratios rise to 28.5 at d=20 and 44.8 at d=50. The corresponding residual bounds are 5.32e-3 and 9.45e-3. Curvature recovery remains intact: restart continuation ratios are 0.980 and 0.974, while the effective sample size of the outer weights is 0.839 and 0.910. The looseness comes from the dual. Compared with the one-asset mean residual bound of about 1.19e-4, the residual bound is about 45 times looser and the radius about seven times wider. The square-root law behaves exactly as advertised.
Huh concedes the limitation in the abstract, in the same sentence as the 50-asset claim. Elsewhere, the pilot is described as conservative and locating the bottleneck is presented as a contribution in itself. Dual approximation breaks down; the residual-to-policy transfer does not. The diagnosis is persuasive and offers little help to the user. A radius 45 times the true error cannot support an accept-or-retrain decision on a real book, even when the failing component has been identified precisely.
The certificate algebra takes 9 seconds for a radius audit, 96 seconds for a face family and 0.091 seconds for each high-dimensional cell. It consumes a tight, exactly feasible dual for a state-dependent high-dimensional problem. This paper does not construct that object and lists it as open.
Conditions for practical use
The scope section says transaction costs, consumption and generic high-dimensional state-dependent dual learning all require additional structure. Its budget identity is derived for a wealth step without costs or consumption. The certificate is also local to the occupancy law of the deployed policy. Huh says it supplies no pointwise or uniform bound, leaving rarely visited states outside its reach.
Inference adds another qualification. The reported confidence levels use asymptotic normal one-sided bounds with Bonferroni allocation. A distribution-free route exists as a theorem, although its required analytic variance bounds are described as substantially looser and are never instantiated. In the Merton replication audit, nominal 95% coverage is observed in 228 of 240 cells, 95.0%. The result is reassuring about the approximation while leaving it approximate.
Two smaller details matter. Three of the ten one-asset seeds are development seeds and are reported separately. The 33.29% radius improvement from the closed-form curvature term is measured on those three. The Spearman correlation of -0.960 between residual tightness and certified fraction uses ten seeds. The paper explicitly declines to treat it as a scaling law.
The wall-clock table for the seven locked one-asset seeds also deserves careful reading. A radius audit takes 9 seconds and a face audit takes 96 seconds. Including 16.5 minutes for primal training, dual fitting and checkpoint selection, the full run takes 18.9 minutes. The fine dynamic program used for validation takes 0.79 seconds. Huh reports this comparison himself and states that inferential validity is the purpose; compute is outside the claim.
A learned state-dependent dual at d=20 with a radius-to-error ratio under about five would change my view. So would a face certificate that survives proportional transaction costs. Until then, this is a rigorous verification layer waiting for a learned dual that nobody has made tight beyond two risky assets. At d=20 and d=50, certification still arrives at 28.5 and 44.8 times the true error. The paper remains worth reading, much as our note on Wasserstein-ball allocation was worth reading: the machinery was sound, while the binding constraint lay elsewhere entirely.