A large swing in model participation need not move net orders. Li and Bandyopadhyay show this in a synthetic market where headline wording determines which language-model agents trade. The financing result is the useful warning for anyone sizing flow from model output: the roster changes sharply, and buy-minus-sell counts barely budge.

Who trades when the headline changes?

The market has one fictional firm, SmallCo, and 15 agents in each run. Five agents use Llama 3.1 8B, five use Qwen 2.5 7B, and five use Mistral 7B, all called through Ollama. A seed assigns cash between 900 and 1,100, zero to five shares, and one of four mandates, including capital preservation and opportunistic trading. Agents act sequentially for 25 rounds. Each sees its position, the last price, top-of-book sizes and the current headline, then returns a JSON buy, sell or hold for at most five shares. News arrives at round 5.

For each text contrast, the authors keep the agent population fixed and alter the announcement. The change bundles interpretation, social cues and length. Each contrast has 30 paired seed configurations, yielding 150 decisions per family and arm at the announcement round. The measure is each family's share of cleaned buy and sell decisions, against its fixed one-third share of agents. Holds cost nothing here. The question is which family enters the order flow.

In the workforce-reduction event, the longer version adds consequences and says traders are splitting. Mistral's active share rises from 17.5% (21 of 120 actions) to 57.4% (112 of 195). The 39.9-point shift has an exploratory bootstrap interval of [32.8, 47.5]; omitting any one configuration leaves a shift between 38.8 and 41.3. Mistral activity rises in 28 seeds and falls in none. Pooled buy share goes from 27.5% to 64.1%. The paper's exact accounting identity attributes 0.3441 of that 0.3660 change to participation and just 0.0219 to the direction participants choose.

Financing changes the traders more than the flow

The financing event supplies the counterexample, which the authors also put in their abstract. Changing the original wording to a de-emotionalized version moves Qwen's share of submitted orders by 48.1 points and Llama's by -61.0. Net buy-minus-sell counts barely move: -29 with the original text and -26 with the factual text. The difference is -0.10 per market, with an exploratory interval of [-0.77, 0.63]. This is not an equivalence test. By comparison, the layoff contrast moves net counts from -54 to 55, or 3.63 per market.

Look at the cells. Under factual wording, Qwen all but vanishes: 66 sells turn into 2 buys. Llama takes its place, while Llama's sell share among active decisions rises from 13 of 23 to 46 of 52. Qwen's sells largely become Llama's sells, and Mistral's buys drop from 40 to 12. In the financing decomposition, composition adds 0.2012 to the change in buy share; conditional direction subtracts 0.1167. The active-family Herfindahl is unhelpful in the layoff contrast as well. Its readings are 0.438 and 0.441 even as the dominant family switches from Qwen to Mistral and net flow changes sign.

Active share is a poor stand-in for signed flow. The authors put it precisely: "Composition can move substantially while aggregate signed counts barely move, because participation scale and conditional direction can offset one another." I agree. The financing decomposition carries more uncertainty than the layoff result, though. Qwen makes only two factual-arm trades, as the authors flag, and 1,256 of 10,000 bootstrap draws have an empty family-arm cell.

Under the original financing text, Mistral and Qwen trade like fixed directional readers; Llama's split changes. Across four repeated batches, active Mistral announcement decisions are 41/0, 36/0, 41/0 and 40/0 buys to sells. Qwen's are 0/65, 0/68, 0/68 and 0/66. A single scorer labels the text of seven of ten sampled Mistral announcement holds unwilling-bearish. The family's buy-only flow comes from a selected subset; agents who abstain often write bearish text. Apparent herding is largely arithmetic: same-family pair agreement is 0.952 observed, against 0.945 predicted from family buy rates alone.

One-sided single-model markets

With the original financing headline, all-Mistral markets send 159 buys and no sells over rounds 5 to 7. All-Qwen markets send 215 sells and no buys. Each is fully one-sided in 30 of 30 seeds. The mixed control sends 52 buys and 81 sells and is one-sided in 4 of 30. Its mean one-sidedness, absolute net over gross orders, is 0.359.

The authors describe this as a total population-assignment effect. Replacing every agent changes model identity and pre-event history along with concentration, so the comparison cannot isolate concentration at constant quality. Their first proposition gives a reason to be careful: if enough independent agents participate and all lean in the same extreme direction, one-sided flow can become nearly certain without coupling. They offer that result as a benchmark, without verifying it for the sequential simulator. At the follow-up round, the mixed market is itself mostly buy-side, and the one-sidedness contrast loses significance (p = 0.125).

Prices stay pinned

Background liquidity consists of 300 shares resting at 99.5 and 100, with no replenishment. The largest cumulative submitted quantity on either side in any run is 94 shares. Every recorded last price lies between 99.5 and 100.0. The authors call the outcomes "submitted decisions rather than price discovery," and there is no fundamental value in the setup against which to score a trade.

Price consequences appear only in the analytical section. In a linear-demand example with λ = 1 and K = 3, three uncorrelated families, each with error variance 1, produce mean squared price error of 0.25. One family with variance 0.5 produces 0.34375. A lone family needs variance no greater than one-third of the diversified variance to break even. The selection proposition extends the point: participation that covaries with valuation error can help or hurt at the same concentration. In the two-family construction, selections with Herfindahl one yield either error a² or zero. Concentration alone is a poor risk gauge in that example. The result remains untested on data.

We could not run any of this ourselves. Reproduction requires the authors' sequential market runner and the three Ollama model tags. Historical price bars record neither holds nor model identities, so a bar backtest would answer a different question. The authors report using the server's default sampling settings; per-call generation seeds and context length were neither set nor recorded.

The participation finding is tightly measured within its setup: two events, three 7 to 8B models, and pegged prices. The authors propose a known-value validation but have not run it. Its result would tell us whether presentation-driven selection creates measurable decision regret, the step this work needs before its flow effects say anything about P&L.