A thing economists sometimes talk about is notches. Back when I lived in the UK buying a house had one, due to stamp duty. You were taxed 1% of the purchase price of a house, up to £250k. But if you were £250,001 you were taxed 3% on the whole lot. Unsurprisingly you got a lot more sales right at the notch than you did just above.
The behavior made sense under the incentives of the system. One thing buyers would do is to buy other things in the house. If 255k is the right price, but you can buy it for 250k and pay 5k for a surprisingly valuable dining-room table the sellers didn’t want to move anyway, everyone is a winner (bar HMRC).
In RL, this kind of behavior is both common, and frowned upon. We talked about Xiaomi’s MiMo 2.6 model the other day, and they published a follow-up on a specific bad behavior just after. The behavior in question:
the model would sometimes issue the same or highly similar tool calls repeatedly, consuming substantial time and context without making meaningful progress.
They had already anticipated excessive tool calls as a problem. If the model issued more than 32 in a turn, the rollout was stopped and the reward set to zero. This penalty, naturally, applied to every token in that turn, chain of thought and so on. Turns out the model would do a fair number of tool calls, and while the penalty capped it, it didn’t discourage it until that point: over a set of test replays that showed the issue the % of rollouts with 10 or more calls went from 11% at step 0 to around 25% by step 20.
They tried reducing the cap to 8, which did seem to work, but took about 20 steps to take effect. This would have cost about $2.3M for their full run. And, when they did a broader evaluation they found the model still sometimes emitted a lot of tool calls, though more rarely.
A cap is something you are incentivized to stay under, but doesn’t particularly drive behavior below it. To fix the model, Xiaomi ended up finding a quite clever approach: they trained a specialist model whose reward was zero if it repeated a single call, and one if it used a tool properly. After 12 steps (about 7k examples) the repetition was 0 on their held-out replays. They then did on-policy distillation between the teacher and the main model. This only cost about 90k, in part because it taught the behavior they actually wanted to incentivize: just stop making pointless tool calls.
On this example, the repaired model had a 99.87% probability of stopping within eight tool calls and showed a strong tendency to stop at every position from five through 20. Before the fix, the cumulative probability of stopping was still only about 50% even after 59 calls.
Possibly relatedly, Fireworks just released Ember-1:
Ember-1 is a new specialized model from Fireworks Research that delivers Kimi K3’s quality with 40% fewer tokens.
Fireworks sells hosted inference, among other things, and customers always want to pay less for a given level of intelligence. Kimi K3 is a very good model, but it’s also a very large model, and at higher reasoning levels creates a lot of thinking tokens.
The K3 technical report describes how they control those reasoning efforts: each problem gets a token budget, and if a rollout goes over that budget too much the reward is set to -1. They train max-effort mode with a large budget, and then lower effort levels with lower ones.
If K3 behaved like MiMo you might expect to see more tokens emitted, during RL, even if it stays below the cap. Moonshot’s own RL charts show tool-call steps rising steadily with RL compute, alongside the scores, which is at least consistent with that.
Fireworks have post-trained a version of K3 that gets similar results with a lot fewer tokens: I have no idea whether it was the same kind of approach as Xiaomi used, but maybe!