Stage 0 · one training example
A response the model must learn to reconstruct.
Prompt: Natalia sold clips to 48 of her friends in April, and then half as many in May. How many clips did she sell altogether? The 27 response tokens are the tiles on the right. Hover any tile for its numbers.
Standard SFT hides a random fraction of them and trains on every hidden token equally. GoldiMask makes both choices from measurements instead.
Stage 1 · over-mask
Hide more than the masking rate asks for.
The rate t = 0.45 calls for K = 12 hidden targets. GoldiMask first masks at rate t + ρ, here 0.75, dropping a mask on a larger candidate set. Which candidates stay masked is now a decision, not a coin flip.
Stage 2 · measure
One forward pass, no gradient.
A measuring light sweeps the grid. Every candidate gets its gold-token probability p, drawn as a pin, and its attention to the other candidates, drawn as threads. Pins glow brightest where the supervision priority peaks.
A token the model gets right about a quarter of the time is the most valuable target: unsure enough to have room, not so lost that a gradient step cannot move it.
Stage 3 · reveal
Lift the mask where it helps the most.
Revealing a token supports every target that attends to it, but forfeits that token as a target. The objective balances the two. It is submodular but not monotone, so we use maximizers with guarantees; greedy is shown, lifting one mask at a time.
Greedy reveal order · marginal gain in F
Stage 4 · weight
Count each target by how much the context helped.
A second forward pass on the revealed input gives a new probability for each masked target. Targets whose probability rose, and that still have room, take more of a fixed loss budget. No target may take more than ten times the uniform share.
Targets · p before → after · loss weight
Stage 5 · after training · inference
A model that commits more tokens per round.
At inference a diffusion decoder starts from a fully masked response, predicts every position at once, and commits those above a confidence threshold. GoldiMask trains exactly this prediction problem: a target recovered from partial context. Watch the masks lift in rounds.
Illustration: confidences and attention are synthetic; the reveal selection and the loss weights are computed live with the paper's rules.