A Million Monkeys (Minus One)

A Million Monkeys (Minus One)

Causal redundancy, and what a measurement can be asked to mean

There's that old joke about a million monkeys at a million typewriters eventually producing Shakespeare. A recent MIT paper suggests a different experiment. Suppose the monkeys have been typing for some time, their pages collected to train a model that generated an output. Now remove one monkey and generate again.

Does much change?

In "Outputs of generative diffusion models are often unattributable," Zheng Dai and David Gifford report that at sufficient scale, often surprisingly little does. They call the effect attribution decay, and their method is subtractive. Generate an image. Remove one defined unit of training data while holding the prompt and the injected noise fixed. Generate again. The largest distance between the original and any of these counterfactual images is the Counterfactual Radius, or CR. A small radius means no single unit could have been withheld to much effect.

This study involved twenty-four diffusion ensembles and seven image distributions, with training sets from 256 to 162,770 images. The authors note that commercially deployed models can be trained on datasets approaching 10⁹ images. Attribution decay therefore appears at a relatively modest experimental scale; applying the finding to the commercial systems where its consequences would matter most requires extrapolating several orders of magnitude beyond it.

The machinery is not plumbing

Doing this by brute force would be ruinous. Every omitted unit would need its own model, retrained from scratch. So the authors built something that can be cut instead of rebuilt: a diffusion ensemble, whose components are trained independently on different slices of the data and then averaged together. To remove a unit's influence, you throw away every component that ever saw it.

That is genuinely clever, and it is also a bigger intervention than "delete one image." The slices are assigned so that each unit appears in a fixed fraction of them, and for the efficiency the method depends on, that fraction is about half. So ablating a single training image can mean discarding roughly half the ensemble. The counterfactual model is not the model minus one picture. It is a noticeably smaller model.

This cuts both ways, which is why it's worth saying plainly rather than deploying as a gotcha. It makes the headline finding conservative — if a cut that deep barely moves the output, honest retraining would presumably move it less. But it also offers a second explanation for the decay itself. An average over many components naturally gets more stable as each component sees more data. Some of what looks like a corpus growing redundant could be an ensemble growing indifferent to losing half its members. The measurement can't tell those apart.

To their credit, the authors saw this coming and ran the control: over a thousand ordinary diffusion models, built the old-fashioned way, with images genuinely withheld. The decay shows up there too. The catch is range. The control runs only at the smallest sizes, a fourfold span against the three orders of magnitude covered by the main experiments.

Which suggests a narrower statement than the title, and one the authors would probably accept: attribution decay is convincingly demonstrated inside this architecture, and reproduced at small scale outside it. Its extension to the large monolithic systems everyone actually cares about is an inference.

The obvious fallback also fails

The paper doesn't rest on subtraction alone, and this is the part most often skipped in summaries.

If you can't trace an image causally, you might try tracing it by resemblance — attribute the output to whichever training image looks most like it. The authors test this and find that as training sets grow, such attributions are increasingly wrong. The image you'd point to is increasingly not the image that mattered.

That closes off the easy reply that we just need a better similarity search. It doesn't close off everything. Against the attribution methods actually under development, which work by tracing gradients and influence rather than resemblance, the paper offers something different in kind: the observation that those methods are rarely evaluated on this sort of leave-one-out test at all. That is a fair complaint about how a field checks its own work. It is not a demonstration that the methods fail.

What the measurement can't see

The authors suggest decay happens because important features end up redundantly scattered across a large corpus, so no single source is holding anything up. That seems right, and it's where the method meets its sharpest limit.

Redundancy can arrive by at least two routes. Many examples might land on the same convention independently. Or many examples might land on it because influence spread outward from somewhere earlier — through imitation, transformation, commercial absorption, quotation, parody — reaching artists who never saw the original. To a leave-one-out test these look identical. Both leave a corpus that shrugs off the deletion. Historically they are not remotely the same thing.

This isn't a boundary between machine learning and art history. It's an identification problem sitting inside the experiment.

It's tempting here to reach for a famous modern artist and ask whether deleting their work would delete their signature from a model's output. That temptation is worth naming and then declining. Artist-level units appear in only two of the twenty-four ensembles, trained on public-domain artwork. Whatever that establishes, it isn't yet evidence about living artists whose influence has travelled through a century of commercial culture.

One word doing two jobs

The paper draws two conclusions from unattributability, pointing in opposite directions, and both deserve the same scrutiny.

Toward copyright, it suggests unattributability might undercut access, an element of infringement. That is a legal claim, made in passing, resting on a single old case note, and it is the most contestable sentence in the paper.

Toward privacy, it argues that models trained on enough data acquire a protective property, since generated faces become causally independent of the particular people in the training data. Worth noting: this rests on the paper's strongest evidence. The person-level results are its best supported.

But the conceptual leap is the larger one. What's been shown is that one output is hard to trace to one subject. Most privacy harm in this setting is measured differently — as whether an adversary can determine that someone was in the training set at all. That question is about what the model's parameters retain, not about what any single generation depends on, and attribution decay is close to beside the point. The paper cites that literature early and doesn't come back to it here.

So "unattributable" is being asked to mean non-infringing in one direction and privacy-preserving in the other, and it isn't obviously entitled to either.

The same stretch shadows the paper's oddest proposal: that you could simply keep sampling until an output with a low enough CR turns up, and thereby generate images guaranteed unattributable. Institutionally, that resembles a scrubbing service. Two things stop it from just being one. It needs an ablatable ensemble, so it's advice about how to build a system rather than something an operator can bolt onto an existing one. And the certificate is circular — low CR means hard to attribute by this method, which says nothing about methods that don't work by subtraction.

A similar deflation applies to the most dramatic-sounding result. A handful of generated samples came back with a Counterfactual Radius of exactly zero: no deletion changed them at all. The authors defend this by noting that digital images live in a finite space, so exact zeroes are possible in principle. Fair enough. But those samples came from black-and-white MNIST digits, one bit per pixel, and the odds of landing on an exact zero collapse as you move toward the millions of colours the paper itself invokes. It's a clean existence proof. It isn't yet a claim about photographs.

The hole

Our monkeys turn out to be badly suited to the original joke, which imagines creativity emerging from a vast number of independent attempts. These monkeys aren't independent. They've been reading each other. Some copied. Some transformed. Some arrived at the same phrase alone. Some inherited conventions from monkeys they never read.

Remove one, and no Shakespeare-shaped hole appears.

That absence is a real measurement, carefully made. What it entitles us to conclude is the harder question, and it's smaller than the absence feels.

Subscribe to The Grey Ledger Society

Don’t miss out on the latest issues. Sign up now to get access to the library of members-only issues.
jamie@example.com
Subscribe