Generate synthetic data to fill gaps in your real data, never to replace it. The distinction decides everything downstream. Accumulating synthetic alongside retained human data behaves very differently from substituting one for the other, and only the second pathway leads to collapse.
The reflex after reading the collapse literature is to avoid synthetic data entirely. That overcorrects. Synthetic generation is the only practical option for privacy-constrained domains, for rare events that occur too infrequently to collect, and for bootstrapping tasks with no labelled corpus at all.
What matters is the pipeline shape, and specifically whether real data survives in it.
At a glance
- Generate for coverage of gaps, not for volume
- Diversity must be engineered into the prompt, not hoped for
- Filtering removes more risk than generation technique adds
- Label provenance on every example, permanently
- Never let synthetic data become the majority without a deliberate decision
- Measure the generated distribution against real data, not against previous batches
When is synthetic data the right tool?
| Situation | Suitability | Note |
|---|---|---|
| Privacy blocks using real records | Strong | Often the only lawful option |
| Rare events under-represented | Strong | Targeted generation fills a genuine gap |
| No labelled data exists yet | Strong | Bootstrapping, with human verification |
| Edge cases for testing | Strong | Adversarial and boundary inputs |
| Real data exists but is inconvenient to collect | Weak | Convenience is not a reason; collect it |
| Scaling a dataset for its own sake | Poor | Volume without coverage adds noise, not signal |
The strong cases share one property: real data for that region genuinely does not exist or cannot be used. The weak cases are substitutions of convenience, which is precisely the pathway the collapse research warns about.
Why does naive generation produce narrow data?
Because a model asked repeatedly for examples returns its most probable outputs, and those cluster hard.
Ask for a hundred customer complaints and you will get a hundred variations on a small number of templates, with similar sentence lengths, similar politeness registers and similar complaint types. The set looks varied on inspection and is statistically narrow.
Figure 1 — Two ways to generate a hundred examples
One prompt, run a hundred times
Every sample is drawn from the same conditional distribution, so outputs concentrate near the mode.
Diversity comes only from sampling temperature, which produces surface variation rather than genuinely different cases.
Narrow
Conditioned on varied attributes
Each generation is conditioned on an explicit combination: customer type, product, severity, tone, length, channel.
The attribute grid forces coverage the model would never produce unprompted, and the gaps become visible.
Engineered coverage
Diversity is a property you construct, not one you sample for. Enumerate the attributes that vary in your real data, build the grid, and generate against cells rather than against a single instruction repeated many times.
What does a defensible pipeline look like?
- Characterise the real distribution first. You cannot fill gaps you have not located. Measure the attribute distribution of the data you already hold.
- Identify the sparse cells in that distribution and target generation at them specifically.
- Condition each generation on an explicit attribute combination rather than a generic instruction.
- Filter aggressively. Deduplicate near-identical outputs, drop anything failing validation, and remove examples that drift from the requested attributes.
- Human-verify a sample, and always verify the whole set for anything high-stakes.
- Label provenance permanently on every example: human, synthetic, or synthetic-and-verified.
- Cap the synthetic proportion as a stated parameter someone owns, rather than letting it emerge from whatever was easy to produce.
Step 4 does more work than step 3. Filtering a large noisy generated set down to a smaller clean one consistently outperforms elaborate generation prompting, because rejection is a cheap way to reintroduce judgment the generator lacks.
Watch for this
Deduplicating on exact text is not enough. Generated examples are frequently distinct as strings while semantically near-identical, which inflates apparent dataset size without adding information. Deduplicate on embedding similarity with a conservative threshold, and expect to discard considerably more than exact matching suggests.
How do you tell whether the generated data is any good?
Compare it against real data, never against previous generated batches.
- Attribute coverage. Does the synthetic set span the cells you targeted, at the proportions you intended?
- Diversity within cells. Sample any single cell and check whether its examples are meaningfully different from each other.
- Distributional distance from a held-out sample of real data, on whatever features matter for your task.
- Downstream effect. Does a model trained with the synthetic data outperform one trained without it, measured on real held-out evaluation?
The last is the only test that ultimately matters, and it is the one teams skip because it requires actually running the training twice.
Where does this connect to model collapse?
Directly, and the connection is the reason for the provenance discipline.
Shumailov and colleagues showed in Nature in 2024 that indiscriminate training on generated content causes irreversible loss of distribution tails. Gerstgrasser and colleagues subsequently showed that accumulating real and synthetic data, rather than replacing real with synthetic, substantially mitigates the effect (arXiv:2404.01413).
So the safe configuration is not “avoid synthetic data.” It is: retain all real data permanently, add synthetic alongside it, keep the proportion deliberate, and never train a generation on the previous generation’s output without human data anchoring the mix.
Where do teams go wrong?
Generating volume instead of coverage
A hundred thousand synthetic examples clustered around the mode is worse than a thousand spanning the space, because the large set actively skews training toward the region already well covered.
Losing provenance in the pipeline
Once synthetic and real examples are merged into one training file with no labels, the distinction is unrecoverable. Every subsequent decision about mixing ratios becomes guesswork, and the collapse risk becomes unmeasurable.
Verifying only the good-looking samples
Human review naturally gravitates toward readable, plausible examples. The ones worth inspecting are the outliers and the boundary cases, which is where generated data goes subtly wrong in ways that matter.
What do experienced teams do differently?
They generate with one model and filter with a different one.
A generator’s own judgment of its output is correlated with the errors it just made, so it cannot reliably reject them. A different model family, or better a deterministic validator plus a human sample, catches failures the generator is structurally blind to.
They also keep the attribute grid as a living artefact. When production reveals an input type nobody anticipated, that becomes a new cell in the grid and a new generation target, which turns the pipeline into something that improves rather than a one-off dataset build.
A short glossary
- Attribute grid
- An enumerated set of attribute combinations used to condition generation, forcing coverage across the space.
- Coverage
- How well a dataset spans the range of real inputs, as distinct from how many examples it contains.
- Semantic deduplication
- Removing near-identical examples by embedding similarity rather than exact string match.
- Provenance label
- A permanent record of whether an example is human-authored, model-generated, or generated and human-verified.
- Accumulation
- Adding synthetic data alongside permanently retained real data, the configuration that mitigates collapse.
Key takeaways
- Generate synthetic data to fill located gaps, never to replace real data for convenience.
- Naive repeated generation clusters near the mode and produces narrow data that looks varied.
- Condition each generation on explicit attribute combinations to engineer coverage deliberately.
- Aggressive filtering contributes more than clever generation prompting.
- Deduplicate on embedding similarity, since generated text is often distinct as strings and near-identical in meaning.
- Label provenance permanently, or mixing decisions become guesswork and collapse risk becomes unmeasurable.
- Judge synthetic data by whether a model trained with it beats one trained without, on real held-out data.
Frequently asked questions
Is synthetic training data safe to use?
Yes, when it supplements permanently retained real data rather than replacing it. The collapse research distinguishes accumulation from substitution, and only substitution across generations produces the irreversible narrowing that erases rare cases from the learned distribution.
How much synthetic data is too much?
There is no universal ratio, but the proportion should be an explicit decision someone owns rather than an accident of what was easy to generate. Keep real data present in meaningful quantity in every training mix, and track the ratio as a monitored parameter.
Why does generated data lack diversity?
Because repeated sampling from one prompt draws from the same conditional distribution, concentrating outputs near the mode. Temperature adds surface variation rather than genuinely different cases. Diversity has to be engineered by conditioning each generation on explicit varied attributes.
Should I deduplicate generated examples?
Yes, and on embedding similarity rather than exact text. Generated examples are frequently distinct as strings while semantically near-identical, which inflates apparent dataset size without adding information. Expect to discard considerably more than exact matching would suggest.
Can I use the same model to generate and filter?
Poorly. A generator’s judgment of its own output correlates with the errors it just made, so it cannot reliably reject them. Use a different model family, deterministic validators, and a human-reviewed sample to catch failures the generator is structurally blind to.
How do I measure whether synthetic data helped?
Train with and without it, then evaluate both on real held-out data. Intermediate metrics such as coverage and diversity are useful diagnostics, but the downstream comparison is the only test that settles the question, and it requires running training twice.
What should I record about each generated example?
Provenance at minimum: whether it is human-authored, model-generated, or generated and human-verified. Also useful are the attribute combination it was conditioned on, the generating model version, and the date. Without provenance, later decisions about mixing ratios become guesswork.
Does synthetic data help with privacy compliance?
It can, since generated records need not correspond to real individuals. The caveat is that models trained on sensitive data can reproduce identifiable details, so synthetic generation is not automatically privacy-safe and needs verification rather than assumption for regulated use.
References
- Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., and Gal, Y. AI models collapse when trained on recursively generated data. Nature, 2024.
- Gerstgrasser, M., et al. Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv:2404.01413.
- Dettmers, T., et al. QLoRA: Efficient Finetuning of Quantized LLMs. arXiv:2305.14314. Source for the finding that small high-quality datasets can outperform much larger ones.
