Even Checking 75% of the Particles Rarely Reproduces a Sample's Real Polymer Mix

A tray of particles sits under an infrared microscope. Each one has been picked out of seawater, or a sediment core, or the gut of a fish, and each one has to be interrogated separately: a spectrum taken, matched against a library, called as polyethylene or polypropylene or a stray cotton fibre. Minutes per particle. Thousands of particles in a study. At some point the arithmetic of a research budget wins, and the analyst characterises a subset and multiplies up. That shortcut is routine, and, as the authors note, it is often applied without prior validation.
Three researchers at the Australian Institute of Marine Science and James Cook University set out to measure the damage. Their paper appeared Aug. 4 in Microplastics and Nanoplastics, open access under a CC BY licence, with M. F. M. Santana as first author and C. A. Motti as senior author. It has two halves: a review of what the field actually does, and a simulation of how that affects the results.
The review covered subtidal marine studies published between 2019 and 2024. Just under half of them, 46%, subsampled. Of those, 50.8% used what the authors call constrained-quota random selection, picking particles at random up to a fixed quota. A third of the studies analysed fewer than 25% of the items they had collected. No standardised method was applied across the literature, not even a shared minimum percentage, and the vocabulary used to describe what had been done varied from paper to paper, which makes two studies hard to compare even when both are transparent.
For the second half, the team needed a sample where the true answer was already known. They used a dataset of 2,137 putative microplastics recovered from eight subtidal matrices: surface water, mid-column water, sediment, and five organisms, namely fish, coral, sponge, sea squirt and sea cucumber. Every particle in it had been characterised. They then drew random subsamples at 25%, 50% and 75% of each matrix, ran 1,000 iterations for every matrix at every threshold, and asked a simple question of each draw: does the subsample reproduce the polymer composition of the whole?
It largely does not. Representativeness improved as the subsample grew, but non-linearly, and in every matrix; even at the 75% threshold, the authors report, reliable representativeness was rarely achieved. The number that carries the finding is easy to misread in the wrong direction. Only 15% of polymer types met the study's effectiveness criterion in at least one of the 1,000 iterations. Not reliably; once, out of a thousand tries. The other 85% never met it in a single draw. And the 15% that occasionally cleared the bar were not consistently the most abundant polymers in the original dataset, so the types that scraped through were not the ones a monitoring programme would care most about.
None of this should be surprising once you look at how polymers are distributed in a real sample. A handful of types dominate, and a long tail of rarer polymers turns up once or twice. Random subsampling captures the dominant types easily and misses the tail almost by construction, so a criterion that asks the subsample to reproduce the whole composition, tail included, is asking for exactly the thing random draws are worst at. Representing how much plastic is present is a different and easier problem than representing which plastics are present.
That asymmetry has been documented before, by a different group with a different dataset. In Chemosphere in 2023, Hannah De Frond, Anna M. O'Brien and Chelsea M. Rochman reported that no standard subsampling protocols existed, that methods varied widely and often lacked evidence of representativeness, and that fewer particles are needed to represent the proportion of plastic present than to represent the diversity of material types. The new paper puts numbers on the harder half of that finding.
What follows from this is narrower than it might first appear, and the distinction matters. The paper does not show that published microplastic counts are wrong. It shows that where a study subsampled, its account of which polymers were present rests on a procedure nobody had validated, and that under simulation that procedure reproduces polymer composition poorly. Total particle counts, and the broad question of how much plastic is in a sample, are not what failed here. Polymer identity is, and polymer identity is what links a particle in a sea cucumber to a source, a product or a regulation.
The authors' own prescription is modest and sits in their recommendations rather than in the abstract. If subsampling cannot be avoided, they write, selecting at least 50% of items represents a practical minimum. Beyond that they ask for something cheaper still: that subsampling outcomes and any extrapolations be reported explicitly and clearly justified, so a reader can tell what was counted and what was multiplied. Given that a third of the reviewed studies worked from under a quarter of their particles, the 50% floor would be a real change in practice.
The stakes are the authors' own framing rather than an extension of it: these are the data used to inform monitoring, mitigation and regulatory decisions. Governments are writing rules about plastic, and monitoring programmes are being designed to check whether the rules work. Both rest on laboratory numbers, and one of the routine steps producing those numbers has now been tested and found wanting for the polymer question. The paper is the final version of record, accepted July 6 and published Aug. 4 after peer review, and it is free to read. For anyone reading a microplastics paper this week, it suggests one question to put to the methods section: how many of these particles did anyone actually look at?
Sources
- Peer-reviewedMicroplastics and Nanoplastics
- Peer-reviewedMicroplastics and Nanoplastics
