What actually happens between the post box and the report
Six things, roughly: your sample is preserved and shipped, DNA is extracted from it, a chosen region of that DNA is copied and read, the resulting sequences are matched to names, those names are converted into proportions, and the proportions are compared with something. Every one of those steps is a decision, and every decision has been shown in published work to change the answer.
That is not a scandal. It is what measurement looks like in a field where the instrument is a chain of laboratory and computational choices rather than a single reading. But it does mean a microbiome result is best understood as this material, processed this way — and it explains why two providers can describe the same stool differently without either of them making a mistake.
Step one — collection, preservation and the post
Home collection is not the weak link people assume. Controlled studies with duplicate samples show technical reproducibility is genuinely good: in 52 healthy volunteers across five collection methods, intraclass correlations between duplicate faecal samples ran from 0.64 to 1.00, and stability at room temperature over 96 hours was high for most methods [4]. In a larger study spanning 132 participants across three distinct populations, reproducibility was excellent — at or above 75% — for every collection method tested [6]. A Bangladeshi cohort of 50 adults found technical reproducibility of 0.79 to 0.99 at both day zero and day four [7].
The caveats are specific rather than general. Agreement was weaker for measures that depend on exact relative abundance than for broad diversity measures, and some preservation chemistries performed poorly: intraclass correlations fell below 0.60 for 95% ethanol on abundance-sensitive metrics [4], and stability was not excellent for no-solution or 70% ethanol collection after four to seven days at ambient temperature [6]. Compared against immediately frozen samples, many correlations were low even where the rank order of a person's taxa was largely preserved [7].
The practical implication is narrow and worth stating: use the preservative supplied, and post within the stated window. The published stability data are specific to particular chemistries and particular numbers of days, so they do not license a general "it will be fine".
Step two — getting the DNA out
Once the tube reaches the laboratory, the bacteria have to be broken open and their DNA recovered. This is the step with the largest documented laboratory-side effect. When 21 representative extraction protocols were applied to the same faecal samples and compared against biological variation within a specimen and within a person over time, DNA extraction had the largest effect on the outcome of metagenomic analysis — larger than library preparation, larger than storage. Extraction choice biased estimates of community diversity and the ratio of Gram-positive to Gram-negative organisms [8].
The clearest demonstration of what that means in practice is uncomfortable. In stool from two breast-fed infants, an extraction kit with no bead-beating step produced sequence data containing no bifidobacteria at all — even with primers chosen to detect them — despite bifidobacteria being the most abundant organisms detected by microscopy in the very same samples [9]. Organisms with tough cell walls resist gentle extraction; if they are not broken open, their DNA is not in the tube, and what is not in the tube cannot be sequenced.
Step three — choosing what to read
Most consumer testing amplifies part of the 16S ribosomal RNA gene: a gene all bacteria carry, containing stretches that are near-identical across the domain (useful for grabbing everything) alongside hypervariable regions that differ between groups (useful for telling them apart). Copying that region by PCR is what makes a tiny amount of DNA readable. It also introduces bias at every sub-step, because PCR-based methods have multiple stages, each susceptible to error and bias [10].
Which region you read changes who you find
There is no universal 16S region, and the choice is consequential. In a controlled comparison of four primer pairs on the same human subgingival plaque samples — oral, not gut — targeting V4–V6 failed to detect the genus Fusobacterium entirely, while V7–V9 primers failed to detect Selenomonas, TM7 and Mycoplasma. The dominant genera differed substantially, though a few were dominant whichever region was read between the three region sets [12]. A separate oral study found that broad between-sample relationships survived a change of region, but richness estimates did not: phylotype richness was systematically higher for V1–V3 than for V3–V4 [13]. In a non-human wastewater system, one region overestimated a group of archaea by more than thirtyfold against a metagenomic reference [22].
The marker itself is imperfect, and this is a limit no protocol choice fixes. Comparing the 16S gene against core-genome phylogeny across human gut core genera, concordance was only about 50.7% within genera and 73.8% between them; the best-performing hypervariable regions reached 60–62.5%. Roughly 690 ± 110 informative positions are needed for 80% concordance, and the 16S gene averages 254 [14]. The same work reports rRNA operon copy number varying from 1 to 27 per genome — meaning a bacterium with many copies contributes more reads than an equally abundant one with few. Correcting for that remains, in the words of the benchmarking paper, an unsolved problem: the available tools explained under 10% of the variance in some cases and disagreed with each other for most communities tested [15].
Chimeras: sequences that were never there
PCR can also fabricate. When amplification stalls partway through one template and resumes on another, the result is a chimera — a hybrid sequence with no corresponding organism. Benchmarked against a mock community of known composition, chimeras formed reproducibly across independent amplifications, and rates exceeded 70% for less-abundant species before filtering, on Sanger and 454 platforms — which is why the same paper introduced a chimera-detection tool now in standard use. They can be falsely interpreted as novel organisms, inflating apparent diversity. Notably, shotgun sequences of the same mock community appeared devoid of 16S chimeras [11].
Step four — sequencing
The sequencing platform itself is a comparatively modest contributor. When short-read and long-read platforms were run on the same extracted DNA from faecal samples, more than 90% of reads on each were classified to genus or species level, and species-level misclassification between the two approaches was smaller than the authors expected [16]. That is a single small study, so it is illustrative rather than definitive — but it is consistent with the ring trials, which place specimen, extraction and analysis ahead of sequencing chemistry as sources of variation [1][8]. A newer sequencer does not repair an extraction, primer or database bias upstream of it.
Step five — turning reads into names
Raw output is millions of short sequences. Two decisions convert them into a list of organisms.
The first is how reads are grouped. Historically they were clustered into operational taxonomic units — bins of reads differing by less than a fixed threshold. Modern error-modelling resolves amplicon sequence variants exactly, down to single nucleotides. The argument for the newer approach is not mainly resolution but portability: sequence variants are consistent labels with intrinsic biological meaning, identified independently of any reference database, which makes results from separately processed datasets comparable in a way that cluster-based results are not [19].
The second is which reference database the sequences are compared against, and this changes how much you learn. In a benchmark using one classifier on one gene region against a bacterial sequence test set, species-level assignment ranged from 10.23% to 24.28% across the three standard databases, and genus-level from 70.88% to 87.20% [21]. A gut-specific database significantly raised assignment rates over general-purpose ones — and the authors of that work noted plainly that even for an ecosystem as well studied as the human intestine, assigning genus and species names to 16S reads remains challenging [20].
Some of what is in a stool sample has no name because no name exists yet. Across 9,428 human metagenomes, 77% of the 4,930 reconstructed species-level genome groups had no genome in public repositories as of that 2019 analysis; reference catalogues have grown since; they were present in 93% of well-assembled samples, and adding them lifted the share of gut reads that could be mapped from about 68% to 88% [23]. Those unknown groups were enriched in non-Westernised populations — an equity caveat worth stating plainly, since it means reference completeness is not the same for everybody.
Shotgun metagenomic sequencing reads all the DNA present rather than one gene, adding strain-level resolution and a catalogue of the genes the community carries [24]. That is genomic potential.Function inferred from a 16S profile is a prediction rather than a measurement, with uncertainty the tool's own authors document [25] — explained in full under gut microbiome and digestion. We compare the approaches in sequencing methods, and the wider gene-to-function distinction is set out in gut microbiome and digestion.
Step six — turning names into numbers
Sequencing produces shares, not counts. The instrument imposes an arbitrary total, so the data are compositional: if one organism's share rises, others must fall whether or not anything about them changed [26]. This is not a subtlety to be filed away — total microbial load differs by up to tenfold between healthy people, and quantitative counting demonstrated that the widely reported trade-off between Bacteroides and Prevotella was an artefact of relative profiling [27]. We treat this at length in relative abundance.
Then something has to decide what counts as different. Fourteen differential-abundance methods benchmarked across 38 datasets identified "drastically different numbers and sets" of significant sequence variants, with results depending on pre-processing and, for many tools, on sample size and sequencing depth; only two were consistent across studies, and the authors recommend agreement across several methods rather than trust in one [28]. Common normalisation shortcuts are not neutral either: both simple proportions and rarefying — randomly subsampling every sample to equal depth — produce high false-positive rates in tests for differentially abundant organisms [29].
Contamination, and why controls exist
Extraction kits and laboratory reagents contain bacterial DNA of their own. This is a documented, reproducible property of the reagents rather than a matter of laboratory tidiness: contaminating DNA is ubiquitous in commonly used kits, varies between kits and even between batches of the same kit, and affects both amplicon and shotgun work [17]. Sequencing is sensitive enough to detect that contaminant DNA as efficiently as real signal [18].
The size of the problem depends entirely on how much genuine material is present. Stool is a high-biomass sample, so reagent DNA is a small fraction of what is read; the acute risk is in low-biomass samples such as tissue or blood, where claims should be treated with corresponding suspicion [18]. The standard safeguard is the same in both cases: concurrent sequencing of negative controls alongside real samples is strongly advised [17].
What this means for reading a report
Three practical consequences follow, none of which requires distrusting the measurement.
Results from different providers are not directly comparable. Different gene region, different reference database, different bioinformatic pipeline — and it is the processing rather than the extraction chemistry that has been shown to drive incomparability; one benchmarking study concluded results can be robust to the extraction and sequencing approach while remaining incomparable once bioinformatically processed. In a study that sent matched intestinal biopsy samples from 32 people to three laboratories, broad group-level signal held up while taxonomic assignment and abundance estimates did not, and the authors concluded that combining differently processed samples is nearly impossible [3]. That is a statement about method, not about quality.
Your own samples are comparable to each other, if the method is held constant. That is the basis for reading change over time, covered in retesting and longitudinal tracking.
Between-person differences remain the largest signal. In a cohort of 308 adult men sampled four times over roughly six months, within-person taxonomic and functional variation was consistently lower than between-person variation — although gene expression profiles were as variable within a person as between people, which is one reason a DNA-based test reports composition and functional potential rather than activity [32].