What a second sample can tell you — and what it cannot
A second sample can tell you how much your profile moves. It cannot tell you whether it moved in a good direction, because no target state has been defined and no validated measure of direction exists — an international expert consensus statement found the evidence supporting the clinical usefulness of microbiome testing to be scarce [13]. That distinction is the whole subject of this article, and everything below follows from it.
The useful framing is not progress but signal and noise. Between any two of your results sit a fortnight of ordinary life, two collection events, two shipments, two extractions, two sequencing runs and two passes through a computational pipeline. Some of the difference you see is you. Some of it is the measurement. Distinguishing them is possible, but only under conditions you have to deliberately create.
The part that stays put
The gut microbiome is genuinely individual and genuinely persistent, and this is the strongest thing the longitudinal literature supports.
Low-error sequencing of 37 US adults sampled for up to five years, combined with whole-genome sequencing of more than 500 cultured isolates, found stability following a power-law function which, on extrapolation, suggests most strains in an individual are residents for decades. Shared strains were recovered from family members but not from unrelated people [1]. In a cohort of 308 adult men given four stool samples each — a pair collected 24 to 72 hours apart, a second pair around six months later — within-person taxonomic and functional variation was consistently lower than between-person variation [2]. Individuality is not confined to bacteria either: longitudinal analysis of faecal viruses found high temporal stability and a numerically predominant, individual-specific persistent personal virome [4].
The part that moves
Persistent membership does not mean fixed proportions, and the two are routinely confused.
The densest early time series — two individuals, four body sites, 396 timepoints — found pronounced variability in an individual's microbiota across months, weeks and even days, with only a small fraction of the taxa at a given site present at all timepoints. The authors' conclusion was carefully hedged: no core temporal microbiome exists at high abundance. Many taxa are persistent but non-permanent members, cycling in and out of detection [3].
An intensive antibiotic time series makes the shape of this clearer. Across 3 individuals sampled 52 to 56 times over ten months, day-to-day temporal variability was evident but constrained around an average community composition that was stable over several months in the absence of deliberate perturbation [11]. That is the right mental model: a stable average with a dynamic regimen around it. A single point on that trajectory is neither wrong nor definitive.
Variability is itself a personal trait
The most useful longitudinal finding for anyone considering a second test is that people differ in how much they differ. In 85 adults sampled weekly for three months across four body sites, gut communities varied mostly in the relative abundances of taxa — and there was a wide range of temporal variability across the study population, with some individuals carrying markedly more variable communities than others. Individuals with more diverse gut communities were associated with greater compositional stability in that cohort [9] — an observational association in 85 adults over three months. It is information about variability, not about health; see understanding gut diversity. The authors concluded that temporal dynamics may need to be considered when attempting to link changes in microbiome structure to changes in health status [9].
Three ways a difference can be technical rather than real
Before reading any difference as biology, it is worth knowing how large the alternative explanations are.
Collection and shipping
A controlled reproducibility study of 19 healthy volunteers across seven collection methods, stored for seven days at room temperature, 4 °C or 30 °C and compared with immediately frozen samples, found significant variation for all collection methods against that reference. Stability was high — at or above 0.75 — for tubes going through standard screening procedures, but with exceptions: the relative abundance of Actinobacteria sat at 0.65, and across different tubes stored at 30 °C the range was 0.41 to 0.90, while at room temperature it ran from 0.06 to 0.94 [18]. An intraclass correlation of 0.06 for a phylum means essentially none of the observed variation for that phylum reflected the donor.
The same study reported the necessary counterweight: interindividual variability was much higher than the variability introduced by the collection method [18]. Handling noise is real without being dominant. It simply has to be held constant if two of your own samples are to be compared.
The laboratory
The laboratory contributes more than most readers expect. Blinded specimen sets sequenced by 15 laboratories and analysed with 9 bioinformatics protocols produced variability driven most by biospecimen type and origin, then DNA extraction, then sample handling, then computation; analysis of artificial communities with known composition revealed genuine differences in extraction efficiency and classification between laboratories [5]. Benchmarking 21 DNA extraction protocols on the same faecal samples found extraction had the largest effect on the outcome, larger than library preparation or storage, and contrasted those effects explicitly against biological variation within a person over time [6].
Those were research laboratories. The consumer services have now been tested directly. A US National Institute of Standards and Technology group ordered three kits from each of seven direct-to-consumer gut microbiome services and inoculated every kit with aliquots of the same standardised, homogenised human faecal material, so that any difference between the returned reports was method rather than biology. Across the nine analyses — the seven services plus two reference workflows run in-house — the number of genera identified ranged from 34 to 906 out of 1,208 identified in total, and only 17 genera appeared in every sample once a single anomalous replicate was set aside: under 2%. For 17 of the 18 genera that every service reported, variation between services either exceeded the variation between eight different donors run on one common workflow, or could not be statistically distinguished from it. All seven services reported on the presence or absence of one particular organism in that identical material: three reported it present and four reported it absent [20]. A separate European exercise sent aliquots of one healthy donor's stool to six testing kits and received three verdicts of "excellent" or "good" bacterial diversity, two of "average" and one of "unfavourable" — from the same sample [21].
Two things that study cannot tell you, and says so. Because the material is a reference standard rather than a known biological truth, it measures precision, not accuracy: the authors state the work cannot determine which result was closest to the real composition, and no service was shown to be wrong. And seven services were tested, not the category — a provider that was not in the study was not evaluated by it [20].
The arithmetic
The last source of false change is the most easily missed, because it survives a perfect laboratory. Sequencing data are compositional — the mechanics are set out under how gut microbiome testing works — and the consequence for a second test is direct: if one organism genuinely rises, every other organism's percentage must fall even if nothing about it changed [15]. Underneath that, total microbial load differs by up to tenfold between healthy people — and the authors of that work state the consequence for exactly this use case: comparative analyses of relative microbiome data cannot provide information about the extent or directionality of changes in taxon abundance or metabolic potential [17]. The same study showed that the widely reported Bacteroides–Prevotella trade-off was an artefact of relative analysis.
Deciding which organisms changed is method-dependent too. Fourteen differential-abundance methods benchmarked across 38 datasets identified drastically different numbers and sets of significant sequence variants, with results depending on pre-processing and, for many tools, on sample size and sequencing depth [16]. If that is true for group comparisons with statistical power behind them, it is more so for a comparison of two samples from one person.
What documented change actually looks like
Set against that noise floor, the changes the literature has captured cleanly are large, and they have identifiable causes.
Daily sampling of two individuals over a year found overall communities stable for months, punctuated by rare life events that changed them rapidly and broadly. In one of those two participants, travel from an industrialised to a developing country produced a nearly two-fold increase in the ratio of Bacteroidetes to Firmicutes, which reversed on return. An enteric infection in the other participant caused a permanent decline of most gut bacterial taxa, which were replaced by genetically similar species. Even during stable periods, changes in fibre intake correlated with next-day abundance changes in about 15% of gut microbiota members [7].
Antibiotics give the clearest timescale. A four-day course of three last-resort antibiotics in 12 healthy men — not a typical prescription — was followed by recovery to near-baseline composition within about 1.5 months, but nine common species present in all subjects beforehand remained undetectable in most of them at day 180 [10]. In the ciprofloxacin time series, communities shifted within 3–4 days, began returning about a week after each course, and often did not return completely; the responses varied between individuals and between two courses in the same individual, ending in a state that was stable but altered [11].
Diet can move things quickly when it is extreme: diets composed entirely of animal or plant products altered community structure within days and overwhelmed inter-individual differences in microbial gene expression [12]. Ordinary eating is a subtler input — composition reflects multiple days of dietary history rather than the last meal, and daily responses to the same foods are highly personalised [8]. The fermentation biology that underlies this is covered in gut microbiome and digestion.
One preliminary finding deserves a line, because it inverts a common assumption. In the same daily-sampling study, in a two-person subgroup, participants consuming only meal-replacement beverages did not become more stable; overall dietary diversity, not monotony, was what tracked with microbiome stability [8].
Designing a retest so it means something
Nothing above argues against sampling more than once. It argues for doing it under conditions that make comparison possible.
- ·Hold the method constant — and treat that as necessary, not sufficient. Same provider, same collection kit, same gene region, same pipeline. Differently processed results are much harder to compare after the fact, and standardising the method up front is what the benchmarking literature recommends [5][6]; sending one standardised material to seven consumer services produced between-service variability on the same scale as the biological variability between eight different donors [20]. Reproducibility inside one locked-down workflow is the part that generally does work — the authors of that evaluation state it "tend[s] to be very good", and setting aside the single service that failed, same-service replicates shared genera accounting for at least 95% of the identified sample composition [20]. But generally is not always: in that same study, three replicates of one material from one service produced two "healthy" verdicts and one "unhealthy" one. A constant method makes two of your results comparable in principle. It does not certify that either one was produced correctly.
- ·Sample more than twice if you can. The studies that characterise a person's own range use many timepoints [3][9]; two points give you no range to read a difference against. In fairness to the opposite case, the authors of one of those cohorts concluded that a single measurement can itself provide long-term information about composition and functional potential [3]. There is no validated number of samples either — this is a practical suggestion, not an evidence-based one. A short run of samples establishes your own range; a single pair does not.
- ·Record the conditions. Recent antibiotics, illness, travel and major dietary change are the inputs the literature shows moving communities [7][10][11][12]. Without them, an unexplained difference stays unexplained.
- ·Expect movement. Day-to-day and week-to-week variation is normal and personal [3][9]. Stability is not a virtue and instability is not a fault.
What repeat testing cannot do
It cannot show that an intervention worked, because the outcome would have to be defined in something other than the microbiome itself. The most sophisticated study in this space illustrates the point precisely: an 800-person cohort with 46,898 meals under continuous glucose monitoring, a 100-person validation cohort and a blinded randomised trial produced genuinely personalised dietary predictions — but the validated outcome was the blood glucose response, with the microbiome serving as one input among many alongside blood parameters, dietary habits and anthropometrics. The intervention did produce consistent alterations in gut microbiota configuration; what it validated was the glucose measurement, not a microbiome trajectory [19].
It also cannot substitute for clinical assessment. An international multidisciplinary consensus statement on microbiome testing found that evidence supporting its clinical usefulness is scarce, and warned that an increasing number of commercial providers offer direct-to-consumer tests without any consensus on regulation or proven value in clinical practice [13]. A policy analysis in Science states that such tests lack analytical and clinical validity [14]. The seven-provider evaluation is the first direct empirical test of the first half of that statement; its authors conclude that analytical performance is a prerequisite for sound clinical recommendations, and that standards are needed to ensure analytical validity and consumer confidence [20]. A European expert panel reviewing six kits reached a compatible judgement from a different direction, considering the interpretations and recommendations in the reports premature for want of robust scientific evidence, and the accompanying analyses of limited clinical utility [21]. Those are the field's own words about the category, and they are the reason this article frames repeat sampling as a way of learning how much your description moves — rather than as a measure of progress toward a state nobody has defined.