Skip to content
Gut Health 15 min read

Retesting and longitudinal tracking

A second sample tells you how much your profile moves, not whether it moved in a good direction. Between any two results sit two collection events, two laboratory runs and two computational passes, alongside whatever actually happened to you. This article sets out how to separate those, and what repeat sampling cannot do at all.

Educational context — not a diagnostic result

Key takeaways

  • Repeat sampling answers a different question from a single sample. One profile describes a moment; a series describes how much a person's profile moves — which is itself informative, and differs markedly from person to person [3][9].
  • The membership of your gut community is remarkably persistent. Over five years in 37 adults, stability followed a power law which, extrapolated, suggests most strains in a person are residents for decades [1].
  • The proportions are not. In an intensive time series across 396 timepoints, only a small fraction of the taxa found at a body site were present at every timepoint, leading the authors to conclude that no core temporal microbiome exists at high abundance [3].
  • A difference between two reports is not automatically a change in you. Sequencing reports shares of a total, and because total microbial load varies, relative data cannot provide information about the extent or directionality of changes in a taxon's abundance [15][17].
  • Handling can manufacture apparent change. In a controlled reproducibility study, agreement for one phylum fell to an intraclass correlation of 0.06 under some room-temperature conditions — meaning almost none of the observed variation for that phylum was biology [18].
  • An international expert consensus statement on microbiome testing concluded that evidence supporting its clinical usefulness is scarce, and that direct-to-consumer tests are sold without proven value in clinical practice [13].
  • Holding the provider constant is necessary, not sufficient. When aliquots of one standardised human faecal material were sent to seven consumer gut testing services, variability between providers was on the same scale as the biological variability between eight different donors — and one service returned a "healthy" assessment for two aliquots of that identical material and an "unhealthy" assessment for the third [20].
  • The other half of that finding has to travel with it: the same authors report that reproducibility within a single locked-down workflow "tend[s] to be very good", and setting that one service aside, replicates from the same provider shared genera accounting for at least 95% of the identified sample composition [20].

What a second sample can tell you — and what it cannot

A second sample can tell you how much your profile moves. It cannot tell you whether it moved in a good direction, because no target state has been defined and no validated measure of direction exists — an international expert consensus statement found the evidence supporting the clinical usefulness of microbiome testing to be scarce [13]. That distinction is the whole subject of this article, and everything below follows from it.

The useful framing is not progress but signal and noise. Between any two of your results sit a fortnight of ordinary life, two collection events, two shipments, two extractions, two sequencing runs and two passes through a computational pipeline. Some of the difference you see is you. Some of it is the measurement. Distinguishing them is possible, but only under conditions you have to deliberately create.

The part that stays put

The gut microbiome is genuinely individual and genuinely persistent, and this is the strongest thing the longitudinal literature supports.

Low-error sequencing of 37 US adults sampled for up to five years, combined with whole-genome sequencing of more than 500 cultured isolates, found stability following a power-law function which, on extrapolation, suggests most strains in an individual are residents for decades. Shared strains were recovered from family members but not from unrelated people [1]. In a cohort of 308 adult men given four stool samples each — a pair collected 24 to 72 hours apart, a second pair around six months later — within-person taxonomic and functional variation was consistently lower than between-person variation [2]. Individuality is not confined to bacteria either: longitudinal analysis of faecal viruses found high temporal stability and a numerically predominant, individual-specific persistent personal virome [4].

The part that moves

Persistent membership does not mean fixed proportions, and the two are routinely confused.

The densest early time series — two individuals, four body sites, 396 timepoints — found pronounced variability in an individual's microbiota across months, weeks and even days, with only a small fraction of the taxa at a given site present at all timepoints. The authors' conclusion was carefully hedged: no core temporal microbiome exists at high abundance. Many taxa are persistent but non-permanent members, cycling in and out of detection [3].

An intensive antibiotic time series makes the shape of this clearer. Across 3 individuals sampled 52 to 56 times over ten months, day-to-day temporal variability was evident but constrained around an average community composition that was stable over several months in the absence of deliberate perturbation [11]. That is the right mental model: a stable average with a dynamic regimen around it. A single point on that trajectory is neither wrong nor definitive.

Variability is itself a personal trait

The most useful longitudinal finding for anyone considering a second test is that people differ in how much they differ. In 85 adults sampled weekly for three months across four body sites, gut communities varied mostly in the relative abundances of taxa — and there was a wide range of temporal variability across the study population, with some individuals carrying markedly more variable communities than others. Individuals with more diverse gut communities were associated with greater compositional stability in that cohort [9] — an observational association in 85 adults over three months. It is information about variability, not about health; see understanding gut diversity. The authors concluded that temporal dynamics may need to be considered when attempting to link changes in microbiome structure to changes in health status [9].

Three ways a difference can be technical rather than real

Before reading any difference as biology, it is worth knowing how large the alternative explanations are.

Collection and shipping

A controlled reproducibility study of 19 healthy volunteers across seven collection methods, stored for seven days at room temperature, 4 °C or 30 °C and compared with immediately frozen samples, found significant variation for all collection methods against that reference. Stability was high — at or above 0.75 — for tubes going through standard screening procedures, but with exceptions: the relative abundance of Actinobacteria sat at 0.65, and across different tubes stored at 30 °C the range was 0.41 to 0.90, while at room temperature it ran from 0.06 to 0.94 [18]. An intraclass correlation of 0.06 for a phylum means essentially none of the observed variation for that phylum reflected the donor.

The same study reported the necessary counterweight: interindividual variability was much higher than the variability introduced by the collection method [18]. Handling noise is real without being dominant. It simply has to be held constant if two of your own samples are to be compared.

The laboratory

The laboratory contributes more than most readers expect. Blinded specimen sets sequenced by 15 laboratories and analysed with 9 bioinformatics protocols produced variability driven most by biospecimen type and origin, then DNA extraction, then sample handling, then computation; analysis of artificial communities with known composition revealed genuine differences in extraction efficiency and classification between laboratories [5]. Benchmarking 21 DNA extraction protocols on the same faecal samples found extraction had the largest effect on the outcome, larger than library preparation or storage, and contrasted those effects explicitly against biological variation within a person over time [6].

Those were research laboratories. The consumer services have now been tested directly. A US National Institute of Standards and Technology group ordered three kits from each of seven direct-to-consumer gut microbiome services and inoculated every kit with aliquots of the same standardised, homogenised human faecal material, so that any difference between the returned reports was method rather than biology. Across the nine analyses — the seven services plus two reference workflows run in-house — the number of genera identified ranged from 34 to 906 out of 1,208 identified in total, and only 17 genera appeared in every sample once a single anomalous replicate was set aside: under 2%. For 17 of the 18 genera that every service reported, variation between services either exceeded the variation between eight different donors run on one common workflow, or could not be statistically distinguished from it. All seven services reported on the presence or absence of one particular organism in that identical material: three reported it present and four reported it absent [20]. A separate European exercise sent aliquots of one healthy donor's stool to six testing kits and received three verdicts of "excellent" or "good" bacterial diversity, two of "average" and one of "unfavourable" — from the same sample [21].

Two things that study cannot tell you, and says so. Because the material is a reference standard rather than a known biological truth, it measures precision, not accuracy: the authors state the work cannot determine which result was closest to the real composition, and no service was shown to be wrong. And seven services were tested, not the category — a provider that was not in the study was not evaluated by it [20].

The arithmetic

The last source of false change is the most easily missed, because it survives a perfect laboratory. Sequencing data are compositional — the mechanics are set out under how gut microbiome testing works — and the consequence for a second test is direct: if one organism genuinely rises, every other organism's percentage must fall even if nothing about it changed [15]. Underneath that, total microbial load differs by up to tenfold between healthy people — and the authors of that work state the consequence for exactly this use case: comparative analyses of relative microbiome data cannot provide information about the extent or directionality of changes in taxon abundance or metabolic potential [17]. The same study showed that the widely reported BacteroidesPrevotella trade-off was an artefact of relative analysis.

Deciding which organisms changed is method-dependent too. Fourteen differential-abundance methods benchmarked across 38 datasets identified drastically different numbers and sets of significant sequence variants, with results depending on pre-processing and, for many tools, on sample size and sequencing depth [16]. If that is true for group comparisons with statistical power behind them, it is more so for a comparison of two samples from one person.

What documented change actually looks like

Set against that noise floor, the changes the literature has captured cleanly are large, and they have identifiable causes.

Daily sampling of two individuals over a year found overall communities stable for months, punctuated by rare life events that changed them rapidly and broadly. In one of those two participants, travel from an industrialised to a developing country produced a nearly two-fold increase in the ratio of Bacteroidetes to Firmicutes, which reversed on return. An enteric infection in the other participant caused a permanent decline of most gut bacterial taxa, which were replaced by genetically similar species. Even during stable periods, changes in fibre intake correlated with next-day abundance changes in about 15% of gut microbiota members [7].

Antibiotics give the clearest timescale. A four-day course of three last-resort antibiotics in 12 healthy men — not a typical prescription — was followed by recovery to near-baseline composition within about 1.5 months, but nine common species present in all subjects beforehand remained undetectable in most of them at day 180 [10]. In the ciprofloxacin time series, communities shifted within 3–4 days, began returning about a week after each course, and often did not return completely; the responses varied between individuals and between two courses in the same individual, ending in a state that was stable but altered [11].

Diet can move things quickly when it is extreme: diets composed entirely of animal or plant products altered community structure within days and overwhelmed inter-individual differences in microbial gene expression [12]. Ordinary eating is a subtler input — composition reflects multiple days of dietary history rather than the last meal, and daily responses to the same foods are highly personalised [8]. The fermentation biology that underlies this is covered in gut microbiome and digestion.

One preliminary finding deserves a line, because it inverts a common assumption. In the same daily-sampling study, in a two-person subgroup, participants consuming only meal-replacement beverages did not become more stable; overall dietary diversity, not monotony, was what tracked with microbiome stability [8].

Designing a retest so it means something

Nothing above argues against sampling more than once. It argues for doing it under conditions that make comparison possible.

  • ·Hold the method constant — and treat that as necessary, not sufficient. Same provider, same collection kit, same gene region, same pipeline. Differently processed results are much harder to compare after the fact, and standardising the method up front is what the benchmarking literature recommends [5][6]; sending one standardised material to seven consumer services produced between-service variability on the same scale as the biological variability between eight different donors [20]. Reproducibility inside one locked-down workflow is the part that generally does work — the authors of that evaluation state it "tend[s] to be very good", and setting aside the single service that failed, same-service replicates shared genera accounting for at least 95% of the identified sample composition [20]. But generally is not always: in that same study, three replicates of one material from one service produced two "healthy" verdicts and one "unhealthy" one. A constant method makes two of your results comparable in principle. It does not certify that either one was produced correctly.
  • ·Sample more than twice if you can. The studies that characterise a person's own range use many timepoints [3][9]; two points give you no range to read a difference against. In fairness to the opposite case, the authors of one of those cohorts concluded that a single measurement can itself provide long-term information about composition and functional potential [3]. There is no validated number of samples either — this is a practical suggestion, not an evidence-based one. A short run of samples establishes your own range; a single pair does not.
  • ·Record the conditions. Recent antibiotics, illness, travel and major dietary change are the inputs the literature shows moving communities [7][10][11][12]. Without them, an unexplained difference stays unexplained.
  • ·Expect movement. Day-to-day and week-to-week variation is normal and personal [3][9]. Stability is not a virtue and instability is not a fault.

What repeat testing cannot do

It cannot show that an intervention worked, because the outcome would have to be defined in something other than the microbiome itself. The most sophisticated study in this space illustrates the point precisely: an 800-person cohort with 46,898 meals under continuous glucose monitoring, a 100-person validation cohort and a blinded randomised trial produced genuinely personalised dietary predictions — but the validated outcome was the blood glucose response, with the microbiome serving as one input among many alongside blood parameters, dietary habits and anthropometrics. The intervention did produce consistent alterations in gut microbiota configuration; what it validated was the glucose measurement, not a microbiome trajectory [19].

It also cannot substitute for clinical assessment. An international multidisciplinary consensus statement on microbiome testing found that evidence supporting its clinical usefulness is scarce, and warned that an increasing number of commercial providers offer direct-to-consumer tests without any consensus on regulation or proven value in clinical practice [13]. A policy analysis in Science states that such tests lack analytical and clinical validity [14]. The seven-provider evaluation is the first direct empirical test of the first half of that statement; its authors conclude that analytical performance is a prerequisite for sound clinical recommendations, and that standards are needed to ensure analytical validity and consumer confidence [20]. A European expert panel reviewing six kits reached a compatible judgement from a different direction, considering the interpretations and recommendations in the reports premature for want of robust scientific evidence, and the accompanying analyses of limited clinical utility [21]. Those are the field's own words about the category, and they are the reason this article frames repeat sampling as a way of learning how much your description moves — rather than as a measure of progress toward a state nobody has defined.

Method boundaries

16S Foundation™

Broad bacterial profiling — taxonomic composition; function is inferred, not measured.

WGS Advanced™

Higher-resolution taxonomy and genomic functional potential (what the community could do).

MetaT Functional™

Expression/activity context at the time of sampling, where supported.

These reports provide sequencing-derived microbiome information and are not diagnostic tests. The method determines which microbial features can be measured; interpretation depends on sample type, analytical method and the strength of the supporting evidence.

What this can help explain

  • Describe the microbial composition detected in the sample
  • Report method-supported diversity and ecological metrics, and what each one measures
  • Compare repeated samples from the same person when collection and analysis are done the same way
  • Set out the limits of the method used, and the questions worth asking next

What it does not mean

  • Not a diagnosis of any disease or medical condition
  • Not a disease-risk estimate or prediction
  • Not a health score — diversity is not a measure of how well you are
  • No universal "normal microbiome": no validated reference range exists for any body site
  • Not proof that a detected organism or pathway caused a symptom (association ≠ cause)
  • Not a reading of your physiology — microbial DNA or RNA does not measure how your body is working
  • A deeper sequencing tier resolves more detail; it does not make a result more clinically meaningful
  • Not a substitute for professional medical advice

Frequently asked questions

How often should I retest?

No established interval exists. We could not verify any published source specifying how often a person should repeat a microbiome test, or defining how large a difference must be before it means something in one individual. Any interval offered by any provider is a practical convention rather than an evidence-based one, and an international expert consensus statement found the evidence supporting the clinical usefulness of microbiome testing to be scarce.

Can a repeat test show whether a dietary change worked?

Not on its own, because the outcome would have to be defined in something other than the microbiome. In the most sophisticated study in this area — an 800-person cohort with 46,898 monitored meals, plus a validation cohort and a blinded randomised trial — the validated outcome was the blood glucose response and the microbiome was one input among many. The intervention did alter gut microbiota configuration, but what was validated was the glucose measurement, not a microbiome trajectory.

How do I tell a real change from measurement noise?

Sample more than twice if you can, and keep the whole method chain identical between samples. The studies that characterise a person's own range use many timepoints; two points give you no range to read a difference against. Keeping the method constant is necessary but not sufficient: when one standardised material was sent to seven consumer services, reproducibility within a single locked-down workflow was reported as generally very good, yet one service still returned two healthy verdicts and one unhealthy verdict from three aliquots of that same material. There is no validated number of samples either — this is a practical suggestion, not an evidence-based one, and no validated retest interval exists.

One organism's percentage fell between my two tests. Did it decrease?

Not necessarily. Sequencing reports shares of an instrument-imposed total, so if one organism genuinely rises, every other percentage must fall even if nothing about those organisms changed. The authors of the study that measured total microbial load state the consequence directly: comparative analyses of relative microbiome data cannot provide information about the extent or directionality of changes in a taxon's abundance.

How long after antibiotics does the gut microbiome recover?

There is no general rule, only specific studies. After a four-day course of three last-resort antibiotics — not a typical prescription — 12 healthy men returned to near-baseline composition within about 1.5 months, but nine common species present in all of them beforehand remained undetectable in most at day 180. In a separate time series, ciprofloxacin shifted communities within 3–4 days, with recovery beginning about a week after each course and often remaining incomplete.

Is my microbiome supposed to be stable?

Some movement is normal, and how much moves is partly a personal trait. In a densely sampled time series across 396 timepoints, only a small fraction of taxa at a body site were present at every timepoint, and the authors concluded that no core temporal microbiome exists at high abundance. In 85 adults sampled weekly for three months, individuals differed widely in how variable their communities were, with more diverse gut communities being more stable in composition.

Can I compare a result from one company with a result from another?

No, and this has now been tested directly rather than inferred. Aliquots of one standardised, homogenised human faecal material were sent to seven direct-to-consumer gut testing services: variability between services was on the same scale as the biological variability between eight different donors, and for 17 of the 18 genera every service reported, between-service variation either exceeded that between-donor variation or could not be distinguished from it. A separate European exercise sent one healthy donor's sample to six kits and received verdicts on bacterial diversity ranging from excellent to unfavourable. Because the test article was a reference material rather than a known biological truth, that work measures precision rather than accuracy — no service was shown to be wrong, and services not in the study were not evaluated by it. Comparing your own samples to each other is only meaningful if the whole method chain was held constant between them, and even then a constant method is necessary rather than sufficient.

Scientific references

  1. Faith JJ, Guruge JL, Charbonneau M, Subramanian S, Seedorf H, Goodman AL, Clemente JC, Knight R, Heath AC, Leibel RL, Rosenbaum M, Gordon JI. The long-term stability of the human gut microbiota Science (2013) ; 341(6141):1237439 .

    Human cohort study DOI: 10.1126/science.1237439 PMID: 23828941

    Microbiota stability followed a power-law function, which on extrapolation suggests most strains in an individual are residents for decades…

  2. Mehta RS, Abu-Ali GS, Drew DA, Lloyd-Price J, Subramanian A, Lochhead P, Joshi AD, Ivey KL, Khalili H, Brown GT, DuLong C, Song M, Nguyen LH, Mallick H, Rimm EB, Izard J, Huttenhower C, Chan AT. Stability of the human faecal microbiome in a cohort of adult men Nature Microbiology (2018) ; 3(3):347–355 .

    Human cohort study DOI: 10.1038/s41564-017-0096-0 PMID: 29335554

    Within-person taxonomic and functional variation was consistently lower than between-person variation over time…

  3. Caporaso JG, Lauber CL, Costello EK, Berg-Lyons D, Gonzalez A, Stombaugh J, Knights D, Gajer P, Ravel J, Fierer N, Gordon JI, Knight R. Moving pictures of the human microbiome Genome Biology (2011) ; 12(5):R50 .

    Review DOI: 10.1186/gb-2011-12-5-r50 PMID: 21624126

    Despite stable differences between body sites and between individuals, there is pronounced variability in an individual's microbiota across months, weeks and even days…

  4. Shkoporov AN, Clooney AG, Sutton TDS, Ryan FJ, Daly KM, Nolan JA, McDonnell SA, Khokhlova EV, Draper LA, Forde A, Guerin E, Velayudhan V, Ross RP, Hill C. The Human Gut Virome Is Highly Diverse, Stable, and Individual Specific Cell Host & Microbe (2019) .

    Human cohort study DOI: 10.1016/j.chom.2019.09.009 PMID: 31600503

    High temporal stability and individual specificity of the faecal virome; a numerically predominant individual-specific "persistent personal virome"…

  5. Sinha R, Abu-Ali G, Vogtmann E, Fodor AA, Ren B, Amir A, Schwager E, Crabtree J, Ma S, Abnet CC, Knight R, White O, Huttenhower C. Assessment of variation in microbial community amplicon sequencing by the Microbiome Quality Control (MBQC) project consortium Nature Biotechnology (2017) ; 35(11):1077–1086 .

    Mechanistic study DOI: 10.1038/nbt.3981 PMID: 28967885

    The paper's framing sentence is the finding: "Achieving sufficient reproducibility in microbiome research has proven challenging…

  6. Costea PI, Zeller G, Sunagawa S, Pelletier E, Alberti A, Levenez F, Tramontano M, Driessen M, Hercog R, Jung FE, Kultima JR, Hayward MR, Coelho LP, Allen-Vercoe E, Bertrand L, Blaut M, Brown JRM, Carton T, Cools-Portier S, Daigneault M, Derrien M, Druesne A, de Vos WM, Finlay BB, Flint HJ, Guarner F, Hattori M, Heilig H, Luna RA, van Hylckama Vlieg J, Junick J, Klymiuk I, Langella P, Le Chatelier E, Mai V, Manichanh C, Martin JC, Mery C, Morita H, O'Toole PW, Orvain C, Patil KR, Penders J, Persson S, Pons N, Popova M, Salonen A, Saulnier D, Scott KP, Singh B, Slezak K, Veiga P, Versalovic J, Zhao L, Zoetendal EG, Ehrlich SD, Dore J, Bork P. Towards standards for human fecal sample processing in metagenomic studies Nature Biotechnology (2017) ; 35(11):1069–1076 .

    Mechanistic study DOI: 10.1038/nbt.3960 PMID: 28967887

    DNA extraction had the LARGEST effect on the outcome of metagenomic analysis — larger than library preparation and larger than storage…

  7. David LA, Materna AC, Friedman J, Campos-Baptista MI, Blackburn MC, Perrotta A, Erdman SE, Alm EJ. Host lifestyle affects human microbiota on daily timescales Genome Biology (2014) ; 15(7):R89 .

    Human cohort study DOI: 10.1186/gb-2014-15-7-r89 PMID: 25146375

    Overall communities were stable for months; but rare life events rapidly and broadly changed them…

  8. Johnson AJ, Vangay P, Al-Ghalith GA, Hillmann BM, Ward TL, Shields-Cutler RR, Kim AD, Shmagel AK, Syed AN, Walter J, Menon R, Koecher K, Knights D. Daily Sampling Reveals Personalized Diet-Microbiome Associations in Humans Cell Host & Microbe (2019) .

    Human cohort study DOI: 10.1016/j.chom.2019.05.005 PMID: 31194939

    Microbiome composition depended on multiple days of dietary history (not the previous meal), and daily microbial responses to diet were highly personalized…

  9. Flores GE, Caporaso JG, Henley JB, Rideout JR, Domogala D, Chase J, Leff JW, Vázquez-Baeza Y, Gonzalez A, Knight R, Dunn RR, Fierer N. Temporal variability is a personalized feature of the human microbiome Genome Biology (2014) ; 15(12):531 .

    Human cohort study DOI: 10.1186/s13059-014-0531-y PMID: 25517225

    Directly rebuts the assumption "temporal variability is negligible for healthy adults." Gut and tongue communities varied mostly in the relative abundances of taxa…

  10. Palleja A, Mikkelsen KH, Forslund SK, Kashani A, Allin KH, Nielsen T, Hansen TH, Liang S, Feng Q, Zhang C, Pyl PT, Coelho LP, Yang H, Wang J, Typas A, Nielsen MF, Nielsen HB, Bork P, Wang J, Vilsbøll T, Hansen T, Knop FK, Arumugam M, Pedersen O. Recovery of gut microbiota of healthy adults following antibiotic exposure Nature Microbiology (2018) ; 3(11):1255–1265 .

    Human cohort study DOI: 10.1038/s41564-018-0257-9 PMID: 30349083

    Initial changes = blooms of enterobacteria and pathobionts, depletion of Bifidobacterium and butyrate producers. Gut microbiota recovered to near-baseline composition within 1…

  11. Dethlefsen L, Relman DA. Incomplete recovery and individualized responses of the human distal gut microbiota to repeated antibiotic perturbation Proceedings of the National Academy of Sciences of the United States of America (2010) ; 108 Suppl 1(Suppl 1):4554–61 .

    Human cohort study DOI: 10.1073/pnas.1000087107 PMID: 20847294

    This is the reference that separates noise from signal biologically. Interindividual variation was the major source of variability between samples…

  12. David LA, Maurice CF, Carmody RN, Gootenberg DB, Button JE, Wolfe BE, Ling AV, Devlin AS, Varma Y, Fischbach MA, Biddinger SB, Dutton RJ, Turnbaugh PJ. Diet rapidly and reproducibly alters the human gut microbiome Nature (2013) ; 505(7484):559–63 .

    Human cohort study DOI: 10.1038/nature12820 PMID: 24336217

    Short-term extreme macronutrient change altered microbial community structure and overwhelmed inter-individual differences in microbial gene expression…

  13. Porcari S, Mullish BH, Asnicar F, Ng SC, Zhao L, Hansen R, O'Toole PW, Raes J, Hold G, Putignani L, Hvas CL, Zeller G, Koren O, Tun H, Valles-Colomer M, Collado MC, Fischer M, Allegretti J, Iqbal T, Chassaing B, Keller J, Baunwall SM, Abreu M, Barbara G, Zhang F, Ponziani FR, Costello SP, Paramsothy S, Kao D, Kelly C, Kupcinskas J, Youngster I, Franceschi F, Khanna S, Vehreschild M, Link A, De Maio F, Pasolli E, Blanco Miguez A, Brigidi P, Posteraro B, Scaldaferri F, Rajilic Stojanovic M, Megraud F, Malfertheiner P, Masucci L, Arumugam M, Kaakoush N, Segal E, Bajaj J, Leong R, Cryan J, Weersma RK, Knight R, Guarner F, Shanahan F, Cani PD, Elinav E, Sanguinetti M, de Vos WM, El-Omar E, Doré J, Marchesi J, Tilg H, Sokol H, Segata N, Cammarota G, Gasbarrini A, Ianiro G. International consensus statement on microbiome testing in clinical practice The Lancet Gastroenterology & Hepatology (2025) ; 10(2):154–167 .

    Consensus statement DOI: 10.1016/S2468-1253(24)00311-X PMID: 39647502

    Verbatim from the retrieved abstract: "evidence supporting its clinical usefulness is scarce…

  14. Hoffmann DE, von Rosenvinge EC, Roghmann MC, Palumbo FB, McDonald D, Ravel J. The DTC microbiome testing industry needs more regulation Science (2024) ; 383(6688):1176–1179 .

    Review DOI: 10.1126/science.adk4271 PMID: 38484067

    The retrieved abstract states, in full: "Tests lack analytical and clinical validity, requiring more federal oversight to prevent consumer harm." Analytical validity = does the test measure what it claims, reproducibly…

  15. Gloor GB, Macklaim JM, Pawlowsky-Glahn V, Egozcue JJ. Microbiome Datasets Are Compositional: And This Is Not Optional Frontiers in Microbiology (2017) ; 8:2224 .

    Mechanistic study DOI: 10.3389/fmicb.2017.02224 PMID: 29187837

    Sequencing datasets are compositional — they carry "an arbitrary total imposed by the instrument," so only relative abundances are observed…

  16. Nearing JT, Douglas GM, Hayes MG, MacDonald J, Desai DK, Allward N, Jones CMA, Wright RJ, Dhanani AS, Comeau AM, Langille MGI. Microbiome differential abundance methods produce different results across 38 datasets Nature Communications (2022) ; 13(1):342 .

    Mechanistic study DOI: 10.1038/s41467-022-28034-z PMID: 35039521

    The tools "identified drastically different numbers and sets of significant ASVs, and results depend on data pre-processing…

  17. Vandeputte D, Kathagen G, D'hoe K, Vieira-Silva S, Valles-Colomer M, Sabino J, Wang J, Tito RY, De Commer L, Darzi Y, Vermeire S, Falony G, Raes J. Quantitative microbiome profiling links gut community variation to microbial load Nature (2017) ; 551(7681):507–511 .

    Human cohort study DOI: 10.1038/nature24460 PMID: 29143816

    Up to tenfold differences in microbial load between healthy individuals. Because standard sequencing reports fractions of a library, "comparative analyses of relative microbiome data cannot provide information about the…

  18. Zouiouich S, Mariadassou M, Rué O, Vogtmann E, Huybrechts I, Severi G, Boutron-Ruault MC, Senore C, Naccarati A, Mengozzi G, Kozlakidis Z, Jenab M, Sinha R, Gunter MJ, Leclerc M. Comparison of Fecal Sample Collection Methods for Microbial Analysis Embedded within Colorectal Cancer Screening Programs Cancer Epidemiology, Biomarkers & Prevention (2022) ; 31(2):305–314 .

    Mechanistic study DOI: 10.1158/1055-9965.EPI-21-0188 PMID: 34782392

    The most usable signal-to-noise statement in this ledger. (a) "When compared with the putative gold standard, we observed significant variation for all collection methods…

  19. Zeevi D, Korem T, Zmora N, Israeli D, Rothschild D, Weinberger A, Ben-Yacov O, Lador D, Avnit-Sagi T, Lotan-Pompan M, Suez J, Mahdi JA, Matot E, Malka G, Kosower N, Rein M, Zilberman-Schapira G, Dohnalová L, Pevsner-Fischer M, Bikovsky R, Halpern Z, Elinav E, Segal E. Personalized Nutrition by Prediction of Glycemic Responses Cell (2015) ; 163(5):1079–1094 .

    Randomised controlled trial DOI: 10.1016/j.cell.2015.11.001 PMID: 26590418

    The strongest example available of an intervention producing personalised, individually-actionable output that involved the microbiome — and it is instructive precisely because of what it did and did not do…

  20. Servetas SL, Gierz KS, Hoffmann D, Ravel J, Jackson SA. Evaluating the analytical performance of direct-to-consumer gut microbiome testing services Communications Biology (2026) ; 9(1) .

    Consensus statement DOI: 10.1038/s42003-025-09301-3 PMID: 41748906

    BETWEEN providers — poor comparability. Verbatim from the abstract: "we found variability between providers was on the same scale as biological variability between different donors…

  21. Rodriguez J, Cordaillat-Simmons M, Badalato N, Berger B, Breton H, de Lahondès R, Deschasaux-Tanguy M, Desvignes C, D'Humières C, Kampshoff S, Lavelle A, Metwaly A, Quijada NM, Seegers JFML, Udocor A, Zwart H, Maguin E, Doré J, Druart C. Microbiome testing in Europe: navigating analytical, ethical and regulatory challenges Microbiome (2024) ; 12(1):258 .

    Consensus statement DOI: 10.1186/s40168-024-01991-x PMID: 39695869

    BETWEEN providers only. From one identical sample: three kits reported "excellent" or "good" bacterial diversity, two "average", one "unfavourable"…

How this article was written

This article is built from 21 peer-reviewed sources, listed in full above. Each was retrieved from PubMed and its identifiers checked. Where a mechanism has been demonstrated in laboratory systems or animal models rather than measured in people, the text says so. Figures are quoted with the study population they came from.

This is educational content about microbial biology and measurement. It describes what sequencing-derived microbiome data can and cannot show. It is not a diagnosis, a treatment recommendation, or a substitute for advice from a qualified healthcare professional.

Explore the tests

Related test

Explore the gut health test that pairs with this topic.

Build a baseline you can actually read

GutX processes every sample through the same laboratory and analysis chain, which removes one source of difference between two of your results. It does not make every difference between them meaningful. GutX describes microbial composition over time; it does not diagnose, monitor or predict any health condition.

Keep reading

More in Gut Health

This article is educational and wellness-oriented. It is not a diagnosis, treatment recommendation, or a substitute for professional medical advice. For symptoms or clinical concerns, speak with a qualified healthcare professional.