Skip to content
Gut Health 13 min read

How gut microbiome testing works

A microbiome test never sees a bacterium. It extracts DNA from your sample, copies and reads chosen fragments, and matches them against a reference database — six steps, each of which has been shown in published work to change the answer. This is what happens at every one of them, and where a result can shift for purely technical reasons.

Educational context — not a diagnostic result

Key takeaways

  • A microbiome test does not look at bacteria. It extracts DNA from a stool sample, copies and reads fragments of it, and matches those fragments against a reference database of sequences. Every organism named on your report is a database match, not a sighting [19][24][25].
  • The single strongest evidence about how much the method matters comes from ring trials: fifteen laboratories sequencing blinded specimen sets with nine bioinformatics protocols produced variability driven mainly by specimen type, then DNA extraction, then handling, then computation [1].
  • The computational step alone is enough. Thirteen laboratories given identical raw sequence files from mock communities of known composition produced different estimates of which organisms were present and in what abundance [2].
  • On the laboratory side, how the DNA is pulled out of the sample matters most. Benchmarking 21 extraction protocols on the same faecal material found extraction had a larger effect on the outcome than library preparation or storage [8].
  • Which stretch of the bacterial marker gene a laboratory chooses to read changes which organisms can appear at all, and the reference database chosen changes how many reads receive a name — species-level assignment ranged from 10.23% to 24.28% across three standard databases using the same classifier and region [12][21].
  • Despite all of that, the largest single source of variation in these studies is still the difference between one person and another [5][21][32].

What actually happens between the post box and the report

Six things, roughly: your sample is preserved and shipped, DNA is extracted from it, a chosen region of that DNA is copied and read, the resulting sequences are matched to names, those names are converted into proportions, and the proportions are compared with something. Every one of those steps is a decision, and every decision has been shown in published work to change the answer.

That is not a scandal. It is what measurement looks like in a field where the instrument is a chain of laboratory and computational choices rather than a single reading. But it does mean a microbiome result is best understood as this material, processed this way — and it explains why two providers can describe the same stool differently without either of them making a mistake.

Step one — collection, preservation and the post

Home collection is not the weak link people assume. Controlled studies with duplicate samples show technical reproducibility is genuinely good: in 52 healthy volunteers across five collection methods, intraclass correlations between duplicate faecal samples ran from 0.64 to 1.00, and stability at room temperature over 96 hours was high for most methods [4]. In a larger study spanning 132 participants across three distinct populations, reproducibility was excellent — at or above 75% — for every collection method tested [6]. A Bangladeshi cohort of 50 adults found technical reproducibility of 0.79 to 0.99 at both day zero and day four [7].

The caveats are specific rather than general. Agreement was weaker for measures that depend on exact relative abundance than for broad diversity measures, and some preservation chemistries performed poorly: intraclass correlations fell below 0.60 for 95% ethanol on abundance-sensitive metrics [4], and stability was not excellent for no-solution or 70% ethanol collection after four to seven days at ambient temperature [6]. Compared against immediately frozen samples, many correlations were low even where the rank order of a person's taxa was largely preserved [7].

The practical implication is narrow and worth stating: use the preservative supplied, and post within the stated window. The published stability data are specific to particular chemistries and particular numbers of days, so they do not license a general "it will be fine".

Step two — getting the DNA out

Once the tube reaches the laboratory, the bacteria have to be broken open and their DNA recovered. This is the step with the largest documented laboratory-side effect. When 21 representative extraction protocols were applied to the same faecal samples and compared against biological variation within a specimen and within a person over time, DNA extraction had the largest effect on the outcome of metagenomic analysis — larger than library preparation, larger than storage. Extraction choice biased estimates of community diversity and the ratio of Gram-positive to Gram-negative organisms [8].

The clearest demonstration of what that means in practice is uncomfortable. In stool from two breast-fed infants, an extraction kit with no bead-beating step produced sequence data containing no bifidobacteria at all — even with primers chosen to detect them — despite bifidobacteria being the most abundant organisms detected by microscopy in the very same samples [9]. Organisms with tough cell walls resist gentle extraction; if they are not broken open, their DNA is not in the tube, and what is not in the tube cannot be sequenced.

Step three — choosing what to read

Most consumer testing amplifies part of the 16S ribosomal RNA gene: a gene all bacteria carry, containing stretches that are near-identical across the domain (useful for grabbing everything) alongside hypervariable regions that differ between groups (useful for telling them apart). Copying that region by PCR is what makes a tiny amount of DNA readable. It also introduces bias at every sub-step, because PCR-based methods have multiple stages, each susceptible to error and bias [10].

Which region you read changes who you find

There is no universal 16S region, and the choice is consequential. In a controlled comparison of four primer pairs on the same human subgingival plaque samples — oral, not gut — targeting V4–V6 failed to detect the genus Fusobacterium entirely, while V7–V9 primers failed to detect Selenomonas, TM7 and Mycoplasma. The dominant genera differed substantially, though a few were dominant whichever region was read between the three region sets [12]. A separate oral study found that broad between-sample relationships survived a change of region, but richness estimates did not: phylotype richness was systematically higher for V1–V3 than for V3–V4 [13]. In a non-human wastewater system, one region overestimated a group of archaea by more than thirtyfold against a metagenomic reference [22].

The marker itself is imperfect, and this is a limit no protocol choice fixes. Comparing the 16S gene against core-genome phylogeny across human gut core genera, concordance was only about 50.7% within genera and 73.8% between them; the best-performing hypervariable regions reached 60–62.5%. Roughly 690 ± 110 informative positions are needed for 80% concordance, and the 16S gene averages 254 [14]. The same work reports rRNA operon copy number varying from 1 to 27 per genome — meaning a bacterium with many copies contributes more reads than an equally abundant one with few. Correcting for that remains, in the words of the benchmarking paper, an unsolved problem: the available tools explained under 10% of the variance in some cases and disagreed with each other for most communities tested [15].

Chimeras: sequences that were never there

PCR can also fabricate. When amplification stalls partway through one template and resumes on another, the result is a chimera — a hybrid sequence with no corresponding organism. Benchmarked against a mock community of known composition, chimeras formed reproducibly across independent amplifications, and rates exceeded 70% for less-abundant species before filtering, on Sanger and 454 platforms — which is why the same paper introduced a chimera-detection tool now in standard use. They can be falsely interpreted as novel organisms, inflating apparent diversity. Notably, shotgun sequences of the same mock community appeared devoid of 16S chimeras [11].

Step four — sequencing

The sequencing platform itself is a comparatively modest contributor. When short-read and long-read platforms were run on the same extracted DNA from faecal samples, more than 90% of reads on each were classified to genus or species level, and species-level misclassification between the two approaches was smaller than the authors expected [16]. That is a single small study, so it is illustrative rather than definitive — but it is consistent with the ring trials, which place specimen, extraction and analysis ahead of sequencing chemistry as sources of variation [1][8]. A newer sequencer does not repair an extraction, primer or database bias upstream of it.

Step five — turning reads into names

Raw output is millions of short sequences. Two decisions convert them into a list of organisms.

The first is how reads are grouped. Historically they were clustered into operational taxonomic units — bins of reads differing by less than a fixed threshold. Modern error-modelling resolves amplicon sequence variants exactly, down to single nucleotides. The argument for the newer approach is not mainly resolution but portability: sequence variants are consistent labels with intrinsic biological meaning, identified independently of any reference database, which makes results from separately processed datasets comparable in a way that cluster-based results are not [19].

The second is which reference database the sequences are compared against, and this changes how much you learn. In a benchmark using one classifier on one gene region against a bacterial sequence test set, species-level assignment ranged from 10.23% to 24.28% across the three standard databases, and genus-level from 70.88% to 87.20% [21]. A gut-specific database significantly raised assignment rates over general-purpose ones — and the authors of that work noted plainly that even for an ecosystem as well studied as the human intestine, assigning genus and species names to 16S reads remains challenging [20].

Some of what is in a stool sample has no name because no name exists yet. Across 9,428 human metagenomes, 77% of the 4,930 reconstructed species-level genome groups had no genome in public repositories as of that 2019 analysis; reference catalogues have grown since; they were present in 93% of well-assembled samples, and adding them lifted the share of gut reads that could be mapped from about 68% to 88% [23]. Those unknown groups were enriched in non-Westernised populations — an equity caveat worth stating plainly, since it means reference completeness is not the same for everybody.

Shotgun metagenomic sequencing reads all the DNA present rather than one gene, adding strain-level resolution and a catalogue of the genes the community carries [24]. That is genomic potential.Function inferred from a 16S profile is a prediction rather than a measurement, with uncertainty the tool's own authors document [25] — explained in full under gut microbiome and digestion. We compare the approaches in sequencing methods, and the wider gene-to-function distinction is set out in gut microbiome and digestion.

Step six — turning names into numbers

Sequencing produces shares, not counts. The instrument imposes an arbitrary total, so the data are compositional: if one organism's share rises, others must fall whether or not anything about them changed [26]. This is not a subtlety to be filed away — total microbial load differs by up to tenfold between healthy people, and quantitative counting demonstrated that the widely reported trade-off between Bacteroides and Prevotella was an artefact of relative profiling [27]. We treat this at length in relative abundance.

Then something has to decide what counts as different. Fourteen differential-abundance methods benchmarked across 38 datasets identified "drastically different numbers and sets" of significant sequence variants, with results depending on pre-processing and, for many tools, on sample size and sequencing depth; only two were consistent across studies, and the authors recommend agreement across several methods rather than trust in one [28]. Common normalisation shortcuts are not neutral either: both simple proportions and rarefying — randomly subsampling every sample to equal depth — produce high false-positive rates in tests for differentially abundant organisms [29].

Contamination, and why controls exist

Extraction kits and laboratory reagents contain bacterial DNA of their own. This is a documented, reproducible property of the reagents rather than a matter of laboratory tidiness: contaminating DNA is ubiquitous in commonly used kits, varies between kits and even between batches of the same kit, and affects both amplicon and shotgun work [17]. Sequencing is sensitive enough to detect that contaminant DNA as efficiently as real signal [18].

The size of the problem depends entirely on how much genuine material is present. Stool is a high-biomass sample, so reagent DNA is a small fraction of what is read; the acute risk is in low-biomass samples such as tissue or blood, where claims should be treated with corresponding suspicion [18]. The standard safeguard is the same in both cases: concurrent sequencing of negative controls alongside real samples is strongly advised [17].

What this means for reading a report

Three practical consequences follow, none of which requires distrusting the measurement.

Results from different providers are not directly comparable. Different gene region, different reference database, different bioinformatic pipeline — and it is the processing rather than the extraction chemistry that has been shown to drive incomparability; one benchmarking study concluded results can be robust to the extraction and sequencing approach while remaining incomparable once bioinformatically processed. In a study that sent matched intestinal biopsy samples from 32 people to three laboratories, broad group-level signal held up while taxonomic assignment and abundance estimates did not, and the authors concluded that combining differently processed samples is nearly impossible [3]. That is a statement about method, not about quality.

Your own samples are comparable to each other, if the method is held constant. That is the basis for reading change over time, covered in retesting and longitudinal tracking.

Between-person differences remain the largest signal. In a cohort of 308 adult men sampled four times over roughly six months, within-person taxonomic and functional variation was consistently lower than between-person variation — although gene expression profiles were as variable within a person as between people, which is one reason a DNA-based test reports composition and functional potential rather than activity [32].

Method boundaries

16S Foundation™

Broad bacterial profiling — taxonomic composition; function is inferred, not measured.

WGS Advanced™

Higher-resolution taxonomy and genomic functional potential (what the community could do).

MetaT Functional™

Expression/activity context at the time of sampling, where supported.

These reports provide sequencing-derived microbiome information and are not diagnostic tests. The method determines which microbial features can be measured; interpretation depends on sample type, analytical method and the strength of the supporting evidence.

What this can help explain

  • Describe the microbial composition detected in the sample
  • Report method-supported diversity and ecological metrics, and what each one measures
  • Compare repeated samples from the same person when collection and analysis are done the same way
  • Set out the limits of the method used, and the questions worth asking next

What it does not mean

  • Not a diagnosis of any disease or medical condition
  • Not a disease-risk estimate or prediction
  • Not a health score — diversity is not a measure of how well you are
  • No universal "normal microbiome": no validated reference range exists for any body site
  • Not proof that a detected organism or pathway caused a symptom (association ≠ cause)
  • Not a reading of your physiology — microbial DNA or RNA does not measure how your body is working
  • A deeper sequencing tier resolves more detail; it does not make a result more clinically meaningful
  • Not a substitute for professional medical advice

Frequently asked questions

What actually happens to my stool sample?

It is preserved and shipped, DNA is extracted from it, a chosen region of that DNA is copied and read, the resulting sequences are matched to names in a reference database, those names are converted into proportions, and the proportions are compared with something. Every one of those steps is a methodological decision, and published benchmarking shows each can alter the reported profile.

Why might two companies report different results for the same stool?

Because they make different choices at each step. In one inter-laboratory study, thirteen laboratories analysed identical raw sequence files with their own pipelines and produced different estimates of which organisms were present and in what abundance. A larger ring trial with fifteen laboratories and nine bioinformatics protocols found variability driven mainly by specimen type, then DNA extraction, then handling, then computation.

Does collecting a sample at home make it unreliable?

Technical reproducibility of home collection is good. Duplicate faecal samples in 52 healthy volunteers gave intraclass correlations of 0.64 to 1.00, and a 132-participant study found reproducibility at or above 75% for every method tested. Agreement is weaker for measures that depend on exact relative abundance, and some preservation chemistries perform poorly after several days at ambient temperature, so using the supplied preservative and posting within the stated window matters.

Why does my report say 'unclassified'?

Usually because the reference catalogue is incomplete rather than because anything is unusual about your gut. Across 9,428 human metagenomes, 77% of reconstructed species-level genome groups had no genome in public repositories; adding them raised the share of gut reads that could be mapped from about 68% to 88%. Those unknown groups were enriched in non-Westernised populations, so reference completeness is not equal for everybody.

Can 16S sequencing identify bacteria to species level?

Only partially, and the names should be treated as provisional. Using the same classifier and gene region, in a benchmark against a bacterial sequence test set, species-level assignment ranged from 10.23% to 24.28% across three standard reference databases, while genus-level assignment ranged from 70.88% to 87.20%. The 16S gene itself agrees with core-genome relationships only about half the time within a genus, and individual hypervariable regions perform worse.

Does a newer sequencing platform make the result more accurate?

Not on its own. When short-read and long-read platforms were run on the same extracted DNA, more than 90% of reads on each were classified to genus or species level, with less species-level disagreement than the authors expected. Published benchmarking places sample handling, DNA extraction and analysis choices ahead of sequencing chemistry as sources of variation, so a newer sequencer does not repair a bias introduced upstream of it.

Do the percentages tell me how much of each bacterium I have?

No. Sequencing reports shares of a total imposed by the instrument, not counts, so one organism's share can rise because another's fell. Total microbial load also differs by up to tenfold between healthy people, and quantitative cell counting showed that the widely reported Bacteroides–Prevotella trade-off was an artefact of relative profiling.

Scientific references

  1. Sinha R, Abu-Ali G, Vogtmann E, et al. Assessment of variation in microbial community amplicon sequencing by the Microbiome Quality Control (MBQC) project consortium Nature Biotechnology (2017) ; 35(11):1077–1086 .

    Review DOI: 10.1038/nbt.3981 PMID: 28967885

    Blinded specimen sets were sequenced by 15 laboratories and analysed using 9 bioinformatics protocols…

  2. O'Sullivan DM, Doyle RM, Temisak S, et al. An inter-laboratory study to investigate the impact of the bioinformatics component on microbiome analysis using mock communities Scientific Reports (2021) ; 11(1):10590 .

    Mechanistic study DOI: 10.1038/s41598-021-89881-2 PMID: 34012005

    Thirteen laboratories analysed the identical FASTQ files using their own pipelines…

  3. Szamosi JC, Forbes JD, Copeland JK, et al. Assessment of Inter-Laboratory Variation in the Characterization and Analysis of the Mucosal Microbiota in Crohn's Disease and Ulcerative Colitis Frontiers in Microbiology (2020) ; 11:2028 .

    Review DOI: 10.3389/fmicb.2020.02028 PMID: 32973734

    Matched samples from each participant were sent to three laboratories, which independently performed DNA extraction, library prep, amplicon sequencing and data processing…

  4. Vogtmann E, Chen J, Amir A, et al. Comparison of Collection Methods for Fecal Samples in Microbiome Studies American Journal of Epidemiology (2016) ; 185(2):115–123 .

    Mechanistic study DOI: 10.1093/aje/kww177 PMID: 27986704

    Technical reproducibility between duplicate faecal samples was high — ICC 0.64 to 1.00. Stability at room temperature for 96 h was high for most methods, though ICC fell below 0…

  5. Sinha R, Chen J, Amir A, et al. Collecting Fecal Samples for Microbiome Analyses in Epidemiology Studies Cancer Epidemiology, Biomarkers & Prevention (2015) ; 25(2):407–16 .

    Mechanistic study DOI: 10.1158/1055-9965.EPI-15-0951 PMID: 26604270

    Microbiome profiles showed "systematic biases according to sample method and time at ambient temperature" — but "the highest source of variation was between individuals…

  6. Byrd DA, Chen J, Vogtmann E, et al. Reproducibility, stability, and accuracy of microbial profiles by fecal sample collection method in three distinct populations PLoS ONE (2019) ; 14(11):e0224757 .

    Mechanistic study DOI: 10.1371/journal.pone.0224757 PMID: 31738775

    Reproducibility ICCs comparing duplicate samples were excellent (≥75%) for all collection methods. Stability after 4–7 days at ambient temperature was excellent (≥75%) for most methods except no-solution and 70% ethanol…

  7. Vogtmann E, Chen J, Kibriya MG, et al. Comparison of Fecal Collection Methods for Microbiota Studies in Bangladesh Applied and Environmental Microbiology (2017) ; 83(10) .

    Mechanistic study DOI: 10.1128/AEM.00361-17 PMID: 28258145

    Duplicate samples had ICCs for technical reproducibility of 0.79 to 0.99 at day 0 and day 4…

  8. Costea PI, Zeller G, Sunagawa S, et al. Towards standards for human fecal sample processing in metagenomic studies Nature Biotechnology (2017) ; 35(11):1069–1076 .

    Mechanistic study DOI: 10.1038/nbt.3960 PMID: 28967887

    "DNA extraction had the largest effect on the outcome of metagenomic analysis" — larger than library preparation or sample storage, and compared explicitly against biological variation within the same specimen and within…

  9. Walker AW, Martin JC, Scott P, Parkhill J, Flint HJ, Scott KP. 16S rRNA gene-based profiling of the human infant gut microbiota is strongly influenced by sample processing and PCR primer choice Microbiome (2015) ; 3:26 .

    Mechanistic study DOI: 10.1186/s40168-015-0087-4 PMID: 26120470

    Using a DNA extraction kit with no bead-beating step resulted in a complete absence of bifidobacteria in the sequence data, even with optimised primers — despite bifidobacteria being the most abundant organisms detected…

  10. Gohl DM, Vangay P, Garbe J, et al. Systematic improvement of amplicon marker gene methods for increased accuracy in microbiome studies Nature Biotechnology (2016) ; 34(9):942–9 .

    Mechanistic study DOI: 10.1038/nbt.3601 PMID: 27454739

    "PCR-based methods have multiple steps, each of which is susceptible to error and bias"; variance also arises from the different next-generation library-preparation methods used…

  11. Haas BJ, Gevers D, Earl AM, et al. Chimeric 16S rRNA sequence formation and detection in Sanger and 454-pyrosequenced PCR amplicons Genome Research (2011) ; 21(3):494–504 .

    Mechanistic study DOI: 10.1101/gr.112730.110 PMID: 21212162

    Chimeras — hybrid PCR products between multiple parent sequences — "can be falsely interpreted as novel organisms, thus inflating apparent diversity…

  12. Kumar PS, Brooker MR, Dowd SE, Camerlengo T. Target region selection is a critical determinant of community fingerprints generated by 16S pyrosequencing PLoS ONE (2011) ; 6(6):e20956 .

    Review DOI: 10.1371/journal.pone.0020956 PMID: 21738596

    Different target regions produced markedly different genus lists from the same samples. Targeting V4–V6 failed to detect the genus Fusobacterium; the V7–V9 primers failed to detect Selenomonas, TM7 and Mycoplasma…

  13. Zheng W, Tsompana M, Ruscitto A, et al. An accurate and efficient experimental approach for characterization of the complex oral microbiota Microbiome (2015) ; 3:48 .

    Review DOI: 10.1186/s40168-015-0110-9 PMID: 26437933

    Beta-diversity structure correlated significantly between V1–V3 and V3–V4 (Procrustes on unweighted UniFrac) — i.e…

  14. Hassler HB, Probert B, Moore C, et al. Phylogenies of the 16S rRNA gene and its hypervariable regions lack concordance with core genome phylogenies Microbiome (2022) ; 10(1):104 .

    Mechanistic study DOI: 10.1186/s40168-022-01295-y PMID: 35799218

    At the intra-genus level the 16S gene showed only ~50.7% average concordance with the core-genome phylogeny; concordance for individual hypervariable regions was lower still…

  15. Louca S, Doebeli M, Parfrey LW. Correcting for 16S rRNA gene copy numbers in microbiome surveys remains an unsolved problem Microbiome (2018) ; 6(1):41 .

    Mechanistic study DOI: 10.1186/s40168-018-0420-9 PMID: 29482646

    Sequence-variant counts are biased towards clades with more 16S copies; copy number could only be accurately predicted for taxa with ≲15% 16S divergence from a sequenced representative…

  16. Wei PL, Hung CS, Kao YW, et al. Characterization of Fecal Microbiota with Clinical Specimen Using Long-Read and Short-Read Sequencing Platform International Journal of Molecular Sciences (2020) ; 21(19):7110 .

    Review DOI: 10.3390/ijms21197110 PMID: 32993155

    Illumina MiSeq (short-read, V3–V4) and Oxford Nanopore MinION (long-read, full 16S) were run on identical genomic DNA…

  17. Salter SJ, Cox MJ, Turek EM, et al. Reagent and laboratory contamination can critically impact sequence-based microbiome analyses BMC Biology (2014) ; 12:87 .

    Review DOI: 10.1186/s12915-014-0087-z PMID: 25387460

    "Contaminating DNA is ubiquitous in commonly used DNA extraction kits and other laboratory reagents, varies greatly in composition between different kits and kit batches," and "critically impacts results obtained from sa…

  18. Eisenhofer R, Minich JJ, Marotz C, Cooper A, Knight R, Weyrich LS. Contamination in Low Microbial Biomass Microbiome Studies: Issues and Recommendations Trends in Microbiology (2019) ; 27(2):105–117 .

    Mechanistic study DOI: 10.1016/j.tim.2018.11.003 PMID: 30497919

    Next-generation sequencing sensitivity "is a double-edged sword" — it detects contaminant DNA and cross-contamination as efficiently as real signal, and this "can confound the interpretation of microbiome data," especial…

  19. Callahan BJ, McMurdie PJ, Holmes SP. Exact sequence variants should replace operational taxonomic units in marker-gene data analysis The ISME Journal (2017) ; 11(12):2639–2643 .

    Mechanistic study DOI: 10.1038/ismej.2017.119 PMID: 28731476

    Historically, reads were clustered into OTUs — groups of reads differing by less than a fixed dissimilarity threshold…

  20. Ritari J, Salojärvi J, Lahti L, de Vos WM. Improved taxonomic assignment of human intestinal 16S rRNA sequences by a dedicated reference database BMC Genomics (2015) ; 16:1056 .

    Mechanistic study DOI: 10.1186/s12864-015-2265-y PMID: 26651617

    "even with well-characterised ecosystems like the human intestinal microbiota it is challenging to assign genus and species level taxonomy to 16S rRNA amplicon reads…

  21. Agnihotry S, Sarangi AN, Aggarwal R. Construction & assessment of a unified curated reference database for improving the taxonomic classification of bacteria using 16S rRNA sequence data Indian Journal of Medical Research (2020) ; 151(1):93–103 .

    Mechanistic study DOI: 10.4103/ijmr.IJMR_220_18 PMID: 32134020

    With the same classifier and the same region, species-level assignment rates ranged from 10.23% to 24.28% across the three standard databases (Greengenes, SILVA, RDP), and genus-level from 70.88% to 87.20%…

  22. Lin L, Ju F. Evaluation of different 16S rRNA gene hypervariable regions and reference databases for profiling engineered microbiota structure and functional guilds in a swine wastewater treatment plant Interface Focus (2023) ; 13(4):20230012 .

    Mechanistic study DOI: 10.1098/rsfs.2023.0012 PMID: 37303742

    Richness captured by different primers decreased V4 > V4–V5 > V3–V4 > V6–V8/V1–V3, and V6–V8 overestimated archaeal methanogens by over 30-fold relative to the metagenomic reference…

  23. Pasolli E, Asnicar F, Manara S, et al. Extensive Unexplored Human Microbiome Diversity Revealed by Over 150,000 Genomes from Metagenomes Spanning Age, Geography, and Lifestyle Cell (2019) .

    Mechanistic study DOI: 10.1016/j.cell.2019.01.001 PMID: 30661755

    154,723 microbial genomes were reconstructed into 4,930 species-level genome bins, 77% of which had no genome in public repositories…

  24. Quince C, Walker AW, Simpson JT, Loman NJ, Segata N. Shotgun metagenomics, from sampling to analysis Nature Biotechnology (2017) ; 35(9):833–844 .

    Mechanistic study DOI: 10.1038/nbt.3935 PMID: 28898207

    What shotgun sequencing adds over amplicon methods (member cataloguing, functional characterisation, strain-level characterisation) and what remains hard — assembly- and mapping-based profiling of "high-complexity sample…

  25. Langille MGI, Zaneveld J, Caporaso JG, et al. Predictive functional profiling of microbial communities using 16S rRNA marker gene sequences Nature Biotechnology (2013) ; 31(9):814–21 .

    Mechanistic study DOI: 10.1038/nbt.2676 PMID: 23975157

    The authors' own framing: profiling the 16S gene "does not provide direct evidence of a community's functional capabilities…

  26. Gloor GB, Macklaim JM, Pawlowsky-Glahn V, Egozcue JJ. Microbiome Datasets Are Compositional: And This Is Not Optional Frontiers in Microbiology (2017) ; 8:2224 .

    Review DOI: 10.3389/fmicb.2017.02224 PMID: 29187837

    Sequencing datasets are compositional — "they have an arbitrary total imposed by the instrument…

  27. Vandeputte D, Kathagen G, D'hoe K, et al. Quantitative microbiome profiling links gut community variation to microbial load Nature (2017) ; 551(7681):507–511 .

    Human cohort study DOI: 10.1038/nature24460 PMID: 29143816

    Up to tenfold differences in microbial load between healthy individuals. Quantitative profiling "bypasses compositionality effects" and reveals that the taxonomic trade-off between Bacteroides and Prevotella is an artefa…

  28. Nearing JT, Douglas GM, Hayes MG, et al. Microbiome differential abundance methods produce different results across 38 datasets Nature Communications (2022) ; 13(1):342 .

    Mechanistic study DOI: 10.1038/s41467-022-28034-z PMID: 35039521

    The tools "identified drastically different numbers and sets of significant ASVs, and results depend on data pre-processing…

  29. McMurdie PJ, Holmes S. Waste not, want not: why rarefying microbiome data is inadmissible PLoS Computational Biology (2014) ; 10(4):e1003531 .

    Review DOI: 10.1371/journal.pcbi.1003531 PMID: 24699258

    Both simple proportions and rarefying (randomly subsampling every sample to equal depth) "result in a high rate of false positives" in tests for differentially abundant species, and rarefying "often discards samples that…

  30. Hoffmann DE, von Rosenvinge EC, Roghmann MC, Palumbo FB, McDonald D, Ravel J. The DTC microbiome testing industry needs more regulation Science (2024) ; 383(6688):1176–1179 .

    Review DOI: 10.1126/science.adk4271 PMID: 38484067

    The published summary states that direct-to-consumer microbiome "tests lack analytical and clinical validity, requiring more federal oversight to prevent consumer harm…

  31. The Lancet Gastroenterology & Hepatology. Direct-to-consumer microbiome testing needs regulation The Lancet Gastroenterology & Hepatology (2024) ; 9(7):583 .

    Review DOI: 10.1016/S2468-1253(24)00163-8 PMID: 38870959

    A second, independent top-tier journal editorial calling for regulation of DTC microbiome testing. No abstract is available in the PubMed record — do not paraphrase its content beyond its title.

  32. Mehta RS, Abu-Ali GS, Drew DA, et al. Stability of the human faecal microbiome in a cohort of adult men Nature Microbiology (2018) ; 3(3):347–355 .

    Human cohort study DOI: 10.1038/s41564-017-0096-0 PMID: 29335554

    Within-person taxonomic and functional variation was consistently lower than between-person variation over time…

How this article was written

This article is built from 32 peer-reviewed sources, listed in full above. Each was retrieved from PubMed and its identifiers checked. Where a mechanism has been demonstrated in laboratory systems or animal models rather than measured in people, the text says so. Figures are quoted with the study population they came from.

This is educational content about microbial biology and measurement. It describes what sequencing-derived microbiome data can and cannot show. It is not a diagnosis, a treatment recommendation, or a substitute for advice from a qualified healthcare professional.

Explore the tests

Related test

Explore the gut health test that pairs with this topic.

Know the method behind the profile

GutX uses published, peer-reviewed laboratory and analysis methods, and states which ones are used at each step. Comparing the approaches shows what each method can and cannot resolve.

Keep reading

More in Gut Health

This article is educational and wellness-oriented. It is not a diagnosis, treatment recommendation, or a substitute for professional medical advice. For symptoms or clinical concerns, speak with a qualified healthcare professional.