What a diversity number is actually measuring
Microbial diversity is a family of statistics that summarise one sample: how many different kinds of organism were detected in it, and how evenly the sequencing reads were spread across them. It describes that sample. It is not a health score — and the clearest demonstration of that comes from a body site where the healthy state is the least diverse one.
Two distinct quantities travel under the single word "diversity", and conflating them is the most common reading error.
Alpha diversity: what is inside one sample
Alpha diversity is a within-sample measure. Richness counts how many distinct types were detected. The Shannon and Simpson indices work differently: they weight that count by evenness, so a community in which one organism overwhelmingly dominates scores lower than one in which the same number of types share the space more equally [1][3].
These are ecological statistics borrowed into microbiology, and the borrowing brings a statistical problem with it. A sequencing run observes a sample, not the community the sample came from, and rare organisms are by definition the ones most likely to be missed. The correct posture is to treat a diversity value as an estimate of an unknown parameter of the environment, with bias and variance attached, rather than as a number the sample simply has [1].
Richness suffers most from this. Working across theory, simulated communities and nine metagenomic datasets, Haegeman and colleagues concluded that one cannot reliably estimate the absolute or relative number of microbial species present without making unsupported assumptions about how abundances are distributed — because the sample carries almost no information about the tail of rare organisms. Applying a widely used richness estimator to simulated communities ranked them incorrectly when rare types were plentiful. Shannon and Simpson diversity could be estimated robustly, and the authors recommend them over species richness [2]. Note what that result is and is not: it is a demonstrated property of the estimators, shown in silico and on largely environmental data, not an observation that two indices have reversed the order of a human cohort.
More detail on the individual metrics sits in our reference entry on alpha diversity.
Beta diversity: distance between samples
Beta diversity is a different object. It is not a property of your sample at all — it is a computed distance between samples. UniFrac, one of the standard methods behind microbiome ordination plots, measures phylogenetic distance between sets of taxa as the fraction of branch length leading to descendants from one community but not the other, and is used to compare many communities at once through clustering and ordination [5][6]. An ordination plot then arranges samples in two dimensions so that more similar communities sit closer together [5][9].
There is a further constraint underneath every plot and every percentage. Sequencing datasets are compositional: they have an arbitrary total imposed by the instrument, so a figure of 8% is a share of that arbitrary whole rather than an absolute quantity [7]. This is why one taxon can appear to rise simply because another fell — the measurement mechanics are set out under how gut microbiome testing works — a point we treat separately under relative abundance.
The counterexample: a site where low diversity is the healthy state
The intuition that more diversity means more health is not a general biological rule. It fails, decisively and in a well-characterised human population, at a body site outside the gut.
In 396 asymptomatic North American women, vaginal communities clustered into five groups — and four of those five were dominated by a single *Lactobacillus* species (L. iners, L. crispatus, L. gasseri or L. jensenii), the fifth carrying fewer lactic acid bacteria and more strictly anaerobic organisms [14]. In a separate comparative study, women without bacterial vaginosis carried Lactobacillus at a median of 96% of all bacteria detected, while women with bacterial vaginosis carried a median of 11 different taxa each present at 1% or more [16]. And in a case-control study of 50 affected and 50 healthy women, bacterial vaginosis was associated with a marked increase in the taxonomic richness and diversity of the vaginal community, with no single organism distinguishing the two groups [15].
That is the strong version of the argument. The honest version has a second half, because the mirror error is just as available. When 32 healthy women were sampled twice weekly for 16 weeks, some communities changed markedly over short periods while others stayed relatively stable — and because every participant was healthy, the authors concluded that neither variation in composition nor a higher observed diversity reading is necessarily indicative of dysbiosis [17]. Diversity alone is not the diagnostic in either direction.
It is worth adding that even the physiology varies between groups. In the same 396-woman cohort, vaginal pH also differed by ethnicity — Hispanic 5.0 ± 0.59, black 4.7 ± 1.04, Asian 4.4 ± 0.59, white 4.2 ± 0.3 [14]. That is population variation to be aware of when reading any reference band, not a personal target.
Where "higher is better" came from, and what it actually showed
The idea did not appear from nowhere. In 292 Danish adults, individuals with low bacterial gene richness — 23% of that study population — a sample deliberately enriched for obesity (169 of 292 participants) — showed more marked adiposity, insulin resistance, dyslipidaemia and a more pronounced inflammatory phenotype, and obese individuals in the low-richness group gained more weight over time [10]. A companion dietary intervention found reduced gene richness in 40% of its much smaller sample, and reported that intervention improved gene richness and clinical measures — while being less effective for inflammation markers in exactly the people with lower gene richness [11].
The wider literature has since made that caution sharper. A meta-analysis that re-processed 28 case-control studies across 10 diseases with standardised methods found that results from individual studies can be inconsistent; that some diseases associate with over 50 genera while most associate with only 10–15; and that about half of the genera flagged in individual studies respond to more than one disease — meaning many such associations are not disease-specific but part of a shared, non-specific response [12].
Diversity also varies with things that are not health at all. Across 531 individuals from Venezuela, rural Malawi and the United States, the microbiome showed a shared functional maturation over the first three years of life, alongside pronounced differences in bacterial assemblages between US residents and the other two populations, evident in infancy and adulthood alike [13]. Low diversity in infancy is normal. A profile typical in one population is not typical in another — and the largest healthy reference set of its era was explicitly a Western one, encountering an estimated 81–99% of the genera and community configurations of the healthy Western microbiome and noting that even healthy individuals differ remarkably from one another [27].
The framing that survives all of this is the ecological one: the relationships between diversity and emergent community properties such as stability, productivity or invasibility are much more nuanced than the popular reading allows, and diversity without context provides limited insight. It belongs at the start of an inquiry rather than at the end of one [4].
What moves the number without anything moving in you
A diversity value is the output of a pipeline, and several stages of that pipeline change it.
- ·Sequencing depth. It is common to find as much as 100-fold variation in the number of 16S sequences recovered across samples within a single study, and the diversity metrics microbial ecologists use are sensitive to differences in sequencing effort [20].
- ·How that unevenness is corrected. This is a genuinely unsettled methodological argument, not a solved one. One influential analysis argued that rarefying counts is statistically inadmissible, produces high false-positive rates and discards usable samples [19]. A later benchmarking across 12 datasets found rarefaction was the only method able to control for uneven sequencing effort across common alpha and beta metrics [20], and a re-analysis of the original work identified 11 factors that could have compromised it — noting that even confusion between the terms "rarefying" and "rarefaction" continues to cloud interpretation [21].
- ·What was sequenced. Targeting 16S variable regions with short-read platforms cannot achieve the taxonomic resolution of sequencing the full, roughly 1,500-base-pair gene; and because bacteria carry multiple slightly different intragenomic copies of that gene, any count of "how many types" must contend with copy-number variation [22].
- ·How it was analysed. Method choice materially changes results — clustering approaches, false-discovery-rate control and taxonomic resolution all differ between pipelines [9], and design, execution and analysis are each independent sources of variation that have to be controlled [8].
- ·How it was collected and extracted, and how fast material was moving through the gut. Both are large effects, and both are covered in detail in our article on gut microbiome and digestion — including why transit time and stool consistency are the single largest measured influence on a stool profile, and why 64 randomised fibre trials produced no change in alpha diversity at all.
Is there a healthy range?
This is the question most readers actually have, and it has a clear, verifiable answer.
A multi-stakeholder workshop convened to determine whether sufficient evidence existed to establish measurable gut microbiome characteristics that could serve as indicators of health concluded that mechanistic links between specific changes in gut microbiome structure and markers of human health are not yet established; that it is not established whether dysbiosis is a cause, a consequence, or both; and that biomarkers and surrogate indicators still need to be determined and validated, along with normal ranges [25]. An authoritative review reaches the same place from another direction: using microbiome-based biomarkers for diagnosis, prognosis or risk profiling requires a definition of a healthy microbiome in different populations, which in turn requires strain-level profiling and better knowledge of variation with age, diet, medication, ethnicity and geography — with many gut viruses, phage, fungi and archaea still uncharacterised [26].
Is it at least stable?
Membership, largely yes — proportions, much less so. In 37 US adults sampled for up to five years, microbiota stability followed a power-law function which, extrapolated, suggests most strains remain resident for decades — the five years are observed, the decades are a model projection [23]. In 308 adult men sampled four times, within-person taxonomic and functional variation was consistently lower than between-person variation over time [24].
Reading a diversity number well
- ·Read it as a description of one sample, processed by one method, on one day.
- ·Note which index it is. Richness, Shannon and Simpson are not interchangeable, and richness is the least reliably estimated of them [1][2][3].
- ·Treat any accompanying band as a comparison with that laboratory's reference cohort, not as a clinical range [25][26].
- ·Do not compare it with a number from a different test [9][20][22].
- ·On an ordination plot, read position relative to the other samples shown — nothing more [5][7].
- ·Where a pattern interests you, follow it over repeat samples taken under similar conditions rather than interrogating a single figure.
Diversity is a real, measurable, informative property of a microbial community. It is simply not a verdict — at any body site, in any direction. If you want to see how a sequencing-derived composition is assembled and compared in practice, our comparison of test types sets out what each method can resolve.