Phylum and genus balance scores are vendor-defined indexes that compress many taxonomic abundances into one “gut balance” or “health score”, often by measuring distance from a reference cohort’s mean profile at phylum or genus level. If your report shows 72% balanced, phylum imbalance detected, or a genus equilibrium index, you are looking at a statistical comparison to other customers or a curated database, not a clinically validated wellness endpoint.
These scores overlap with dysbiosis labels and F:B ratio lines but add proprietary weighting that varies by company. Without published validation, a “low balance” flag is sample-specific context, not a diagnosis.
How balance scores are constructed
Exact algorithms are trade secrets, but common patterns include:
| Step | Typical approach | Limit |
|---|---|---|
| Taxonomic profiling | 16S or shotgun → phylum/genus table | Resolution and database affect all downstream scores |
| Reference cohort | Mean or median profile of “healthy” customers or public data | Geography, diet, and age skew the reference |
| Distance metric | Euclidean, Bray-Curtis, or weighted index vs reference | Sensitive to rare taxa and sequencing depth |
| Thresholding | Flag if distance exceeds percentile cut-off | Cut-offs are internal, not guideline-backed |
| Composite label | May merge balance + diversity + opportunist rules | Double-counts the same taxonomic signal |
Some panels weight “beneficial” genera (Bifidobacterium, Lactobacillus, Faecalibacterium) against “undesirable” buckets (Proteobacteria, opportunists). That embeds value judgments that research cohorts do not treat as universal optima, high Bifidobacterium after a prebiotic trial is not the same construct as high Bifidobacterium in a untreated IBS sample (Duvallet et al., 2017).
Cross-study re-analysis (Duvallet 2017) undercuts the logic further: on average ~51% of genus-level associations in individual disease datasets were genera linked to more than one disease. Clostridiales (especially Lachnospiraceae and Ruminococcaceae, home to Faecalibacterium, Roseburia, and related butyrate producers) were depleted across multiple sick cohorts; Lactobacillales were enriched across multiple diseases as a shared sickness/transit signal. A balance score that penalises Proteobacteria while rewarding Lactobacillus can double-count one ecological story (diarrhoea, antibiotics, fast transit) as both “bad” and “good” depending on the bucket.
| Vendor bucket | Research pattern (Duvallet meta-analysis) | Score risk |
|---|---|---|
| “Beneficial” Firmicutes | Often depleted in IBD and general sickness, not high = healthy | False reassurance when Clostridiales are low |
| “Opportunistic” Proteobacteria | ↑ in diarrhoea; also post-antibiotic | True bloom vs transit overlap |
| “Beneficial” Lactobacillus | ↑ non-specifically in multiple diseases | Penalising low Lactobacillus misses context |
| Composite balance | Mixes disease-specific (e.g. CRC enrichment) with shared responders | One number hides conflicting biology |
Disease-specific enrichment (e.g. Fusobacterium in colorectal cancer across studies) is rarer than shared depletion of health-associated Clostridiales. Balance algorithms that treat all deviation as one “imbalance” cannot separate those patterns.
Overlap with dysbiosis labels
Reports often show balance score, dysbiosis index, and low diversity together. They are related but not redundant:
| Label | What it usually captures |
|---|---|
| Balance score | Distance from reference at chosen taxonomic rank |
| Dysbiosis flag | Often balance + opportunist rules + diversity (vendor-specific) |
| Alpha diversity | Richness/evenness within sample, one axis of ecology |
| F:B ratio | Two-phylum legacy summary (F:B ratio) |
You can have “balanced” phylum scores with low diversity after antibiotics, or “imbalanced” genus scores driven by one dietary shift (e.g. sudden fiber increase) that is physiologically benign. See Dysbiosis for how the ecological term is used on panels versus in research.
High vs low on reports
| Report display | Likely meaning | Common misread |
|---|---|---|
| High imbalance / low score | Deviation from vendor “healthy” centroid | Proof of illness or need for intervention |
| In-range / balanced | Within vendor tolerance band | Guaranteed optimal microbiome |
| Improving trend on retest | Profile moved toward reference | Clinical improvement without symptom change |
Reference cohort effects dominate interpretation. A profile typical in East Asia may read “imbalanced” against a US-customer reference and vice versa (Rinninella et al., 2019). Retests are comparable only with same lab and pipeline (Retesting over time).
What cannot be inferred
Balance scores do not measure:
| Cannot show | Better tool or page |
|---|---|
| Mucosal inflammation | Calprotectin, Gut inflammation markers |
| Barrier permeability | Clinical workup; Intestinal barrier |
| Small-intestinal overgrowth | Breath testing (SIBO) |
| Food intolerance | Elimination diet trials (FODMAPs) |
| Infection requiring antibiotics | Stool culture/PCR; Opportunistic lists |
A surprising pattern on some reports is high balance score alongside persistent symptoms, the index tracks taxonomic similarity to a reference, not symptom mechanism. Functional disorders can coexist with a “balanced” stool profile.
What not to conclude from balance scores
| If the report says… | Do not conclude… |
|---|---|
| Low balance / imbalanced | You have clinical dysbiosis requiring antimicrobials |
| High balance | No need to evaluate symptoms clinically |
| Score below “optimal” | Vendor threshold equals medical reference range |
| Balance improved after supplement | Symptom endpoint proven without separate tracking |
| Genus-level imbalance | Specific genus is pathogenic at that read level |
Related pages
- Reading your microbiome report, hub for report lines
- Dysbiosis, enrichment vs depletion framing
- Duvallet 2017, non-specific responders and balance-score limits
- Alpha diversity, within-sample complexity
- Firmicutes/Bacteroidetes ratio, legacy phylum summary
- Multi-marker synthesis, reconciling conflicting scores
- Retesting over time, valid comparisons