How to Read a Phylogenetic Tree: Branch Lengths, Support Values, Rooting, and How One Gets Built

What a phylogenetic tree states and what it cannot, how to read one in order, what the scale bar and the numbers at the nodes are telling you, how a tree is inferred from sequence data, and what a journal requires when the tree becomes a figure.

Scientific Figure Team
Condividi

A phylogenetic tree carries more information than a cladogram and each extra piece can be misread. This is what branch length, the scale bar, the root and the support values are each claiming, where those claims come from, and how to draw the result so it survives a journal’s reduction to 89 mm.

A sequencing facility at night: a row of benchtop instruments, and a researcher at a wall display showing a circular tree carrying no labels — the point at which a genome becomes a hypothesis about ancestry.
A sequencing facility at night: a row of benchtop instruments, and a researcher at a wall display showing a circular tree carrying no labels — the point at which a genome becomes a hypothesis about ancestry. Immagine generata con Scientific Figure

Quick answer

Read a phylogenetic tree in four passes. Find the root and read outward, because that direction is time. Read the nesting, never the tip order — everything beyond a node shares a common ancestor not shared with anything outside it. Find the scale bar, which says whether branch length is substitutions per site, elapsed time, or nothing. Then read the numbers at the nodes, after checking the caption for whether they are bootstrap percentages or posterior probabilities, since the same number means different things.

What does a phylogenetic tree show?

A phylogenetic tree represents the evolutionary relationships among a set of organisms or groups of organisms, called taxa. That definition is the one the UC Museum of Paleontology uses in its Understanding Evolution material, and it is worth taking literally: the tree shows relationships, not organisms and not history as such.

Reading a phylogeny is close to reading a family tree. The root represents the ancestral lineage and the tips represent the descendants of that ancestor, and as you move from the root to the tips you are moving forward in time. Where a lineage splits, a single ancestral lineage has given rise to two or more daughter lineages. Each lineage has a part of its history unique to it alone, and parts it shares with other lineages.

Two things follow that catch people out. The first is that the tips are not necessarily species. Depending on how much of the tree you are looking at, the descendants at the tips might be different populations of one species, different species, or different clades each containing many species. A tree labelled with genus names and a tree labelled with strain identifiers are doing the same job at different scales.

The second is that a published tree is a hypothesis with an error bar, not a photograph of the past. Everything below — the branch lengths, the support values, the position of the root — exists to tell you how much of that hypothesis the data actually carried.

The parts of a phylogenetic tree

The vocabulary shared with every branching diagram is quick. The tips are the taxa. Each node is the common ancestor of the lineages leaving it. Two descendants that split from the same node are sister groups, and they are each other’s closest relatives on that tree. A clade is a group containing a common ancestor and all its descendants — Understanding Evolution offers the useful physical test of imagining a single branch clipped off the tree, where everything on the pruned branch makes up a clade. Clades nest inside clades, forming a nested hierarchy.

What a phylogenetic tree adds beyond that vocabulary is where the reading actually gets difficult, because each addition is a separate quantitative claim made in a separate visual language.

GibbonOrangutanGorillaHumanChimpanzee100 / 10096 / 9971 / 880.02 substitutions per site1234561Root — the base, and the direction of time2Node — a common ancestor3Branch length — a quantity, read against the scale4Support value — how well the data back that split5Scale bar — what one unit of branch length means6Outgroup — the taxon that gives the tree a root
The six things a phylogenetic tree carries that a cladogram does not. Note that the branches end at different depths and are joined to the aligned label column by dotted leaders — a phylogram’s tips do not line up, and the leaders are how every tree viewer keeps the labels readable anyway. The support values are given here as SH-aLRT / ultrafast bootstrap; the node joining the two closest tips fails both of IQ-TREE’s recommended thresholds. Topology after the standard great-ape phylogeny; the branch lengths are illustrative geometry, not estimates from any analysis.

Only the horizontal axis carries information here. The vertical order of the tips is a drawing convention — iTOL, for instance, sorts leaves by default so that clades with fewer leaves sit towards the top, purely because it produces a tidier stair-like display. Rotating the branches at a node changes the picture and changes nothing about the hypothesis, which is the single most common misreading of any tree and is worked through in detail in our guide to reading a cladogram.

What branch length means, and why the scale bar decides

Branch length is the difference between a phylogenetic tree and a cladogram, and it is the one element whose meaning the figure has to declare because the drawing cannot.

Tree typeWhat branch length meansWhat the figure must carry
CladogramNothing. Horizontal extent is whatever the layout needed to line the tips upNo scale bar, because there is nothing to scale
PhylogramAmount of evolutionary change along that lineage, conventionally substitutions per siteA scale bar in those units
Chronogram / time-calibrated treeElapsed timeAn axis, usually in millions of years

The distinction is not academic, because the same file produces all three. iTOL displays any tree containing branch length information as a phylogram by default, and turning that same tree into a cladogram is a single setting — toggle Branch lengths to Ignore and the lengths are discarded from the display. Nothing about the underlying tree changed. Only the claim the picture is making did.

So the scale bar is load-bearing. A tree with no scale bar and no time axis is making no quantitative claim through its branch lengths at all, and reading one lineage as more evolved because its branch is longer is reading the layout rather than the data. When you are the one drawing the figure, a missing scale bar is not a cosmetic omission — it removes the reader’s ability to interpret the axis you spent the whole analysis producing.

Rooted, unrooted, and what an outgroup is for

Most trees you see are rooted. Most trees as they come out of standard inference are not.

This is a real property of the models, not a software limitation. IQ-TREE 1 could only infer unrooted trees, because the substitution models it used assume time reversibility and a time-reversible model cannot tell which direction along a branch is forward. Rooted inference arrived in IQ-TREE 2 through non-time-reversible models, together with a root search that moves the root to neighbouring branches and keeps the position with the highest likelihood. Until you supply direction from somewhere, an unrooted tree states relatedness without stating ancestry.

The usual way to supply it is an outgroup: a taxon outside the group of interest, chosen so that all the members of the group of interest are more closely related to each other than any of them is to the outgroup. The outgroup therefore stems from the base of the tree, and it also tells you roughly where on the wider tree of life your group sits.

Where the root goes is not a presentational choice.

a. Unrooted — relationships, no directionABCDEbcCircles mark where panels b and c place the rootb. Rooted on EEABCDA+B and C+D are sister pairs; E is the outgroupc. Rooted on the branch to AABECDE is now the sister of C+D; A is the outgroup
One unrooted tree, two root placements, two different biological claims. Panels b and c contain exactly the same edges as panel a — nothing was added, removed or re-estimated. Rooting on the branch to E makes A+B and C+D sister pairs; rooting on the branch to A makes E the sister of C+D instead. The circles on panel a mark where each root was placed.

When there is no defensible outgroup, tree viewers offer midpoint re-rooting — iTOL has the function on its node menu, applicable regardless of which node you clicked. Treat it as a display convenience rather than a result. It places the root by branch-length geometry alone, so a single long branch, which is precisely what a distant or fast-evolving taxon produces, can drag the root onto the wrong edge and hand you the arrangement in panel c.

What the number at a node actually means

The numbers next to the nodes are the most-quoted and least-understood part of a published tree, mostly because two very different quantities get printed in the same place with no visual difference between them.

Bootstrap proportions come from resampling the alignment columns with replacement, rebuilding the tree from each resampled dataset, and counting how often each clade reappears. The resampling has a consequence worth holding onto: because sites are drawn with replacement, on average only about two thirds of the original sites appear in any one bootstrap alignment. The number is a statement about how consistently the data support a grouping, not a p-value.

Posterior probabilities come from Bayesian inference, where an MCMC run samples trees in proportion to their posterior probability and the frequency of a clade in that sample becomes its support. MrBayes annotates the consensus tree with these split and clade frequencies, along with node times and branch rates, through its sumt command.

Then there is the 70 percent rule, repeated everywhere and almost never with its conditions attached. It traces to Hillis and Bull’s 1993 simulation study in Systematic Biology, and their actual finding was narrower than the folklore. Under conditions of equal rates of change, symmetric phylogenies, and internodal change of 20 percent of the characters or fewer, bootstrap proportions of 70 percent or more usually corresponded to a probability of 95 percent or more that the clade was real. Outside those conditions the paper is blunt: where rates of internodal change are very high, or rates among taxa are highly unequal, bootstrap proportions above 50 percent are overestimates of accuracy. The same study found that as a measure of repeatability, any given bootstrap proportion is unbiased but so imprecise as to be virtually useless.

Two practical consequences follow.

A percentage means nothing until you know which method produced it. IQ-TREE’s documentation states directly that you should not compare standard bootstrap percentages with ultrafast bootstrap percentages, because ultrafast bootstrap values are less biased — 95 percent UFBoot corresponds roughly to a 95 percent probability that a clade is true, and the guidance is to rely on a branch only from 95 percent upward. Standard bootstrap is more conservative at the same number. The recommendation is to run the SH-aLRT test alongside it and treat a clade as well supported when SH-aLRT is 80 percent or more and ultrafast bootstrap is 95 percent or more.

Those thresholds only apply to single-gene trees. IQ-TREE notes that in a concatenation analysis across many genes, ultrafast bootstrap and even the more conservative Felsenstein bootstrap tend towards 100 percent everywhere, and recommends computing concordance factors instead for any phylogenomic analysis. A phylogenomic tree with 100 percent support at every node is not a strong result; it is a saturated statistic.

Support measureRead a branch as supported atSource of the threshold
Standard bootstrap (parsimony simulations)≥ 70%, under equal rates, symmetric phylogenies and ≤ 20% internodal changeHillis & Bull 1993, Syst. Biol. 42(2)
Ultrafast bootstrap (UFBoot)≥ 95%IQ-TREE documentation
SH-aLRT, alongside UFBoot≥ 80% with UFBoot ≥ 95%IQ-TREE documentation
Any of the above, on a concatenated phylogenomic treeNot applicable — use concordance factorsIQ-TREE documentation

A polytomy — a node with more than two lineages leaving it — is the other thing a tree can say about uncertainty, and it says it structurally rather than numerically. Understanding Evolution notes that it usually means there was not enough data to work out how those lineages are related, and that by leaving the node unresolved the authors are telling you not to draw conclusions from it. Occasionally it means the opposite: several speciation events genuinely happening at once, in which case the paper should say so. Both readings exist, so the caption is the only place to settle which one you are looking at.

How do you read a phylogenetic tree?

In this order, and the order matters because each step changes how the next one reads.

  1. Find the root. Everything is relative to it, and root-to-tip is time.
  2. Read only the nesting. At every node, everything beyond it shares a common ancestor not shared with anything outside it. Two lineages leaving one node are equally related to everything else on the tree.
  3. Find the scale bar or the time axis. This tells you whether the horizontal axis is substitutions per site, elapsed time, or nothing.
  4. Check the caption for what the node numbers are. Bootstrap and posterior probability are not interchangeable, and neither are standard and ultrafast bootstrap.
  5. Locate the outgroup, and confirm it actually sits outside the group being discussed.
  6. Look for polytomies, and read them as the statements of uncertainty they usually are.
  7. Ignore the vertical order of the tips. Adjacent tips are not necessarily close relatives, and a tip near the top is not primitive.

How a tree gets built from sequences

Between an alignment and a figure sit four steps, and knowing which one a paper skipped is often how you judge the tree.

Alignment. Homologous positions have to be placed in the same columns first, because every downstream method reads a column as a set of characters sharing an ancestor. An alignment error is a phylogenetic error that no amount of bootstrap replication will reveal.

Model selection. A substitution model describes how characters change, and model choice is a step in its own right — IQ-TREE 2 supports more than 200 time-reversible models alone, plus partitioned, mixture and heterotachy models for phylogenomic data, and its ModelFinder component exists to choose among them. MrBayes takes the alternative route of integrating over model uncertainty during the run, sampling across all 203 possible time-reversible rate matrices according to their posterior probability rather than committing to one in advance.

Tree search. Maximum likelihood searches for the topology and branch lengths that make the observed alignment most probable under the chosen model; Bayesian inference samples trees in proportion to their posterior probability via MCMC. Distance methods such as neighbour-joining collapse the alignment to pairwise distances first, which is fast and is why they survive as a starting point rather than a result.

Support. Bootstrap replicates, ultrafast bootstrap, SH-aLRT, or posterior probabilities from the same MCMC run that produced the tree.

For phylogenomic data there is a fifth step, because gene trees are inferred per locus and then reconciled. IQ-TREE 2 will infer individual locus trees for exactly this purpose, feeding either coalescent or concordance analyses downstream.

Newick: the whole tree as one line of text

Almost every tree you will ever handle is stored in Newick format, which represents a tree as nested parentheses — a correspondence Felsenstein credits to Arthur Cayley, who noticed it in 1857. The rules are short enough to learn in a sitting.

ElementWritten asExample
End of treeA semicolon(B,(A,C,E),D);
Interior nodeA pair of matched parentheses(A,C,E)
Tip nameAny printable characters except blanks, colons, semicolons, parentheses and square bracketsHomo_sapiens
A blank inside a nameAn underscoresea_lion
Branch lengthA colon and a real number after the node(A:5.0,C:3.0,E:4.0):5.0
Interior node nameText immediately after the right parenthesis(A:5.0,C:3.0)Ancestor1:5.0

Two properties of the format cause more confusion than the syntax does.

Newick is not a unique representation. The left-right order of a node’s descendants changes the string without changing the tree, so (A,(B,C),D);, (A,(C,B),D);, (D,(C,B),A); and ((C,B),A,D); are all the same tree to a biologist. Comparing two Newick strings character by character will tell you nothing about whether two trees agree.

The format is defined for rooted trees, and unrooted trees are stored by arbitrarily rooting them. So (B,(A,D),C); and (A,(B,C),D); represent the same unrooted tree. A root position in a file is not evidence that anybody chose it deliberately.

There is also a genuine ambiguity the standard never resolved, and it is worth knowing about because it silently corrupts figures. Felsenstein’s specification puts interior node names after the right parenthesis. Software conventionally puts support values in that same slot — iTOL documents (A:0.1,(B:0.1,C:0.1)90:0.1)98:0.3); with 90 and 98 as bootstrap values. The file itself cannot tell you which is which, or whether a given number is a bootstrap percentage, a posterior probability, or a label someone typed. That is a real limitation rather than an oversight on our part: as Felsenstein records, there has never been a formal publication of the Newick Standard. It was adopted by an informal committee that convened during the 1986 Society for the Study of Evolution meetings and took its name from the restaurant in Dover, New Hampshire, where the second session met. Where a number’s meaning matters, the caption is the only authority.

For richer metadata, iTOL also reads Nexus, PhyloXML and Jplace, and parses MrBayes and NHX annotations where they are present.

A gene tree is not a species tree

Different genes routinely give different trees, and the usual first instinct — that somebody made a mistake — is usually wrong. Evolutionary histories are genuinely discordant across the genome, and that discordance has to be accounted for rather than averaged away.

Incomplete lineage sorting is the ubiquitous cause. Under the multi-species coalescent model, the branches of a species tree are populations, and gene lineages coalesce inside them; lineages that fail to coalesce before a branch ends are carried up into the parent branch, where they may then coalesce in an order the species tree never took.

Species tree, with one gene traced through itA and C coalesce firstABCThe resulting gene treeACBA and C are sisters — the species tree says otherwise
How a gene tree comes to contradict the species tree with nobody making an error. The pale bands are populations, the thin lines are one gene’s ancestry traced backwards. Lineages A and B enter the ancestral population without coalescing, so A is free to coalesce with C first — producing a gene tree in which A and C are sisters, while the species tree says A and B are. Schematic; the geometry is illustrative.

The scale of this in real data is easy to underestimate. In the 48-genome avian dataset used to benchmark ASTRAL-III, neoavian relationships show extremely high levels of gene tree discord across 14,446 loci, attributed to a rapid radiation in their ancestors. Summary methods such as ASTRAL address it directly by searching for the species tree sharing the maximum number of quartet topologies with the input gene trees.

One counter-intuitive finding from that same work is worth carrying into your own analyses: contracting branches with very low support in the input gene trees — below roughly 3 to 10 percent — improved species-tree accuracy, while aggressive filtering at 50 or 75 percent made results worse. Discarding weakly supported signal is not automatically the safe choice.

Drawing a phylogenetic tree for a paper, poster or thesis

A phylogeny is the figure type with the most small lettering and the thinnest lines per square centimetre of any figure in biology, and journals reduce figures to fit their columns. That combination is why trees come back from production more often than almost anything else.

Nature’s final-submission guide is specific about the target, and the numbers are a reasonable proxy for most journals even though each one sets its own.

SpecificationNature’s requirementWhy a tree hits it first
Figure width89 mm single column, 183 mm double column, 120–136 mm for a column-and-a-halfTip labels are sized for the width you drew at, not the width it prints at
Full page depth247 mmA tree with 60 tips needs the depth budgeted before the labels are set
LetteringSans-serif, preferably Helvetica or Arial, the same font throughout all figuresTree viewers default to whatever the system offers
Type size5 pt minimum, 7 pt maximum for all text other than panel labels; panel labels 8 pt boldTip labels and support values are usually the smallest text in the paper
Line weight0.25–1 pt at final size; lines thinner than 0.25 pt may vanish in printBranches scale down with the figure, and hairline branches disappear
FormatVector — Illustrator, PostScript, EPS or PDF for line art; do not rasterize or convert text to outlinesA tree is line art, and a rasterized tree cannot be relabelled at proof

Note what the guide does not say. There is no rule for how many tips a tree figure may carry, no requirement about where support values sit, and no instruction on whether to show a scale bar — those are conventions of the field, not journal policy, and no journal we have checked publishes a tree-specific specification. What is binding is the type size, the stroke weight and the format.

Four things follow for a tree specifically:

  • Set the width first, then the type. Draw at 89 mm or 183 mm from the start. Exporting a screen-sized tree and scaling it down multiplies every violation at once.
  • Count your tips against the depth. At 5 pt minimum with sensible leading, a single column of 247 mm holds roughly 100 tips before labels collide. Past that, collapse clades — iTOL will collapse by average branch length, by support value threshold, or by node class.
  • Thin the support values before you thin the type. Showing support only where it falls below your threshold, and stating that convention in the caption, removes most of the smallest text on the figure without removing information.
  • Export vector, and check at final size. Nature’s own advice is to verify at the smallest size the figure might be printed that lettering is still readable and lines still print clearly.

The same discipline applies to every figure in the paper, and the general rules are collected in our guide to making scientific figures. For a worked example of one journal’s numbers end to end, see our Scientific Reports figure requirements breakdown.

Frequently asked questions

What is a phylogenetic tree? A branching diagram of the evolutionary relationships among a set of taxa. The tips are descendant taxa, the nodes are their common ancestors, two descendants splitting from one node are sister groups, and the root is the ancestral lineage. Moving from root to tips is moving forward in time, and any node with all of its descendants is a clade.

How do you read a phylogenetic tree? Find the root and read outward. Read the nesting rather than the tip order: everything beyond a node shares a common ancestor not shared with anything outside it. Then find the scale bar or time axis, check the caption for what the node numbers are, and locate the outgroup. Ignore the vertical order of the tips.

What do branch lengths mean? Whatever the figure declares. Substitutions per site on a phylogram, elapsed time on a chronogram, nothing at all on a cladogram. The scale bar or axis is the conversion, and a tree carrying neither is making no quantitative claim through its branch lengths.

What is an outgroup? A taxon outside the group of interest, chosen so every member of that group is more closely related to the others than to it. It stems from the base of the tree and its job is to place the root — which matters because standard time-reversible models infer unrooted trees, and where the root goes changes which taxa the tree says are sisters.

What does a bootstrap value of 70 mean? It comes from Hillis and Bull 1993 and carries conditions: under equal rates of change, symmetric phylogenies and internodal change of 20 percent or less, bootstrap proportions of 70 percent or more usually corresponded to at least a 95 percent probability that the clade was real. Under very high or very unequal rates, proportions above 50 percent overestimate accuracy. It is a conservative rule of thumb from parsimony simulations, not a significance test.

Is 95 percent ultrafast bootstrap the same as 95 percent bootstrap? No. IQ-TREE’s documentation says the two should not be compared directly. Ultrafast bootstrap is less biased, so 95 percent corresponds roughly to a 95 percent probability that the clade is true and is the point at which to start relying on a branch; standard bootstrap is more conservative at the same number. Running SH-aLRT alongside and requiring 80 percent SH-aLRT with 95 percent UFBoot is the documented recommendation.

Why do different genes give different trees? Because histories genuinely differ across the genome. Incomplete lineage sorting means two lineages entering an ancestral population without coalescing can coalesce in an order the species tree never took. In the 48-genome avian dataset behind ASTRAL-III, neoavian relationships show extremely high gene tree discord across 14,446 loci. Summary methods infer the species tree while accounting for it.

Why is my tree unreadable when the journal shrinks it? Because trees carry more small text and thinner lines than any other figure type. Nature sets 89 mm and 183 mm standard widths, 5 pt minimum lettering, and 0.25–1 pt strokes, warning that lines thinner than 0.25 pt may vanish. Draw at final width, export vector, and check every label at that size.

Where to go next

  • Cladogram vs phylogenetic tree — what the same diagram claims when the branch lengths are removed, plus how a character matrix becomes a topology.
  • How to make scientific figures — the general resolution, type and format rules that every figure in the paper has to meet.
  • Phylogenetic tree maker — describe the group or paste a Newick string and get a phylogram, time-calibrated tree or annotated-clade figure you can export at publication size.

Riferimenti

  1. Understanding Evolution — Reading trees: A quick reviewUniversity of California Museum of Paleontologyhttps://evolution.berkeley.edu/phylogenetic-systematics/reading-trees-a-quick-review/Consultato il 18 ago 2026
  2. Understanding Evolution — Understanding phylogeniesUniversity of California Museum of Paleontologyhttps://evolution.berkeley.edu/evolution-101/the-history-of-life-looking-at-the-patterns/understanding-phylogenies/Consultato il 18 ago 2026
  3. Understanding Evolution — Phylogenetic pitchforksUniversity of California Museum of Paleontologyhttps://evolution.berkeley.edu/phylogenetic-systematics/reading-trees-a-quick-review/phylogenetic-pitchforks/Consultato il 18 ago 2026
  4. The Newick tree formatJoseph Felsenstein / PHYLIPhttps://phylipweb.github.io/phylip/newicktree.htmlConsultato il 18 ago 2026
  5. iTOL — Help and documentationEuropean Molecular Biology Laboratoryhttps://itol.embl.de/help.cgiConsultato il 18 ago 2026
  6. IQ-TREE — Frequently asked questionsIQ-TREE developershttps://iqtree.github.io/doc/Frequently-Asked-QuestionsConsultato il 18 ago 2026
  7. Hillis & Bull 1993 — An empirical test of bootstrapping as a method for assessing confidence in phylogenetic analysis, Systematic Biology 42(2):182–192Oxford University Press / Society of Systematic Biologistshttps://doi.org/10.1093/sysbio/42.2.182Consultato il 18 ago 2026
  8. Minh et al. 2020 — IQ-TREE 2: new models and efficient methods for phylogenetic inference in the genomic era, Molecular Biology and Evolution 37(5)Oxford University Press / PubMed Centralhttps://pmc.ncbi.nlm.nih.gov/articles/PMC7182206/Consultato il 18 ago 2026
  9. Zhang et al. 2018 — ASTRAL-III: polynomial time species tree reconstruction from partially resolved gene trees, BMC Bioinformatics 19(S6):153BioMed Central / PubMed Centralhttps://pmc.ncbi.nlm.nih.gov/articles/PMC5998893/Consultato il 18 ago 2026
  10. Ronquist et al. 2012 — MrBayes 3.2: efficient Bayesian phylogenetic inference and model choice across a large model space, Systematic Biology 61(3):539–542Oxford University Press / PubMed Centralhttps://pmc.ncbi.nlm.nih.gov/articles/PMC3329765/Consultato il 18 ago 2026
  11. Nature — Final submission and figure preparationSpringer Naturehttps://www.nature.com/nature/for-authors/final-submissionConsultato il 18 ago 2026

I nomi di riviste, congressi e prodotti sono marchi dei rispettivi titolari. Questa pagina è contenuto editoriale indipendente e non è affiliata, autorizzata né approvata da essi. Le specifiche cambiano: verifica ogni requisito sulle fonti ufficiali collegate sopra prima di inviare o stampare.

Continua a leggere

La tua prossima figura è a 30 secondi di distanza

Inizia gratis con 50 crediti: nessuna carta di credito, nessuna competenza di design richiesta.