Ancestrify field guide
Glossary
The vocabulary this field uses, defined plainly — and for the terms people routinely read too much into, what they do not mean.
Global25
Global25 and coordinate methods
- Global25G25
- A set of 25 numbers that locates one genome in a genetic reference space built from ancient and modern samples. Coordinates come from the independent Eurogenes Global25 service; comparing them means measuring distances and fitting mixtures inside that space.
- What it is not. Not an ancestry test and not a set of ethnicity percentages. A coordinate is a position, and it only means something relative to the other positions you compare it against.
- Coordinate row
- One sample's 25 Global25 values, written as a single comma-separated line beginning with the sample's label. It is the input format every coordinate tool takes.
- Scaled and unscaled coordinates
- Two published forms of the same 25 values. The scaled form multiplies each dimension by that dimension's share of the variance, so the earlier dimensions carry more weight in a distance calculation; the unscaled form treats all 25 equally.
- What it is not. Not interchangeable. Mixing scaled and unscaled rows in one comparison produces distances that mean nothing, and the arithmetic will not warn you — the numbers still come out looking reasonable.
- Euclidean distance
- The straight-line distance between two coordinates across all 25 dimensions: square the difference in each dimension, add them, take the square root. It is how closest-population rankings are produced.
- What it is not. Not a measure of relatedness or of shared ancestors. Two groups can sit close together because they genuinely share ancestry, or because they are both mixtures that happen to average out to a similar point.
- Admixture model
- An estimate of your coordinate as a weighted mixture of chosen source populations, searched for by Monte-Carlo methods in the nMonte tradition and reported with a fit distance.
- What it is not. Not a test. A coordinate fit cannot reject a model — it always returns percentages, including for a source panel containing nobody your ancestors met.
- Fit distance
- How far the fitted mixture's combined coordinate still sits from your own after the search finishes. Smaller means the mixture lands nearer your point.
- What it is not. Not a p-value and not a confidence measure. A small fit distance says the arithmetic closed the gap, not that the historical story behind the sources is right.
- Calculator
- A curated, era-scoped set of source populations an admixture model is fitted against. Which calculator you choose shapes the answer more than the arithmetic does.
- Era
- A dated window that scopes which reference populations a comparison uses. Distances run across six: Late Bronze Age (3000–1200 BC), Pre-Classical Iron Age (1200–0 BC), Imperial Antiquity (0–600 AD), Middle Ages (600–1400 AD), Early Modern (1400–2000 AD) and the Modern Era (2000 AD onward).
- PCAprincipal component analysis
- A projection that flattens many dimensions into a plot you can read, placing samples so that the largest axes of variation become the axes of the chart.
- What it is not. Not a map, and the axes are not ancestries. Distance on a two-dimensional scatter can hide separation that exists in the dimensions the plot dropped.
- Population average
- A coordinate built by taking the arithmetic mean of each dimension across the individuals in a group. Our reference set holds 1,535 curated populations averaged from 30,386 individual samples.
- What it is not. Not a person. An average describes the centre of a group, and individuals scatter widely around it — which is why matching an average closely is not the same as matching anyone who lived.
qpAdm
Formal statistics and qpAdm
- qpAdm
- A formal admixture test that estimates a target's ancestry as proportions of chosen source populations, working from allele-frequency statistics and a set of outgroups. It returns a weight, a standard error and a z-score per source, plus a p-value for the model as a whole.
- What it is not. Not a coordinate fit dressed up in different language. The defining difference is that qpAdm can reject a model outright, which no distance-based method can do.
- ADMIXTOOLS 2
- The R package these methods are implemented in, used in published ancient-DNA research. Our paid analysis runs its qpadm() function; the free Lab exposes f-statistics, qpWave, qpAdm and admixture-graph fitting against a reference panel.
- p-value
- The probability of seeing data at least this far from the proposed model if the model were true. A high value means the model is compatible with the data; a low one means it is not.
- What it is not. Not the probability that the model is correct, and not a measure of how much ancestry came from anywhere. A model can pass with a comfortable p-value and still be historically wrong, because a different model can pass too.
- Standard errorSE
- The uncertainty around one estimated weight. It depends far more on how many markers your file carries after merging than on how long anyone searched for the model.
- What it is not. Not something extra effort can shrink. Coverage sets the floor, which is why a sparse file cannot reach the tightest statistical bars no matter what is paid for it.
- Z-score
- How many standard errors a weight sits from zero. A source needs a large |Z| before its contribution can be called distinguishable from nothing at all.
- Outgroupthe right set
- A population used as a reference point rather than as a candidate ancestor. The outgroups are what give qpAdm its power to discriminate between models.
- What it is not. Not a formality to be filled in. A right set that is too small, or too closely related to your sources, will accept almost any model you propose — the commonest way to produce a confident and meaningless result.
- Left set and right set
- The two halves of a qpAdm model: the left set is the target plus the candidate sources, the right set is the outgroups it is measured against.
- Rotation
- An automated strategy that moves populations between the left and right sets to search many models at once.
- What it is not. Not what we publish from. Rotating qpAdm has a high false-discovery rate, so automated rotation was built here and then deliberately removed — every published model is composed and checked by hand.
- f-statisticsf2, f3, f4, D
- Summaries of shared genetic drift between sets of populations. They are the raw material qpAdm, qpWave and admixture graphs are all computed from.
- AADRAllen Ancient DNA Resource
- The public compilation of published ancient and modern genotypes that formal modelling here is run against — version 66, roughly 23,265 samples across roughly 6,015 distinct population labels.
- Nested model
- A simpler model contained inside a more complex one — for example the same model minus one source. If the simpler version also passes, the extra source is not earning its place.
Segments
Segments, matches and lineages
- IBDidentity by descent
- A stretch of DNA two people share because both inherited it from a common ancestor.
- What it is not. Not what a scan of ancient samples can establish. Demonstrating descent requires evidence a genotype comparison alone does not carry — which is why we describe our results as IBS.
- IBSidentity by state
- A stretch of DNA where two samples simply match, whether or not the match came from a shared ancestor. It is what a comparison against ancient genomes actually detects.
- What it is not. Not proof of a family connection. A shared segment with an excavated individual is a genuine signal of shared ancestry in a population, never a claim that the person was your ancestor or relative.
- Haplogroup
- A branch on the single-line tree of either the Y chromosome (paternal) or the mitochondrial genome (maternal), defined by the mutations that mark it.
- What it is not. Not an ethnicity, and not a summary of your ancestry. A haplogroup traces one thread — father's father's father, or mother's mother's mother — while almost all of your ancestry sits in the rest of your genome.
- Clade
- A haplogroup branch together with everything descending from it. A deeper clade is a more specific placement on the tree.
- Coverage
- How many usable markers a raw DNA file actually contains, and where they fall. It governs which analyses a file can support and how tight their standard errors can be.
- What it is not. Not fixed per company. Files from the same provider differ by chip generation, and some products carry no Y-chromosome rows at all — so a file can work for one analysis and be unusable for another.
- Genotype merge
- Combining your markers with a reference panel so the two can be compared position by position. Only the positions present in both survive, which is why the merge — not the panel — sets the marker count a model is computed from.
