Surname Frequency and Distribution: How to Read the Data

Short Answer

Every surname article quotes a ranking — here is how to judge one, map one, and turn both into family-history progress Sooner or later, every surname researcher meets the numbers. You have decoded the name (our guides to Occupational Surnames and Patronymic and Matronymic Surnames Explained will have helped); now you find a table announcing […]

Every surname article quotes a ranking — here is how to judge one, map one, and turn both into family-history progress

Sooner or later, every surname researcher meets the numbers. You have decoded the name (our guides to Occupational Surnames and Patronymic and Matronymic Surnames Explained will have helped); now you find a table announcing that your surname is “the 27th commonest in Britain” — or the 11,432nd. What, honestly, does that tell you? And then someone produces a map, with epicentres and gradient arrows and a confident county verdict. What can that prove?

Quite a lot — if you read the data as data. Frequency tables and distribution maps are the most powerful research instruments in surname studies, but they fail silently when read carelessly, because every number in surname studies answers a different question than the one you are asking. The tables measure the surname (how many people use it, where its bearers cluster); they never measure your family. As throughout this series, the distinction between surname origins and the history of a particular family governs everything: frequency and distribution are how we interrogate origins; documents are how we prove a particular family.

This guide walks through both data types using the actual sources genealogists use — with their citations, their strengths and their published flaws.

Key takeaways

  • Frequency ≠ distribution. “How many?” (ranks, percentages, cumulative shares) and “where?” (maps, epicentres, diaspora clusters) are different datasets with different uses.
  • Every table has a vintage, a geography, a unit and a caveat. An 1853 registration list, a 1986 phone survey, a 1991–2000 NHS register extract and a 2021 census table will disagree — by design.
  • Concentration is the headline. In Wales, the ten commonest surnames covered 55.85% of the population in 1813–37 — against 5.15% for England’s top ten in 1853 (Rowlands; Registrar General). Regionally, the Welsh top-ten share ranged from 27.32% (Maelor hundred, Flintshire) to 90.69% (Uwchgwyrfai, Caernarfonshire).
  • Variant handling can reorder the whole league table. Combine Schmidt/Schmitt/Schmitz/Schmid and smith-names beat Müller (German data); split Price/Pryce/Preece and each drops.
  • Maps make research hypotheses; documents test them. The Rowlands’ probability method used four surnames from an emigrant family to locate its origin — and was right 97% of the time across their tests.
  • Modern Britain contains 378,782 surnames with two or more bearers (Webber 1997 data) — the “long tail” is long, and surname death shapes what survives.

1. Two questions, two datasets

Ask, first, which question you are actually asking:

  • “How common is my surname?” — a frequency question, answered by counting bearers (or events: births, marriages, deaths, phones, electors) in some covered population, then ranking.
  • “Where does my surname come from?” — a distribution question, answered by plotting that same count by place and reading the pattern.

Genealogy needs both: frequency tells you how hard the variant-search battle will be; distribution tells you where to fight it. But the two are computed from different sources with different biases, and mixing them is the commonest reading error.

The core datasets you will actually meet:

DatasetWhat it countsCoverage & vintageReading cautions
Registrar General’s lists, 1853 (quoted in Rowlands, Nomina 29)Registration events, 1837–52England & Wales onlyPercentages interpreted against population; pre-dates industrial mixing partially
Rowlands’ Welsh marriage-register survey, 1813–37Marriages (two surnames per record)All Welsh parishes; ~270,000 occurrences, ~5,500 surnamesEstimated 30–40% of the potential population; variants deliberately combined
1881 national census (Steven Archer’s Surname Atlas)Persons present on census nightAll Britain (and maps widely reused)Census-night residence; transcription errors inflate rare-name counts
Telephone directories, 1986 (Michael Williams’ survey, via the Guild)Phone-holding households, ~six BT Welsh directoriesWalesOwnership and household bias; heads of households counted
NHS Central Register survey, 1991–2000 (ONS, published via Guild/FOI)Registered individuals (~60m names)England & Wales & Isle of Man onlyThe Guild’s analyst himself concluded the results were “somewhat suspect” against the electoral register — read his caveat
Electoral register data, 1996/2002 (Richard Webber; Technoleg Taliesin/ONS derivations)Adult electorsUK (separate Scottish & NI registers)Excludes non-registrants; commercial versions only cover large names
Census Wales 2021 tables (BBC/WalesOnline via Forebears)Reported bearersWales onlyTwo 2021-based tables differ slightly (Jones 170,633 vs “>170,000”) — a lesson in definitions
US census 2010 surname filePersons (names with 100+ occurrences)United StatesImmigrant country: reflects adoption and translation, not just descent (Smith 2,442,977)
DIY indices: FreeBMD, FreeCEN, FreeREG, 1939 RegisterEvents/personsAs digitisedVariant handling is your job — wildcards and spelling control
FaNUK surname inventory (Webber list)378,782 names with 2+ bearersBritain, 1997An inventory, not a frequency table — but it defines the universe

Rule one of reading: before trusting any figure, identify its vintage (year), geography (Wales? England & Wales? the whole “UK”?), unit (persons, events, households, phone lines, electors), and the compiler’s variant policy (combined or split).


2. Reading a frequency table

Ranks, shares, cumulative totals. A rank without a share is trivia; a share without a denominator is noise. Three exhibit-numbers to train the eye:

  • In the NHS-register survey, the top 500 names cover 39.33% of the register — a compact national concentration worth memorising as Britain’s “normal.”
  • In 1853, England’s top ten covered a mere 5.15% of the population (Smith leading at 1.37%) — while in Wales in 1813–37 the top ten covered 55.85% (Jones alone 13.84%). Same language, two naming demographies.
  • Wales’ extreme regional top-ten shares (27.32%–90.69%) show that even “concentration” is a distributional fact in disguise — frequency tables are maps wearing a suit.

Stability across sources builds trust. Compare Smith across the historical series the Guild tabulates: 1.36% (1853 registration) — 1.76 (1944 Treasury sample) — 1.37 (1966 National Insurance) — 1.36 (1975, the demographer Gabriel Lasker’s study) — 1.37 (1984–97 GRO deaths) — 1.20 (1996 electoral register) — 1.15 (1991–2000 NHS register). The rank never moves: #1 in every source across 140 years. When ranks disagree between sources, suspect the units (events vs persons) or the coverage, not the surname.

Variant handling can reorder everything. Our occupational guide showed the German case: Schmidt ranks second, but combine Schmidt/Schmitt/Schmitz/Schmid and smith-names out-rank Müller to become the largest occupational name-stock in German. The same levers apply in Britain: Roberts vs Robert vs the McKinley-documented post-medieval -s additions, Davies/Davis, and the Price/Pryce/Preece/Breese cluster as our Welsh spelling guide showed. When a league table doesn’t state its variant policy, treat its comparisons as decorations.

Watch the inclusion rules. The Guild’s Wales table (1986 phone survey) explicitly excluded English-language surnames — and notes that had Smith been included, it would have ranked 13th in Wales and Green 38th. Frequency tables are curated objects; read the curatorial note before the list.

The “perpetual incognito” problem. Frequency’s most honest lesson is negative: commonness destroys identifiability. The Registrar General George Graham wrote in 1856 that Welsh surnames’ frequency “almost defeats the primary object of a name” — and that “the name of John Jones is a perpetual incognito in Wales.” For research, the counter-move is not a better surname search but a better combination: forename + surname + place + occupation, or spouse-pairs (below) — the same principle as our spelling-variations pillar’s cluster method.


3. Reading a distribution map

A distribution map is a frequency table with geography restored. Learn its anatomy:

  • The epicentre: the zone of maximum concentration (Ravenshaw/Ramshaw → near Bishop Auckland, Durham; Greatorex → Wormhill, Derbyshire — the FaNUK cases in our place-names pillar).
  • The gradient: how fast frequency falls from the epicentre — Jones from 30.71% in Penllyn (Merionethshire) to 1.06% in Dewisland hundred (Pembrokeshire).
  • The clusters: multiple zones of strength, often betraying multiple origins — the classic case is Gwynne (south) versus Wynne (north), split almost cleanly along a line through mid-Wales in the 1813–37 survey (our Welsh pillar details the figures).
  • Corridors and diaspora tails: migration lines, like the Meredith axis from Radnorshire across to Monmouthshire the Rowlands mapped, or Wales-to-Glamorgan industrial flows.
  • Index values: the Welsh Government diaspora study scores concentration where 100 = a country’s average, 50 = half, 200 = twice — the same convention behind “South Carolina 9.5%, North Dakota 1.1%” for Welsh-origin names in the US.

Three reading disciplines follow:

  1. Add time depth. The FaNUK team warns that “the epicentre itself may move from century to century”; compare 1881 (Surname Atlas) against modern data (gbnames profiler) before pronouncing an origin.
  2. Let the map overrule the dictionary. The methodological lesson of our place-names pillar in statistical form: Pardoe and Pardey etymologise identically (the oath-nickname par Dieu), yet the map shows Pardoe glued to Staffordshire and Pardey to Dorset — the FaNUK team concludes the two are probably genealogically unrelated. Resemblance proposes; distribution disposes.
  3. Map precisely, then search locally. Distribution converts a national surname problem into a county or hundred problem — the conversion that made the Rowlands’ method work.

4. Diaspora data: reading emigration fingerprints

National frequency tables end at borders; diaspora studies extend them, and they repay careful reading:

  • The Welsh Government’s diaspora study (Richard Webber, UCL) found nearly 35% of Wales’ population bears a Welsh-origin family name, versus 5.3% in the rest of the UK, 4.7% in New Zealand, 4.1% in Australia and 3.8% in the US; 6,461 names were classified Welsh, and 16.3 million people across the studied countries bear one.
  • Its US maps show the concentration pattern an 1850s–90s emigration wave would predict: 9.5% in South Carolina down to 1.1% in North Dakota, mid-Atlantic states, the Carolinas, and Appalachia — plausibly coal-linked, as the study itself suggests.
  • Australia/New Zealand map the earlier-settlement signature: Tasmania 17% above the Australian average; the New Zealand provinces Marlborough and Taranaki at twice the national average.
  • The study’s own sober note — that English neighbourhoods with high Caribbean-descended populations show elevated Welsh and Scottish surnames, “a likely reason… the tendency for slaves to have adopted the family names of slave masters” — is essential reading for diaspora researchers: surname frequency abroad measures naming inheritance, including adoption, not necessarily ancestry. The same logic underlies the American Smith explosion noted in our occupational guide.

And a home-market corollary for family historians: two 2021-based tables of Welsh surnames, both circulating online, differ in the details (Jones 170,633 versus “more than 170,000”; Davies 111,559 versus 112,000). Both are legitimate; neither is identical; the difference is the lesson — definitions and sources, not sloppiness.


5. What the data can actually do for your research

Frequencies and maps are not decoration; deployed as a pair they become the research engine.

(a) They locate. The Rowlands’ probability method remains the exemplar: knowing an emigrant family’s surnames, compute where each is regionally frequent and intersect the maps. Their test family — Simon “Davis” (the e lost in migration) of Meigs County, Ohio, wife Richards, co-emigrants named Oliver and Evans — resolved decisively to one area of Wales, and the family indeed came from mid-Cardiganshire (Llanrhystud). Across a hundred such tests, the method placed families exactly or almost exactly right 97% of the time. Marriage registers are the researcher’s gift here: each record carries two surnames, doubling the intersecting evidence.

(b) They size and scope one-name studies. The Guild’s study categories and variant-registration rules (about five primary variants; more by exception) presuppose knowing a name’s frequency — a rare name (under, say, a hundred modern bearers) is a very different project from a Wilson. The FaNUK inventory’s structure (43,877 “frequent established” names vs 14,452 with 20–100 bearers) is the working scale.

(c) They discipline expectations before searching. A top-100 surname mandates wildcard/variant/DNA strategies (our preceding pillars); a genuinely rare name invites different errors — spelling-garble hunting and surname death. Sturges and Haggett’s model implies over half of single-bearer surnames die out within ~600 years; the name in a 1377 poll tax roll may simply be extinct, so distribution work for rare names must chain medieval, 17th-century and modern maps to reveal breaks in transmission.

(d) They identify individuals — in reverse. For the perpetually incognito common names, the statistical reflex is to increase specificity: rare forename–surname combinations (an “Ephraim Jones” is a much smaller population than a “John Jones”), spouse-surname pairs, address-plus-occupation anchors. This is the data-reading version of the cluster method from our spelling pillar.

(e) Modern data is obtainable. The ONS publishes surname tables under FOI (the “500 most common surnames” and “twenty most common surnames at birth, 2014 and 2024” requests are public precedents); FreeBMD/FreeCEN/FreeREG let you compute your own variant-aware counts; Archer’s Surname Atlas and the free gbnames profiler cover the mapping. Do not underestimate how much one careful self-made table can replace a dozen internet rankings.


6. Case studies

1. The Ohio Davis family — statistics that found a parish. As above: four surnames, three intersecting Welsh distribution maps, one hundred (Llanrhystud), plus the documentary confirmation in family papers and local history. Every step is auditable from the sources cited in our Welsh pillar (Rowlands, Nomina 29). Lesson: distribution plus variant-correction (Davis → Davies) plus multi-surname intersection — not any single number — produced the answer.

2. Pardoe/Pardey: one etymology, two families. Same linguistic explanation, utterly separate geographies (Staffordshire vs Dorset), and the FaNUK team’s conclusion that distribution suggests no genealogical connection. Lesson: frequency/mapping data exists precisely because etymology alone cannot separate independent coinages — the pillar-series distinction (name origin vs family history) wearing numbers.

3. Lee/Li: a frequency table that hides two worlds. The Oxford Family Names launch materials: English Lee (“clearing in a wood”) has 84,000 British bearers — of whom nearly half, the FaNUK team found, are of Chinese ancestry; Li itself had 9,000+ bearers in 2011 before spellings are consolidated. A national ranking like “Lee = 35th” is therefore a composite of unrelated naming systems. Lesson: behind every frequency figure stands an unstated answer to “which bearers?” — read the compiler’s category logic before the rank.


FAQ

My surname is ranked Xth — so what? Rank without share is trivia. What matters: the share, the vintage, the geography and the variant policy. Use rank only to choose strategies (common: variant-plus-DNA plus mapping; rare: continuity-hunting plus surname-death awareness).

Why do websites disagree about rankings? Different units (events, electors, phone lines), different geographies (E&W vs UK), different vintages and different variant-collapsing decisions. The Guild’s own NHS-register table carries its analyst’s “somewhat suspect” caveat — a model of honest presentation.

Can I locate my family’s origin from maps alone? You can generate a strong hypothesis — the 97%-accurate Rowlands method is evidence — but distribution is about surname populations, not your line. Confirm with documents; where two places compete, DNA.

What is “cumulative share” and why care? The proportion of the population covered by the top N names — a concentration index. Wales 1853/1813–37’s 55.85% versus England’s 5.15% explains at a glance why Welsh genealogy is a variant-and-probability problem and English genealogy is not.

How rare is rare? Below ~100 modern bearers (the FaNUK threshold), statistics wobble; below ~20, every count is fragile; and surname death means some historically attested names have no modern bearers at all.


Sources and further reading

  • John & Sheila Rowlands, “The Distribution of Surnames in Wales,” Nomina 29 (2006) — the 1813–37 survey (270,000+ occurrences), the 55.85%/5.15% comparison, hundred-level extremes (27.32%–90.69%), Jones 30.71%–1.06%, Gwynne/Wynne, Meredith’s axis, ap-name and Old-Testament clusters, the Meigs County, Ohio case, and the 97% accuracy record; their The Surnames of Wales (1996; updated 2013).
  • Guild of One-Name Studies: “Top 500 names in England and Wales” (ONS NHS Central Register survey, 1991–2000, ~60m names; top 500 = 39.33%; with the Guild’s own “somewhat suspect” caveat); “Wales: the top hundred surnames” (Michael Williams’ 1986 BT-directory survey — Smith 13th if included; Green 38th); “Comparisons” table (Registrar General 1853; Treasury 1944; National Insurance 1966; Lasker 1975; GRO deaths 1984–97; 1996 electoral register).
  • Registrar General George Graham’s Sixteenth Annual Report (1856) — the “perpetual incognito” passage, quoted in Rowlands, Nomina 29.
  • Patrick Hanks, Richard Coates & Peter McClure (eds), Oxford Dictionary of Family Names in Britain and Ireland (OUP, 2016), plus the FaNUK methodology paper — the 378,782-surname inventory; 43,877/14,452 structure; epicentre doctrine; Pardoe/Pardey; surname inconstancy; the IGI-scale warnings; the Rochester, Harmison and Greatorex cases (detailed in our place-names pillar).
  • Richard Webber, The Welsh Diaspora: Analysis of the Geography of Welsh Names (report for the Welsh Assembly Government) — 34.9%/5.3%/4.7%/4.1%/3.8% figures; 6,461 Welsh names; 16.3m bearers; index-value maps (South Carolina 9.5%–North Dakota 1.1%; Tasmania; Marlborough/Taranaki); the Caribbean-descended neighbourhoods finding.
  • Steven Archer, The British 19th-Century Surname Atlas (Archer Software); gbnames public profiler (Webber-based); H.B. Guppy, Homes of Family Names (1890).
  • BBC News, “Jones, Davies and Williams…” (26 Dec 2023) — 2021 census Wales figures; WalesOnline (2023) Forebears-based table — the two slightly differing 2021 tables.
  • ONS FOI responses: “500 most common surnames in England and Wales” (2022); “Twenty most common surnames for births in England and Wales, 2014 and 2024” (released Jan 2026).
  • Big Think/Frank Jacobs (after Marcin Ciura), “How the Smiths took over Europe” (2018) — US 2010 census Smith count (2,442,977) and the German variant-collapsing result.
  • Sturges & Haggett, Inheritance of English Surnames (1987) — surname-death model; McKinley, A History of British Surnames (1990) — inorganic -s and variant behaviour.
  • Cross-links in this series: Welsh Surnames: Patronymics, Spelling Changes, and History; Surname Spelling Variations: How to Search Historical Records; Occupational Surnames; Patronymic and Matronymic Surnames Explained; Locational and Topographic Surnames.

Individual surname profiles on Select Surname List list each name’s modern frequency tiers, regional epicentres and variant sets — the practical outputs of the method taught here — together with the records needed to test any origin hypothesis for a particular family.

Leave a Reply

Your email address will not be published. Required fields are marked *