Common Surnames in the United States: A Guide to 2020 Census Data

Short Answer

Smith’s fifth straight century on top, Garcia at number six, and a surname table that finally includes 2020 — how the Census Bureau’s surname data works, what it says, and what it can (and cannot) tell your family history research For surname researchers, the U.S. Census Bureau quietly maintains the most complete surname-frequency dataset in […]

Smith’s fifth straight century on top, Garcia at number six, and a surname table that finally includes 2020 — how the Census Bureau’s surname data works, what it says, and what it can (and cannot) tell your family history research

For surname researchers, the U.S. Census Bureau quietly maintains the most complete surname-frequency dataset in the world — and the 2020 decennial census has now added the newest installment. Frequently Occurring Last Names in the 2020 Census (Census Brief C2020BR-14, by Joshua Comenetz, April 2026) is the fifth data product in the Bureau’s surname series, and the first to report 2020 results. Reading it properly teaches nearly every lesson this series has built up — including the standing distinction: the table tells you about a surname’s population; only documents tell you about your family. Smith’s 70.9% white share and 23.1% Black share (2010 data) are a statement about hundreds of unrelated founding lines — not a statement about any Smith in particular.

Here we walk through the whole lineage — 1790, 1990, 2000, 2010, 2020 — with the sources cited at each step, then the practical ways to use the new tables.

Key takeaways

  • The 2020 product exists and is specific: Last Name Data From the 2020 Census (C2020BR-14, April 2026), by the Bureau’s Joshua Comenetz, with a full public table of 156,621 last names (those with 100+ occurrences) and a companion first-name product (53,615 names).
  • Smith is #1 in every list ever published — 1790 (tabulated in 1909), 1990, 2000, 2010 and 2020 — and the top five (Smith, Johnson, Williams, Brown, Jones) have been identical since 2000.
  • Coverage shrank in 2020: surnames recorded for 299 million people = 90.2% of the 331,449,281 enumerated (2010: 95.5%). Read 2020 rates, not just counts.
  • The demographic columns are the treasure: Garcia 91% Hispanic, Washington 83.6% Black, Yoder 96.7% White, Xiong 97.3% Asian (all 2020). Surnames are predominantly one thing — and often several.
  • The fastest-growing names of 2010–2020 are Asian: Zhang +74.1%, Liu +62.4%, Patel +34.5% (to 309,327), Chen +37.9% (to 233,798) — plus, top of all, the data caveats themselves.
  • The 2020 file has its own quirks researchers must navigate: a 2010-frequency filter (brand-new names can’t appear), confidentiality “noise infusion” (±2 at 95% probability), a documented DOE inflation, and the conversion of PEÑA to PENA.
  • Institutional users apply these tables in Bayesian surname geocoding (BISG/BIFSG) to estimate race/ethnicity — a lesson in what surname statistics are, and are not.

1. Where the numbers come from: five products and a 1909 ancestor

The Census Bureau has now issued surname tables four times in the modern era — plus one antique:

ProductBasisKey facts to read it by
1790 tabulation (published 1909, A Century of Population Growth)The first U.S. census; white population onlyThe original SMITH count “includes SCHMIDT, SMYTH, and six other names” — variant-grouping is as old as the series
1990 Census product (U.S. Census Bureau, 1995)Post-enumeration survey, not the census itselfDifferent method; not directly comparable to 2000+
Census 2000 product (Word, Coleman, Nunziata & Kominski, 2007)Full census responses, heavily editedIntroduced the race/Hispanic-origin columns and the 100-occurrence threshold
2010 product (Comenetz, October 2016)Full census; 2010 methodology builds on 20006.3 million distinct surnames; 162,253 names with 100+ cover 90.1% of surname-recorded people; 95.5% coverage (295m of 308,745,538)
2020 product (Comenetz, C2020BR-14, April 2026)Full census; 2010 methods revised + expanded9.4 million raw entries edited down to 7.8 million distinct names; 156,000+ names published; 90.2% coverage (299m of 331,449,281)

The coverage figure is the single most consequential difference: 2020 recorded surnames for about five percentage points fewer people than 2010, so the brief’s own guidance is that 2010–2020 growth rates “were likely higher than what is shown.” Two editions are also filtered at the source in 2020: (a) the 2010-frequency rule — 2020 last-name data “were limited to names also found in the 2010 Census Names Table that had a frequency of at least 100 in both years,” so a name crossing the 100-occurrence line for the first time in 2020 cannot appear; and (b) noise infusion of the race/Hispanic counts (±2 at 95% probability, so a true 1,000 shows as 998–1,002; a research file retains the raw negatives).


2. The 2020 table: structure and the columns that matter

Download the files from the Census Bureau’s names data page (census.gov/topics/population/genealogy/data/2020_names.html): a top-1,000 table, the complete 53,615 first-name and 156,621 last-name datasets, and research-purpose files (with raw negative counts). Each surname row carries:

  • Rank among all reported names;
  • Frequency — the count of people who reported the name in 2020;
  • Occurrences per 100,000 people — note carefully: the denominator is the surname-recorded population (about 299 million in 2020; 295 million in 2010), not the headline census count. That is why Smith’s 2020 rate is 792.8 rather than 2,369,644 ÷ 331.4m;
  • Six mutually exclusive race/Hispanic columns (identical to 2000 and 2010): Non-Hispanic White alone; Non-Hispanic Black; Non-Hispanic American Indian and Alaska Native alone; Non-Hispanic Asian and Native Hawaiian and Other Pacific Islander alone; Non-Hispanic Two or More Races; and Hispanic origin (an identity overlaid on the race columns — the columns are “mutually exclusive” only within the scheme as tabulated).

The top of the national list — 2020 with 2010 for comparison (official product data, as tabulated with earlier censuses):

RankName2020 count2020 rate/100k2010 count2010 rate/100k
1Smith2,369,644792.82,442,977828.2
2Johnson1,858,234621.71,932,812655.2
3Williams1,561,395522.41,625,252551.0
4Brown1,386,083463.71,437,026487.2
5Jones1,383,250462.81,425,470483.2
6Garcia1,149,510384.61,166,120395.3
7Miller1,129,186377.81,161,437393.7
8Rodriguez1,085,838363.31,094,924371.2
9Davis1,074,448359.51,116,357378.5
10Martinez1,039,848347.91,060,159359.4

Three structural facts to lodge: the top ten cover only 4.7% of the population (2020; 4.9% in 2010); 258 names are needed to cover a quarter of Americans (2010: 239); and the top 100 surnames — 50,086,264 people, or 15.1% of the population — include eleven names reported over a million times each, which together cover just 5.0%. Compare with Britain in our fifth pillar, where Smith alone was 1.15–1.36% of England & Wales: American surnames are diluted by the world’s immigration — Smith is relatively rarer in the US than in Britain, though no less dominant.


3. The demographic columns: what they teach about surnames

The race/Hispanic tables are the series’ most under-read feature, and the 2020 brief’s Table 2 gives the extremes within the top 1,000:

  • Predominantly Hispanic (2020): Barajas 95.7%, Velazquez 95.2%, Vazquez 95.1% — and the marquee Garcia, 91% Hispanic, sixth overall. In 2000 the top 50 contained ten names at least 90% Hispanic; by 2020, thirteen.
  • Predominantly Black (2020): Pierre 86.2%, Washington 83.6%, Jefferson 71.6%, Mohamed 67.7%, Booker 62.7%, Jackson 51.2%.
  • Predominantly White (2020): Yoder 96.7%, Friedman 94.0%, Schwartz 93.7%, Schneider and Klein both 92.7–92.9%, Mueller 92.9% — note how German-heritage names, including Jewish-heritage ones like Friedman, run high but rarely top 96%.
  • Predominantly Asian and NHPI (2020): Xiong 97.3% (Hmong), Zheng/Zhu/Zhao/Zhou/Xu/Zhang/Huang 96.2–97.1% (Chinese), Kaur 96.3% (Sikh).
  • The American Indian and Alaska Native column behaves differently: even its “top” names score low — John 9.6%, Hunt 3.7%, Hunt/James/Sampson/Lucero all under 10% — because most AIAN-identifying people bear surnames shared across America. The Bureau’s own data underlines how imprecise surname→ancestry inference is.
  • Two or More Races column (2020): Washington (6.0%), Wong, Booker, Jefferson, Banks, Mosley, Gaines, Ware, Blackwell, Dorsey — a list of the “predominantly Black” names, doubling as the multiracial cluster.

Two deeper readings:

“Most common surname among Black Americans is Williams” — check the arithmetic. Applying the 2010 percentages to the 2020 counts: Black Williams ≈ 0.477 × 1,561,395 ≈ 745,000; Black Johnson ≈ 0.346 × 1,858,234 ≈ 643,000; Black Smith ≈ 0.231 × 2,369,644 ≈ 547,000. The Census-based tabulations of the surname products reach the same headline. It’s a beautiful, teachable demonstration that rank among a group ≠ rank overall — and it is an estimate, computed across noisy 2020 data; treat it as one decimal place, not one digit.

Names as population statistics — and how institutions use them. The race-by-surname table is the standard input for Bayesian Improved Surname Geocoding (BISG) and its first-name extension (Elliott et al. 2009; Tzioumis 2018; and Census-surname-based validation studies such as Morrison’s health-plan work, which reported a 0.76 correlation with self-reported race/ethnicity). Statisticians use it to estimate race/ethnicity where records lack it. The Morrison work’s caution is the family historian’s caution: “Imputing Native American and multiracial identities from surname and residence remains challenging.” For genealogy, the lesson is the inverse of institutional: the column tells you the population a surname comes with (for BISG’s purposes), and tells you nothing about any individual bearer. Your Smith line may belong to any of Smith’s groups — and only parish-level and household documents can say which.

(The Census Bureau adds its own guardrail, worth quoting: names tables “are not suitable for producing aggregate statistics such as the total U.S. population of a given race or Hispanic origin group.”)


4. Change over the decades — and the data-quality field notes

Stability at the top; motion beneath. The Census Bureau’s own “Hello my name is…” infographic (2016) tracks the top 15 across 1990/2000/2010: Garcia, Rodriguez and Martinez first appeared in the top 15 in 2010 (shown in red). The 2020 brief extends the series: the top fifteen is unchanged 2010→2020, and the top five has been Smith–Johnson–Williams–Brown–Jones since 2000. Meanwhile Garcia’s overall climb — 8th (2000) → 6th (2010) → 6th (2020) — tracks the Hispanic population’s growth.

The growth table. The 2020 brief’s Table 3 lists the fifteen fastest-growing names among the top 1,000 from 2010 to 2020: Zhang +74.1% (70,125→122,053), Liu +62.4%, Wang +54.7%, Ahmed +54.1%, Kaur +53.9%, Li +48.3% (111,786→165,790), Lin +46.9%, Ali +40.9%, Singh +40.1%, Chen +37.9% (169,580→233,798), Wu and Huang +36.0%, Patel +34.5% (229,973→309,327), Khan +33.6%, Shah +31.9%. All fifteen are predominantly Asian or Hispanic — except Ali, “a name with exceptionally broad… diversity.” With the 2020 coverage dip acknowledged, true growth was, per the brief, “likely higher than what is shown.”

The 2020 file’s special traps — read before you cite:

  1. DOE is inflated by design compromises. The brief is candid: “John DOE and JANE DOE” entries in mixed forms were deleted, but simple John/Jane Doe entries stayed — “therefore, the frequency of DOE listed in the last name table is likely inflated.” That explains the 2020 table’s eye-catching Doe (262,774) with no 2010 comparison value. The same class of issue affects a handful of other names where generic and genuine uses mixed (PERSON, CHILD, SON, YOUNG).
  2. Raw data was messy: ~1 million generic-word entries (ELDERLY, TENANT, HOUSEHOLDER, TWENTY-TWO, MRS SMITH); Spanish Ñ arriving as ? and corrected to N — PEÑA → PENA. Nonresponse and transcription of handwritten forms generated 9.4 million “unique” raw names; editing left 7.8 million, of which ~98% occur fewer than 100 times.
  3. Privacy processing: noise infusion (±2, 95%) plus the 2010-frequency filter. The Bureau documents a Disclosure Review Board approval (CBDRB-FY26-058).
  4. Historical comparability: the 1990 lists stand apart (post-enumeration survey); the 1790 tables counted only the white population and grouped similar names (SMITH included Schmidt, Smyth and six others) — the ancestor of every variant-collapsing decision we’ve discussed in this series.

For your surname: cite the edition (C2020BR-14), pair the 2020 row with the 2010 row, use rates for trend statements, and report single counts as approximations within the noise envelope.


5. From the table to your family history

The census surname data does three jobs for a family historian — and the third is the decisive one:

  1. It sizes your problem. A count (and rate) tells you how many variant-spelling descendants of a 2.4-million-strong Smith pool you might be indexing against; how a Patel line’s growth (+34.5%, and +113% since 2000) signals recent-cohort arrival, with the documentation implications of immigrant research (ship manifests, naturalization files — see our second pillar, including the corrected “Ellis Island” story).
  2. It shapes hypotheses, geographically. Pair national counts with state-level and diaspora reads: our Welsh pillar showed Welsh-origin names at 3.8% of the US population overall — and 9.5% in South Carolina — an index-value way of thinking borrowed directly from the diaspora studies; surname-geocoding logic (BISG) is, used cautiously, the same instinct: population statistics can narrow a search, never prove a line.
  3. It forces the distinction this series exists to teach. The six demographic columns demonstrate at national scale what our case studies demonstrate at family scale: surnames do not sort people. Smith is 70.9% white and 23.1% Black; Washington is 83.6% Black and 6.0% two-or-more-races; Lee (693,023 in 2010) split 36.0% white, 42.2% Asian, 16.3% Black — one spelling, a crossroads of three naming systems (English patronymic; Chinese; and African-American adoptions). Your family’s thread begins where the table ends: in the census returns themselves — the 1790–1950 censuses, with their households and neighbours — plus naturalization, vital and church records. A surname row tells you where to look; the records tell you who you are descended from.

A working sequence: decode the meaning and origin type (pillars 3–4); measure the name’s national 2020/2010 footprint here; map state and county concentrations (Surname Atlas equivalents for the US: census surname tables by state where published, plus diaspora index maps); build the variant list (pillar 2: Soundex/Daitch-Mokotoff apply directly to US index systems); then take proven ancestors document by document — census schedules first, because they attach names to households, addresses and neighbours.


FAQ

Is Smith really the most common surname in every census? In every surname product — 1790 (as tabulated in 1909), 1990, 2000, 2010 and 2020 — yes. In 2020, 2,369,644 people reported it (792.8 per 100,000 of the surname-recorded population; roughly 0.7% of the total population enumerated).

Why do other websites’ numbers differ from these? They usually count different things: headline census population vs surname-recorded population; unedited vs edited datasets; or they simply copy 2010 into “2020” templates. The Bureau’s own numbers are the standard; third-party tables restate them (helpfully, with 2000/2010/2020 columns) or diverge (unhelpfully, without sources).

If my surname is “predominantly” something, does that describe my ancestor? No — and the columns prove their own uselessness for that: Smith splits 70.9/23.1 across centuries-old founding lines. The column describes the name’s population. Your ancestor’s identity is a document question.

What is BISG and should genealogists use it? BISG/BIFSG is a demographic method (Elliott et al. 2009; Tzioumis 2018) that combines surname (and forename) data with geography to estimate race/ethnicity in records lacking it. Genealogists should understand it — it’s the main institutional use of the table you’re reading — but apply the same caution its authors do: strong for Hispanic and Asian groups, weak for Native American and multiracial identities, and never a substitute for a document.

Where do I get the data? The Bureau’s names data page hosts the April 2026 brief (C2020BR-14 “Last Name Data From the 2020 Census”; companion first-name brief), the top-1,000 table, and the complete 156,621-name last-name and 53,615-name first-name datasets (plus research files). The 2010 brief (Comenetz, October 2016), the 2000 methodology (Word et al., 2007) and the 1995 1990-methodology paper remain the comparison baselines.


Sources and further reading

  • Joshua Comenetz, Last Name Data From the 2020 Census, U.S. Census Bureau, Census Brief C2020BR-14 (April 2026) — coverage (299m/90.2%; 331,449,281 enumerated), 7.8m distinct names after editing, 156,000+ names ≥100, Table 1 frequency distribution, Table 2 group extremes (Yoder 96.7%; Pierre 86.2%; Washington 83.6%; Xiong 97.3%; Garcia 91% Hispanic), Table 3 fastest-growing (Zhang +74.1%; Patel +34.5%), Table 4 clustering (top-10 = 4.7% total; 35 names = 25% Hispanic), DOE/generic-word/Ñ/PEÑA edits, noise infusion, 2010-frequency filter.
  • Joshua Comenetz, Frequently Occurring Surnames in the 2010 Census (U.S. Census Bureau, October 2016) — 294,979,229 surname-recorded people (95.5%); 6.3m surnames; 162,253 names ≥100 covering 90.1%; Table A-1 top-50 counts and race/Hispanic shares (Smith 2,442,977; 70.9% White/23.1% Black; Williams 47.7% Black; Lee 36.0/42.2/16.3); Tables 2–4; Garcia #6.
  • U.S. Census Bureau, “Frequently Occurring First Names and Last Names in the 2020 Census” (April 14, 2026) — data page and file sets (53,615 first names; 156,621 last names).
  • David L. Word, Charles D. Coleman, Robert Nunziata & Robert Kominski, Demographic Aspects of Surnames from Census 2000 (2007); U.S. Census Bureau, Documentation and Methodology for Frequently Occurring Names in the U.S.—1990 (1995); Bureau of the Census, A Century of Population Growth in the United States 1790–1900 (1909) — the 1790 white-population tabulation and variant-grouped SMITH count.
  • U.S. Census Bureau, “HELLO my name is…” infographic (December 15, 2016) — top-15 by rank 1990/2000/2010, with Garcia, Rodriguez, Martinez marked as first-time entrants.
  • Wikipedia, “List of most common surnames in North American countries” (US section) — tabulation of official 2020/2010/2000 counts and rates for the top 100 (Smith 2,369,644/792.8 in 2020, etc.); cite the underlying Census products when quoting.
  • Marc N. Elliott et al., “Using the Census Bureau’s Surname List to Improve Estimates of Race/Ethnicity…” (2009); Konstantinos Tzioumis, “Demographic Aspects of First Names,” Scientific Data 5 (2018); Oregon Criminal Justice Commission race-correction technical documentation (2018) — BISG/BIFSG methodology, 0.76 correlation validation, and the Native-American/multiracial imputation caveat; Frank Nuessel’s Names notes (2017) on the series.
  • Frank Jacobs (after Marcin Ciura), “How the Smiths took over Europe,” Big Think (2018) — Smith 2,442,977 in the 2010 census; adopted-Smith phenomenon; C.M. Matthews’ History Today (1967) explanation.
  • Cross-links in this series: Surname Spelling Variations (variant searching, Ellis Island reality check); Occupational Surnames (Smith’s origin); Locational and Topographic Surnames; Patronymic and Matronymic Surnames Explained; Surname Frequency and Distribution (British comparisons: Smith 1.15–1.36% in E&W; concentration statistics); Welsh Surnames (diaspora index values: US 3.8%; South Carolina 9.5%).

Individual American surname profiles on Select Surname List carry each name’s 2010/2020 census counts and rates, demographic shares with sources, origin notes — and the records needed to take any line back the document way.

Leave a Reply

Your email address will not be published. Required fields are marked *