Skip to content
BabyNames

Methodology

How every number on this site is produced, and where it can be wrong.

Source data

All figures derive from the US Social Security Administration’s national and territory baby name files, covering births from 1880 to 2024. Each record is a name, a sex, a year and a count. The SSA includes a name only if it was given to at least five children of that sex in that year, so very rare names are absent rather than shown as small.

Exclusions

Two categories are removed before anything is counted:

  • Administrative placeholders. Entries such as “Unknown”, “Infant”, “Notnamed” and similar are filing artifacts rather than names.
  • Suffix artifacts. Names ending in “jr” where the base name exists separately are treated as filing errors. The check is deliberately narrow: a name is only removed if stripping the suffix leaves a name that exists in the data, so genuine names that happen to end in those letters are kept.

Grouping spellings by pronunciation

This is the step that makes the numbers here differ from the official rankings. Spellings that share a pronunciation are combined into one name, with the most common spelling used as the label and the others listed as variants.

  1. Pronunciations come from the CMU Pronouncing Dictionary, an ARPABET phonetic dictionary of around 126,000 entries.
  2. Names absent from it — which is most invented and modern names — are transcribed with the g2p_en neural grapheme-to-phoneme model.
  3. Names the model also handles poorly fall back to recursive subword splitting.
  4. A curated override file corrects specific names where all of the above produce the wrong pronunciation, and a forced-merge file joins groups that should be together but are not.

A name may have several valid pronunciations. Grouping uses all of them, so two spellings merge if any pronunciation is shared.

Counts, ranks and percentages

  • Total births is the sum across every spelling in the group and every year in the record, including territories.
  • Rank is by total births within a gender, using dense ranking — tied names share a rank and the next rank is not skipped.
  • Peak year is the single highest year for the primary spelling.
  • Unisex share is the minority gender’s percentage of the total, reported only where both genders reach a meaningful count.

The year-by-year series

The chart on each name page is built by summing the group’s spellings for each year separately. It is generated from the same raw files as the totals, and every name’s series sums exactly to the total births shown next to it — this is checked automatically whenever the data is regenerated, because a chart that disagrees with the number beside it is worse than no chart.

Which names get a page

The SSA files list about 90,000 distinct names once spellings are grouped. Most of them are a handful of births in a handful of years, and a page built from three numbers is not worth anyone’s time to read or ours to publish. So the cut is at 5,000 recorded births: roughly 3,600 names, each with enough of a record to plot a curve and say something about it.

Anything below that is absent rather than thin. Browse and search cover exactly the same set, so nothing on the site links to a name that has no page.

Known limitations

  • The five-child threshold. Rare names are undercounted, and a name given to four children a year for a century does not appear at all.
  • Pronunciation is not universal. Regional and family variation means the grouping will sometimes merge names people consider distinct, or separate names they consider the same.
  • Sex is binary in the source. The SSA files record M or F only. That is a property of the source data, not a claim about people.
  • US only. Nothing here describes naming outside the United States.
  • Recorded name, not used name. The data reflects what was filed at birth, which is not always what a person goes by.

If you find an error in the grouping, the contact page explains how to report it.