Method
How the order is decided
The whole site hangs off one claim: that this is the real order of English by frequency. That claim is only worth anything if the working is visible, so here it is.
The corpora
SPOKEN
OpenSubtitles 2018
Film and television subtitles. Excellent for conversational English, and the reason a list built on it alone over-ranks interjections.
NEWS
Leipzig Corpora — news 2023
English-language journalism. Brings in the vocabulary of public life — government, economy, law — that dialogue barely touches.
ENCYCLOPEDIC
Leipzig Corpora — Wikipedia 2016
Formal, descriptive prose, and the source that keeps abstract and technical vocabulary in proportion.
One corpus is a genre, not a language. Subtitles alone would teach you to follow a film and leave you lost in a newspaper; Wikipedia alone would do the reverse. Each word is ranked on how it behaves across all three.
How they are combined
01
Rates, not counts
The corpora are different sizes, so raw counts cannot be compared. Each is converted to occurrences per million words first.
02
Geometric mean
The three rates are averaged in log space. This rewards a word that is moderately common everywhere over one that spikes in a single register.
03
What is removed
Proper nouns, detected by how often the word appears capitalised in the case-preserving sources. Plus abbreviations and web debris. A word must appear in at least two corpora to rank at all.
04
What is added
Pronunciation, from CMUdict. Translations, part of speech, level and example sentences are being written by hand and are not here yet.
What English made us change
- Pronunciation is stored, not derived. Spanish spelling predicts sound well enough to compute it; English does not, so every transcription comes from CMUdict. It is General American and broad, so a British learner's vowel in “bath” is not what this list shows.
- The sources disagree about apostrophes — Leipzig keeps them, the subtitle list strips them — so only plain alphabetic words are ranked. The stripped forms fall out of the 5,000 on their own. The exception is “its”, where the subtitle list folds the possessive and “it's” together, so it ranks slightly high.
- “I” is always capitalised, so the proper-noun filter discarded it until it was named as an exception. Loosening the filter instead would have let real names back in.
- British and American spellings rank as separate entries — centre and center, defence and defense — because words are ranked as they are written. This is honest about the data and unhelpful to a learner, and it may change.
What is not finished
- Translations, part of speech, level and example sentences are being written. Until they are here, this is a ranked list with pronunciation and nothing more — which is why there is no vocabulary test yet.
- Three corpora are still three corpora. Spoken English that never becomes a subtitle, regional vocabulary and specialist registers are all under-represented.
- The coverage figures are measured against these corpora, with proper nouns, numerals and everything else the ranking excludes left in the denominator. That is why they are lower than the numbers usually quoted for this: the first 1,000 words cover 73.1% of subtitle dialogue but only 59.6% averaged across dialogue, news and encyclopedic prose. Quoting the first figure alone is how “1,000 words is 75% of English” became folklore.