Learners ask how many words they need and get answers ranging from a few hundred to tens of thousands. Both are correct, because they answer different questions. The arithmetic underneath is worth understanding, since it explains the most common frustration in language learning.
Frequency is brutally lopsided
Words are not evenly distributed. A small number of them do most of the work in any language, and the drop-off is steep. Approximate coverage of everyday text, which varies by language and by the corpus measured:
| Words known | Approx. coverage of running text |
|---|---|
| 100 | ~50% |
| 1,000 | ~80% |
| 2,000 | ~85–90% |
| 3,000 | ~95% |
| 5,000 | ~98% |
| 10,000 | ~99% |
The first thousand words buy you eighty per cent. The next nine thousand buy you nineteen. That shape is why beginners feel fast and intermediate learners feel stuck — the returns genuinely collapse, and it is not a failure of method.
The gap between coverage and comprehension
Here is the part that surprises people, and it is where the frustration comes from.
Ninety-five per cent coverage sounds close to complete. It means one unknown word in twenty — roughly one per two lines of text, several per paragraph. Every one interrupts you, and worse, the unknown words are disproportionately the ones carrying the specific meaning, because the frequent words you do know are largely grammatical machinery.
Research on reading comprehension has repeatedly landed on around ninety-eight per cent as the threshold for comfortable independent reading. Getting from ninety-five to ninety-eight is not a small final push — on the table above it is roughly a doubling of vocabulary.
This is the intermediate plateau, quantified. The plateau is not psychological and it is not a sign you have stopped learning. It is the point where each new word is genuinely worth less than the last one was, while the amount of text you can read comfortably has not yet crossed the threshold where reading becomes self-sustaining. Knowing the shape of the curve makes it easier to keep going through it.
Spoken language is much smaller than written
The numbers above describe text. Speech uses a considerably narrower vocabulary — conversation reuses a small core relentlessly, and topics are constrained by context in a way that writing is not.
This is why learners who can hold a conversation are so often unable to read a newspaper. They are not two levels of the same skill. Conversation rewards a small well-drilled vocabulary and rapid recall; reading rewards breadth. If your goal is talking to people, the first two thousand words go a very long way. If your goal is reading, they do not.
What counts as one word
Every figure here depends on a definition, and this is where published numbers diverge most.
- Word families group a base with its predictable relatives — help, helps, helped, helpful, unhelpful as one item. Most coverage research counts this way.
- Individual forms count each separately, which inflates the total substantially.
The distinction matters enormously for heavily inflected languages. A Finnish noun has fifteen cases in singular and plural; a Russian verb has dozens of forms. Counting forms rather than families makes those languages look like they need five times the vocabulary, when the learner's actual task is a base plus a system.
What the numbers suggest doing
- Do the first thousand by frequency, deliberately. This is the one stage where a frequency list beats anything else, and the payoff is the largest you will ever get.
- Then switch to material you care about. After roughly two thousand words, general frequency lists start giving you words you rarely meet. Vocabulary drawn from things you actually read or watch is more efficient from that point.
- Accept partial knowledge. Recognising a word when you meet it is most of its value and arrives far sooner than being able to produce it. Chasing production for every word slows you down.
- Make lookup free. The single strongest predictor of vocabulary growth is how often you check a word you half-know, and that is governed by friction. If a lookup costs a loading spinner or a data charge, you will guess instead — which is how half-learned words fossilise.
The honest summary: a thousand words to start talking, three thousand to stop drowning, five thousand before reading feels like reading. The first of those takes weeks and the last takes years, and the curve between them is the reason.
Look words up without friction
Vocabulary grows fastest when checking a word costs nothing. NDT Studio's offline dictionaries make lookup instant, with no connection and no data cost.
Open the dictionary appFrequently asked questions
How many words do I need to have a conversation?
Roughly a thousand of the most frequent words covers a large majority of everyday spoken language — commonly estimated around 80 per cent of the words people actually say. That is enough to hold a simple conversation, though not enough to follow one between two native speakers talking normally.
Why do I still not understand anything at 2,000 words?
Because coverage and comprehension are different things. At 95 per cent coverage you meet an unknown word roughly every twentieth word, which is about one per two lines of text. That is frequent enough to break the thread continuously. Comfortable reading generally needs closer to 98 per cent, which takes several times more vocabulary.
Is it better to learn frequent words or words about my interests?
Frequent words first, then your own field. The frequency list gets you the machinery of the language — the words that appear regardless of topic. After the first couple of thousand, general frequency lists give diminishing returns and vocabulary from material you actually care about is more efficient.