The Plateau · Train

Can you learn Spanish with the 1,000 most common words?

The ninth most common Spanish word is se. Try putting that on a flashcard. How many Spanish words you really need, and exactly where the shortcut stops.

THE TOP OF A SPANISH FREQUENCY LIST 01 de 02 la 03 que 04 el 05 en 06 y 07 a 08 los 09 se 10 del Ten out of ten are grammatical words. The first noun on the list sits at rank 47. COVERAGE OF SPOKEN SPANISH, BY THOUSAND 1st thousand · 87.8% +2nd thousand: 4.9 +3rd thousand: 2.3 6% still unknown A list ranked by usefulness puts the hardest grammar in the first fifty cards. se (9) · lo (21) · le (27) · none of them fits on the back of a flashcard

You have seen the thumbnail. A number, a stopwatch, a promise: learn the 1,000 most common Spanish words and you will understand 80% of everything. Sometimes it is 300 words and 65%. Sometimes 2,000 and 90%. The number moves. The pitch does not.

It is a good pitch, and the reason it works is that the underlying data is real. Somebody did count. The trouble starts in the gap between the corpus paper and the video: a coverage statistic gets quietly reclassified as a study plan, and nobody stops to ask whether the thing being counted is the thing being sold.

So here is an audit. Where the number comes from, who turned it into a method, the strongest version of their case, the four places it breaks, and what the people defending it say back. We build a Spanish system for people stuck at the B1 plateau, so we are not a neutral party. We will say where we land at the end rather than pretending we have no position.

The short answer

Can you learn Spanish with the 1,000 most common words? No, but you can understand a surprising amount of it. The coverage claim is roughly accurate: a thousand of the most frequent Spanish lemmas account for about 88% of spoken Spanish and 76 to 80% of written Spanish. What that buys you is recognition, sitting still, with time to think.

It does not buy you speech, for four reasons this note works through in order. Nobody agrees what one word is, so the promised thousand quietly includes fifty conjugated forms per verb. The top of the list is not vocabulary at all: se is the ninth most common word in Spanish, lo is twenty-first, le is twenty-seventh, and all ten of the top ten are grammatical words that will not fit on a flashcard. Coverage is not comprehension; the two rise together in a straight line with no magic threshold. And roughly half of real conversation is made of multi-word chunks that no list of single words contains.

The honest version of the method: use frequency to decide what order you meet words in, and something else entirely to decide what you do with them once you have met them.

Where the 80% claim comes from

Not from language influencers. From corpus linguistics, and it is a serious, careful literature.

For Spanish, the reference point is Mark Davies, who built a 20-million-word corpus split evenly between spoken transcripts, fiction and non-fiction, from Spain and Latin America, tagged it for part of speech and lemma, and then measured exactly this question. His coverage paper (Davies, 2005), written alongside the Routledge Frequency Dictionary of Spanish, gives the honest table:

TEXT COVERAGE BY FREQUENCY BAND · SPANISH (DAVIES, 2005) 1ST THOUSAND +2ND +3RD TOTAL Spoken 87.8% 4.9 2.3 94.0% Fiction 79.6% 6.5 3.5 89.6% Non-fiction 76.0% 8.0 4.2 88.2% The first thousand does almost all the work. Everything after it is a long tail. Davies: "There clearly is a law of diminishing returns in terms of vocabulary learning."

So the popular claim, on its own terms, is roughly right. A thousand Spanish lemmas cover about 76 to 80 percent of written text and close to 88 percent of conversation. Doubling to two thousand buys you five to eight points more. The third thousand buys two to four. Davies says it plainly: there is a law of diminishing returns.

The same shape holds in English. Paul Nation’s survey of vocabulary size (2006) found the first thousand word families covering around 78 to 82 percent depending on the text, with each later band adding a few points, and concluded that you need 8,000 to 9,000 word families to read unassisted and 6,000 to 7,000 to follow speech. Adolphs and Schmitt’s study of spoken discourse (2003) found 2,000 word families still falling short of 95% coverage of real conversation.

Notice what the papers are measuring: the percentage of running words on a page or in a transcript. That is a fact about texts. It was never a claim about what a learner can do.

Who turned the frequency list into a study method

Teaching from a frequency list is not new and did not start on YouTube. Michael West’s General Service List of roughly 2,000 English words appeared in 1953 and shaped textbook design for decades. Frequency dictionaries for Spanish go back to the 1920s, and Davies built his because the existing ones were drawn from literary corpora and produced a top-500 containing poeta, marqués and dama.

The modern popular version has a different flavour. Tim Ferriss’s 12 rules put the frequency argument in front of a mass audience with a Pareto framing: “in English just 300 words make up 65% of all written material,” so front-load the 20% of vocabulary that does 80% of the work, and return to “academic material and grammar books” afterwards to tidy up. Benny Lewis built Fluent in 3 Months partly on the same logic. Gabriel Wyner’s Fluent Forever opens with a fixed 625-word base list, learned as images rather than translations and fed through spaced repetition. Around all of it grew the Anki and sentence-mining ecosystem, and then the video genre: 1000 Most Common Spanish Words, four hours long, one word every few seconds, millions of views.

Worth saying: the more careful voices inside that same community are already sceptical. Olly Richards, who sells language courses for a living, argues that “knowing a word and being able to use it appropriately in conversation are two completely different things,” and that isolated list learning leaves you “not really knowing those words particularly well.” He quotes Steve Kaufmann’s point that if you are genuinely spending time in the language, the top thousand words arrive on their own without any deliberate effort at all.

The gap that matters is not between researchers and practitioners. It is between what the careful proponents say and what the thumbnail says.

The strongest case for learning the most common words first

Before the objections, the steelman, because this method is not a scam and its defenders are not cranks.

Frequency is the best single predictor of usefulness there is. If you can only choose the order in which you meet 5,000 words, ranking them by frequency and range beats every other heuristic, including intuition, including a textbook’s topic order. That is not controversial.

Nation, the researcher most often cited against the method, is on their side for the first stretch. His position across Learning Vocabulary in Another Language is that high-frequency words, roughly the first 2,000 to 3,000, are worth deliberate teaching time precisely because their coverage is so lopsided, and that mid- and low-frequency words are the ones better left to extensive reading and learner strategies. Deliberate attention to the top of the list is the mainstream recommendation, not a fringe one.

Deliberate list learning produces real knowledge, not inert knowledge. The common objection that flashcard words never become usable turns out to be too strong. Irina Elgort’s experiment (2011) taught 48 invented words through word cards over a week, then probed them with masked repetition and semantic priming. The deliberately learned items behaved like genuine lexical entries: they primed, they were accessed fluently, they had been integrated into the mental lexicon. Per hour spent, deliberate vocabulary study is faster than waiting for words to turn up.

So: frequency-ordered, deliberately studied vocabulary is a legitimate component of a serious course. That is the true claim. The oversold claim is that it is the course.

Problem one: nobody agrees what one Spanish word is

Ask three people for “the 1,000 most common Spanish words” and you get three different lists, because word means three different things.

A word form is a string. Tengo, tienes, tuve and tendría are four words.

A lemma is a dictionary headword with its inflections folded in. Those four collapse into tener, one word.

A word family goes further and gathers derivations too: pintar, pintura, pintor and pintoresco under a single head.

The difference is not academic. Take the Real Academia Española’s CREA frequency list, 160 million words of Spanish from 1975 to 2004, half from Spain and half from the Americas. It counts orthographic forms. On that list the top 1,000 items cover 66% of running text, not 80%. Davies reaches 76 to 88 percent for the same nominal thousand because he counts lemmas, and Nation’s English figures run higher again because he counts families.

Which means the popular claim, unpacked, says something it never admits. “Learn 1,000 words and understand 80% of Spanish” holds only if your thousand entries are lemmas, and a lemma is a promissory note. Behind tener stand roughly fifty conjugated forms across indicative and subjunctive. The flashcard says tener = to have. The transcript says tuviéramos. You cannot cash the coverage figure without doing the morphology, and the morphology is exactly what the list leaves out. It is the same confusion we wrote about in conjugation tables don’t make you conjugate, arriving from a different direction.

Problem two: the most common Spanish words are grammar, not vocabulary

This is the objection that finishes the argument, and you can check it yourself in ten minutes.

Take the CREA list and look at the first hundred entries. Not the first thousand. The hundred that carry the most weight, covering 48% of all running text between them.

WHAT THE 100 MOST FREQUENT SPANISH FORMS ACTUALLY ARE 51 function words 15 verb 14 adv 12 det 8 n. The 15 verb forms belong to only 7 verbs: ser, estar, haber, tener, hacer, poder, decir. The 8 nouns: años, vez, parte, tiempo, vida, gobierno, día, país. WHERE THINGS SIT ON THE LIST 01 de 09 se 21 lo 27 le 47 años first noun on the list 100 Every one of the first ten entries is a grammatical word. Source: RAE CREA, 160M words.

Fifty-one of the hundred are pure function words: articles, prepositions, conjunctions, pronouns, clitics. Twelve more are determiners, quantifiers, numerals and prenominal modifiers like gran. Fourteen are adverbs. Fifteen are verb forms, and those fifteen belong to just seven verbs, every one of them irregular, with fourteen of the fifteen appearing as conjugated forms rather than dictionary entries. Which leaves eight nouns in the entire top hundred: años, vez, parte, tiempo, vida, gobierno, día, país. The first noun does not show up until rank 47.

Now look at where the hard cases sit. Se is the ninth most common word in the Spanish language. Lo is twenty-first. Le is twenty-seventh.

Try writing a flashcard for se. What goes on the back?

It is the reflexive clitic in se levanta. It is reciprocal in se conocen. It is impersonal in se dice que. It is passive in se venden pisos. It marks the accidental in se me cayó, which quietly reassigns blame from you to the plate. It is the lexical part of pronominal verbs where it changes the meaning outright: ir is to go, irse is to leave. And it is the stand-in for le whenever an indirect object collides with a direct one, which is why le lo di is impossible and se lo di is what people actually say.

Seven jobs, one form, no shared meaning between them, ranked ninth. A card reading se = oneself is not a piece of knowledge. It is a placeholder for a chapter you have not read.

ONE CARD, SEVEN JOBS se oneself, himself… RANK 9 OF 160M WORDS se levanta reflexive se conocen reciprocal se dice que… impersonal se venden pisos passive se me cayó accidental ir / irse meaning shift se lo di stands in for le and the ones you meet later No meaning is shared across the seven. Frequency ranked the form. It cannot explain it. Same problem, smaller: lo (21) against le (27) is case, not vocabulary.

Lo and le are that problem in miniature. A list gives you lo = it, him and le = to him, to her, and that gloss will not tell you why le goes with gustar and lo with ver, why Madrid says le he visto and Seville says lo he visto, or why lo also fronts a whole clause in lo que pasa es que. The distinction is case and argument structure. It is grammar wearing a two-letter costume.

Here is the structural point, and it is not an accident of Spanish. Frequency lists rank by usefulness, and grammatical words are the most useful words in any language, so a frequency list is guaranteed to put the least memorisable material at the very top. The method’s own ranking principle loads the hardest grammar into the first fifty cards and then labels it vocabulary.

Problem three: 94% coverage is not 94% comprehension

Suppose you get past all of that. You know the top 3,000 lemmas and their inflections. Davies says that is 94% coverage of spoken Spanish. Ninety-four sounds like a pass.

It is not. Ninety-four percent means roughly one unknown word in every seventeen you hear. In conversation that is a gap every sentence or two, arriving at speaking speed, with no pause button.

The experimental work is blunter. Hu and Nation’s study of unknown word density (2000) replaced fixed proportions of a fiction text with nonsense words and measured comprehension. At 80% coverage nobody reached adequate comprehension. At 90%, a small minority did. They proposed 98% as the threshold for reading unassisted, and that figure is where the 8,000 to 9,000 word family estimate comes from.

Then Schmitt, Jiang and Grabe ran the largest test of the idea (2011), with 661 participants across eight countries, and found something more uncomfortable for everybody. The relationship between coverage and comprehension is essentially linear. There is no threshold, no cliff, no percentage at which understanding switches on. Every point of coverage buys a proportional slice of comprehension, and no amount of it buys a step change.

That kills the shape of the promise. “Learn 1,000 words and unlock 80% of Spanish” implies a door. There is no door. There is a ramp, and the first thousand words leave you at the bottom of it with a very good view.

And all of this is about understanding, sitting still, with time to think. It says nothing at all about your ability to say anything.

Problem four: real Spanish is built from chunks, not single words

The deepest problem is not that the list is too short. It is that the list is made of the wrong things.

Erman and Warren’s count of prefabricated language (2000) found that a little over half of both spoken and written discourse consists of multi-word units speakers select whole rather than assemble word by word. Conversation is not built from a dictionary. It is built from a phrasebook you cannot see.

In Spanish, the units that carry a real conversation are almost all made from top-100 words, and almost none of them are derivable from the individual entries. Es que. O sea. Lo que pasa es que. De hecho. A ver si. Me da igual. Qué va. Ya te digo. You can know lo, que, pasa and es perfectly and still not have lo que pasa es que, which is not a sum of its parts but a single move meaning roughly “here comes my excuse.” We came at this from the listening side in why native speakers sound fast and from the speaking side in filler words and discourse markers.

Two smaller problems in the same family. Davies notes that raw frequency has to be corrected for range: a word can look frequent overall because it is dense in three articles about cardiology and absent everywhere else. And every list carries the fingerprints of its corpus. CREA is 90% written and pan-Hispanic. A subtitle-derived list is dubbed television. Neither one is a Tuesday afternoon in a Spanish gestoría. That is why vosotros never rises in a pan-Hispanic frequency list even though it is unavoidable in Spain, and why a list built for one continent quietly miscalibrates you for the other.

Can you learn Spanish without grammar? What the defenders say

Three arguments come up, and they are not equally good.

“Grammar comes free from exposure. Get the words in and let the input do the rest.” This is Krashen’s position filtered through a decade of comprehensible-input advocacy, and there is real substance in it: nobody gets fluent without volume. But the strong version does not survive contact with the meta-analytic record. Norris and Ortega’s synthesis of L2 instruction studies (2000) found that focused instruction produces large, durable gains and that explicit approaches outperform implicit ones. Input is necessary. Input alone is slow, and it leaves stable holes precisely where the forms are hard to notice.

“Sentence mining handles the grammar. See se lo di five hundred times and you internalise it.” For salient, meaning-bearing forms this works reasonably well. For non-salient, redundant ones it works badly, and Bill VanPatten’s input processing research explains the mechanism. Under his Lexical Preference Principle, learners take meaning from content words and skip the grammatical marker carrying the same information, so the marker never gets processed however often it appears. Under his First Noun Principle, learners assign the agent role to the first noun or pronoun they meet, which means an intermediate hearing Lo saluda María reliably decodes it as “he greets María” rather than “María greets him.” Every word in that sentence sits in the top thirty of the frequency list. Knowing all of them produces the wrong reading. That is the cleanest answer available to “grammar will come later”: for a subset of forms, quite reliably, it will not. We take that argument apart properly in can you learn Spanish without grammar?, including the decades of immersion evidence that settles it.

“It’s a starting point, not the whole plan.” This one is correct, and it is what the serious proponents actually believe. Richards says it. Kaufmann says it. Wyner’s 625 words are explicitly a base for pronunciation and imagery before grammar work begins, not a replacement for it. The trouble is that this version does not fit in a thumbnail, so the version that travels is the other one.

So how many Spanish words do you actually need?

Here is the number, with the caveats attached. For conversational comfort in Spain, aim at 2,000 to 3,000 high-frequency items, which is where Adolphs and Schmitt put the edge of 95% coverage of real speech. For reading a newspaper without a dictionary, the honest figure is 8,000 to 9,000 word families, because that is what 98% coverage costs. And for saying any of it out loud, the count matters far less than what state each item is in, which is the part no list tracks.

Strip out the overclaim and something genuinely useful is still standing.

Frequency should decide order, not method. It is the right way to sequence what you meet and the wrong way to describe what you do with it.

The list splits in two, and the halves need opposite treatment. Content words, mostly further down the ranking, respond well to deliberate spaced retrieval, and Elgort’s evidence holds for them. Grammatical words, which is most of the top two hundred, need teaching, contrast and production practice, because a gloss on a card is not a description of what they do.

The unit should be the chunk more often than the word, because half of real speech is prefabricated.

And deliberate study should be a slice, not the whole. Nation’s four strands (2007) argue for a balanced course in four roughly equal quarters: meaning-focused input, meaning-focused output, deliberate language-focused learning, and fluency development. On that arithmetic a pure frequency-list regime is one quarter of a course being sold as the entire thing, and the missing three quarters are the ones that produce speech.

What we do with this at Suelto

We took the frequency argument seriously and then refused to stop where it stops.

Vocabulary in Suelto is frequency-informed and Peninsular-weighted, but no item is ever a single card with a gloss. Every item moves along a five-state ladder: recognised, retrieved, produced, spoken, spontaneous. Nothing counts as learned until it appears unprompted in something you produced, which is the distinction we set out in passive vs active vocabulary. Scheduling runs per item and per state, because recognising a word and reaching for it under pressure decay on different curves.

The grammatical top of the list is taught as grammar, not stocked as vocabulary. Se gets a chain of lessons, not a card. Clitics get contrast drills where the meaning you want forces the choice between lo and le. Verbs get their own daily step, one verb in one tense, all six persons, timed, because a form you can only produce in four seconds is a form you will avoid in conversation.

Territory exists because coverage statistics lie. It maps your Spanish on two axes, Grammar (can you produce the structures?) and Function (can you do things with the language?), and every category sits in one of three states: fog, partial, or clear. Fog is not “you got it wrong.” Fog is “we have never seen you try this,” which is exactly the state a green progress bar is designed to hide. Journey puts that map beside your course, your verb matrix and your vocabulary ladder in one view, so the four things that can independently stall are visible at the same time.

None of this is invented here. The method page lists the sources we build on and where each one shows up in the product: retrieval practice, spaced repetition, the lexical approach, comprehensible input, automaticity through timed practice. We would rather be checkable than clever. If you want the ten-minute version of your own map, the gap test samples it without a signup.

The one Spanish vocabulary shortcut that is real

Here is the irony. There is a genuine free-vocabulary shortcut in Spanish for English speakers, it is larger than the frequency list, and almost nobody sells it, because it is a rule rather than a number.

English and Spanish share upwards of twenty thousand cognates through Latin, and a recent Applied Linguistics study of academic spoken vocabulary found roughly half the list to be Spanish-English cognates, with false friends under one percent. Better still, cognates come in patterns. -tion becomes -ción, and that single correspondence hands you situación, educación, información, condición, hundreds of items at once, with the meaning already installed.

That is what a real shortcut looks like: generative, not enumerated. You learn a rule and get a family. Our English Advantage walkthrough runs eight minutes, needs no signup, and covers six of these patterns plus the -ate to -ar verb bridge, three sentence frames to hang them on, and the five false friends most likely to embarrass you first. Most people finish it holding several hundred words they did not know they had.

It has honest limits too, and we would rather state them than let you discover them. Cognates skew Latinate, formal and abstract. They will hand you la situación es complicada and they will not hand you venga, tío, qué va. They are a launchpad, not a course, which is precisely what we would say about the thousand most common words.

Frequency data is good data. It tells you which words to meet first and it is right about that. It simply cannot tell you what any of them do, and in Spanish the words it puts at the very top are the ones that need the most explaining.

Common questions

How many Spanish words do you need to know to be fluent?

Around 1,000 lemmas covers roughly 88% of spoken Spanish and 76 to 80% of written Spanish, and about 3,000 gets you to 94% of speech (Davies, 2005). But coverage is not comprehension: research on unknown word density puts the threshold for unassisted understanding near 98%, which is 8,000 to 9,000 word families. For conversational comfort, 2,000 to 3,000 high-frequency items plus the grammar that holds them together is the realistic target.

Is learning the 1,000 most common Spanish words enough to speak Spanish?

No. It is enough to follow a lot of what you hear and almost never enough to produce it. Recognising a word and retrieving it under time pressure are separate skills on separate schedules, and a frequency list only ever trains the first one.

What percentage of Spanish do the 1,000 most common words cover?

About 87.8% of spoken Spanish, 79.6% of fiction and 76.0% of non-fiction, measured on lemmas in a 20-million-word corpus. The second thousand adds only 5 to 8 points and the third adds 2 to 4, so returns fall off sharply after the first band.

Can you learn Spanish without studying grammar?

Not efficiently. The most frequent Spanish words are grammatical words: se ranks 9th, lo 21st, le 27th, and every one of the top ten is a function word. Meta-analytic evidence also favours explicit instruction over purely implicit exposure, and input-processing research shows learners can know every word in a sentence and still parse it wrong.

Does the 80/20 rule work for learning Spanish vocabulary?

For deciding the order in which you meet words, yes. Frequency is the single best predictor of usefulness. As a description of what to do with those words it fails, because the first 20% of the list is mostly grammar and roughly half of real speech is made of multi-word chunks rather than single words.

The method · Train

The reflex is trained, not learned.

Notes like this one name the mechanic. The nightly drill is how it becomes a reflex.

See the method how Suelto actually closes the loop