Ask a plateaued learner to read a Spanish article and they glide through it. Two thousand words, maybe more, all understood or close enough. Ask them to explain the same event out loud without preparation and watch the vocabulary shrink to bien, vale, es que, no sé, and básicamente deployed as a stalling word.
This is not forgetting, and it is not a vocabulary shortage. It is the distance between passive and active knowledge, and it is both the most under-measured fact in language learning and one of the better documented ones.
The short answer
Knowing a word is not one thing that you either have or do not have. It is at least four things, arranged in a fixed order of difficulty, and almost every learning tool tests only the easiest one. Recognising a word, producing it from a cue, reaching for it while writing under pressure and saying it unprompted in live speech are four different states of the same item.
Your vocabulary is therefore not a number. It is four numbers, and they are very far apart. The first one is what apps report. The last one is what you actually have when you open your mouth, and it is usually a small fraction of the first.
The four rungs are measured, not a metaphor
The ladder is not our invention. Batia Laufer and Zvia Goldstein built a test around it and validated it on 435 learners in Testing Vocabulary Knowledge: Size, Strength, and Computer Adaptiveness. Their four degrees of knowing a word, ordered from easiest to hardest, are passive recognition, active recognition, passive recall and active recall.
Two findings from that paper matter here. The hierarchy held at every word-frequency level tested, so it is a property of how words are known rather than an artefact of which words were chosen. And the harder rungs predicted actual classroom performance better than the easy one, which is a polite way of saying that the thing cheap tests measure is the thing that matters least.
The counts in that figure are illustrative rather than measured, and the shape is not. Every study of the receptive and productive split finds the same steep fall.
A concrete example
Take imprescindible, meaning essential or indispensable. You have seen it dozens of times and understand it instantly. That is rung one.
Asked how to say “indispensable” in Spanish, you can produce it. Rung two, and note that the question did most of the work: it handed you the meaning and asked only for the form.
Now you are writing an email and something needs to be described as essential. Do you reach for imprescindible or do you type muy importante? Most people type muy importante, because precision costs retrieval effort and under time pressure the system takes the cheap option. That is rung three, and most people fail it quietly and never find out.
And in live conversation, asked for an opinion, do you produce es imprescindible que lo tengamos resuelto antes del viernes, or es muy importante? Almost everyone says the simpler one. Not from ignorance. The production path has never been built to conversational speed.
Four states, one word, entirely different levels of knowing. Every app in the market would record all four as a single green tick after you picked it correctly out of four options.
The example is deliberately unglamorous, and that is the point. Imprescindible is not a rare or literary word. It is ordinary adult Spanish, the register you already operate in when you speak English, and it is sitting in your head one rung too low to arrive when you need it. Multiply that by every precise word you have met and never said, and you have a fairly complete description of why speaking feels like a downgrade of your personality.
The gap widens as you improve
The intuition is that these numbers converge with study. The data says the opposite.
Laufer’s longitudinal work on passive and active vocabulary tracked the ratio of active to passive size across a year of instruction and watched it fall from 89% to 73%. The learners were improving throughout. Passive vocabulary was simply growing faster.
Stuart Webb found the same widening along the frequency axis: the rarer the word, the bigger the distance between recognising it and producing it. Which is exactly backwards from what you need, because the rare words are the ones that would make you sound like yourself.
This is why more input can make the experience worse. Every hour of series and reading pushes items onto rung one. Nothing about that hour pushes anything onto rung four. The pile grows at the bottom and the feeling of being stuck grows with it, and the same asymmetry drives understanding more than you can use.
Why apps collapse the ladder
A typical exercise shows por mucho que among distractors and asks you to identify it. Correct answer, item marked learned, item drifts into revision or out of scope.
That tested rung one. The other three were never probed, so their absence was never recorded. The progress bar fills on recognitions while the part of your Spanish that produces sentences stays where it was.
The collapse also hides the shape of your knowledge. You might be at rung three for half the subjunctive triggers, writing them fine under pressure, and rung one for all of them in speech, having never said one aloud without reading it. No dashboard built on right and wrong can represent that difference, because its exercises only sample one state.
What the collapse costs you
The cost is not abstract, and it is not really about vocabulary. It is about what you can and cannot do on a given afternoon.
You feel fluent enough on the strength of your comprehension, then freeze somewhere recognition is not enough: a doctor’s appointment, a work meeting, a disagreement that matters. You cannot explain why you are stuck, because every metric you have says you are doing well. You study harder, which pushes more items onto rung one and makes the gap feel wider. And you conclude that the problem is failing to apply what you know, when the truth is that applying it was never trained.
Merrill Swain’s output hypothesis names the reason input cannot fix this on its own. Comprehension lets you succeed semantically without processing the form, because context and content words carry the meaning. You can understand imprescindible perfectly for years without ever being forced to retrieve it, and nothing in the understanding builds the retrieval.
The strongest case for the collapse
There is a real argument on the other side and it deserves stating properly.
Recognition is not a consolation prize. It is the precondition for everything else, it is what makes input comprehensible, and a large recognition vocabulary is the thing that lets you read a newspaper and follow a film. A tool that grows it efficiently is doing something genuinely valuable, and the spaced repetition systems that do this are among the best-evidenced pieces of technology in education.
There is also an honest argument that rungs three and four cannot be forced. On this view production ripens: build enough recognition and enough exposure and the words move up on their own when they are ready, so measuring the upper rungs is measuring something you cannot act on anyway.
That position has a respectable version and a practical one. The respectable version is that forcing production before a word is well enough known produces errors that stick, so patience is protective rather than lazy. The practical version is that most learners have far more recognition than they can use, so the fastest route to better Spanish really is to keep reading and let the surplus drain upward. Both are worth taking seriously, and neither is obviously wrong on a single afternoon.
Why it does not hold
The ripening story predicts that the gap closes with proficiency, and the measurements say it widens. Laufer’s ratio moved the wrong way over a year of instruction, and Webb’s widened with rarity. Whatever moves words up the rungs, it is not simply more time and more input.
The mechanism for the stall is avoidance. Elaine Schachter’s 1974 paper found learners producing almost no relative clauses rather than producing them badly, which made an error count read them as strong. The Spanish version is muy importante standing in for imprescindible forever. The substitute is correct, nothing flags it, and the item never leaves rung one.
Speed is the other half. Norman Segalowitz separates three kinds of fluency, and the one that binds here is cognitive fluency: how fast the machinery can reach a word and assemble an utterance. Rung four is not a knowledge state at all. It is an access-speed state, and adding more words to a system that cannot reach the words it has does nothing for it.
What moving items up the ladder takes
Each transition needs its own kind of task, and none of them is a recognition exercise.
Rung one to two. Retrieval from a meaning cue rather than a word bank: a Spanish definition, a situation, a paraphrase. Never an English translation, which builds a dependency you do not want. Roediger and Karpicke’s test-enhanced learning work is the reason to insist on retrieval rather than review: rereading wins on an immediate test and loses badly on a delayed one.
Rung two to three. Writing under time pressure where you choose your own vocabulary and nothing tells you which word to use. Your meaning demands a certain precision and you either reach for it or you do not.
Rung three to four. Speaking in real time with no planning window. Answering questions, narrating, arguing a problem with your landlord. This is where the last path gets built, and it is the only rung that cannot be faked.
Paul Nation’s four strands is the useful frame for the diet: meaning-focused input, meaning-focused output, language-focused study and fluency development at roughly a quarter each. A learner living entirely on input is running one strand and wondering why rung four is empty.
What we do with this at Suelto
Suelto tracks this ladder per item rather than storing one score, and that is the main structural difference between it and a flashcard app.
An item is never “known”. It sits at a rung, each rung has its own schedule, and moving up a rung is the unit of progress rather than answering correctly again at the rung you already hold. Territory shows the shape that produces: fog where we have never seen you attempt something, partial where you hold it in one modality and not another.
Free use is the only graduation. A phrase counts as yours when it surfaces unprompted in something you produced spontaneously, which is why the conversation step records what you reached for and what you swapped out. The verb step is timed, because rung four is an access-speed state and untimed practice cannot detect it.
The gap test samples where your vocabulary actually sits, in about ten minutes with no signup, and the method page lists what each part rests on.
Your vocabulary is not two thousand words. It is two thousand, then eight hundred, then four hundred, then two hundred, and the only one of those four numbers that has ever spoken out loud is the last one.