There is a version of the plateau story you have probably heard. You got comfortable. You stopped pushing. The app could not help you because apps are for beginners anyway. The implied fix is character: hire a tutor, move to Spain, want it more.
Most of that story is wrong, and the part that is right is right for the wrong reason. There is a mechanical account of what happens around B1, it is not flattering to the tools rather than to you, and it predicts the experience much better.
The short answer
You did not run out of discipline at B1. You ran out of signal. Up to that point, learning Spanish means adding material you have never seen, which is easy for you to feel and trivial for software to score. From B1 onward the remaining work is depth on material you have already met, and depth is invisible to any instrument that only records whether you answered correctly.
So the measurements keep coming back clean while the Spanish stops moving. You experience that as stagnation and reach for the only explanation available, which is yourself. The accurate description is narrower: the feedback loop broke, and no amount of additional effort aimed through a broken loop produces a different result.
The plateau is a named stage, not a personal event
It helps to know that this has a literature and a name. Jack Richards wrote a Cambridge monograph called Moving Beyond the Plateau about exactly this transition, and his framing is that the learner arriving at the intermediate stage does not simply need more grammar and more words. They need new uses of forms they already hold, more complex resources, and the features of natural speech that no beginner syllabus contains.
The arithmetic underneath is also public. Cambridge publishes guided learning hours as cumulative estimates: roughly 180 to 200 hours to reach A2, 350 to 400 for B1, 500 to 600 for B2, 700 to 800 for C1. Read as steps rather than totals, the early moves cost around 100 to 150 hours and the B1 to B2 move costs 150 to 200.
So some of the flatness is real and expected, and we put numbers on that in B1 to B2: how long it takes. But the curve alone does not explain why learners at this stage cannot say what is wrong. That part is the instrument.
What actually changes at B1
Beginner courses are dense with novelty. Every lesson introduces visible new material, so progress is easy to feel and easy to score. The measurement model works because everything you need is on the menu. You learn ser and estar. You learn the present, then the past. Each unit adds a tool, and you can feel the belt getting heavier.
Around B1 the menu runs out. Not because Spanish is finished, but because the type of work changes. What remains is depth on material you have met: the ser and estar subtleties that were not in the first lesson, the subjunctive in contexts you half-learned, the register difference between two correct sentences, the social functions you have never once been asked to produce.
That depth was always there. Beginner exercises never touched it, so your tools never measured it, and their silence read as coverage.
From that point on, four things happen at once. You answer correctly, so you are scored as knowing. You keep producing the same safe constructions, so production never upgrades. Your understanding keeps growing through series and reading, which makes the gap between input and output wider rather than narrower. And the app reports streaks and XP, which were always measuring engagement.
The feedback loop is the thing that broke
This is the part worth being precise about, because “the app is bad” is not the claim.
Feedback is one of the most powerful influences on learning, and also one of the most variable. Hattie and Timperley’s synthesis on feedback makes the distinction that matters here: feedback works when it closes the distance between where a learner is and where they are going, and it does very little when it only confirms the current state. A green tick on a question you were always going to answer correctly is the second kind. It is information about the question, not about you.
It is not that you stopped climbing. The ladder stopped having visible rungs.
The strongest case for the motivation story
The motivation framing deserves its best version, because it is not stupid and it is not always wrong.
Plenty of people genuinely do coast. Two years into a language, the practical pressure drops off. You can order, work, socialise and handle the bank. The marginal return on effort falls at exactly the moment the marginal cost rises, and a lot of learners quietly stop trying while continuing to open the app. For those people, “you got comfortable” is an accurate description, and being told to push harder is correct advice.
The framing also survives because it is falsifiable in a satisfying way: people who take a tutor and go hard often do improve. That looks like proof.
And there is a real mechanism behind that improvement, which is worth conceding rather than explaining away. A good tutor is a measurement device. They hear the sentence you did not say, they ask the follow-up question you were hoping to avoid, and they refuse to accept the version of your answer that was merely acceptable. What looks like the effect of motivation is often the effect of finally being observed by something that can tell good enough from better.
Why it is still the wrong diagnosis for most people
The tell is what happens to learners who are demonstrably still trying.
If effort were the variable, sustained effort would move the needle. What actually happens is that people study for another year and produce the same Spanish, because more practice inside the same loop generates more of the same data: correct answers to questions drawn from a fixed tree. Volume without diagnosis deepens the grooves that are already there, including the avoidance grooves.
And avoidance is the mechanism that makes effort invisible. Elaine Schachter’s 1974 paper on avoidance found learners producing almost no relative clauses rather than producing them wrongly. Counting errors made them look strong. In Spanish this is exactly why an hour of aunque practice does nothing for the por mucho que you route around daily, and why subjunctive avoidance survives years of subjunctive exercises.
The end state has a name too. Larry Selinker described interlanguage in 1972 as the learner’s own systematic version of the language, with fossilisation as the point where it stops developing. His observation was that the large majority of adult learners stop somewhere short and stay there. Stable, fluent, finished. That is what a broken feedback loop produces given enough time, and no quantity of willpower routed through it will produce anything else.
What this looks like in daily life
You live in Spain. You have been here two years. Your Spanish is good enough: you handle bureaucracy, you socialise, you follow the news. You also notice that you avoid phone calls when you could text, that at group dinners you understand everything and contribute in fragments, that fast or emotional conversation collapses you back into simple sentences, and that serious discussions with your partner happen in English because your Spanish feels too blunt for nuance.
You have been B1-ish for a year and nothing moves. None of that is a discipline failure. It is what good-enough Spanish looks like from the inside when nothing can tell you what specifically is holding the next level back.
What the stage actually requires
Three things, all diagnostic before they are instructional.
Coverage probing. Systematic checks across structures and real-life functions, aimed specifically at what was never asked. Not “do you know the subjunctive” but “produce a polite refusal, a nuanced disagreement, a hypothetical regret, without dropping into simple structures”. This is what DELE B2 speaking exposes, and why the result so often startles confident B1 speakers.
Avoidance tracking. Noticing when a correct sentence used a simpler construction than the speaker’s level implies. This is the hard one, because the sentence was right. The question is whether it was the best version available or the safest one you always reach for.
Modality separation. Treating recognition, cued production, written production and spoken production as different states of the same item, each with its own schedule. Laufer and Goldstein validated that hierarchy across 435 learners in Testing Vocabulary Knowledge, and found the harder rungs predicted real performance better than the easy one. You know the word. Whether you can say it live at speed is a different question and needs a different score.
None of this is exotic pedagogy. It is simply expensive to measure, which is why products optimised for engagement do not measure it.
Why the fix is diagnosis rather than more hours
Once the three signals exist, the instruction that follows them is well understood and fairly boring.
Retrieval beats review. Roediger and Karpicke’s test-enhanced learning work found that restudying wins on an immediate test, which is why rereading feels productive, and loses badly on a delayed one. Producing from memory is not how you check that you learned something. It is the learning event.
Deliberate work has to be a portion of the diet rather than all of it or none. Paul Nation’s four strands put meaning-focused input, meaning-focused output, language-focused study and fluency development at roughly a quarter each. The plateau is what a diet of one strand looks like after two years, whichever strand it was.
And the target has to be the production column specifically, because that is the one that stalls. Batia Laufer’s work on passive and active vocabulary found the ratio between them moving the wrong way over a year of instruction: learners improving, passive vocabulary outrunning active, the gap widening as proficiency rose. More input aimed at a system that already recognises more than it can say makes the number worse.
What we do with this at Suelto
Suelto is a diagnostic engine with a course attached, and that ordering is the whole argument of this note.
Territory maps your Spanish on two axes, Grammar and Function, and marks every category as fog, partial or clear. Fog means we have never seen you attempt it, and it exists as a state because of Schachter: any system that scores only attempts will report a learner who avoids everything as a learner with no problems.
The conversation step records avoidance as a first-class signal rather than discarding it as a correct answer, and lessons get scheduled off that signal instead of off a syllabus position. Recognition and production are never merged into a single score, because the distance between them is the thing we are actually trying to move. Verbs get a timed step, because a form you can only produce in four seconds is a form you will avoid.
If you are stuck at B1, you do not have a motivation defect. You have an instrument problem, and that is better news than the alternative: the problem has a shape and a fix rather than being a vague fact about your character.
The gap test takes about ten minutes, needs no signup, and shows you what your current instruments cannot. The method page sets out the research the rest of it rests on.
You did not stop climbing. Somebody took the rungs off.