The Plateau · Diagnose

The B1 plateau is not a motivation problem

You did not lose discipline at B1. You ran out of signal. Why the intermediate plateau is a measurement failure, and what the stage actually requires.

PROGRESS, AS MEASURED B2 B1 A2 the menu runs out the plateau starts here scores keep saying “known”. production says otherwise.

There is a version of the plateau story you have probably heard. You got comfortable. You stopped pushing. The app could not help you because apps are for beginners anyway. The implied fix is character: hire a tutor, move to Spain, want it more.

Most of that story is wrong, and the part that is right is right for the wrong reason. There is a mechanical account of what happens around B1, it is not flattering to the tools rather than to you, and it predicts the experience much better.

The short answer

You did not run out of discipline at B1. You ran out of signal. Up to that point, learning Spanish means adding material you have never seen, which is easy for you to feel and trivial for software to score. From B1 onward the remaining work is depth on material you have already met, and depth is invisible to any instrument that only records whether you answered correctly.

So the measurements keep coming back clean while the Spanish stops moving. You experience that as stagnation and reach for the only explanation available, which is yourself. The accurate description is narrower: the feedback loop broke, and no amount of additional effort aimed through a broken loop produces a different result.

The plateau is a named stage, not a personal event

It helps to know that this has a literature and a name. Jack Richards wrote a Cambridge monograph called Moving Beyond the Plateau about exactly this transition, and his framing is that the learner arriving at the intermediate stage does not simply need more grammar and more words. They need new uses of forms they already hold, more complex resources, and the features of natural speech that no beginner syllabus contains.

The arithmetic underneath is also public. Cambridge publishes guided learning hours as cumulative estimates: roughly 180 to 200 hours to reach A2, 350 to 400 for B1, 500 to 600 for B2, 700 to 800 for C1. Read as steps rather than totals, the early moves cost around 100 to 150 hours and the B1 to B2 move costs 150 to 200.

EACH LEVEL COSTS MORE THAN THE ONE BEFORE CUMULATIVE GUIDED LEARNING HOURS, CAMBRIDGE ESTIMATES A2 180 to 200 B1 350 to 400 B2 500 to 600 C1 700 to 800 The B1 to B2 step alone is roughly 150 to 200 hours. Same weekly effort, less visible movement. The stall is partly just the shape of the curve.

So some of the flatness is real and expected, and we put numbers on that in B1 to B2: how long it takes. But the curve alone does not explain why learners at this stage cannot say what is wrong. That part is the instrument.

What actually changes at B1

Beginner courses are dense with novelty. Every lesson introduces visible new material, so progress is easy to feel and easy to score. The measurement model works because everything you need is on the menu. You learn ser and estar. You learn the present, then the past. Each unit adds a tool, and you can feel the belt getting heavier.

Around B1 the menu runs out. Not because Spanish is finished, but because the type of work changes. What remains is depth on material you have met: the ser and estar subtleties that were not in the first lesson, the subjunctive in contexts you half-learned, the register difference between two correct sentences, the social functions you have never once been asked to produce.

That depth was always there. Beginner exercises never touched it, so your tools never measured it, and their silence read as coverage.

From that point on, four things happen at once. You answer correctly, so you are scored as knowing. You keep producing the same safe constructions, so production never upgrades. Your understanding keeps growing through series and reading, which makes the gap between input and output wider rather than narrower. And the app reports streaks and XP, which were always measuring engagement.

The feedback loop is the thing that broke

This is the part worth being precise about, because “the app is bad” is not the claim.

Feedback is one of the most powerful influences on learning, and also one of the most variable. Hattie and Timperley’s synthesis on feedback makes the distinction that matters here: feedback works when it closes the distance between where a learner is and where they are going, and it does very little when it only confirms the current state. A green tick on a question you were always going to answer correctly is the second kind. It is information about the question, not about you.

THE SAME LOOP, BEFORE AND AFTER THE MENU RUNS OUT A1 TO A2 · LOOP CARRIES SIGNAL task can genuinely fail failure names the gap next task changes progress is visible to you and the tool B1 ONWARD · LOOP RETURNS NOTHING task is inside your range correct answer names nothing next task is the same the dashboard still looks healthy Nothing about the learner changed between these two panels. Hattie and Timperley: feedback that only confirms the current state does very little.

It is not that you stopped climbing. The ladder stopped having visible rungs.

The strongest case for the motivation story

The motivation framing deserves its best version, because it is not stupid and it is not always wrong.

Plenty of people genuinely do coast. Two years into a language, the practical pressure drops off. You can order, work, socialise and handle the bank. The marginal return on effort falls at exactly the moment the marginal cost rises, and a lot of learners quietly stop trying while continuing to open the app. For those people, “you got comfortable” is an accurate description, and being told to push harder is correct advice.

The framing also survives because it is falsifiable in a satisfying way: people who take a tutor and go hard often do improve. That looks like proof.

And there is a real mechanism behind that improvement, which is worth conceding rather than explaining away. A good tutor is a measurement device. They hear the sentence you did not say, they ask the follow-up question you were hoping to avoid, and they refuse to accept the version of your answer that was merely acceptable. What looks like the effect of motivation is often the effect of finally being observed by something that can tell good enough from better.

Why it is still the wrong diagnosis for most people

The tell is what happens to learners who are demonstrably still trying.

If effort were the variable, sustained effort would move the needle. What actually happens is that people study for another year and produce the same Spanish, because more practice inside the same loop generates more of the same data: correct answers to questions drawn from a fixed tree. Volume without diagnosis deepens the grooves that are already there, including the avoidance grooves.

And avoidance is the mechanism that makes effort invisible. Elaine Schachter’s 1974 paper on avoidance found learners producing almost no relative clauses rather than producing them wrongly. Counting errors made them look strong. In Spanish this is exactly why an hour of aunque practice does nothing for the por mucho que you route around daily, and why subjunctive avoidance survives years of subjunctive exercises.

The end state has a name too. Larry Selinker described interlanguage in 1972 as the learner’s own systematic version of the language, with fossilisation as the point where it stops developing. His observation was that the large majority of adult learners stop somewhere short and stay there. Stable, fluent, finished. That is what a broken feedback loop produces given enough time, and no quantity of willpower routed through it will produce anything else.

What this looks like in daily life

You live in Spain. You have been here two years. Your Spanish is good enough: you handle bureaucracy, you socialise, you follow the news. You also notice that you avoid phone calls when you could text, that at group dinners you understand everything and contribute in fragments, that fast or emotional conversation collapses you back into simple sentences, and that serious discussions with your partner happen in English because your Spanish feels too blunt for nuance.

You have been B1-ish for a year and nothing moves. None of that is a discipline failure. It is what good-enough Spanish looks like from the inside when nothing can tell you what specifically is holding the next level back.

What the stage actually requires

Three things, all diagnostic before they are instructional.

Coverage probing. Systematic checks across structures and real-life functions, aimed specifically at what was never asked. Not “do you know the subjunctive” but “produce a polite refusal, a nuanced disagreement, a hypothetical regret, without dropping into simple structures”. This is what DELE B2 speaking exposes, and why the result so often startles confident B1 speakers.

Avoidance tracking. Noticing when a correct sentence used a simpler construction than the speaker’s level implies. This is the hard one, because the sentence was right. The question is whether it was the best version available or the safest one you always reach for.

Modality separation. Treating recognition, cued production, written production and spoken production as different states of the same item, each with its own schedule. Laufer and Goldstein validated that hierarchy across 435 learners in Testing Vocabulary Knowledge, and found the harder rungs predicted real performance better than the easy one. You know the word. Whether you can say it live at speed is a different question and needs a different score.

WHAT THE STAGE NEEDS THAT SCORING CANNOT GIVE IT Coverage probing asks for what you did not choose, so the never-attempted becomes visible finds absence Avoidance tracking reads a correct sentence and asks whether it was the safest one available finds the dodge Modality separation scores recognise, write and say as three states of one item, never as one finds the ceiling

None of this is exotic pedagogy. It is simply expensive to measure, which is why products optimised for engagement do not measure it.

Why the fix is diagnosis rather than more hours

Once the three signals exist, the instruction that follows them is well understood and fairly boring.

Retrieval beats review. Roediger and Karpicke’s test-enhanced learning work found that restudying wins on an immediate test, which is why rereading feels productive, and loses badly on a delayed one. Producing from memory is not how you check that you learned something. It is the learning event.

Deliberate work has to be a portion of the diet rather than all of it or none. Paul Nation’s four strands put meaning-focused input, meaning-focused output, language-focused study and fluency development at roughly a quarter each. The plateau is what a diet of one strand looks like after two years, whichever strand it was.

And the target has to be the production column specifically, because that is the one that stalls. Batia Laufer’s work on passive and active vocabulary found the ratio between them moving the wrong way over a year of instruction: learners improving, passive vocabulary outrunning active, the gap widening as proficiency rose. More input aimed at a system that already recognises more than it can say makes the number worse.

What we do with this at Suelto

Suelto is a diagnostic engine with a course attached, and that ordering is the whole argument of this note.

Territory maps your Spanish on two axes, Grammar and Function, and marks every category as fog, partial or clear. Fog means we have never seen you attempt it, and it exists as a state because of Schachter: any system that scores only attempts will report a learner who avoids everything as a learner with no problems.

The conversation step records avoidance as a first-class signal rather than discarding it as a correct answer, and lessons get scheduled off that signal instead of off a syllabus position. Recognition and production are never merged into a single score, because the distance between them is the thing we are actually trying to move. Verbs get a timed step, because a form you can only produce in four seconds is a form you will avoid.

If you are stuck at B1, you do not have a motivation defect. You have an instrument problem, and that is better news than the alternative: the problem has a shape and a fix rather than being a vague fact about your character.

The gap test takes about ten minutes, needs no signup, and shows you what your current instruments cannot. The method page sets out the research the rest of it rests on.

You did not stop climbing. Somebody took the rungs off.

Common questions

Why do so many Spanish learners get stuck at B1?

Because the kind of work changes at B1 and the measurement does not. Up to that point every lesson adds visible new material, so progress is easy for you and your tools to see. From B1 the remaining work is depth on material you have already met: register, precision, the structures you understand but never produce. Nothing in a right-or-wrong exercise can detect depth, so the instrument goes blind exactly where the work moves.

Is the intermediate plateau real or just an excuse?

It is a named, studied stage. Jack Richards wrote a Cambridge monograph on it in 2008, and the guided-learning-hour estimates published by exam boards show the underlying non-linearity: roughly 100 to 150 hours to reach A2, and 150 to 200 more to move from B1 to B2. Each level costs more than the one before, so the same weekly effort produces visibly less movement. That is arithmetic, not a character flaw.

Why does my app say I am doing well when I feel stuck?

Because it is measuring what it asked you, and it stopped asking anything that could fail. Correct answers to questions drawn from a fixed tree produce a clean progress bar regardless of whether your Spanish is moving. Streaks and XP were always measuring engagement. At the plateau, engagement and learning come apart, and only one of them is on the dashboard.

Will a tutor or moving to Spain fix the B1 plateau?

Both help and neither is sufficient on its own, for the same reason more app practice is not. A tutor who does not systematically probe what you avoid will mostly exercise the Spanish you already bring, and living in Spain supplies enormous input while supplying very little pressure to produce the structures you route around. What closes the gap is aim, not volume or location.

What is the difference between a motivation problem and a measurement problem?

A motivation problem shows up as less practice. A measurement problem shows up as the same practice producing the same result. If you are still studying, still using the app, still getting correct answers, and nothing is moving, effort is not the variable. The feedback loop is: your tools can no longer tell the difference between good enough and actually improving.

How do I find out what is actually holding my Spanish back?

You need three things that ordinary practice does not give you: coverage probing that asks for structures and social functions you did not choose, avoidance tracking that notices when a correct sentence was the safest one rather than the best one, and separate scores for recognising, writing and saying the same item. All three are diagnostic before they are instructional.

The gap test · Diagnose

Measure the gap, not your motivation.

This note describes the gap. The gap test finds where yours lives — in about ten minutes, no signup.

Find your gap free · no signup · twelve prompts