The Plateau · Train

Memorisation lessons shouldn't be lectures

Retrieval and spacing are the only two high-utility study strategies in the evidence. Forty list-heavy lessons rebuilt around that, and why it matters.

STATIC SLIDE → 10 ACTIVE DRILL TYPES OLD: STATIC LESSON 2 examples 1 check FLASH RECALL 47 MATCH PAIRS 🔗 cinco ↔ 5 ✓ ✓ ✓ · 12s SPEED TYPE hacer, yo → hice ✓ BINARY CHOICE Es lista? ser estar GAP FILL 🕐 ___ la una. Es ERROR FIX Mi tio vive… Mi tío vive… ✓ CONTEXT MATCH "something shocking" ¡Qué fuerte! GRID TAP 3 7 12 SEQUENCE FILL 10, 11, ?, 13 doce ✓ TILDE CHOICE ___ eres mi amigo. tu 40 LESSONS CONVERTED teach 30s → drill immediately → repeat × chunks → quiz to pass

Open any Spanish learning app. Navigate to “Numbers 0-100.” You’ll get a slide that says something like: Numbers from 11 to 20 have unique forms you memorise. From 21 onward, you combine tens and units with “y.” Two example sentences. A multiple-choice check. Move on.

That’s the numbers lesson. All of it. You just “learned” a hundred number words by reading a paragraph about them.

This isn’t teaching. It’s a PDF with a progress bar.

The short answer

Explanation is the wrong tool for a list, and the evidence on this is unusually one-sided. John Dunlosky and colleagues reviewed ten common study techniques across more than 200 studies and rated exactly two as high utility: practice testing and distributed practice. Rereading and highlighting, which is structurally what a teach-then-quiz lesson gives you, came out low.

So a numbers lesson that explains the pattern, shows two examples and asks eight questions has used the weakest strategy available and skipped both strong ones. The fix is not more explanation or a prettier table. It is to break the teaching into thirty-second pieces and put a retrieval attempt after every one, then to space the whole thing across days.

The architecture problem

Traditional lesson systems are designed for grammar rules. Explain the rule, show examples, test comprehension. That works when the content has internal logic: ser vs estar, por vs para, subjunctive triggers. You understand the principle, then you apply it.

But forty of the topics in a Spanish curriculum aren’t principles. They’re lists:

  • Cardinal numbers (cero through novecientos)
  • Days, months, seasons (all lowercase, unlike English)
  • Irregular verb stems (tuv-, estuv-, hic-, pud-, pus-, quis-)
  • Pronoun paradigms (me/te/le/nos/os/les × direct/indirect/reflexive)
  • False friends (embarazada ≠ embarrassed, sensible ≠ sensible)
  • Tilde pairs (tú/tu, él/el, sí/si, same spelling, different meaning)
  • Administrative vocabulary (empadronamiento, NIE, certificado digital, nothing to derive, everything to memorise)

These have nothing to reason through. Quince means fifteen because it does. Dieciséis fuses “dieci” and “seis” into one word because that’s the convention. You can explain the pattern in thirty seconds. Understanding isn’t the bottleneck. Remembering is.

And for remembering, a slide with two example sentences does approximately nothing.

What the research says

Hermann Ebbinghaus showed in 1885 that you forget most of what you read within twenty-four hours. A century of replication has confirmed it. The forgetting curve is exponential, and it doesn’t care how many bullet points were on the slide.

What slows the decay? Active retrieval: trying to produce the answer before seeing it, not reading it again. Dunlosky et al. (2013), in a meta-analysis of over 200 studies, rated retrieval practice and spaced repetition as the only two “high utility” learning strategies. Re-reading, highlighting, and re-watching were rated “low utility.” The gap isn’t marginal. It’s a factor of three to five.

Elizabeth and Robert Bjork’s work on desirable difficulties makes the mechanism precise: memory gets stronger when recall is effortful but successful. Too easy, with the answer in front of you, wastes the opportunity. Too hard, with the item completely gone, produces frustration without encoding. Their list of difficulties worth creating deliberately is short and specific: vary the conditions of practice, interleave separate topics, space the sessions, and use tests as learning rather than as measurement.

Roediger and Karpicke’s test-enhanced learning work is the cleanest demonstration of that last one. On an immediate test, restudying wins, which is precisely why rereading feels productive. On a delayed test the result reverses and retrieval wins by a wide margin.

And the spacing half has been quantified. Nicholas Cepeda and colleagues synthesised 839 assessments of distributed practice across 317 experiments and found spacing reliably beating massed repetition, with the optimal gap growing as the retention interval grows.

Robert DeKeyser’s skill acquisition framework adds the speed dimension: knowing something is not the same as knowing it fast. Skills move from declarative through procedural to automatic, and what automatises is whatever you actually practised. Setenta y cinco produced after four seconds of visible effort is technically correct and conversationally useless.

Norman Segalowitz’s cognitive fluency, the efficiency of the machinery that reaches a form and assembles an utterance, is the variable being trained. An untimed exercise cannot see it, which is why a lesson can report full marks on content you will never produce in time. That is the same wall behind conjugation tables not making you conjugate.

None of this is obscure. It’s mainstream cognitive science. And yet the default lesson architecture for memorisation-heavy content in every app we’ve audited is: show a table, explain the pattern, quiz with eight questions. A format designed for grammar rules, applied to content that needs something completely different.

TEN STUDY TECHNIQUES, TWO THAT WORK DUNLOSKY ET AL, REVIEW OF OVER 200 STUDIES HIGH UTILITY Practice testing Distributed practice retrieve it, then wait, then retrieve again LOW UTILITY Rereading Highlighting Summarising what a teach-then-quiz lesson is The standard lesson format sits entirely in the right-hand box. Not because anyone chose badly. Because it was designed for rules, then reused for lists.

What we built instead

We replaced the explain-then-quiz format for all forty memorisation-heavy lessons with a teach-drill interleave, short teaching chunks alternating with immediate practice drills, so you never read more than thirty seconds of explanation before you’re actively retrieving what you just saw.

Each lesson is broken into three to eight chunks. Each chunk has a brief teach card (a reference grid you can tap to hear each item, a pattern-discovery question, or a short explanation) followed by one or more drill exercises drawn from ten different types:

Flash recall. You see “47”; a countdown bar gives you a few seconds scaled to the answer length; you try to produce cuarenta y siete before the timer runs out. The answer reveals, you hear the pronunciation, you move on. Twelve items per round, shuffled.

Pair matching. Two columns; tap one item on the left, its match on the right. Five pairs, timed. The competitive element is your own speed.

Speed type. A prompt appears (poder, yo, preterite) and you type pude. Instant feedback, latency invisible but shaping the pace. If you don’t know, the button says “I don’t know” and shows you the answer.

Binary choice. A sentence appears: María ___ lista para el examen. Two buttons: ser and estar. You tap. Explanation follows. Fifteen items in rapid succession, building the discrimination habit that ser/estar adjective pairs require.

Gap fill. A sentence with a blank, context around it, sometimes a clock emoji showing 🕐 1:00 as a hint; you type the missing word. Grammar-focused: Es or Son? La or las? Not “guess the number.”

Error correction. A sentence with a mistake pre-filled in the input. You edit it. Mi tio vive en el centroMi tío vive en el centro. The act of finding and fixing the error builds a different retrieval path from producing from scratch.

Context match. A situation described (“You’re at a bar and your friend says something shocking”) pick the right expression from four options. ¡Qué fuerte!, not desde luego.

Grid tap. A four-by-four grid of numbers. A prompt appears; you tap the correct cell. Rapid-fire, no typing, pure recognition-to-action mapping.

Sequence fill. A sequence with a gap: 10, 11, ___, 13, 14; type doce. Pattern reinforcement within ordered sets.

Tilde choice. A sentence appears. Two buttons: with tilde and without tilde. You decide whether the sentence needs or tu, él or el. The accent mark is the entire point, not a decoration, a meaning-changer.

Ten types, not because ten is a magic number, but because different memorisation targets need different retrieval paths. Matching a pronoun to its person is not the same skill as producing a number under time pressure, which is not the same as deciding whether listo needs ser or estar. One drill type cannot train all three.

The interleave that matters

The critical design choice isn’t the variety of drills. It’s where they sit relative to the teaching.

A traditional lesson teaches everything first, then tests at the end. By the time you reach the quiz, you’ve forgotten the first chunk. Ebbinghaus measured this decay at forty percent within twenty minutes.

The teach-drill interleave eliminates that delay. Chunk one: numbers 0-10, reference grid with audio for each. Immediately: flash recall drill on those same eleven items. Chunk two: numbers 11-15, the irregular forms. Immediately: flash recall plus a sequence fill. Chunk three: the 16-19 pattern. Immediately: pattern discovery plus flash recall mixing all items seen so far.

By the time you’ve seen thirty-three base forms (the only ones that require real memorisation, everything else is compositional), you’ve already retrieved each one at least twice in active drills. The quiz at the end is confirmation, not the learning event.

This is not a minor UX improvement. It is the structural change that retrieval practice research has been asking for since Ebbinghaus: practice the recall while the memory is fresh but not free, at the point of maximum encoding benefit.

WHERE THE RETRIEVAL SITS TEACH EVERYTHING, THEN TEST explanation, examples, more explanation quiz by the quiz, chunk one is already decaying TEACH, DRILL, TEACH, DRILL teach drill teach drill teach drill quiz every item retrieved at least twice before the quiz, which becomes confirmation

Silent mode and the no-audio path

Every drill works without sound. If you’re on a train, in a library, or simply prefer reading and writing to speaking and listening, every audio-dependent element has a text fallback. Flash recall shows the answer as text instead of playing it. Grid tap shows the prompt as written Spanish instead of speaking it. The Listen button is there for those who want it; its absence doesn’t break the exercise.

This isn’t a compromise. It’s a design principle. Memorisation practice should work wherever you are, however you learn.

DIFFERENT LISTS NEED DIFFERENT RETRIEVAL CONTENT WHAT THE DRILL HAS TO ASK FOR Numbers, dates, question words production against a clock Pronouns, possessives slot filling inside a sentence Irregular verb stems infinitive to form, at speed Fixed expressions, markers situation to expression Ser/estar meaning shifts binary discrimination, repeated Accent and tilde rules error correction, not recognition One drill format cannot train recognition, production and discrimination.

Forty lessons, not one

We didn’t build a drill prototype for numbers and call it done. Forty lessons, spanning six distinct memorisation types, were converted:

Discrete item lists (numbers, dates, question words, false friends): reference grids, flash recall, matching, speed type, grid tap, sequence fill.

Paradigm tables (pronouns, possessives, demonstratives): paradigm grids with audio, sequence fill for rows, sentence-slot drills for contextual use.

Irregular verb stems (present stem-changers, preterite irregulars, future stems, participles, imperatives): stem-flash mapping drills, conjugation speed type, matching infinitive to irregular form.

Fixed expressions (tener expressions, gustar family, verbal periphrases, connectors, discourse markers, reactions, repair strategies, idioms, collocations, slang): expression flash, context match, gap fill with situational prompts, register-choice drills.

Meaning-shift pairs (ser/estar adjective shifts, preterite/imperfect meaning changes, reflexive meaning changes): rapid binary choice, contrast flash showing both meanings, translation pair drills.

Rule-plus-exception sets (accent rules, tilde diacrítica): speed type where you add the missing accent, binary tilde choice, error correction of accent mistakes.

Each category uses a different subset of drills because each demands different cognitive work. Matching a false friend to its real meaning is a recognition task. Producing the correct accent under time pressure is a production task. Choosing ser or estar for listo is a discrimination task. One format cannot serve all three.

The strongest case for the lecture format

Worth stating, because the explain-then-quiz format is not the product of laziness.

Explanation genuinely is the right tool for a rule. Ser against estar, por against para, mood selection: these have internal logic, and a learner who grasps the principle can generate cases they were never shown. DeKeyser’s own framework insists the declarative stage is real rather than a detour, so teaching the rule first is correct sequencing, not a shortcut.

The format also scales, which matters more than it sounds. A teach card and eight multiple-choice questions can be authored in an afternoon and graded for nothing. Ten drill types with timers, audio and per-item latency cost far more to build and to run, and a product that spends that everywhere ships less curriculum. Choosing breadth over depth is a defensible trade when a learner’s main risk is running out of material.

Why it still fails on lists

The trade stops being defensible when the content has no logic to grasp.

There is no principle underneath the fact that eleven is once. Nothing generalises from dieciséis to the question words or the false friends. For a list, understanding is not the bottleneck and never was, so an architecture that spends its budget on explanation is spending it on the one thing that was never the difficulty. Paul Nation’s four strands put deliberate language-focused study at roughly a quarter of a balanced diet, and this is that quarter being spent on the wrong operation.

It also produces a specific, familiar failure. The learner passes, feels they have covered numbers, and discovers six weeks later at a till that they cannot produce setenta y cinco. Not a knowledge gap: a rung gap, in the sense Laufer and Goldstein validated across 435 learners, where recognition and active recall are separate states of the same item and only the easy one was ever tested. It is the passive and active split built into the lesson format itself.

The quiz that earns the pass

Every drill lesson ends with the same ten-question quiz used across all grammar lessons. Same grading, same eighty-percent pass threshold, same hybrid scoring (deterministic matching plus AI grading for ambiguous free-text answers). Same FSRS integration; wrong answers create review objects that enter the spaced repetition queue and shape tomorrow’s conversation scenarios.

The drill lessons aren’t a separate track. They’re embedded in the curriculum, same prerequisite chains, same progress tracking, same daily loop. Pass a drill lesson in the morning, and the conversation module that afternoon will prefer scenarios that use those structures. The connection is automatic.

What this changes

If you’re stuck at B1 and you can’t remember whether quinientos agrees in gender, or whether supe means “I knew” or “I found out,” or whether it’s a la una or a las una, or you’re staring at CCSE exam prep and the constitutional vocabulary feels like a wall, you haven’t failed at memorisation. You’ve been given explanation where you needed retrieval, and a table where you needed ten different ways to practise until it stuck.

None of this is a claim that memorisation is the interesting part of learning Spanish. It is the least interesting part, which is exactly why it deserves an architecture that gets it over with efficiently instead of one that spreads it thinly across months of rereading. The interesting work is what you can do with the language, and lists are the toll on the way there.

Reading the table again won’t help. What works is being asked for the answer before you see it, under just enough time pressure to make the recall effortful, in enough variety that your brain can’t game the pattern, and then being asked again two days later, just as you’re about to forget.

That’s what these forty lessons do now. If you want to feel the difference, open Suelto and work through a drill lesson. The first flash card will show you a number. Try to say it before the bar runs out. That small act of effortful recall (not reading, not highlighting, not watching) is where the memory forms.

Common questions

Why can't I remember Spanish numbers even after studying them?

Because reading a table is the study strategy the evidence rates lowest. Dunlosky and colleagues reviewed over 200 studies and rated only retrieval practice and distributed practice as high utility, with rereading and highlighting rated low. A numbers lesson that explains the pattern and quizzes you once has used the weak method and skipped the strong one.

What is retrieval practice and why does it beat rereading?

Trying to produce an answer from memory before you see it. Roediger and Karpicke found rereading winning on an immediate test and losing badly on a delayed one, which is exactly why rereading feels productive and does not last. The act of failing to recall, then recalling, is the encoding event rather than a check on whether encoding happened.

How should Spanish vocabulary lists actually be taught?

In short teaching chunks alternating with immediate drilling, so nothing is explained for more than about thirty seconds before you have to produce it. That keeps the recall attempt inside the window where the memory is fresh but no longer free, which is where Bjork and Bjork locate the encoding benefit of a desirable difficulty.

Does spaced repetition really work for language learning?

It is one of the two best-evidenced strategies there is. Cepeda and colleagues synthesised 839 assessments of distributed practice across 317 experiments and found spacing reliably beating massed repetition, with the optimal gap widening as the retention interval grows. Three short sessions across three days beat one long session, and the long session will feel far more productive.

Why do timed drills matter more than getting the answer right?

Because correct and automatic are different states and only one is usable in conversation. Robert DeKeyser's skill acquisition account has knowledge move from declarative through procedural to automatic, and Segalowitz's cognitive fluency is the speed at which the machinery actually reaches a form. Setenta y cinco produced after four seconds is right and conversationally too late.

Are grammar lessons and memorisation lessons different?

They need different architectures. A grammar rule has internal logic, so explaining it then applying it works. A list has no logic to grasp: numbers, question words, false friends, irregular stems. There is nothing to understand, only something to retrieve, and explanation is the wrong tool for a retrieval problem.

The method · Train

The reflex is trained, not learned.

Notes like this one name the mechanic. The nightly drill is how it becomes a reflex.

See the method how Suelto actually closes the loop