You study for three months. You do practice tests. You handle the reading and manage the listening. Then you sit down, are asked to evaluate a set of proposals, and what comes out is the same three phrases rearranged, creo que es bueno, por otro lado, en mi opinión, while your head fills with opinions you cannot assemble fast enough to use.
The DELE B2 oral is the plateau with a score attached. Everything else on this site, the gap between understanding and producing, the avoidance of complex structures, the retreat to safe vocabulary, turns up in twenty minutes with two examiners watching.
The short answer
It tests whether you can sustain and interact, not whether you know things, and the official rating scale says so in the plainest language you will find anywhere. The Instituto Cervantes publishes the DELE B2 exam guide with the scales the examiners use, and the top coherence band describes a candidate who uses varied organisational structures, connectors and cohesion mechanisms without apparent effort and follows turn-taking naturally.
Not one word about conjugation. The exam is a test of discourse management, and most B1 speakers have never practised discourse management even once.
Which is why the usual preparation misfires. Candidates revise verb tables and vocabulary lists, arrive able to produce every form the examiner might want, and lose points on whether their argument survived ninety seconds of being disagreed with. That is the same mismatch behind filling a conjugation table and freezing in conversation: the thing practised and the thing assessed are only distantly related.
The structure, corrected
A great deal of what is written about this exam online describes an older or invented format, with four tasks and forty minutes appearing regularly. That is not the exam. It is worth stating from the source, because turning up having rehearsed a structure that does not exist is a poor use of three months.
The oral paper is three tasks in twenty minutes, with twenty minutes of preparation beforehand, and it is worth 25 of the 100 points.
Task two is the one that quietly demands the most grammar, because describing an imagined situation is speculation, and speculation in Spanish runs on parece que, podría ser que, puede que and es posible que. Several of those take the subjunctive, which is exactly the structure most candidates have spent two years routing around.
The twenty minutes of preparation are worth understanding properly too. They are not twenty minutes to write a script, because you will not have one in front of you, and a memorised opening is audible and collapses the moment the examiner interrupts it. They are twenty minutes to decide what you think, in what order, and which two or three structures you intend to use. Candidates who spend the time drafting sentences arrive with a fragile monologue. Candidates who spend it deciding a position arrive with something that survives being argued with. Robert DeKeyser’s skill acquisition account is the reason: what you can execute under pressure is what you proceduralised beforehand, and you cannot proceduralise a script in twenty minutes.
The mark scheme is the trap
The four papers are worth 25 points each, and they are not simply added.
For scoring purposes the guide groups them: Grupo 1 is reading comprehension plus written expression and interaction, Grupo 2 is listening comprehension plus oral expression and interaction. Each group is worth 50, and the guide states the pass condition plainly: a global grade of Apto requires a score at or above the minimum for each group of tests, and that minimum is 30 points in Grupo 1 and 30 in Grupo 2.
Note where speaking sits. It is bundled with listening, not with writing, so the two skills you perform live and under time pressure are weighed as one, and the two you do with time to think are weighed as the other. That grouping is not accidental, and it is why a learner who has spent two years reading is more exposed than they feel.
It also means the arithmetic people do in their heads is wrong. The instinct is to treat the exam as a hundred points where a strong paper compensates for a weak one, and to plan revision accordingly: shore up the strengths, hope the speaking is survivable. Under the real scheme, a brilliant reading paper cannot buy a single point of headroom in Grupo 2. The only revision that changes your Grupo 2 outcome is revision aimed at listening and speaking, which for most candidates is the revision they enjoy least and postpone longest.
What the scale actually rewards
This is the part worth reading closely, because the Cervantes guide publishes the analytic scale and it is far more concrete than most preparation advice.
Coherence is scored in bands from 0 to 3. The descriptors, paraphrased from the guide:
Band 3. Coherent, cohesive discourse with appropriate and varied organisational structures, connectors and other cohesion mechanisms. Converses with ease, using the right resources without apparent effort, and follows turn-taking naturally.
Band 2. Clear and coherent, with adequate but limited cohesion. May start to lose control of the discourse if the turn runs long.
Band 1. Linear sequences of related ideas as brief, simple statements, linked by common connectors, and the guide gives the examples: es que, por eso, además.
Band 0. Limited discourse made of word groups and simple connectors, with the examples y, pero, porque.
That is unusually actionable. The difference between the two lowest bands is described, in the official document, as a list of six words. If your Spanish links ideas with y, pero and porque, the scale has a place for you and it is the bottom one. We went through why these markers are the cheapest fluency available, and here is an examining board putting them in a rubric.
Britt Erman and Beatrice Warren’s study of the idiom principle found formulaic sequences making up 58.6% of spoken discourse, which is why a scale about cohesion is really a scale about how much prefabricated material you have. Fluency is scored separately, and Norman Segalowitz’s three kinds of fluency is the useful frame: the examiner can only observe perceived and utterance fluency, both of which are downstream of the cognitive fluency you have or have not built.
Why B1 speakers fail it
Not through error. Through running out.
Elaine Schachter’s 1974 paper found learners producing fewer attempts rather than more mistakes, which made an error count read them as strong. Under exam conditions this inverts: an examiner is not counting errors, they are watching whether the discourse holds up for the length of a turn, and a candidate who avoids everything difficult produces something short, flat and correct. Band 1.
The specific failures are predictable. The turn dies at ninety seconds because there was never a plan beyond the opening claim. Pushback produces agreement rather than a defended position, because conceding is easier to say. Speculation collapses into description, because parece que sea costs more than hay un hombre. And register stays uniform across three tasks that call for three different ones.
The strongest case for not preparing specifically
Worth stating, because exam-focused preparation has a bad reputation it partly earns.
A B2-level speaker passes a B2 exam. If your Spanish genuinely operates at that level, learning the format is an afternoon of admin, and time spent on exam technique is time not spent on Spanish. There is also a real risk of producing a candidate optimised for twenty minutes in a room and no better at ordering a coffee, which is the standard criticism of teaching to the test and it is often fair.
And the pass threshold is not high. Sixty out of a hundred with thirty per group is a competent performance, not an outstanding one, so a learner comfortably inside B2 has slack.
Why the exam is still the right forcing function
The counter-argument holds if you already know where you sit. Most people do not.
The value of this exam is not the certificate, it is that the mark scheme refuses the trade your daily life keeps offering you. Every day, your comprehension is allowed to cover for your production: people accommodate you, conversations move on, and the gap never gets priced. The grouped minimum prices it, once, in public. That is the same reason we treat B1 to B2 as a question about the composition of your hours rather than the number of them.
Preparing properly for this paper is not exam technique. It is sustained production, interaction under pushback and speculation with the structures you avoid, which is the work regardless of whether you ever book a date.
How to actually prepare
Practise sustaining, not answering. Set a timer for two minutes, take an unfamiliar prompt, prepare for twenty minutes without writing a script, then talk. The failure is running out, and it is invisible until you force yourself to continue past the point where you would normally stop.
Get pushback. A partner whose job is to disagree with you is worth more than one who is being kind. Conceding is the reflex under pressure, and holding a position while conceding a point is a separate skill.
Drill the connectors as production, not recognition. Laufer and Goldstein’s four degrees of word knowledge, validated across 435 learners, is why recognising por otra parte does not put it in your mouth. Roediger and Karpicke’s test-enhanced learning work found restudying losing badly to retrieval on any delayed test, so say them rather than read them.
Keep the balance. Paul Nation’s four strands put deliberate study at roughly a quarter alongside input, output and fluency work. A month of nothing but past papers produces a worse speaker than a month of speaking.
Read the guide. It is free, it is on the Cervantes site, and it contains the scales the two people in the room will be using on you. Most candidates never open it, which is a strange thing to be true of a document that lists the criteria.
What we do with this at Suelto
We are not an exam course, and this paper happens to test precisely what the rest of the system is built around.
The conversation step scores sustained discourse rather than single answers, and records avoidance rather than marking the safe substitute correct, since Band 1 is made entirely of correct sentences. Discourse markers are their own coverage area in Territory because the official scale treats them as one. And latency is stored alongside accuracy, because twenty minutes with two examiners is a cognitive fluency test wearing a grammar costume.
The gap test shows where your production sits in about ten minutes with no signup, and the method page sets out the rest. If you are still deciding which certificate you need at all, DELE, SIELE or CCSE covers that first.
Three tasks, twenty minutes, and a published rubric that names the six words separating its bottom two bands. For an exam with a reputation for being opaque, it is remarkably willing to tell you what it wants, and almost nobody preparing for it goes and reads the thing that says so.