Here's a pattern almost every learner hits: the vocabulary keeps growing and the speaking doesn't. Eight thousand words in, and you're still producing the same handful of sentences.
That isn't a discipline problem. You're storing the wrong unit — and linguistics has measured this.
The number: 58.6%
In 2000, Erman and Warren hand-analysed 19 stretches of real language to work out how much of it was prefabricated — retrieved whole rather than built word by word from grammar.
Their answer: 58.6% of spoken English, and 52.3% of written.
Later corpus work using different methods gets different figures, but the direction holds: conversation is the most prefabricated register there is, with some analyses putting it close to 70%.
The phenomenon has a name — lexical bundles — introduced by Biber and colleagues in the 1999 *Longman Grammar of Spoken and Written English*. The recurring sequences they found in conversation were mostly fixed combinations about four words long.
Why the number matters to you
If more than half of native speech is retrieved whole, then "vocabulary + grammar = speaking" was never an accurate model of what's happening.
Compare the two production paths:
- A native speaker: intention → retrieve a whole chunk → say it. One step.
- You right now: intention → find it in your first language → find the English words → work out the structure → check the article and tense → say it. Five steps, each one a place to stall.
The issue isn't that you're slow. You're running four extra steps. And more vocabulary only widens the menu at step three — it doesn't remove any steps.
This is also the whole explanation for "I understand everything but can't say anything": you've stored components and the task requires assemblies.
What to store instead
Switch the unit from word to chunk. Four categories, in order of usefulness:
1. Sentence openers (highest frequency, highest priority) I don't know if… · Do you want to… · I was going to… · Are you going to… · What do you think about… · Is it okay if… These dominate conversational corpora precisely because they're beginnings — and once a sentence has started, the rest is far easier to finish.
2. Fillers and hedges I mean… · You know… · sort of · kind of · or something like that · Let me think Learners often treat these as sloppy padding to be eliminated. The opposite is true: they're the devices native speakers use to buy themselves thinking time. Without them you either freeze or force yourself to a speed you can't sustain.
3. Reactions Oh nice. · That's rough. · No way. · I know, right? · Fair enough. Short enough to require no thought, and they decide whether the other person feels listened to.
4. Functional lines (one or two per situation) Could you hold for a moment? · Sorry, I didn't catch that. · I'd better let you go.
How to get them in
Three rules, all plain, none skippable:
- Store whole, don't decompose. Learn
That reminds me — I need to…, notremind = to make someone remember. Whatever shape it goes in as, it comes out as. - Out loud, not silently. A chunk's value is retrieval speed, and retrieval speed is only built by saying it. Three times aloud beats ten times read.
- Put it in your own sentence. Say the chunk again inside something you'd genuinely say. A chunk attached to an example sentence drags that example along with it.
Finding the chunks that are actually yours
The list above is a starting point. A better method is to let your own gaps decide what you store:
Think of something you want to say → say it at your current level, awkward and all → look at how a native would put it → treat that as one chunk and say it out loud three times.
The order is the point. Struggle first, answer second. That gap is precisely where your missing chunk is, and what you learn at a gap sticks. Going straight to the model answer leaves almost nothing behind.
MirrorSay is built in that order: you say your own sentence, it shows the native version and reads it aloud, you shadow it, and it compares your rhythm word by word. Every chunk you accumulate comes from something you actually wanted to say.
Two other angles on the same problem: You Understand English but Can't Speak It and Does Shadowing Work?
References
- Erman, B., & Warren, B. (2000). The idiom principle and the open choice principle. *Text*, 20(1), 29–62.
- Biber, D., Johansson, S., Leech, G., Conrad, S., & Finegan, E. (1999). *Longman Grammar of Spoken and Written English.* (origin of the term *lexical bundles*)
*MirrorSay is a free AI speaking-practice app: say a sentence, see how a native would say it, then shadow it. No sign-up, and your voice stays on your phone.*
