MirrorSay ← Speaking Guides
Speaking Guides · Published 2026-09-28

Connected Speech: Why You Know Every Word but Still Can't Understand

You pause the show, read the subtitle, and every single word is one you learned years ago. You press play again, and it dissolves back into noise.

That gap is not vocabulary, and it is not speed. It is this: the pronunciation in the dictionary and the pronunciation in a real sentence are two different things. You learned words standing alone. Natives speak them squeezed together, with the edges worn off.

Here are the three things that happen to those edges — and a drill for getting them into your ear and mouth.

1. Linking: the end of one word joins the start of the next

The core rule is simple: consonant ending + vowel beginning = one blended chunk.

This is the default, not laziness. Saying "pick / it / up" as three clean words actually takes deliberate effort.

2. Reduction: the small words get crushed

English is stress-timed. Content words (nouns, verbs, adjectives) get the beat; function words (prepositions, articles, auxiliaries, pronouns) shrink almost to nothing.

The pair worth memorising on its own is can vs. can't. In a positive sentence, *can* reduces to a weak /kən/ and rushes past. *Can't* does the opposite: it takes the stress, keeps a full vowel, and its final t is often silent — just a stopped mouth shape. So when you can't hear whether there's a *not* in there, don't listen for the t. Listen for which word is louder and longer. "I can SWIM" vs. "I CAN'T swim."

3. Dropped and flapped sounds: some sounds just don't show up

The counterintuitive part: listening is fixed with your mouth

Your brain recognises speech by matching it against your own stored pronunciation. If your internal copy of *want to* is forever two crisp words, then *wanna* arrives and matches nothing — which is exactly that feeling of "I heard it clearly, I just don't know what it was."

Flip it around: once *wanna* is comfortable in your own mouth, you recognise it instantly. Practising pronunciation is really installing a decoder for your ears.

One common misunderstanding while we're here: *gonna* and *wanna* are not slang and not sloppy. Americans use them in meetings and on client calls. You don't have to use them — plenty of people sound perfectly natural without them — but you absolutely have to understand them. And they live in speech only: don't write them in an email.

The 3-minute drill

1. Pick one short sentence — 5 to 12 words, conversational, content you understand 100%. 2. Listen once and mark the changes: which two words fused, which word got crushed, where the stress landed. 3. Say it out loud. Out loud is non-negotiable — silent repetition builds no muscle memory. 4. Compare. Which word did you stretch, which did you swallow? The more specific the gap, the more useful the next rep. 5. Three reps of the same sentence. By the third it's usually noticeably smoother. Then move on.

Don't drill a whole TED talk, and don't drill material you can't follow. The rule for shadowing is *understand the content completely, train only the mouth* — anything you have to decode belongs in intensive listening, not here.

The full shadowing routine and its common failure points are in How to Practice English Speaking Alone.

Doing steps 3–5 in one move

MirrorSay covers exactly that part of the loop: you say a sentence, it shows how a native would say it and reads it aloud, you press and hold to shadow, and it compares your rhythm to the reference word by word — so which word you dragged and which you swallowed is on the screen instead of a guess.

The material is your own sentence, so "too difficult" never comes up, and one sentence takes about thirty seconds.


*MirrorSay is a free AI speaking-practice app: say a sentence, see how a native would say it, then shadow it. No sign-up, and your voice stays on your phone.*

Say your next sentence into the mirror

Free · No sign-up · Your voice never leaves your phone