Learning Tipsenglishpronunciationlisteningconnected speechintermediate

Connected Speech in English: Why Fast Speech Sounds Like One Word

By the Conversa team · · 15 min read

Paper cards pegged evenly along a washing line, with the last few slid together into one overlapping clump.

"J'eat jet?"

Four words are written there: Did you eat yet? Out of an American mouth they arrive as two syllables with no boundaries you could point at. The phrase comes from a list of reduced forms that nine native speakers of North American English, from California, New York City, Boston and the Midwest, all agreed they personally produce.

You know all four words. You have known them for years. You caught none of them.

That gap has a name: connected speech in English. Keith Johnson went through 88,000 words of recorded interviews and found that more than 60% of them deviate from their dictionary form on at least one sound. One word in twenty loses an entire syllable.

Below: why fluent English feels fast, the three things that happen to a word plus the fourth that happens at the seam, the pair where getting it wrong reverses the sentence, how to tell which mechanism ate the word you missed, and the drill. Transcriptions are General American throughout.

Why fluent English feels too fast

John Levis and Kate Challis name the mechanism in the Iowa State open textbook on teaching pronunciation: learners "often do not recognize known words when they occur in speech", and so they "think that fluent speech is too fast." Speed is what the failure feels like from inside.

Sven Mattys and colleagues, reviewing speech in adverse conditions, list syllable deletion, elision and reduction as the things that degrade intelligibility, "and possibly faster speech rate (though this is a disputed fact)". A review paper put that one in print as disputed and gave the blame to the compression.

The most useful finding is about timing. Marco van de Ven and colleagues found that reduced words are not lost, they arrive late: reduced primes failed to prime low-frequency targets until the experimenters stretched the interval between them, at which point the priming came back. You are not failing to understand. You are understanding a beat behind.

Which is what it feels like at a coffee counter. You catch the front of the sentence, spend a beat working out that the noise in the middle was what do you want, and by the time it resolves the person has asked the follow-up and is waiting on you with their hand on the cup.

What connected speech is, and what it isn't

Connected speech is what English does to words when they touch: sounds move across the boundary between two words, unstressed vowels thin out, and consonants at the ends of syllables go missing or change into something else.

It is not a shortcut and not a regional habit. Connected speech processes "are not improper English, slang, or sloppy speech," Levis and Challis write. "In fact, they occur in even the most formal and careful registers of spoken English." A newsreader links and reduces. So does a judge.

Cauldwell's ladder, quoted in the same chapter, names the gap you fell into: very careful speech is "greenhouse listening," because every word is planted in its own pot, classroom audio is "garden listening," and real speech between two people busy communicating is "jungle listening."

Three things happen to a word, and a fourth happens at the seam

Danxin Liang played 25 native-spoken sentences to 50 Chinese English majors and scored a gap-fill listening test item by item. The four worst items were but you (90%), take it away (88%), thinking of (86%) and office is (84%).

Look at what those are. Weak forms and linking, with cross-word fusion joining them further down the list, and not one of the sixteen elision items anywhere on it. That is fifty listeners in a single study, so take it as a hint about where to look first and not as a league table.

Three of those get a repair move below. Deletion does not, and inventing one would be dishonest: Jennifer Jenkins' point, quoted in the same chapter, is that learners "cannot easily reconstruct words when sounds and syllables are missing." Once the sound is gone there is only context, which is why deletion is the most expensive of the four, why the drill at the end of this post exists, and why how to ask someone to repeat in English matters more for this mechanism than for the other three.

On the pagerun into
In the airru‿nintothe n changed address
On the pagefirst aid
In the airfir‿staidthe whole cluster moved
On the pageI went in to work
In the airI went ən to workin, on and and all land here

In the first two rows nothing was dropped and the consonant simply moved house, while in the third the vowel collapsed.

Rule 1: linking sounds move the consonant onto the next word

Run into comes out of a native mouth as ru‿ninto. The /n/ has changed address, and into now begins with a consonant.

This is the regular pattern. English wants consonant + vowel, so Levis and Challis describe linking as the thing that "ensures that words beginning with vowels will actually sound like they begin with a consonant." Their table runs some ofso‿muv, check inche‿kin, first aidfir‿staid, post officepo‿stoffice. When the first word ends in a cluster, the whole cluster travels.

Turn it off hides a second change. Unstressed it in General American is /ət/, not /ɪt/, so the middle of that phrase is a schwa. If you are listening for the vowel in sit, you will listen straight past it. Phrasal verbs lose the most to this, because the particle is usually the unstressed piece: it is why get over it arrives as one blurred word.

Vowels link too, with a glide in the seam: see‿y‿it, go‿w‿out.

The trap in this rule is that linking can build a word nobody said. Take the chapter's own example: it's kind of nice arrives as it's kayn‿dof nice, which parses equally well as "it's kine dove nice." Native listeners are unlikely to notice, because kind of is frequent and kine has been archaic for centuries. It is the same machinery behind ice cream and I scream.

Rule 2: the small words collapse into one sound

"If you say 'I went in to work on Thursday and Friday,'" write Levis and Challis, "the words in, on, and will all be reduced so that they are pronounced in almost the same way… in connected speech, the differences disappear."

Their weak-form table: and → [ən], in → [ən], on → [ən], of → [ə], to → [tə], can → [kn̩], will → [əl]. Wiktionary's unstressed entries extend it with for → /fɚ/, you → /jə/, have → /(h)əv/, him → /əm/ and was → /wəz/.

Three different spellings, one sound in the air. That is the reason you cannot find the boundary, and it is why this bucket is worth your attention before any of the others: Liang's four worst items were all weak forms and linking, and none of them were rare words.

The /h/ is the cruellest of them, because it erases whole pronouns. Levis and Challis note that /h/ drops most often at the start of unstressed function words, which is to say he, him, his and her. It goes from have, has and had as well, leaving [əv], [əz] and [əd]. So should have surfaces as shou‿duv, which is probably why a generation of native writers types "should of."

Rule 3: t, d, s and z fuse with "you" into a new sound

What did you think? comes out as wha‿dja‿think, and the /dʒ/ in the middle belongs to neither word on its own.

Four lines cover it, and once you have the rule you do not need the list:

You hearIt was
/tʃ//t/ + /j/
/dʒ//d/ + /j/
/ʃ//s/ + /j/
/ʒ//z/ + /j/

Wiktionary transcribes the results as though they were words: gotcha /ˈɡɑtʃə/, didja /ˈdɪdʒə/, wouldja /ˈwʊdʒə/, dontcha /ˈdoʊntʃə/, whatcha /ˈwʌtʃə/. Wikipedia, in its history of English consonant clusters, confirms the fusion crosses word boundaries and offers gotcha and whatcha as its examples.

Run miss you through the /s/ line and out comes the "mishu" that learners hear as vocabulary they have never met.

Sorry, what does 'mishu' mean? I can't find it in the dictionary.
Miss you.

/s/ meeting /j/ makes /ʃ/, so miss you arrives as a word you have never met.

Test any unfamiliar /tʃ/ or /dʒ/ on a boundary this way before you go hunting for new vocabulary. Brazilian Portuguese speakers meet the same fusion inside single words, which is why "team" comes out sounding like "cheam".

Is "gonna" a real word?

Yes. Gonna, wanna, gotta and dunno have headwords in the major learner dictionaries, labelled "not standard" and glossed by Cambridge as "spelled the way it is often spoken."

Now compare hafta, the identical process applied to have to. It gets into none of them: Oxford Learner's returns a 404 and Dictionary.com sends it to a spellcheck suggestion, while gotta sits there with a full entry. Both reductions are equally real in equally many mouths. Whether one has been granted a spelling tells you nothing about whether you will hear it, so do not use the dictionary as your list of what to expect.

None of this is new, either. In 1687 a grammarian called Christopher Cooper complained in The English Teacher, under the chapter heading "Of Barbarous Speaking," that people were saying gim me for give me. The complaint is older than the United States.

The pair where getting it wrong reverses the sentence: can vs can't

"I can meet Tuesday" and "I can't meet Tuesday" differ by one vowel, and the /t/ you are listening for often isn't released and sometimes isn't there at all. It is the same vanishing /t/ that makes past-tense -ed so hard to hear, except there it costs you a tense instead of a negative. Miss a linked consonant and you lose a word. Miss this one and you are told the opposite of what was said.

Rachel of Rachel's English is blunt about the /t/: "most Americans when they're speaking everyday speech don't release final T's."

The cue is the vowel and the stress. Can has a weak form, /kən/. Can't has no weak form at all, and Wiktionary's usage note spells out the consequence: "can't is never reduced, so its /æ/ vowel clears up the ambiguity." Baruch College's Tools for Clear Speech frames the same contrast as stress: "Can is unstressed at the beginning and in the middle of a sentence, but can't is stressed in these positions."

canI kən meet tomorrow afternoon.schwa, hurried past
can'tI kænt meet tomorrow afternoon.full vowel, and it lands hard

Listen for the vowel and the stress, because the t is often not there at all.

At the end of a sentence or thought group both words are stressed and the cue is gone. Baruch's own advice at that point is to "try using the context of the situation to guess if the speaker is using can or can't." There is a position where the acoustic evidence genuinely runs out and everyone in the room is inferring, natives included.

One more wrinkle: British can't is /kɑːnt/, a long back vowel General American never uses in this word. If you trained your ear on British audio, that particular pair does not transfer.

Try it in Conversa

Practice with AI characters who adapt to your level and give real-time feedback.

Try Conversa Free

When you miss something, which one was it?

Laura Dilley and Mark Pitt slowed the speech around a function word and native listeners stopped hearing the word entirely: "leisure or time" was heard as "leisure time." Speed the rate up around a gap where no word existed and the same listeners heard one that was never spoken. Native listeners, deleting and hallucinating on cue.

So the question after a miss is not whether your ear is bad. It is which mechanism got you.

What it felt likeProbablyWhat to put back
The words ran together and I couldn't find the boundarylinkingmove the consonant onto the previous word
There was a sound I don't know a word forfusion with yourestore the you
A word I know was simply missingweak form, or deletionpredict it from the grammar
I heard a word that wasn't therelinking built a new onecheck whether two words made a third
I heard a completely different wordyour first language got there firstthe competing word from your L1

That last row has a measurement behind it. Andrea Weber and Anne Cutler tracked Dutch listeners' eyes during English sentences and found that at the word desk they looked at a picture of a lid, which is deksel in Dutch. Their conclusion: "lexical competition is greater for non-native than for native listeners." Hearing the wrong word is not carelessness. There are more candidates loaded in your head than in a native's.

Where your first language points you wrong

Japanese listeners hear vowels that were never spoken. Emmanuel Dupoux and colleagues found they perceive an illusory /u/ inside acoustically vowel-less consonant clusters, so the word you said and the word they heard are different lengths before anyone has made a mistake. Spanish speakers have the opposite gap: "the Spanish vowel space has an empty central area with no central vowel categories", so schwa, one of the commonest vowels in reduced English, lands somewhere they have no category for.

Holger Mitterer and Annelie Tuinman found the shape of it in German learners of Dutch: where the two languages' reduction patterns line up, learners performed close to native; where they diverge, they fell back on grammatical cues. So this is hard in the specific places your first language did not prepare you for, which is why the fix differs for a Spanish speaker adding an e to s-words and a French speaker putting English stress in the wrong place.

The drill: write down what you heard, then diff it

Brown and Hilferty ran four weeks of reduced-forms dictation with their own students in the early eighties, and comprehension of reduced-form sentences went from 35% to 61% between the pretest and the posttest. A pretest-posttest comparison in one program, forty years ago. It is still the closest thing to a number anyone has on this drill.

The kind of practice matters more than the hours. Elizabeth Kissling taught 116 novice learners of Spanish about stress and rhythm and then had them practise hearing it rather than saying it, and the target language came out more intelligible to them than to controls. Different language, same split. As Celce-Murcia and colleagues put it in the same chapter, the listening goal is "fast, messy, authentic speech," which is "much more varied and unpredictable than what they need to produce in order to be intelligible." Shadowing trains your mouth. Your ear needs its own hour.

So run it as a diff:

  1. Take 20 to 30 seconds of unscripted audio. An interview, not a news read.
  2. Write down what you actually hear, nonsense syllables and all. Leave the gaps as gaps.
  3. Compare it against the transcript.
  4. Run every miss through the table above and name the mechanism.
  5. Listen one more time, now that you know what was there.

Step four is the one people skip, and I think it is the one that does the work. Naming the mechanism turns a single miss into a pattern you will catch next week.

Natasha Warner at the University of Arizona has built the ready-made version, and it is jungle listening rather than garden: real spontaneous American speech, narrowly transcribed, with each clip playable both out of context and in context. In one clip, what are you doing is transcribed [iɪɾ̞ɪɨ̰ɪ̰ɛ̰]. Out of context it is noise. In context it is a sentence, and hearing it flip is the entire skill in one click.

This drill works if you already know the words. When the transcript turns up vocabulary you have never met, that is a vocabulary problem and no amount of dictation will touch it. If you can read the transcript comfortably but could not have produced the sentence yourself, that is a third problem again.

What the finish line looks like

Warner's own note on those recordings sets the target honestly: in context, she writes, a native listener "doesn't even notice that there's anything unusual about the reduced part." Not noticing is the finish line, and it is a good deal further out than four weeks.

J'eat jet is closer than that, though. Its first /dʒ/ is did you, the same fusion that gives you didja, and its second sits on exactly the kind of seam Rule 3 describes, where the end of one word meets the y at the start of the next. You will not rebuild that from first principles the first time you hear it. You will recognise it the second time, which is the whole trick, ready for the next time somebody behind a counter asks you.

Share this article

Related Posts

Ready to start speaking?

Join thousands learning with AI-powered conversations

Get Started Free