Advanced connected speech in American English is the system of sound modifications that native speakers apply automatically across word boundaries to produce fluid, natural-sounding speech. If you have ever heard a native speaker say “didja eat yet?” instead of “did you eat yet?” or “gonna” instead of “going to,” you have heard connected speech at work. These are not lazy habits or regional quirks. They are predictable phonological processes that govern how sounds link, blend, disappear, or get inserted in fluent American English.
The four core processes you need to master are:
- Linking (catenation): Final sounds connect to initial sounds of the next word, creating unbroken syllable chains.
- Assimilation: A sound changes its phonetic identity to match or partially match a neighboring sound.
- Intrusion: A glide sound (/y/, /w/, or /r/) is inserted between two words to smooth a vowel-to-vowel transition.
- Elision: A sound, usually a consonant, is deleted entirely to reduce articulatory effort.
The gap between a word’s citation form (how it sounds in isolation, as in a dictionary) and its normal spoken form (how it sounds in real conversation) is where most advanced learners lose comprehension. Research confirms that advanced learners often score similarly to lower-proficiency peers when listening to fluent connected speech, despite stronger vocabulary and grammar. Closing that gap requires explicit, structured training in each of these processes.
Explore Curriculum | Book Your Sample Class | Read Reviews

How linking works in American English connected speech

Linking, also called catenation, is the process of connecting the final sound of one word to the initial sound of the next. Native speakers do this constantly to maintain speech flow without pausing between every word. The result is that word boundaries become nearly invisible to the ear.
There are three main types of linking in American English:
Consonant-to-vowel (C-V) linking is the most frequent type. When a word ends in a consonant and the next begins with a vowel, the consonant moves forward and attaches to the following vowel. “Pick it up” becomes /pɪ.kɪ.tʌp/, sounding like one continuous word. “Turn it off” flows as /tɜr.nɪ.tɔf/. This is why “an apple” sounds like “a napple” to untrained ears.
Consonant-to-consonant (C-C) linking occurs when two identical or similar consonants meet at a word boundary. Instead of pronouncing both, native speakers produce one slightly lengthened consonant. “Big game” becomes /bɪg:eɪm/, and “hot tea” becomes /hɒt:iː/. This gemination effect is subtle but consistent.

Vowel-to-vowel (V-V) linking requires a glide to bridge two adjacent vowels. This is technically intrusion (covered in Section 4), but it originates from the linking need. “Go away” naturally picks up a /w/ glide: /goʊ.wə.weɪ/.
Key linking rules and patterns to practice:
- Final /t/ or /d/ before a vowel-initial word links and often flaps: “get it” → /ɡɛ.ɾɪt/
- Final /r/ in American English links strongly to following vowels: “for a while” → /fɔ.rə.waɪl/
- Linking does not depend on spelling; it depends on the actual sounds at the boundary
- American English links more aggressively than British English, particularly with /r/ and flapped /t/
Pro Tip: Practice C-V linking by rewriting phrases as single phonetic words. “I need it” becomes /aɪ.niː.dɪt/. Drill it until the boundary disappears.
For a structured approach to American English pronunciation practice, working through linking patterns with phonetic transcription is one of the fastest ways to close the gap between your reading fluency and your speaking fluency.
How assimilation changes sounds in American English
Assimilation is the process by which a sound changes to become more like a neighboring sound. It reduces articulatory effort by letting the mouth anticipate or carry over the position of an adjacent sound. Native speakers do this without thinking. For learners, recognizing and producing assimilation is one of the most challenging aspects of connected speech in American English.
There are three directions assimilation can take:
Progressive assimilation means the preceding sound influences the following one. The plural suffix “-s” is a classic example: it becomes /s/ after voiceless consonants (“cats” /kæts/) and /z/ after voiced ones (“dogs” /dɒgz/).
Regressive assimilation means the following sound influences the preceding one. This is far more common in connected speech. “In case” becomes /ɪŋ.keɪs/ because the /n/ anticipates the velar /k/ and shifts to the velar nasal /ŋ/.
Coalescent (reciprocal) assimilation is where two sounds merge into a third. This is the most dramatic form and the one that most surprises advanced learners.
The two most important coalescent assimilations in American English are:
- /t/ + /j/ → /tʃ/: “Did you eat?” becomes “Didja eat?” (/dɪ.dʒə.iːt/). “Don’t you?” becomes “Doncha?” (/doʊn.tʃə/).
- /d/ + /j/ → /dʒ/: “Would you?” becomes “Wouldja?” (/wʊ.dʒə/). “Did your” becomes /dɪ.dʒər/.
Place assimilation with /n/ is equally common. Before a bilabial consonant like /p/ or /b/, the alveolar /n/ shifts to the bilabial /m/: “ten people” is often pronounced /tɛm.piː.pl/. Before a velar like /k/ or /g/, it shifts to /ŋ/: “in case” → /ɪŋ.keɪs/, “one goal” → /wʌŋ.goʊl/.
Assimilation types and their phonetic triggers:
- Voicing assimilation: voiced/voiceless environment determines suffix pronunciation
- Place assimilation: /n/ shifts to match the place of articulation of the following consonant
- Coalescent assimilation: /t/ or /d/ before /j/ produces /tʃ/ or /dʒ/ respectively
- Regressive assimilation: the most frequent direction in American English connected speech
What are intrusion glides and when do they appear?
Intrusion is the insertion of a glide sound between two words where no such sound exists in the citation forms. It happens automatically when two vowels meet at a word boundary, because the vocal tract needs a transitional movement to avoid an abrupt vowel-to-vowel shift. This insertion is a natural and frequent feature of American English connected speech.
Three glides appear in intrusion:
The /j/ (y) glide appears after front vowels like /iː/, /ɪ/, /eɪ/, or /aɪ/ when the next word begins with a vowel. “See it” becomes /siː.jɪt/. “They asked” becomes /ðeɪ.jæskt/. “I am” becomes /aɪ.jæm/. The tongue naturally rises toward the palate as it transitions from the high front vowel to the next sound.
The /w/ glide appears after back rounded vowels like /uː/, /oʊ/, or /ɔː/. “Go on” becomes /goʊ.wɒn/. “Do it” becomes /duː.wɪt/. “How are” becomes /haʊ.wɑːr/. The lips, already rounded for the back vowel, push slightly forward and then release into the following vowel.
The /r/ glide is specific to American English and is called intrusive /r/. It appears after central vowels like /ə/ or /ɑː/ when the next word begins with a vowel. “The idea of” becomes /ðə.aɪ.diː.ər.əv/. “Law and order” becomes /lɔː.rən.ɔːr.dər/. This feature is more common in some American dialects than others, but it appears across many regional varieties.
Common intrusion patterns and exceptions:
- /j/ intrudes after /iː/, /ɪ/, /eɪ/, /aɪ/, /ɔɪ/ before any vowel-initial word
- /w/ intrudes after /uː/, /ʊ/, /oʊ/, /aʊ/ before any vowel-initial word
- /r/ intrudes after /ə/, /ɑː/, /ɔː/ in American English (not in non-rhotic British English)
- Intrusion does not occur when a pause separates the two words
- Intrusion is unconscious for native speakers; forcing it artificially sounds unnatural at first
How elision removes sounds in fluent American English speech
Elision is the deletion of a sound, most often a consonant, that would be present in the citation form of a word. Deletions occur more commonly at syllable endings and help native speakers speak more quickly and efficiently without losing meaning. For learners, elision is often the hardest process to accept because it feels like something is being “left out” that should be there.
The most common elision environments in American English:
Consonant cluster reduction is the deletion of a stop consonant when it is sandwiched between two other consonants. “Last night” becomes /læs.naɪt/ (the /t/ disappears). “Next day” becomes /nɛks.deɪ/. “Acts quickly” loses the /t/: /æks.kwɪk.li/. The rule is consistent: when a stop sits between two consonants, it is a strong candidate for deletion.
Word-final /t/ and /d/ deletion is extremely common before consonant-initial words. “Good morning” often sounds like /gʊ.mɔːr.nɪŋ/. “Kept talking” becomes /kɛp.tɔː.kɪŋ/. “Old friend” becomes /oʊl.frɛnd/.
Schwa deletion occurs in unstressed syllables. “Family” is often pronounced /fæm.li/ rather than /fæ.mɪ.li/. “Every” becomes /ɛv.ri/. “Comfortable” is frequently /kʌmf.tər.bl/.
Function word reduction removes sounds from grammatical words under low stress. “And” becomes /ən/ or even /n/. “Of” becomes /ə/. “Him” becomes /ɪm/. “Her” becomes /ər/.
Typical elision scenarios:
- Stop deletion in three-consonant clusters: /t/ or /d/ between two consonants
- Final /t/ and /d/ before consonant-initial words in rapid speech
- Schwa deletion in unstressed syllables of polysyllabic words
- Function word reduction under low prosodic stress
- /h/ deletion in unstressed pronouns: “tell him” → /tɛ.lɪm/
Elision directly affects word recognition. When you hear /læs.naɪt/, your brain must reconstruct “last night” from context and phonological knowledge. Without explicit training in elision patterns, even advanced learners misparse these sequences and lose comprehension.
Why mastering connected speech matters for advanced learners
Lack of connected speech knowledge is a major reason why advanced learners struggle with spoken English comprehension despite strong vocabulary and grammar. You can know every word in a sentence and still fail to understand it when the sounds have been linked, reduced, and modified beyond recognition. This is not a vocabulary problem. It is a phonological one.
Connected speech shapes three critical dimensions of communication:
Naturalness and rhythm. American English has a stress-timed rhythm, meaning stressed syllables occur at roughly regular intervals regardless of how many unstressed syllables fall between them. Connected speech processes, especially elision and reduction, are what make this rhythm possible. Without them, speech sounds choppy and foreign, even when every word is grammatically correct.
Listener perception. Native speakers process connected speech automatically. When a non-native speaker produces citation-form speech in conversation, native listeners often perceive it as accented or effortful, even if every sound is technically correct. Producing natural connected speech signals fluency at a level that grammar and vocabulary alone cannot.
Listening comprehension. Production and perception are deeply linked. Research shows that many advanced learners mistakenly perceive connected speech difficulties as listening problems, when the root cause is actually a production gap. When you cannot produce a connected form, you often cannot recognize it either.
Strategies and benefits of mastering connected speech:
- Train each process explicitly: linking, assimilation, intrusion, and elision require separate focused practice
- Use phonetic transcription to see the spoken form, not just the spelling
- Listen to authentic American English at natural speed, not slowed-down recordings
- Shadow native speakers phrase by phrase, focusing on sound boundaries, not individual words
- Record yourself and compare your connected forms to a native model
- Study American English intonation alongside connected speech, since rhythm and reduction work together
How visual-physical training accelerates connected speech mastery
Traditional pronunciation instruction has a fundamental limitation: it asks you to copy sounds you cannot fully see or feel. Repeating after a recording or watching a teacher’s mouth gives you only partial information. You cannot see tongue placement, the degree of lip rounding, or the precise moment a sound transitions to the next. This is why so many advanced learners plateau despite years of listening practice.
Phonetic-physical training, including the use of Interactive 2D Sound Video Simulators, addresses this directly. The 2D Sound Motion Technology developed by Prof. Alex, Ph.D., makes the internal movements of the speech organs visible. You can watch exactly how the tongue, lips, jaw, and velum move through a connected speech sequence, frame by frame if needed. Sound becomes visible, and doubt becomes clarity.
Here is how the training sequence works in Prof. Alex’s 1-on-1 program:
- Speech-organ awareness: You learn the anatomy of the vocal tract and how each organ contributes to sound production before attempting any connected speech pattern.
- Sound retraining: Each American consonant and vowel is retrained with correct organ placement, using the 2D simulator to confirm the movement.
- Connected speech application: Once individual sounds are stable, you practice linking, assimilation, intrusion, and elision in phrases and sentences.
- Phonetic exercises: Structured drills build muscle memory for connected transitions until they become automatic.
- Paragraph-level practice: You apply connected speech in extended speech until the patterns are habitual, not deliberate.
Pro Tip: The goal is to stop consciously managing every sound. Once the speech-organ movements are retrained, you produce connected forms automatically, the same way you do in your native language.
Watch how this method works in practice:
“Thiago came to Myaccentway as a Portuguese speaker who understood English well but felt his speech was holding him back professionally. After structured phonetic retraining with the 2D Sound Motion Technology, his connected speech patterns became measurably more natural, and his confidence in professional conversations improved significantly.”
Watch
to see what structured connected speech training produces.
Individual factors like motivation, native language transfer, and prior phonetic exposure all influence how quickly connected speech patterns take hold. Research confirms that recognizing these individual differences and tailoring instruction accordingly produces better outcomes than one-size-fits-all methods. That is exactly why Prof. Alex’s program is built around 1-on-1 coaching rather than group classes or self-study apps.
How connected speech affects listening comprehension
Connected speech is the primary reason fluent English sounds so different from what you studied in a classroom. When sounds link, reduce, and blend, word boundaries disappear. A phrase like “I’m gonna get outta here” contains five connected speech processes in seven words. For a learner who has only practiced citation forms, parsing that phrase in real time is genuinely difficult.
The comprehension challenge is structural, not a matter of vocabulary. Research shows that teaching connected speech processes significantly improves listening comprehension for advanced non-native learners in professional contexts. The key is moving from passive exposure to active, explicit training.
Practical strategies to improve listening comprehension through connected speech awareness:
- Dictation with connected speech focus: Listen to short authentic clips and transcribe what you hear phonetically, not what you expect the words to be. This forces you to process the actual sounds.
- Reduced-form recognition drills: Practice identifying common reduced forms: “wanna,” “gonna,” “hafta,” “shoulda,” “coulda,” “wouldja,” “didja.” Build a mental lexicon of these forms.
- Chunking by prosodic phrase: Native speakers group words into breath groups with a single intonation contour. Train yourself to hear these chunks rather than individual words.
- Speed variation practice: Start with slightly slowed authentic speech, then gradually increase to natural speed. The goal is to process connected forms in real time, not just recognize them when slowed down.
- Shadowing with connected speech markers: Shadow native speakers while consciously tracking where linking, elision, and assimilation occur. This trains both production and perception simultaneously.
The connection between production and perception is direct. When you can produce “last night” as /læs.naɪt/, you immediately recognize it when you hear it. Improving your connected speech production is one of the most efficient ways to improve your listening comprehension.
How does American English connected speech differ from other varieties?
American English connected speech has several features that set it apart from British, Australian, and other varieties. Knowing these differences prevents you from applying the wrong patterns and helps you target specifically American speech.
Rhotic /r/ and intrusive /r/. American English is rhotic, meaning /r/ is pronounced in all positions, including after vowels. This creates strong consonant-to-vowel linking with /r/: “for a moment” links as /fɔ.rə.moʊ.mənt/. Non-rhotic varieties like standard British English drop the /r/ in those positions entirely, which changes the linking environment completely.
The flap /ɾ/. American English converts intervocalic /t/ and /d/ to a flap /ɾ/ between vowels, especially when the following vowel is unstressed. “Better” becomes /bɛ.ɾər/, “water” becomes /wɔ.ɾər/, “city” becomes /sɪ.ɾi/. British English retains the full /t/ in these positions. This flapping also applies across word boundaries: “get it” → /gɛ.ɾɪt/, “put it on” → /pʊ.ɾɪ.tɒn/.
Vowel reduction patterns. American English reduces unstressed vowels to schwa /ə/ more aggressively than many other varieties. “Can” in an unstressed position becomes /kən/, “and” becomes /ən/, “of” becomes /ə/. These reductions are more consistent and more extreme in American speech than in, for example, Australian English.
Coalescent assimilation frequency. The /t/+/j/ → /tʃ/ and /d/+/j/ → /dʒ/ assimilations are particularly prominent in American English. While they occur in other varieties, American speakers apply them more consistently in casual and even semi-formal speech.
Glottal stop usage. British English uses the glottal stop /ʔ/ as a /t/ variant far more frequently than American English. In American speech, /t/ between vowels typically flaps rather than glottaling. This is a consistent and diagnostically useful difference.
If you have trained with British English materials, you will need to consciously retrain several of these patterns for American professional contexts. The connected speech techniques differ enough that cross-variety confusion is a real and common problem for advanced learners.
Common challenges non-native speakers face with connected speech
The challenges are predictable, and so are the solutions. Understanding where you are most likely to struggle lets you target your practice precisely.
Challenge 1: Hearing word boundaries where there are none. Learners trained on citation forms expect to hear each word as a separate unit. In connected speech, “did you eat?” has no clear boundary between “did” and “you.” The solution is explicit boundary-dissolution training: practice listening to phrases as phonetic strings, not word sequences.
Challenge 2: Native language transfer. Motivation, anxiety, and native language transfer all affect the pace of acquiring connected speech skills. Spanish speakers, for example, tend to produce more syllable-timed speech, which resists the reduction patterns of stress-timed American English. Mandarin speakers often pronounce every syllable with full vowel quality, resisting schwa reduction. The solution is identifying your specific L1 transfer patterns and targeting them directly.
Challenge 3: Producing assimilation feels “wrong.” Many learners resist saying “wouldja” or “doncha” because it feels incorrect or informal. In reality, these are the standard spoken forms in American English at normal conversational speed. The solution is reframing: these are not errors. They are the correct phonological forms for fluent speech.
Challenge 4: Elision makes speech feel incomplete. Dropping the /t/ in “last night” or the /d/ in “old friend” feels like leaving something out. The solution is understanding that elision is phonologically governed, not random. Learning the rules removes the anxiety.
Challenge 5: Inconsistent production under pressure. Learners often produce connected speech correctly in drills but revert to citation forms in real conversations. This is a muscle memory issue, not a knowledge issue. The solution is extended paragraph-level practice until connected forms become the default, not a deliberate choice.
Techniques to build natural connected speech into everyday speaking
Knowing the rules of connected speech is not the same as producing them automatically. The gap between knowledge and habit closes through specific, consistent practice techniques.
Phrase-level drilling before sentence-level practice. Start with two-word phrases and master the connection before adding more words. “Pick it” → “pick it up” → “can you pick it up?” Each step adds one more connection. This builds the neural pathways for connected transitions without overwhelming working memory.
Phonetic journaling. Write out common phrases you use at work or in daily life in phonetic transcription, marking where linking, assimilation, intrusion, and elision should occur. Then practice those phrases aloud until the connected form feels natural. This technique works particularly well for everyday American expressions that you use repeatedly.
Shadowing with connected speech annotation. Choose a short clip of authentic American English, 20–30 seconds. Transcribe it phonetically, mark every connected speech process, then shadow it repeatedly. The annotation step forces conscious awareness; the shadowing builds automaticity.
Prosodic phrase practice. American English groups words into prosodic phrases, each with one primary stress and reduced unstressed syllables. Practice producing these phrases as single phonetic units: “I’m gonna go to the store” is one prosodic phrase, not six words. Training at the phrase level naturally produces connected speech.
Real-time self-monitoring. Record yourself in a real conversation or presentation, then listen back specifically for connected speech. Are you linking consonants to vowels? Are you reducing function words? Are you flapping intervocalic /t/? Targeted self-monitoring accelerates improvement faster than undirected practice.
Structured phonetic retraining. For learners who have plateaued, the most effective route is working with a trained accent coach who can assess your specific connected speech gaps and build a targeted program. The science-backed approach of retraining speech organs first, then applying connected speech patterns, produces faster and more durable results than surface-level imitation.
Myaccentway gives you a structured path to connected speech fluency
Most connected speech training stops at awareness. You learn what linking and elision are, you recognize them in examples, and then you are left to figure out production on your own. That gap between understanding and speaking is exactly where Myaccentway works.

Prof. Alex’s program is built around one principle: you cannot reliably produce what you cannot physically feel and see. The Interactive 2D Sound Video Simulators show you the exact speech-organ movements behind every American sound and every connected speech transition, something no audio recording or classroom exercise can do. Students who have tried years of self-study and group classes consistently report that seeing the movement is what finally makes the sound click.
The program begins with speech-organ awareness, moves through structured retraining of American consonants and vowels, and then applies those retrained sounds directly to connected speech patterns in phrases, sentences, and extended speech. Every session is 1-on-1 with Prof. Alex, which means your specific L1 transfer patterns, your professional speaking context, and your individual pace all shape the training. This is not a generic curriculum. It is a personalized phonetic program backed by 20+ years of linguistics research and coaching experience.
If you are ready to move from knowing about connected speech to producing it automatically, start with a sample class and experience the method directly.
FAQ
What is an example of linking in American English connected speech?
“Turn it off” is a clear example: the final /n/ of “turn” links to the initial vowel of “it,” and the /t/ of “it” links to the vowel of “off,” producing /tɜr.nɪ.tɔf/ as a single phonetic unit rather than three separate words.
What are the different types of connected speech processes?
The five main types are catenation (linking), assimilation, elision, intrusion, and geminates. Each modifies sounds at word boundaries in predictable ways, and all five are essential for natural American English fluency.
When do /t/ and /d/ assimilate before /y/ in American English?
When /t/ or /d/ immediately precedes the /j/ sound (spelled “y”), they coalesce into /tʃ/ and /dʒ/ respectively. “Don’t you” becomes /doʊn.tʃə/ and “did you” becomes /dɪ.dʒə/, a process called coalescent assimilation that is especially consistent in American English.
What is the difference between American and British English linking?
American English links /r/ across word boundaries because it is rhotic (“for a moment” → /fɔ.rə.moʊ.mənt/) and flaps intervocalic /t/ to /ɾ/ (“get it” → /gɛ.ɾɪt/). British English drops post-vocalic /r/ and retains a full /t/ in those positions, producing distinctly different connected speech patterns.
How does Myaccentway help with connected speech production?
Myaccentway’s 1-on-1 coaching with Prof. Alex uses 2D Sound Motion Technology to make speech-organ movements visible, allowing students to physically retrain the transitions behind linking, assimilation, elision, and intrusion rather than guessing from audio alone.
Key Takeaways
Mastering advanced connected speech in American English requires explicit training in four phonological processes: linking, assimilation, intrusion, and elision, supported by physical speech-organ retraining rather than auditory imitation alone.
| Point | Details |
|---|---|
| Citation form vs. spoken form | The gap between how words sound in isolation and in fluent speech is where most advanced learners lose comprehension. |
| Four core processes | Linking, assimilation, intrusion, and elision each follow predictable phonological rules that can be explicitly taught and trained. |
| American English specifics | Rhotic /r/ linking, intervocalic /t/ flapping, and coalescent assimilation of /t/+/j/ distinguish American connected speech from British and other varieties. |
| Production drives perception | Improving connected speech production directly improves listening comprehension, because the two skills share the same phonological representations. |
| Myaccentway’s approach | Prof. Alex’s 1-on-1 program uses 2D Sound Motion Technology to make speech-organ movements visible, accelerating connected speech mastery beyond what auditory-only methods achieve. |