Learning Tipsjapanesepronunciationlisteningdevoicingbeginner

Why Japanese desu sounds like 'des': the vowels you can't hear

By the Conversa team · · 9 min read

Why Japanese desu sounds like 'des': the vowels you can't hear

ありがとうございます (arigatou gozaimasu) ends in a su that arrives with no voice in it. Listen to anyone thank you in Japanese and the phrase lands as "arigatō gozaimas," the last vowel still there in the timing and gone from the sound. It's the same reason です is pronounced "des" and not "de-su." You've heard both a thousand times, in anime and in shop doorways, and you've probably still been saying "go-za-i-ma-su."

The vowel is in the kana. It's in the romaji you first learned the phrase from. It's in the timing too. What it isn't in is your voice, and it never was.

Most articles about this treat it as a polish problem: stop over-enunciating and blend in. That gets the cost backwards. Long before it makes you sound stiff, this is why fast Japanese turns into mush in your ears. Here's where the vowel drops, why you sometimes still hear it, and a drill that works on your ear as well as your mouth.

です (desu) has two beats and one voice

です (desu) is written with two kana and lands with one audible vowel. Same with ます (masu). 好き (suki) comes out "ski." 少し (sukoshi) comes out "skoshi." 学生 (gakusei) comes out "gaksei."

That's standard Tokyo Japanese, and it's what your textbook's own audio was doing while the spelling insisted otherwise. Linguists call it vowel devoicing: the vowel loses its voice while everything around it carries on as normal. The sci.lang.japan FAQ states the rule plainly, and it turns on the consonants standing either side.

Two of the five Japanese vowels do this often enough to matter: i and u. A, e and o you can treat as safe.

Three questions tell you if a vowel will drop

好き (suki) is the cleanest case to test. The u sits between s and k. Put your hand on your throat and say "sss," then "kkk." No buzz on either. The vowel trapped between them has voiceless neighbours on both sides, and it gives up its own voice to match.

That's the whole test, and you can run it on a word you've never seen:

  1. Is the vowel i or u? If it's a, e or o, stop. You can treat it as safe.
  2. Is it either squeezed between two voiceless consonants, meaning k, s, sh, t, ts, ch, h, f or p, or the last thing in the phrase after one of them?
  3. Is it the last beat before the pitch drops? If so, it tends to keep its voice.

好き passes question 2: the u comes after s and before k. Question 3 is the brake, and 好き is where you can watch it work. That final i sits at the end of the word behind a k, so question 2 says it should drop too, which would leave you with "sk." It doesn't drop, because that beat is the one the pitch falls after. So the word arrives as "ski," with one vowel gone and one protected. です has nothing to protect its u, and out it goes.

Run the same test on a word you hear rather than one you read. If a Japanese word seems to open with a consonant cluster Japanese isn't supposed to have, the "sk" of "ski" or the "ts" of "tsu," that's a devoiced i or u, and the test tells you which vowel to write back in.

Writtensu-ki, two beats
Heardskiu caught between s and k
Writtende-su
Hearddesu ends the phrase after s

i and u go voiceless between two voiceless consonants, or at the end of a phrase after one.

The word you can't hear is usually the one that lost a vowel

好きです (suki desu) is four written vowels and two voiced ones. It arrives as "ski-des," and if you learned those two words separately, from a page, there's a good chance you've never once matched what you read to what you heard.

Here's how I think about it. If you came to Japanese through romaji, and almost every English speaker does, you filed each word by the vowels you saw. Gakusei has four of them. The audio gives you three with any voice in them. So your ear goes looking for a match, finds nothing shaped like that, and files the whole thing under "too fast."

You're searching for a word that was never going to arrive in the shape you stored it in. That model is mine, not a research finding, but the underlying effect is measurable. A 2025 study in Applied Psycholinguistics found that English speakers hear Japanese devoiced-vowel sequences differently from true consonant clusters, and that English-speaking learners of Japanese show measurable sensitivity to the vowel that isn't voiced. That's the encouraging half. The ear does adapt. What it adapts to is whatever you feed it, and if you're still reading four vowels where the audio gives three, you're feeding it a mismatch every time. Earlier work by Cutler, Otake and McQueen found devoicing changes how words get picked out of a stream of speech, though that study tested native Japanese listeners rather than learners.

Su-ki de-su.
Ski des.

The first u sits between s and k. The second ends the phrase after s. Both go voiceless.

Devoiced is not deleted, and the length is the tell

少し (sukoshi) is three beats whether or not you can hear the u. This is where learners overcorrect: they discover devoicing, decide the vowel is gone, and start saying a two-beat word. Japanese counts beats, and dropping one changes the shape of a word as surely as adding one does.

The vowel keeps its slot. Your mouth still moves into position for it. The vocal cords just don't switch on.

Tanner, Sonderegger and Torreira put numbers on this in Frontiers in Psychology: across a corpus of spontaneous Tokyo Japanese, a consonant-plus-high-vowel beat runs about 25% shorter before a voiceless consonant, while the same beat built on a, e or o shortens by about 3%. That gap is the tell. Devoicing is a switch Japanese throws on purpose, not the blur any language picks up at speed. Their data stops short of showing the timing comes out perfectly preserved in natural conversation, so treat the beat as keeping its weight rather than staying untouched.

If holding a beat while dropping its voice sounds familiar, it should. It's the mirror image of Japanese long vowels and double consonants, where the beat is the thing being added or held. Same rhythmic grid, opposite failure mode: there you lose a word by shortening a beat, here you lose one by deleting a beat you were only supposed to quieten.

Try it in Conversa

Practice with AI characters who adapt to your level and give real-time feedback.

Try Conversa Free

Why you sometimes do hear the Japanese silent u

靴下 (kutsushita) is where the rule appears to break. Socks. Three high vowels in a row, all sitting in positions that should devoice: ku, tsu, shi. Japanese doesn't quieten all three. In the commonly cited rendering, the ku goes voiceless, the tsu keeps its vowel, and the shi goes voiceless again.

That particular word is dictionary lore rather than lab result, so don't build a system on it. The underlying fact is solid: when devoiceable vowels stack up next to each other they don't all give way, which is the case Tsuchida studied in the Journal of East Asian Linguistics. This is also the real answer to "but sometimes I definitely hear the u." You were hearing correctly.

Be careful how hard you lean on it. Teaching sites like to state a tidy alternating rule, first one drops, second one survives, but the research describes consecutive environments as variable and speaker-dependent. A University of Edinburgh thesis on the mechanisms behind devoicing found that in single environments devoicing is close to compulsory, and that the blocking factors only really come into play once two of them collide. Which vowel wins can shift with the speaker and how fast they're talking.

Pitch is the other factor, and it's question 3 doing its work again. A beat carrying the accent tends to hang onto its voice, and in 靴下 the most commonly listed accent pattern puts it on that surviving tsu, which is at least consistent with the tendency even if one word proves nothing on its own. If you've already worked through Japanese pitch accent, this is where the two systems start talking to each other. If you haven't, leave it. Devoicing first.

Kansai holds the vowel Tokyo drops, so check who you're shadowing

The same です (desu) from an Osaka speaker keeps much more of its vowel. Devoicing is a feature of Tokyo standard speech, and it weakens noticeably in Kansai, a point Migaku's Kansai dialect guide makes in passing. I'd read that line three times before it occurred to me it explained why one podcast host sounded nothing like my textbook.

The practical version: if you've been shadowing a Kansai YouTuber for six months and drilling against NHK news, you've been training two different ears. Neither is wrong. Nobody has published a hard number for the gap, so treat it as a tendency and not a percentage. Just know which one you picked, because "that's not how my podcast host says it" usually means you found a real regional difference.

The drill: whisper one beat, then go find it

好き (suki) gets two taps. Tap the table once per beat and say the word at normal volume, then say it again, whispering only the beat that actually devoices. The tap stays where it was. Whisper the beat your three-step test flagged, not every i and u in the word.

です (desu): tap, tap, whisper the second. 少し (sukoshi): three taps, whisper the first. 学生 (gakusei): four taps, whisper the second. ありがとうございます (arigatou gozaimasu): ten taps, whisper the last. Leave 靴下 out of this until the single-vowel cases are automatic, because it has two devoicing beats and one that refuses.

Whispering isn't quite what a native speaker does. It's close enough to stop you deleting the beat, which is the mistake waiting on the other side of learning this rule.

Then run the drill backwards, because this started as a listening problem. Play a clip and write down two things: what you heard, and how many beats it took. Look at the kana last. "Des" plus two beats means your ear has both halves. "Des" plus one beat means you caught the sound and lost the rhythm. The beat is there in the length of that final s, and the rhythm is the half that decides whether you recognise the word next time.

Next time a sentence dissolves on you, don't rewind it and blame the speed. Rewind and ask which vowel went missing. Start with anything ending in ます, because you already know exactly what should be there.

Share this article

Related Posts

Ready to start speaking?

Join thousands learning with AI-powered conversations

Get Started Free