Say the word ベッド out loud. Beddo. Two clean syllables, both vowels straight out of the Japanese set. Now say the English word it came from: bed. One syllable, and a vowel that sits somewhere your mouth has never been asked to go. The katakana version already made a decision for you: it took the English /ɛ/, filed it under エ, bolted an extra /o/ onto the end, and handed you back something easy to say and slightly wrong.
That is the trap. Japanese has exactly five vowels: あ い う え お, /a i u e o/, and vowel length is the only other thing that changes (Wikipedia: Japanese phonology). English has around eleven to twelve vowel qualities, plus a stack of diphthongs, so somewhere between fifteen and twenty vowel sounds depending on whose accent you count. When you learned English words as katakana, every one of those got rounded off to the nearest of five. This post is about which English vowels got merged in your pronunciation, why sit and seat can sound nearly identical to your ear, and how to pull them back apart.
Why "bed" becomes ベッド
Say bus and English lets it stop on the /s/. Japanese won't. It props a /u/ underneath and hands you バス, basu. That inserted vowel is the first of two things that happen every time an English word gets katakana-ized, and they happen together.
The first is the extra vowels, because Japanese syllables are mostly consonant-plus-vowel and the only sounds allowed to end one are ん and the little pause written ッ (Wikipedia: Loanwords in Japanese). That geminate ッ is the same mora-length machinery behind long vowels and double consonants in Japanese. Milk becomes ミルク, miruku, two extra vowels of its own. Strike, one syllable in English, balloons into ストライク, sutoraiku, five beats long.
The second thing is quieter and more damaging: the vowel that was already there gets swapped for the nearest Japanese one. Bed's /ɛ/ is close enough to エ that nobody notices the substitution. Cup's /ʌ/ becomes the /a/ in カップ, kappu. The English vowel doesn't survive the trip. And because you say these katakana words every day, the wrong vowel is the one that feels natural when you switch into English.
One honest caveat before the map: katakana isn't perfectly consistent. The /æ/ in bag shows up as バッグ with a plain あ-vowel, but the /æ/ in cat gets the ya-glide キャット. Don't treat any single English vowel as having one fixed katakana home. Katakana English carries a vocabulary trap too, the made-in-Japan words that aren't real English; this post is only about the vowels. The point isn't the exact spelling. It's that the five-vowel filter throws information away.
How katakana collapses English vowels
Each Japanese vowel swallows a whole set of English ones, and あ alone takes four (Tofugu: Katakanization):
| Japanese vowel | English vowels that collapse into it | Example words |
|---|---|---|
| あ /a/ | /æ/, /ʌ/, /ɑː/, /ə/ | cat, cut, cart, about |
| い /i/ | /ɪ/, /iː/ | sit, seat |
| う /u/ | /ʊ/, /uː/ | full, fool |
| え /e/ | /ɛ/ | bed |
| お /o/ | /ɔː/, /oʊ/, /ɒ/ | bought, boat, hot |
Look at the top row. Four separate English vowels, all landing on あ. That is the pile-up that makes cat, cut, and cart sound interchangeable. The い row is the famous one, the sit / seat problem. え has it easy: /ɛ/ maps to エ cleanly, and it's the one vowel you rarely need to fix. Spend your effort where the merges are worst, which means あ first and い second.
Sit or seat: the い trap
Sit and seat both become シート-ish sounds in katakana, so the only difference katakana preserves is length, the long ー mark. That trains you to think the contrast is short-versus-long. In English it isn't, or not only. /iː/ in seat is tense: your tongue is high and your lips spread into a slight smile. /ɪ/ in sit is lax: the tongue drops a little, the lips relax, the sound is looser (7ESL: /iː/ vs /ɪ/ minimal pairs).
The pairs that catch people: sit / seat, ship / sheep, live / leave, fill / feel, chip / cheap (EnglishClub minimal pairs). Say sheep and hold the smile. Now say ship and let your face go slack on the vowel. If you only shorten sheep without relaxing the mouth, a listener still hears sheep, just clipped. Length alone won't rescue you.
Cat, cut, cart: the あ pile-up
Cat /kæt/, cut /kʌt/, and cart /kɑːt/ are three words with three different vowels, and katakana gives you one bucket for all of them. This is the merge that does the most damage.
The pairs to drill: cat / cut, cap / cup, bad / bud, ran / run, hat / hut. Three targets to keep separate. /æ/ in cat is wide and bright, mouth open, tongue pushed forward, almost a complaining sound. /ʌ/ in cut is short and central, the vowel you make when a doctor says "say ah" but only halfway. /ɑː/ in cart is long and pulled to the back of the mouth, the vowel in father. If cat and cut come out identical, you're most likely making both of them as /ʌ/, the comfortable middle one. The fix is to exaggerate cat until it feels rude, then dial back.
Full or fool, bought or boat
The う and お rows have the same length-versus-quality trap as い. Full /fʊl/ and fool /fuːl/, pull and pool: katakana keeps the length difference and drops the rest. /ʊ/ is lax and loose; /uː/ is tense, with the lips rounded and pushed forward. Round harder on fool and pool.
お hides a different problem. Bought /bɔːt/ is a single steady vowel. Boat /boʊt/ is a diphthong, a glide that starts at /o/ and slides toward /u/, two vowel positions inside one sound. Katakana flattens both into a plain お or a long オー, so the glide in boat disappears and it comes out sounding like bought. Let the boat vowel move. Your lips should travel.
Try it in Conversa
Practice with AI characters who adapt to your level and give real-time feedback.
Try Conversa FreeHear it before you say it
If sit and seat sound identical going into your ear, no amount of mouth-position advice fixes them coming out. You cannot reliably produce a difference you cannot hear, and that is the step most pronunciation guides skip. Perception first, production second.
The drill is minimal-pair listening. Have someone say one word from a pair, sit or seat, without showing you which, and you guess. A partner works, or an app, or a text-to-speech tool. That is the whole drill. You're not trying to say anything yet. You're teaching your ear that these are two categories, not one blurry one. Only once you can hear the split reliably does saying it become worth practicing.
This retraining is slow and a little boring, and it is the part that actually moves the needle. Shadowing helps too: play a short clip of real English, not the katakana reading in your head, and copy it immediately, vowel for vowel. This is also where an AI conversation partner earns its keep, because it reacts to the vowel you actually produced instead of the one you meant.
A ten-minute daily version
Pick one collapse and stay on it for the day. Three minutes of listening discrimination, guessing words from a minimal pair. Four minutes reading the pair list aloud, watching your mouth in a mirror or your phone camera. Three minutes shadowing a short clip. Rotate the vowel each day, and give あ two days a week, because it's carrying four English vowels and one day won't clear it.
None of this erases your accent, and that was never the goal. The goal is narrower and more useful: to stop four English words from arriving as one. The next time you order a コーヒー, notice how automatic that vowel is, and then aim your English somewhere your katakana never let it go.
