/v/as in van
You likely let the voicing drop out completely at the end of a word and keep the vowel before it short, producing what is effectively an f.


What you probably do
You likely let the voicing drop out completely at the end of a word and keep the vowel before it short, producing what is effectively an f. Hungarian doesn't rely on vowel length to signal a consonant's voicing, so this substitution feels natural.
How natives do it
Americans keep the buzz going right into the final closure of v and stretch the vowel just before it noticeably longer than they would before f. The lip position itself never changes.
Why it matters
Since vowel length carries much of the signal, dropping it can turn 'leave' into 'leaf' or 'save' into 'safe' in a listener's ear — a small timing habit with a real effect on being understood correctly.
Hear the difference
Words this touches
Hungarian speakers often substitute /f/.
How the mouth differs
Both languages form v identically, with the lower lip lightly touching the upper teeth and the vocal folds vibrating throughout. English marks word-final v versus f mainly through a longer preceding vowel plus residual voicing; without the habit of using vowel length this way, Hungarian speakers often devoice the final v fully, merging it with f.
Listen for the pattern
A devoiced final v can make 'leave' sound like 'leaf' to American ears, since the main cue distinguishing them — vowel length before the consonant — gets lost.
Practise it
Hear the Difference: V vs. B
sourced- Put on headphones in a quiet space.
- Listen to each word pair, 'van/ban', 'vote/boat', 'very/berry', one at a time.
- Before the recording reveals the answer, decide which word you heard.
- Replay any pair you got wrong and listen specifically for a buzzing hum (V) versus a short popping sound (B) right before the vowel.
- Repeat the full set three times, tracking your score each round.
- Stop once you correctly identify at least 9 out of 10 pairs in a row.
Success check: You can correctly tell 'van' from 'ban' (and similar pairs) at least 9 times out of 10 without seeing the spelling.


Why this works. Minimal-pair listening trains categorical perception by forcing the learner to attend to the single acoustic cue that separates labiodental /v/ from bilabial /b/ — continuous frication noise with a gradual amplitude onset versus the sharp stop-burst of /b/. Because Spanish treats these as variants of one category, repeated forced-choice identification reshapes the learner's phonemic boundary, a process documented for other L1-categorization mismatches such as Japanese /l/–/r/.
Sources (2)
- Phonological Interference: How Native Language Habits Affect Pronunciation in a New Language – The English Nook — 2025
- Substitution occurs when learners replace L2 phonemes with the closest available sound in their native language, often causing meaning changes (e.g., Spanish /v/ becoming /b/).
- Omission involves dropping sounds that do not exist in the learner's L1 or are difficult to articulate due to L1 constraints.
- Insertion (epenthesis) happens when learners add vowels to break up consonant clusters that their native language does not support.
- Distinguishing universal and language-dependent levels of speech perception: Evidence from Japanese listeners' perception of English “l” and “r” — Virginia A. Mann, 1986
- Japanese listeners successfully discriminate the acoustic properties of English /l/ and /r/ when presented with non-English contrasts, ruling out a universal hearing deficit.
- The primary source of error is language-dependent categorization, where listeners map English sounds onto their existing L1 phonological categories (e.g., classifying both as liquids).
- Performance improves significantly when the acoustic contrast between /l/ and /r/ is exaggerated or presented in contexts that highlight their distinctiveness.
Drill Sentences for V
generated- Read this sentence slowly, marking every 'v': 'Victor loves to drive his van to visit seven villages.'
- Before each 'v', pause briefly and check that your lower lip is rising toward your upper teeth, not closing both lips.
- Record yourself reading the sentence at a natural pace.
- Play it back and listen to each 'v' — does it buzz continuously, or does it sound like a 'b'?
- Mark any words where the 'v' sounded like 'b' and repeat just those words five times each.
- Read the full sentence again at normal speed, aiming for every 'v' to buzz clearly.
Success check: On playback, every 'v' in the sentence has a continuous buzzing quality and none of them sound like 'b'.
Why this works. Embedding the target sound in varied, meaningful sentence contexts promotes transfer from isolated articulatory control to connected speech, where coarticulation with neighboring vowels and consonants can pull the gesture back toward the L1 bilabial habit if not actively monitored.
Exaggerate V, Then Dial It Back
sourced- Say 'vvvvvvan' holding the initial 'v' buzz for a full three seconds, deliberately too long and too loud.
- Record this exaggerated version and listen for a clear, sustained buzz with no trace of a lip-closing pop.
- Repeat with 'vvvvvvery' and 'vvvvvvote', again holding the 'v' far longer than normal.
- Now say the same words with the 'v' at only half that length, still clearly buzzing but closer to normal speed.
- Finally, say the words at a natural conversational pace, keeping the buzz but shortening it further.
- Compare your natural-pace recording to the exaggerated one and confirm the buzzing quality survived the shrink.
Success check: Your natural-speed 'v' still has an audible, brief buzz (not a pop), and you can clearly recall the exaggerated version as a reference if the sound slips back toward 'b'.
Why this works. Temporarily overshooting the contrast — holding the labiodental contact and voiced buzz far longer and more forcefully than natural speech requires — widens the perceptual and motor distance from the L1 bilabial substitute, making the category boundary unmistakable before the learner dials the gesture back down to a natural, brief duration; this overshoot-then-fade strategy mirrors temporal exaggeration techniques shown to sharpen categorical perception in L2 training.
Sources (1)
- The Role of Temporal Acoustic Exaggeration in High Variability Phonetic Training: A Behavioral and ERP Study — Bing Cheng, Xiaojuan Zhang, Siying Fan et al., 2019
- The HVPT-E group showed greater improvement in natural word identification performance compared to the standard HVPT group.
- Training with temporal acoustic exaggeration induced native-like categorical perception based on spectral cues.
- MMN responses demonstrated training-induced changes at pre-attentive neural levels, suggesting enhanced brain plasticity.
Say It Right: Producing English V
sourced- Stand in front of a mirror and say 'ffff' first, noticing your lower lip touching your upper teeth.
- Keep that exact lip-teeth position, then turn on your voice to make a buzzing 'vvvv' sound without moving your lips.
- Say 'van' slowly, holding the 'v' for two full seconds before releasing into the vowel.
- Check in the mirror that your lips never fully close together during the 'v'.
- Record yourself saying 'van, vote, very' and compare to the model audio.
- Repeat until your lip-teeth contact is visible and consistent in every repetition.
Success check: In your recording, the 'v' sound hums continuously for about a quarter second before the vowel, and your mirror shows the lower lip against the teeth with the upper lip never closing against it.


Why this works. Production practice on true minimal pairs forces the learner to commit to one articulatory gesture per trial rather than an ambiguous bilabial-labiodental blend. Recording and self-comparison against a model gives feedback on visible lower-lip-to-teeth contact and continuous frication, the two cues American listeners rely on most to separate /v/ from /b/.
Sources (2)
- Phonological Interference: How Native Language Habits Affect Pronunciation in a New Language – The English Nook — 2025
- Substitution occurs when learners replace L2 phonemes with the closest available sound in their native language, often causing meaning changes (e.g., Spanish /v/ becoming /b/).
- Omission involves dropping sounds that do not exist in the learner's L1 or are difficult to articulate due to L1 constraints.
- Insertion (epenthesis) happens when learners add vowels to break up consonant clusters that their native language does not support.
- Wells, J.C. (1982). Accents of English.
Use Your 'F' to Find 'V'
sourced- Say a long 'ffffff', the same sound as in Spanish 'fácil', and notice exactly where your lower lip touches your upper teeth.
- Keep your lips frozen in that exact position.
- Without moving your lips, turn on your voice so the sound becomes a buzzing 'vvvvv' instead of the hissy 'ffff'.
- Alternate ffff-vvvv-ffff-vvvv several times, changing only the voicing, not the lip position.
- Once the switch feels automatic, attach it to a vowel: 'ffff...vvvv-an' to produce 'van'.
- Check in the mirror that your lip position truly does not move between the f and v versions.
Success check: You can flip between 'ffff' and 'vvvv' by only turning your voice on and off, with zero visible change in lip position.


Why this works. English /v/ shares its exact place of articulation (labiodental) with /f/, a voiceless labiodental fricative that already exists in the Spanish sound inventory (e.g., 'fácil'). Using /f/ as a proxy motor template lets the learner borrow an already-automatized lip-teeth gesture from the orbicularis oris / lower-lip musculature, then simply add vocal fold vibration to convert it into /v/, bypassing the need to learn a brand-new articulatory placement.
Sources (2)
- Phonological Interference: How Native Language Habits Affect Pronunciation in a New Language – The English Nook — 2025
- Substitution occurs when learners replace L2 phonemes with the closest available sound in their native language, often causing meaning changes (e.g., Spanish /v/ becoming /b/).
- Omission involves dropping sounds that do not exist in the learner's L1 or are difficult to articulate due to L1 constraints.
- Insertion (epenthesis) happens when learners add vowels to break up consonant clusters that their native language does not support.
- Wells, J.C. (1982). Accents of English.
Slow-Motion V: Step by Step
consensus- Start with your lips relaxed and slightly open, not touching.
- Very slowly, raise only your lower lip until it lightly touches the bottom edge of your upper front teeth.
- Hold that contact for two seconds without making any sound.
- Now add your voice, creating a steady buzzing hum while keeping the same lip-teeth contact.
- Slowly release the lip and let the hum flow directly into a vowel, as in 'vvv-an'.
- Repeat the four steps three times, gradually speeding up until it feels like one smooth motion.
- Try the full-speed word 'van' and check it still has the buzz, not a pop.
Success check: You can feel the lower lip touch only the upper teeth (never the upper lip) at slow speed, and the buzzing sound appears before any sense of a 'popped' release.



Why this works. Breaking the gesture into slow-motion stages isolates the single articulatory parameter that differs from the L1 habit — lip rounding/closure versus labiodental contact — letting the learner consciously rehearse the unfamiliar motor sequence before reintegrating it into normal-speed speech, a standard technique for installing a new place of articulation.
Sources (1)
- Wells, J.C. (1982). Accents of English.
Reader ratings and feedback are coming soon.
Sources (1)
- Wells, J.C. (1982). Accents of English.