/v/as in van
Because Spanish spelling uses both 'b' and 'v' for the same single sound, you probably pronounce English 'v' the same way you pronounce 'b' — bringing both lips together or close together — rather than using your lower…


What you probably do
Because Spanish spelling uses both 'b' and 'v' for the same single sound, you probably pronounce English 'v' the same way you pronounce 'b' — bringing both lips together or close together — rather than using your lower lip and upper teeth, since your ear and mouth have never needed to separate these into two different sounds.
How natives do it
Americans keep these as two completely separate sounds: for 'v', the lower lip rises to touch the edge of the upper front teeth, leaving the upper lip out of the picture entirely, and the sound continues steadily rather than stopping the air the way 'b' does.
Why it matters
This merger is one of the highest-impact consonant issues for Spanish speakers because English uses this contrast to separate common, unrelated words like 'van/ban', 'vote/boat', and 'very/berry' — collapsing them into one sound can genuinely confuse listeners, not just mark an accent.
Hear the difference
Words this touches
Spanish speakers often substitute /b, β/.
How the mouth differs
Spanish treats the letters b and v as spelling variants of the same single phoneme, realized as a bilabial stop or approximant depending on position, with the lips fully or partly closing; English v is a distinct labiodental fricative made with the lower lip against the upper teeth, never involving both lips together, and contrasts meaningfully with b in many word pairs.
Listen for the pattern
Merging v into b (or its approximant variant) causes genuine lexical confusion for American listeners, since English treats 'van/ban', 'vote/boat', and 'very/berry' as entirely different words distinguished primarily by this single consonant.
Practise it
Hear the Difference: V vs. B
sourced- Put on headphones in a quiet space.
- Listen to each word pair, 'van/ban', 'vote/boat', 'very/berry', one at a time.
- Before the recording reveals the answer, decide which word you heard.
- Replay any pair you got wrong and listen specifically for a buzzing hum (V) versus a short popping sound (B) right before the vowel.
- Repeat the full set three times, tracking your score each round.
- Stop once you correctly identify at least 9 out of 10 pairs in a row.
Success check: You can correctly tell 'van' from 'ban' (and similar pairs) at least 9 times out of 10 without seeing the spelling.


Why this works. Minimal-pair listening trains categorical perception by forcing the learner to attend to the single acoustic cue that separates labiodental /v/ from bilabial /b/ — continuous frication noise with a gradual amplitude onset versus the sharp stop-burst of /b/. Because Spanish treats these as variants of one category, repeated forced-choice identification reshapes the learner's phonemic boundary, a process documented for other L1-categorization mismatches such as Japanese /l/–/r/.
Sources (2)
- Phonological Interference: How Native Language Habits Affect Pronunciation in a New Language – The English Nook — 2025
- Substitution occurs when learners replace L2 phonemes with the closest available sound in their native language, often causing meaning changes (e.g., Spanish /v/ becoming /b/).
- Omission involves dropping sounds that do not exist in the learner's L1 or are difficult to articulate due to L1 constraints.
- Insertion (epenthesis) happens when learners add vowels to break up consonant clusters that their native language does not support.
- Distinguishing universal and language-dependent levels of speech perception: Evidence from Japanese listeners' perception of English “l” and “r” — Virginia A. Mann, 1986
- Japanese listeners successfully discriminate the acoustic properties of English /l/ and /r/ when presented with non-English contrasts, ruling out a universal hearing deficit.
- The primary source of error is language-dependent categorization, where listeners map English sounds onto their existing L1 phonological categories (e.g., classifying both as liquids).
- Performance improves significantly when the acoustic contrast between /l/ and /r/ is exaggerated or presented in contexts that highlight their distinctiveness.
Drill Sentences for V
generated- Read this sentence slowly, marking every 'v': 'Victor loves to drive his van to visit seven villages.'
- Before each 'v', pause briefly and check that your lower lip is rising toward your upper teeth, not closing both lips.
- Record yourself reading the sentence at a natural pace.
- Play it back and listen to each 'v' — does it buzz continuously, or does it sound like a 'b'?
- Mark any words where the 'v' sounded like 'b' and repeat just those words five times each.
- Read the full sentence again at normal speed, aiming for every 'v' to buzz clearly.
Success check: On playback, every 'v' in the sentence has a continuous buzzing quality and none of them sound like 'b'.
Why this works. Embedding the target sound in varied, meaningful sentence contexts promotes transfer from isolated articulatory control to connected speech, where coarticulation with neighboring vowels and consonants can pull the gesture back toward the L1 bilabial habit if not actively monitored.
Exaggerate V, Then Dial It Back
sourced- Say 'vvvvvvan' holding the initial 'v' buzz for a full three seconds, deliberately too long and too loud.
- Record this exaggerated version and listen for a clear, sustained buzz with no trace of a lip-closing pop.
- Repeat with 'vvvvvvery' and 'vvvvvvote', again holding the 'v' far longer than normal.
- Now say the same words with the 'v' at only half that length, still clearly buzzing but closer to normal speed.
- Finally, say the words at a natural conversational pace, keeping the buzz but shortening it further.
- Compare your natural-pace recording to the exaggerated one and confirm the buzzing quality survived the shrink.
Success check: Your natural-speed 'v' still has an audible, brief buzz (not a pop), and you can clearly recall the exaggerated version as a reference if the sound slips back toward 'b'.
Why this works. Temporarily overshooting the contrast — holding the labiodental contact and voiced buzz far longer and more forcefully than natural speech requires — widens the perceptual and motor distance from the L1 bilabial substitute, making the category boundary unmistakable before the learner dials the gesture back down to a natural, brief duration; this overshoot-then-fade strategy mirrors temporal exaggeration techniques shown to sharpen categorical perception in L2 training.
Sources (1)
- The Role of Temporal Acoustic Exaggeration in High Variability Phonetic Training: A Behavioral and ERP Study — Bing Cheng, Xiaojuan Zhang, Siying Fan et al., 2019
- The HVPT-E group showed greater improvement in natural word identification performance compared to the standard HVPT group.
- Training with temporal acoustic exaggeration induced native-like categorical perception based on spectral cues.
- MMN responses demonstrated training-induced changes at pre-attentive neural levels, suggesting enhanced brain plasticity.
Say It Right: Producing English V
sourced- Stand in front of a mirror and say 'ffff' first, noticing your lower lip touching your upper teeth.
- Keep that exact lip-teeth position, then turn on your voice to make a buzzing 'vvvv' sound without moving your lips.
- Say 'van' slowly, holding the 'v' for two full seconds before releasing into the vowel.
- Check in the mirror that your lips never fully close together during the 'v'.
- Record yourself saying 'van, vote, very' and compare to the model audio.
- Repeat until your lip-teeth contact is visible and consistent in every repetition.
Success check: In your recording, the 'v' sound hums continuously for about a quarter second before the vowel, and your mirror shows the lower lip against the teeth with the upper lip never closing against it.


Why this works. Production practice on true minimal pairs forces the learner to commit to one articulatory gesture per trial rather than an ambiguous bilabial-labiodental blend. Recording and self-comparison against a model gives feedback on visible lower-lip-to-teeth contact and continuous frication, the two cues American listeners rely on most to separate /v/ from /b/.
Sources (2)
- Phonological Interference: How Native Language Habits Affect Pronunciation in a New Language – The English Nook — 2025
- Substitution occurs when learners replace L2 phonemes with the closest available sound in their native language, often causing meaning changes (e.g., Spanish /v/ becoming /b/).
- Omission involves dropping sounds that do not exist in the learner's L1 or are difficult to articulate due to L1 constraints.
- Insertion (epenthesis) happens when learners add vowels to break up consonant clusters that their native language does not support.
- Wells, J.C. (1982). Accents of English.
Use Your 'F' to Find 'V'
sourced- Say a long 'ffffff', the same sound as in Spanish 'fácil', and notice exactly where your lower lip touches your upper teeth.
- Keep your lips frozen in that exact position.
- Without moving your lips, turn on your voice so the sound becomes a buzzing 'vvvvv' instead of the hissy 'ffff'.
- Alternate ffff-vvvv-ffff-vvvv several times, changing only the voicing, not the lip position.
- Once the switch feels automatic, attach it to a vowel: 'ffff...vvvv-an' to produce 'van'.
- Check in the mirror that your lip position truly does not move between the f and v versions.
Success check: You can flip between 'ffff' and 'vvvv' by only turning your voice on and off, with zero visible change in lip position.


Why this works. English /v/ shares its exact place of articulation (labiodental) with /f/, a voiceless labiodental fricative that already exists in the Spanish sound inventory (e.g., 'fácil'). Using /f/ as a proxy motor template lets the learner borrow an already-automatized lip-teeth gesture from the orbicularis oris / lower-lip musculature, then simply add vocal fold vibration to convert it into /v/, bypassing the need to learn a brand-new articulatory placement.
Sources (2)
- Phonological Interference: How Native Language Habits Affect Pronunciation in a New Language – The English Nook — 2025
- Substitution occurs when learners replace L2 phonemes with the closest available sound in their native language, often causing meaning changes (e.g., Spanish /v/ becoming /b/).
- Omission involves dropping sounds that do not exist in the learner's L1 or are difficult to articulate due to L1 constraints.
- Insertion (epenthesis) happens when learners add vowels to break up consonant clusters that their native language does not support.
- Wells, J.C. (1982). Accents of English.
Slow-Motion V: Step by Step
consensus- Start with your lips relaxed and slightly open, not touching.
- Very slowly, raise only your lower lip until it lightly touches the bottom edge of your upper front teeth.
- Hold that contact for two seconds without making any sound.
- Now add your voice, creating a steady buzzing hum while keeping the same lip-teeth contact.
- Slowly release the lip and let the hum flow directly into a vowel, as in 'vvv-an'.
- Repeat the four steps three times, gradually speeding up until it feels like one smooth motion.
- Try the full-speed word 'van' and check it still has the buzz, not a pop.
Success check: You can feel the lower lip touch only the upper teeth (never the upper lip) at slow speed, and the buzzing sound appears before any sense of a 'popped' release.



Why this works. Breaking the gesture into slow-motion stages isolates the single articulatory parameter that differs from the L1 habit — lip rounding/closure versus labiodental contact — letting the learner consciously rehearse the unfamiliar motor sequence before reintegrating it into normal-speed speech, a standard technique for installing a new place of articulation.
Sources (1)
- Wells, J.C. (1982). Accents of English.
Reader ratings and feedback are coming soon.
Sources (1)
- Wells, J.C. (1982). Accents of English.