What a filler actually is
"Um", "uh", "er", "like", "you know", "basically", "actually", "sort of", "I mean" — these are the sounds that appear when you have committed to speaking but haven't yet worked out what the next words are. That gap between commitment and content is the whole mechanism. Speech production runs slightly ahead of speech planning, and when planning falls behind, something has to occupy the airtime. A filler is that something.
This changes what the fix is. If fillers were a habit — a verbal tic you'd absentmindedly acquired — noticing them and deciding to stop would work. It doesn't, as anyone who has tried mid-answer can confirm; concentrating on not saying "um" reliably produces more of them, because you've added a monitoring task on top of the planning task that was already behind. Fillers are a symptom of planning load, so they respond to reducing that load, not to policing the output.
Veracue reports fillers as a rate per 100 words rather than a raw count, because a long answer will naturally contain more of them. Roughly: at or under about 2 per 100 words the feedback treats your filler use as natural, up to about 5 as moderate, and above that as heavy enough that a listener starts noticing the fillers instead of the content. Those bands are the product's own thresholds, not a universal standard — the useful signal is your own trend across sessions.
Why they cluster
Fillers are almost never spread evenly through an answer. They bunch at specific points, and the pattern is diagnostic — where they cluster tells you what part of the answer wasn't ready.
- At the very start. The most common cluster by far. The question has landed, you've begun speaking to avoid the silence, and you're choosing your example while already talking. "Um, so, yeah, I guess, like, one time..." is the sound of a decision being made out loud.
- At transitions. Moving from context into what you did, or from what you did into the outcome. These joins are where an unrehearsed answer discovers it doesn't know what comes next.
- Before anything specific. Numbers, dates, names, technical terms. You're retrieving a detail, and retrieval takes time you're filling with sound.
- At the end, when the answer doesn't know how to stop. "...so yeah, that was, um, basically it, I guess." This trailing cluster is usually a sign there was never a planned ending, so the answer runs out rather than concluding.
Read your transcript with those four buckets in mind and the fix usually becomes obvious. Heavy clustering at the start means you're starting too soon. Clustering at transitions means the middle of the answer isn't structured. Trailing fillers at the end mean you need a closing line.
The silent-pause substitution
The single technique that works is replacing the filled pause with a silent one. Not removing the pause — you genuinely need that time to think — but taking it without making a noise.
This is not intuitive, because silence feels enormous when you're the one producing it. A one-second gap can feel like five. Played back on a recording, that same pause is almost invisible, and frequently reads as composure. The asymmetry is reliable enough that you can simply trust it: the pause is always shorter to the listener than it is to you.
How to practise the substitution, in order of how well it works:
- Pause before you start, not after. The biggest single win. Take two or three seconds after the question to pick your example and decide how it ends, then begin. Nearly all of the opening cluster disappears, because you're no longer choosing while speaking.
- Close your mouth at the end of each sentence. Physically. Fillers need an open mouth and moving air; closing it makes an "um" mechanically awkward and turns the gap into silence by default. It sounds trivial and it is unusually effective.
- Slow down slightly. Fillers and pace are linked — speaking faster means less planning time, which widens the gap that fillers fill. Bringing your pace back towards the 140–170 range often reduces fillers without any direct work on them.
- Rehearse transitions specifically. If your fillers cluster at the joins, practise just the joins: the sentence that moves from situation to action, and the one that moves to the result. Two sentences, said aloud a few times, removes the cluster that a whole extra run-through wouldn't.
Fillers are close to impossible to count in yourself while speaking. Veracue gives you the rate per 100 words, the most common filler, and the full transcript — so you can see where they cluster rather than guessing.
Why zero is the wrong target
It's worth saying plainly: trying to eliminate fillers entirely will make you a worse interviewee, not a better one.
Fillers are a normal feature of spontaneous spoken English. Everyone uses them, including excellent communicators, and listeners don't consciously register them at ordinary rates. They also do useful work — a brief "um" signals that you're mid-thought and haven't finished your turn, which is part of how conversation manages itself. Speech with no fillers at all doesn't sound polished; it sounds recited, because the only reliable way to produce it is to read or recite from memory.
The cost of chasing zero is real, too. Monitoring your own speech for banned words consumes the attention you need for the content, which is the thing being evaluated. An interviewer will forgive a dozen "um"s in an answer with a concrete example and a clear outcome. They won't be won over by a filler-free answer that said nothing. The substance of the answer outranks its surface every time.
So the honest target is not elimination. It's getting from "the fillers are the thing you notice" to "the fillers are there and nobody cares" — and, more usefully, breaking the clusters at the start and at the transitions, because those are the ones that make an otherwise good answer sound uncertain. If your rate drops from 9 per 100 words to 3, you have made the entire available improvement. Going after the last 3 is effort spent on the wrong problem.