What email verification accuracy is actually achievable?

Consider two claims that are both true at once. ZeroBounce publishes 99.6%, NeverBounce 99.9%, Clearout 99.73%. Hunter ran 40,000 verifications across the market and found real-world accuracy between 63 and 70%. Nobody is lying, and the two are measuring different things.

The vendor number answers a narrow question: of the addresses we returned a definitive verdict on, how many were right. Strip out every address the engine declined to judge, and what remains is the easy portion. A verifier could reach 100% on that measure by answering only the questions it found simple.

Hunter's number answers the question a buyer is actually asking: of the addresses I gave you, how many did you get right. On that denominator the unknowns count, the catch-alls count, and the score falls to something that describes the job rather than the sample.

So two numbers are needed and neither works alone. Coverage is the share of a list a verifier returned a definitive verdict on. Accuracy is how often those verdicts were correct, with the false-positive rate stated separately, because a wrong valid costs a bounce and a wrong invalid costs a customer.

When you next read an accuracy claim, ask what it was divided by. If the answer is not stated on the page, the number is decorative. Ask also what the test list contained: a corpus of consumer Gmail addresses is a much easier examination than a business list where a quarter of the domains accept everything.

One list, two scorecards

Say you give a verifier 10,000 addresses. It returns a clear verdict on 7,000 and declines to judge the other 3,000. Later you learn the truth about every address, and 6,860 of those 7,000 verdicts were right.

On the vendor's scorecard, that's 6,860 out of 7,000, or 98%. It goes on the home page. Score it the way you'd score it as the buyer and the math changes: 6,860 right answers out of the 10,000 addresses you paid attention to, or 68.6%. That sits inside the range Hunter measured across the market.

Nothing about the engine changed between those two scores. Only the bottom of the fraction moved. Whenever you see a single accuracy figure, picture the 3,000 that were left out and ask where they went.

A wrong valid and a wrong invalid don't cost the same

There are two ways for a verdict to be wrong, and you feel them differently. A wrong valid sends mail to a dead mailbox. You get a bounce, you see it in a report, and enough of them will hurt your sender reputation.

A wrong invalid is quieter. Somewhere a real person gets deleted from your list and never hears from you again. No report shows it. Say your list has 50,000 people and one verdict in a hundred is a false invalid. That's 500 real subscribers gone, and if each is worth a few dollars a year to you, the cost repeats every year.

So ask a vendor for both error rates, stated apart. A single accuracy figure can hide a tool that rarely lets a bad address through because it throws out good ones freely. For a store with paying customers on the list, that trade is often the more expensive one.

Is 99% accuracy possible?

On the subset of domains that answer honestly, yes, roughly 95 to 98%. Across a whole business list including catch-alls and unknowns, no service reaches it.

What is a good coverage figure?

It depends entirely on the list. A consumer Gmail list resolves far more readily than a B2B list where a quarter of domains accept every address.

Why does Hunter's benchmark differ so much from vendor claims?

It counts every address submitted, including the hard ones. Vendor figures are computed after excluding what the engine could not judge.

What should I ask a verifier before buying?

What share of my list will you decline to judge, and what was your accuracy figure divided by. A vendor that cannot answer the first has answered the second.

Check this against your own list

100 free credits a month, no card. Unknown results come back labeled and are never billed.