The benchmark, published method first
There is no result on this page. That is deliberate: the method goes up before the numbers, so we cannot quietly move the standard once we see how we did.
The short version
- Hunter tested 40,000 verifications across the field, disclosed that their own data shaped the sample, and got cited everywhere. That is the standard worth meeting.
- We will report coverage and accuracy together. Either number alone can be improved by making the product worse.
- Results by provider, including the ones we handle badly. A benchmark that only reports wins is marketing in a lab coat.
- Until it runs, this site publishes no ZapBounce accuracy figure at all.
What will be measured
| Metric | Definition | Why both |
|---|---|---|
| Coverage | Share of the test list that produced a definitive verdict. | Alone it is gameable: resolve every ambiguous address to valid and it hits 100%. |
| Accuracy | Of those verdicts, the share that held when mail was actually sent. | Alone it is gameable the other way: judge less and it approaches 100% on what is left. |
| False positives | Addresses called valid that hard bounced. | The costly error. Reported separately rather than averaged away. |
| False negatives | Addresses called invalid that accepted mail. | The quiet error: customers you deleted. |
| Per provider | The same four figures for Gmail, Microsoft 365, Yahoo, and self-hosted domains. | A single average hides that Yahoo is close to unmeasurable. |
How it will run
Build a list with known outcomes
Addresses whose real state is established by actually sending to them and recording what happened, not by asking another verifier. A verifier scored against a verifier measures agreement, not truth.
Describe the sample honestly
Size, how it was assembled, the mix of providers, and the B2B to consumer split. Any sample has a bias; the useful move is naming it, as Hunter did with theirs.
Verify without special handling
The same pipeline a customer's list goes through. No retries we would not normally do, no domains excluded for being difficult.
Send, and record what bounced
The ground truth is delivery, not a second opinion. This is the slow part and the reason the benchmark does not already exist.
Publish everything, including the losses
Method, dataset description, date, and per-provider results. Where we score badly, that number goes up too.
What the benchmark will not be able to settle
Questions
Why publish a method with no results?
Because a method written after the results is a method chosen to flatter them. Committing first is the only way the number means anything later.
Will you benchmark competitors too?
Comparing on a dataset we built and control would flatter us by construction. Hunter's test is the better reference, and we cite it including where it puts them ahead of the field.
When?
When the send-and-record phase has run properly. Publishing early would contradict the only thing this company is arguing.
What if the numbers are bad?
They go up as they are. The page that only appears when the result is flattering is worth nothing to a reader.
Measure us yourself in the meantime
100 free checks a month. Run a sample and count the unresolved share.