The benchmark, published method first

There is no result on this page. That is deliberate: the method goes up before the numbers, so we cannot quietly move the standard once we see how we did.

The short version

  • Hunter tested 40,000 verifications across the field, disclosed that their own data shaped the sample, and got cited everywhere. That is the standard worth meeting.
  • We will report coverage and accuracy together. Either number alone can be improved by making the product worse.
  • Results by provider, including the ones we handle badly. A benchmark that only reports wins is marketing in a lab coat.
  • Until it runs, this site publishes no ZapBounce accuracy figure at all.

What will be measured

MetricDefinitionWhy both
CoverageShare of the test list that produced a definitive verdict.Alone it is gameable: resolve every ambiguous address to valid and it hits 100%.
AccuracyOf those verdicts, the share that held when mail was actually sent.Alone it is gameable the other way: judge less and it approaches 100% on what is left.
False positivesAddresses called valid that hard bounced.The costly error. Reported separately rather than averaged away.
False negativesAddresses called invalid that accepted mail.The quiet error: customers you deleted.
Per providerThe same four figures for Gmail, Microsoft 365, Yahoo, and self-hosted domains.A single average hides that Yahoo is close to unmeasurable.

How it will run

  1. Build a list with known outcomes

    Addresses whose real state is established by actually sending to them and recording what happened, not by asking another verifier. A verifier scored against a verifier measures agreement, not truth.

  2. Describe the sample honestly

    Size, how it was assembled, the mix of providers, and the B2B to consumer split. Any sample has a bias; the useful move is naming it, as Hunter did with theirs.

  3. Verify without special handling

    The same pipeline a customer's list goes through. No retries we would not normally do, no domains excluded for being difficult.

  4. Send, and record what bounced

    The ground truth is delivery, not a second opinion. This is the slow part and the reason the benchmark does not already exist.

  5. Publish everything, including the losses

    Method, dataset description, date, and per-provider results. Where we score badly, that number goes up too.

What the benchmark will not be able to settle

Questions

Why publish a method with no results?

Because a method written after the results is a method chosen to flatter them. Committing first is the only way the number means anything later.

Will you benchmark competitors too?

Comparing on a dataset we built and control would flatter us by construction. Hunter's test is the better reference, and we cite it including where it puts them ahead of the field.

When?

When the send-and-record phase has run properly. Publishing early would contradict the only thing this company is arguing.

What if the numbers are bad?

They go up as they are. The page that only appears when the result is flattering is worth nothing to a reader.

Measure us yourself in the meantime

100 free checks a month. Run a sample and count the unresolved share.