Hunter published a benchmark across 40,000 verifications that measured the field at 63% to 70%, including themselves. They are an interested party selling a competing product, and the benchmark is still cited everywhere, because they described what they did.
That is the lesson. A self-interested benchmark with a disclosed method beats an independent-sounding claim with no method at all. The method is what makes it checkable.
So here is ours before the numbers exist. The dataset is a list of addresses with known outcomes, established by actually sending to them and recording what came back, rather than by asking another verifier what it thought. Anything else measures agreement between vendors rather than truth.
The dataset has to include the hard cases in their real proportion. A benchmark run on domains that answer honestly measures the easy 65% of the problem and produces the same inflated figure everyone else publishes.
What the results will have to include
Coverage and accuracy as a pair, with the false positive rate stated separately. A false positive is an address we called valid that bounced, and it is the error that costs you something. A false negative costs you a contact and no reputation.
The breakdown by provider, because an aggregate hides everything interesting. Our numbers on Yahoo will be poor. Our numbers on a domain that rejects unknown recipients cleanly should be good. Publishing only the average would conceal both.
The cases where a competitor did better. A benchmark that happens to rank its author first is not evidence, and everyone reading it knows that.
Where this argument costs us something
The short version
- Ask any vendor citing a benchmark whether the method is published.
- Check whether the dataset includes catch-all domains and Yahoo in realistic proportions.
- Treat a benchmark with no losing results as a marketing document.
Questions people ask
Why publish the method before the results?
Because the method is the part that can be argued with. Publishing results first invites everyone to assume the method was chosen to fit them.
When will the numbers be available?
When the benchmark has run against a dataset with known outcomes. Until then we publish no accuracy or coverage figure, which is the only position consistent with the rest of this site.