Methodology · Edition 3 · 2026
How we scored thirty address verification APIs
How a 113-product field was catalogued from G2 and Capterra and narrowed to 30 scorable vendors, then the complete rubric: six criteria, fixed weights, explicit scoring bands, an evidence hierarchy, and a written account of what this method cannot tell you. Published so the ranking can be checked rather than trusted.
The fieldHow 113 products became 30
A ranking is only as honest as its candidate list, so the field was assembled mechanically before any judgement was applied.
- Catalogue both portals in full. Every listing in the G2 Address Verification category, 103 products across seven pages, and every listing in the Capterra Address Verification category, 25 products. No sampling, and no filtering by rating or popularity.
- Deduplicate. Several vendors maintain multiple listings, sometimes in both portals and sometimes several times within one. Experian alone holds seven G2 listings, WinPure and Fetchify two each, and Melissa several under different product names. Collapsing these left 113 distinct products.
- Remove non-members of the category. Twenty listings are identity, KYC, AML or fraud tools that verify people rather than addresses, and two appear to be outright categorisation errors. All are named with reasons in the exclusions table on the main report rather than dropped silently.
- Test for scorability. A product was scored only where accessible public documentation supported a verifiable claim set across the six criteria. Thirty passed. The rest are catalogued in the directory alongside whatever portal rating exists.
Why we did not score all 113
Roughly half the catalogued products carry zero reviews on both portals and publish no technical documentation, pricing or certification detail. There is no honest way to place a number on coverage, accuracy or reliability for a product about which nothing is published. Filling those cells with estimates would make the table look more complete and be worth less. Every unscored product is still listed, so nothing is hidden, and a reader can see exactly how thin the long tail of this category is.
Inclusion criteria
A product was eligible for scoring if it offers address verification, validation or standardisation as a named, generally available commercial product with public documentation, and serves at least one of the US, UK, EU, Canadian or Australasian markets. Identity-verification platforms that check addresses as part of KYC were excluded, since they solve a different problem. Pure geocoding providers were excluded unless they market address validation specifically, which is why Radar, Geocodio and Woosmap are in and general mapping APIs are not.
One vendor outside the portals
Radar is scored despite appearing in neither category listing, because it markets an address verification API directly and competes for the same budget. Its profile says plainly that the portals do not classify it here. No other unlisted vendor was added.
PrinciplesFour rules we held to
- Weights before scores. The six weights were written down and published before any vendor was assessed. This is the single most important protection against a rubric being reverse-engineered to justify a preferred winner. If the weights had been chosen after scoring, the ranking would be an opinion dressed as arithmetic.
- Published claims only. A capability counts if a buyer can find it stated on the vendor's own documentation, product pages, trust pages or pricing pages without a sales call. Capabilities that exist but are undisclosed do not earn credit, because a buyer cannot rely on what they cannot find.
- Attribution over assertion. Where a figure comes from a vendor, we say so. Vendor-supplied accuracy percentages and case-study outcomes are reported as vendor claims, never as independently verified results.
- Publish the limitations. The limitations section below lists what this method cannot establish. A comparison that only advertises its strengths is marketing.
CriteriaThe six weights
- Coverage and data depth20%
- Developer experience20%
- Accuracy and certification20%
- Reliability and speed15%
- Pricing and value15%
- Compliance and security10%
Why these weights
The three 20% criteria are the ones that determine whether the product solves the problem at all. An API that does not cover your markets, that your team cannot integrate, or whose output you cannot trust, has failed regardless of how cheap or fast it is.
Reliability and pricing sit at 15% because they are real but substitutable concerns. A slower API can be worked around with caching and async processing. An expensive one can be negotiated or used more sparingly. Neither is fatal in the way a coverage gap is.
Compliance sits at 10% not because it is unimportant but because it is usually binary and market-specific. For most buyers it is a pass or fail gate applied before scoring, not a gradient. Where it is a gradient, as with EU data residency, that shows up in the sub-score.
Scoring bandsWhat each number means
Every criterion uses the same 0 to 10 scale, with band definitions written before scoring began.
1. Coverage and data depth. 20%
Measures how much of the world the vendor resolves, and how precisely. Country count alone is close to meaningless, so granularity is weighted more heavily than breadth.
| Band | Definition |
|---|---|
| 9 to 10 | 200+ countries with delivery-point or premise-level granularity in major markets; multiple authoritative reference sources per region; non-Latin script handling. |
| 7 to 8 | Broad international coverage but uneven granularity, or excellent coverage limited to two or three major markets. |
| 5 to 6 | Meaningful coverage in one region with shallow or autocomplete-only reach elsewhere. |
| 3 to 4 | Single-country product, however deep within that country. |
| 0 to 2 | Partial coverage of a single country. |
2. Developer experience. 20%
Measures time from landing page to a working, production-shaped integration. Documentation quality, SDK breadth, self-serve access, sandbox availability, error-message clarity and prebuilt platform integrations.
| Band | Definition |
|---|---|
| 9 to 10 | Self-serve signup, no card; keys in minutes; SDKs in 8+ languages; complete public reference with runnable examples; clear error semantics. |
| 7 to 8 | Good public docs and several SDKs, but signup friction, gated sandbox, or gaps in the reference. |
| 5 to 6 | Documentation exists but is thin or partly gated; few or no official SDKs; integration is mostly manual. |
| 3 to 4 | API is secondary to a GUI or desktop product; reference is incomplete or sales-gated. |
| 0 to 2 | No meaningful public developer documentation. |
3. Accuracy and certification. 20%
Because we did not run a bake-off, this criterion scores the strongest available proxies: formal postal certification, directness of the reference-data relationship, refresh cadence, and the specificity of any published accuracy evidence. Held certifications carry the most weight because they are externally audited.
| Band | Definition |
|---|---|
| 9 to 10 | Multiple postal certifications across regions (e.g. CASS + SERP + AMAS or NCOA); direct reference-data relationships; documented refresh cadence; specific published match-rate evidence. |
| 7 to 8 | At least one major postal certification plus credible published accuracy evidence. |
| 5 to 6 | One certification or authoritative licence, with little published accuracy evidence. |
| 3 to 4 | Authoritative data sources claimed but no external certification and no published evidence. |
| 0 to 2 | No certification, no named sources, marketing claims only. |
4. Reliability and speed. 15%
Rewards published, specific and contractually meaningful commitments. A vendor with a modest but stated SLA scores above one with no SLA and better marketing copy, because only one of those is something you can hold them to.
| Band | Definition |
|---|---|
| 9 to 10 | Published latency SLA with a percentile, published uptime target, public status page, documented throughput ceilings. |
| 7 to 8 | Published uptime target or typical latency range, plus a status page. |
| 5 to 6 | General performance claims without percentiles or contractual commitment. |
| 3 to 4 | No published performance information, no evident reliability problems. |
| 0 to 2 | No published information plus documented reliability complaints. |
5. Pricing and value. 15%
Combines cost per verified address with pricing transparency. Transparency is scored explicitly because an unpublished price is a real cost to the buyer in evaluation time and negotiating position.
| Band | Definition |
|---|---|
| 9 to 10 | Fully published pricing, competitive unit cost, meaningful free tier, no mandatory sales contact. |
| 7 to 8 | Published entry pricing or a substantial trial; enterprise tiers quoted but the on-ramp is self-serve. |
| 5 to 6 | Trial available but no published pricing at any tier. |
| 3 to 4 | Quote-only, with reported large gaps between trial and production cost. |
| 0 to 2 | No pricing information and no trial without a sales process. |
6. Compliance and security. 10%
Audited security certifications, sector-specific compliance, data residency options, retention policy clarity and access-control features.
| Band | Definition |
|---|---|
| 9 to 10 | SOC 2 Type 2 or ISO 27001, plus sector compliance (HIPAA, PCI-DSS), plus published retention and residency controls, plus SSO and RBAC. |
| 7 to 8 | At least one audited certification plus a clear published privacy and retention position. |
| 5 to 6 | GDPR compliance asserted, no audited certification published. |
| 3 to 4 | Privacy policy only, no certifications, no retention detail. |
| 0 to 2 | No published security or compliance position. |
CalculationWorked example
The overall score is a weighted arithmetic mean, rounded to one decimal place. Using the top-ranked vendor:
PostGrid
Coverage and data depth 9.6 x 0.20 = 1.920
Developer experience 9.5 x 0.20 = 1.900
Accuracy and certification 9.4 x 0.20 = 1.880
Reliability and speed 9.2 x 0.15 = 1.380
Pricing and value 9.3 x 0.15 = 1.395
Compliance and security 9.4 x 0.10 = 0.940
-------
Weighted overall 9.415 -> 9.4
To recompute the table under your own priorities, replace the weight column with your own values, keep the published sub-scores, and renormalise so the weights sum to 1.0. Every sub-score needed is in the scorecard.
Ties. Where two vendors round to the same overall score, the higher unrounded value ranks first. No tie occurred in this edition.
EvidenceHierarchy and inclusion rules
Evidence hierarchy, strongest first
- Postal authority registers. USPS CASS and NCOALink, Canada Post SERP, Royal Mail PAF licensing, Australia Post AMAS. Externally audited, so treated as fact.
- Independent security audits. SOC 2 Type 2 and ISO 27001 attestations. Treated as fact where the vendor publishes the certification.
- Vendor primary documentation. Public API reference, product pages, trust pages, status pages, published pricing. Treated as a vendor claim, reported with attribution.
- Third-party review portals. G2, Capterra, The CTO Club, LinkedIn product directory. Used to sanity-check support and reliability signals, never to set a score directly, because review volume tracks marketing spend more than product quality. Both portal ratings are published side by side on the main report so readers can see the disagreement rather than inherit ours.
- Vendor case studies. Lowest weight. Discounted heavily where the vendor supplies both the metric and its measurement.
What we excluded and why
- Brand recognition and analyst placement. Not a product property.
- Review counts as a score input. A vendor with four reviews and strong documentation outranks one with ninety reviews and no published SLA.
- Roadmap and beta features. Only generally available, documented capability as of September 2026 counts.
- Sponsored or paid placement. No ranking position is available for purchase.
LimitationsWhat this method cannot tell you
We did not run an accuracy bake-off
The obvious way to rank verification APIs is to send all of them the same corpus of addresses and count correct results. We did not do this, and you should know why. A credible test needs licensed ground-truth data for every market in scope, and results would be dominated by corpus composition rather than vendor quality. A corpus heavy in US suburban addresses produces one ranking; one heavy in UK new-builds or Japanese addresses produces another. Rather than run a cheap test and present it as authoritative, the accuracy criterion scores externally auditable proxies and says so.
- Coverage counts are vendor-stated and not directly comparable. One vendor's "250 countries" may mean delivery-point precision in 40 and city-level in the rest. We weighted granularity where it was published, but vendors publish this unevenly.
- Pricing comparisons are incomplete by construction. Ten of the thirty scored vendors publish no list pricing at any tier. Their value scores reflect that opacity, which is a fair penalty for a buyer's evaluation cost but is not the same as knowing they are expensive.
- Support quality is largely unmeasured. We have no reliable way to compare response times or escalation paths across thirty vendors without being a paying customer of each.
- Portal ratings are noisy and sometimes contradictory. The same product can hold 4.6 from 339 reviews on one portal and 4.9 from 91 on another. Several vendors have a strong rating on one portal and zero reviews on the other, including Addressfinder and Woosmap. Ratings built on fewer than ten reviews are reported but carry essentially no information.
- The directory reflects portal categorisation, not ours. Products appear in it because G2 or Capterra placed them in the category. Their presence is not an endorsement, and in twenty-two cases we say explicitly that the listing does not belong there.
- This is a snapshot. Postal certifications lapse and renew, pricing changes, coverage expands, and portal listings are added and withdrawn. Every figure is as of September 2026.
- Sub-scores involve judgement. The bands constrain it, but two careful analysts applying this rubric would not produce identical numbers. They should produce a similar ordering. If they would not, the rubric is the problem.
Corrections
If a figure here is wrong or out of date, it should be corrected. Vendor documentation changes frequently and this evaluation is only as good as its last review. Corrections are applied at the next quarterly review, or immediately where a factual error is material to a ranking position.