Your data provider is lying to you.
Not maliciously. It just doesn't know what it doesn't know.
Apollo doesn't have complete coverage of European SMBs. Hunter misses personal email addresses on certain domains. Datagma has excellent mobile numbers but patchy titles. Dropcontact is GDPR-compliant but only relevant for French and European contacts. No single source has everything — and every GTM engineer who has shipped a live campaign knows this from the first bounce report.
The question isn't which provider to pick.
It's how to architect a system that uses all of them intelligently, in sequence, without burning your credit budget on redundant lookups.
That system is called a waterfall.
The Coverage Gap Problem
Here is what happens when you rely on one provider.
You import 1,000 target companies into Clay. You run Apollo enrichment for verified work emails. Apollo returns results on 680 rows. That means 320 companies — 32% of your list — have no contact data. You're launching a campaign against two-thirds of your target market.
Now run a waterfall.
Apollo covers 680. Hunter runs only on the 320 rows Apollo missed — adds 140 more. Datagma runs only on the remaining 180 — adds 60 more. You're now at 880 contacts with verified emails. 88% coverage from the same list.
The delta between 68% and 88% isn't a nice-to-have. It's 200 accounts you were invisible to. It's the difference between a campaign that generates 12 replies and one that generates 18.
At scale — 10,000 accounts — that's 2,000 contacts you were leaving entirely untouched. Because you picked one provider.
How a Waterfall Actually Works
A waterfall is a sequenced series of enrichment steps with conditional logic at each stage.
The logic is simple: run provider 1. If provider 1 returns data, stop. If provider 1 returns empty, run provider 2. If provider 2 returns empty, run provider 3. Continue until you have data or exhaust your provider stack.
In Clay, this is built with "Run only if" conditions.
Each enrichment column gets a condition: only execute this column if the previous column is empty. This is the mechanism that prevents you from running five providers against a row that already has a valid verified email from provider 1.
Without "Run only if," every row runs against every provider. On 5,000 rows with 4 providers, that's 20,000 enrichment calls — most of them redundant. Conditional logic cuts this to the minimum required calls. The difference between a $50/month enrichment bill and a $500/month one is often nothing more than whether you bothered to set the condition.
Provider order matters. Start with your lowest-cost provider that has acceptable hit rate. Move to more expensive providers only for rows the cheap one missed. The goal is to let cheap data cover most of your list, with expensive data filling the edges.
Output format consistency matters. If provider 1 returns emails in one format and provider 2 in another, downstream columns will break. Normalize before you continue. A formula column that unifies format is worth writing once, using forever.
Provider Selection Guide
No provider is universally best. Each has a strength and a weakness. The waterfall is how you exploit their strengths while hiding their weaknesses.
Apollo — High data accuracy. Best for US tech companies, especially Series A and above. Slows noticeably under heavy filtering (10+ combined filters). Good as your primary sourcing layer for building the initial list. Less reliable for European SMBs and non-English-speaking markets.
Prospio — Superior speed and filtration. Preferred for enrichment workflows that need to run fast at scale. Better than Apollo in filtering-heavy scenarios. Use it as your primary enrichment layer when Apollo sourced the list.
Hunter — Email finding and verification. Strong for domain-based email discovery — you give it a domain and a name, it returns the most likely email format and verifies it. Weaker on mobile numbers. Good second-layer provider for email when Apollo/Prospio miss.
Datagma — Contact data enrichment with strong mobile coverage. Use it when you need phone numbers or when the first two layers miss. Weaker geographic coverage outside major markets.
Dropcontact — European-focused, GDPR-compliant. If any part of your campaign targets France, Germany, or broader EU contacts, Dropcontact belongs in your stack. It's built for compliance in ways the US-centric providers are not. Non-negotiable for EU outbound.
Clearbit — Company-level firmographic enrichment. Strong for enriching account-level data: funding stage, employee count, technology category. Use it at the company level, not the contact level.
Coresignal — Professional data, LinkedIn-adjacent. Useful for title verification and professional history. More accurate than scraped LinkedIn data on fresh role changes.
BuiltWith — Technology detection from public-facing website analysis. Identifies CRM, marketing automation, analytics tools, hosting infrastructure. Free tier covers most common lookups. Use it for tech stack qualification.
HG Insights — Deeper technology verification. Research-based verification of internal software — what the company actually uses vs what their website happens to load. More accurate than BuiltWith for enterprise accounts and internal tools. More expensive. Use for Tier 1 accounts where tech stack is a qualification criterion.
Where does Google go? At the bottom. Google is a search platform, not a proprietary data provider. It doesn't have a contact database — it indexes what others publish. Positioning it high in a waterfall wastes time and returns inconsistent results. If you use it at all, it's as a last-resort fallback for domain finding when all proprietary sources have failed.
The Technographic Layer
Firmographic data tells you who the company is. Technographic data tells you how they operate.
Knowing a company is a "B2B SaaS with 150 employees" is useful. Knowing they use Salesforce + Outreach + Bombora is better. That tech stack tells you their budget (they're spending on premium tooling), their sophistication (they have a dedicated outbound motion), and their likely pain points (they're probably buying more enrichment data than they know what to do with).
Technographic data also qualifies accounts in ways firmographic filters can't.
You can't filter Apollo for "companies that use HubSpot but not Salesforce and are evaluating a CRM switch." But you can enrich your list with BuiltWith, identify the ones running HubSpot, and cross-reference with Koala or 6sense for intent signals around CRM category research.
BuiltWith handles public-facing technology — what loads in the browser when you visit the site. Adequate for identifying marketing automation, analytics, chat tools, and hosting. Fast, cheap, widely available.
HG Insights verifies internal software through research-based methods — not just scraping the tech that loads publicly. More accurate for identifying what the company actually runs internally: ERP systems, internal BI tools, compliance software. More expensive. Reserve it for enterprise accounts where the technology detail justifies the cost.
Sumble provides account-level technology mapping that goes beyond website detection. Better for complex organizations with multiple products and systems.
The "AI-native" problem. If your ICP includes "AI-native companies" — firms built around AI infrastructure rather than bolt-on AI features — no database filter catches this directly. The terminology is too new, too inconsistently used, and too contextual. The solution is secondary AI validation: run a Claygent or Claude column that researches each company and validates whether they qualify against your definition. Expensive, but the only reliable approach for criteria that databases can't operationalize.
Data Decay: The Hidden Enemy
Enrichment is not a one-time event.
Email addresses decay at approximately 25% per year. A list you enriched 18 months ago has lost validity on roughly 37% of its contacts. People change jobs. Companies restructure. Domains change.
Phone numbers decay faster than emails — people change personal numbers less often, but direct-line business numbers evaporate when someone leaves a role.
Title and role data has the shortest shelf life. The average professional tenure is 2.1 years. That means in any given 24-month window, roughly half your list has moved to a different role, division, or company entirely.
The practical implication: enrichment data has an expiry date.
For warm accounts — companies you're actively working — re-enrich every 90 days. For cold accounts — the broader universe — re-enrich before any major campaign. A list built from data that is 12 months old is not a lead list. It's a collection of historical records.
Set freshness schedules. In Clay, you can build a formula that flags rows where the enrichment date exceeds a threshold. Any row older than 90 days gets flagged for re-enrichment before inclusion in active sequences. This runs automatically. You set it once.
Deduplication: Before, Not After
When you combine three data sources, you will have duplicates.
Apollo found the VP of Marketing at a company. Hunter also found the same person from domain-based lookup. Datagma found them via LinkedIn data. You now have three rows for the same contact — potentially with slightly different email formats, titles, or company names.
Deduplication must happen before enrichment, not after.
If you deduplicate after, you've already run enrichment against three rows for the same person. You've burned three times the credits for one contact. You've also created downstream chaos in your CRM — three records, three sequences, three sets of activity for the same person.
The deduplication logic needs to distinguish between two different problems:
Company-level deduplication — same company appearing under different names ("Acme Corp," "Acme Corporation," "Acme Corp Ltd"). Merge these before processing contact-level data.
People-level deduplication — same person appearing from multiple sources. Match on email as the primary key (most reliable) and fall back to name + company domain combination when email is missing.
Claude handles deduplication at scale well. Give it a structured prompt that defines the matching criteria, run it as a formula column before any enrichment starts, and let it flag probable duplicates for review. For large lists (10,000+), this step alone saves meaningful credit budget and prevents sequence duplication downstream.
Documentation Discipline
This is the step most engineers skip. It compounds in ways they later regret.
For every enrichment waterfall you build, document three things:
Provider performance — which providers returned data on what percentage of rows, for which types of companies and geographies. Apollo hit rate might be 68% on US Series B SaaS and 31% on European seed-stage. This matters for future list planning.
Formula logic — every "Run only if" condition, every normalization formula, every output mapping. Document it in a shared sheet linked to the Clay table. When a client comes back in 6 months, you're not rebuilding from memory.
Failure patterns — which rows returned empty from all providers, and why. Were they bootstrapped companies with no LinkedIn presence? Companies under 10 employees where contact data is sparse? Government entities with non-standard domain formats? These patterns tell you how to refine your sourcing filters upstream.
This dataset becomes proprietary knowledge. After 6 months of documenting provider performance across different ICP types, you have benchmarks no public resource offers. You know that for European fintech companies with 50-200 employees, the optimal waterfall is Dropcontact → Datagma → Hunter — in that order — because Dropcontact's GDPR-compliant European coverage is strongest and cheapest there, and the other two fill edge cases.
That knowledge is the difference between a GTM engineer who is learning and one who ships.
Building Your Waterfall: The Practical Sequence
Start here. Do not overthink the provider selection. Optimize after you have data.
Step 1 — Source the list. Apollo or Prospio for initial list building. Apply ICP filters, export to Clay, deduplicate at company level before proceeding.
Step 2 — Run primary email enrichment. Your best-performing provider for the target geography. Set a freshness threshold — only enrich rows where the enrichment date is null or older than 90 days.
Step 3 — Set "Run only if" on all subsequent providers. Every secondary provider runs only on rows where the primary returned empty. Non-negotiable.
Step 4 — Add technographic enrichment. BuiltWith for all rows (cheap, fast). HG Insights only for rows that pass a qualification threshold (save budget).
Step 5 — Deduplicate at contact level. Match on email first, then name + domain. Flag duplicates, review, merge or remove.
Step 6 — Document everything. Provider hit rates, formula logic, failure patterns. Store it per client, per ICP type.
Step 7 — Set freshness schedule. Flag rows for re-enrichment at 90-day intervals. Build this as a formula column so it runs automatically.
The Test to Run This Week
Take 100 rows from your current lead list. Any 100 rows.
Run your primary provider. Note how many return a valid verified email.
Now run two secondary providers on the rows the first one missed. Note how many they fill.
Calculate your coverage gap. That percentage — the contacts your current setup cannot reach — is what you're leaving on the table in every campaign you run.
If your gap is under 10%, your waterfall is working. If it's 20-30%, you're running a single-provider setup against a multi-provider problem.
Fix it once. It runs forever.