Join the waitlist

Let us know how we should get in touch with you.

Thank you for your interest! We’re excited to show you what we’re building very soon.

Close
Oops! Something went wrong while submitting the form.

How to Build a Waterfall Enrichment Workflow (Step-by-Step)

Austin Hughes
·
Updated on: August 3, 2026
TL;DR: A waterfall enrichment workflow queries B2B data providers in priority order, calling the next source only when the last one returns a blank or low-confidence value. For RevOps and sales teams, five decisions, provider stack, cascade logic, conflict resolution, cost control, and monitoring, determine whether match rates hold and bounce rates stay low, as CandorIQ's fell 87% after consolidating onto one waterfall.

Key Facts at a Glance

The numbers below are the only quantitative claims in this article, pulled into one place so you don't have to hunt for them section by section.

Claim Value Source and date
Unify's contact and company database 1.1B+ contacts, 65M+ companies Unify, B2B Company & Contact Data product page, 2026
Signal and intent data sources feeding Unify 40+ Unify, B2B Company & Contact Data product page, 2026
Email and phone vendors waterfalled inside Unify 11+ Unify, B2B Company & Contact Data product page, 2026
Bounce rate change after consolidating onto Unify's waterfall and deliverability layer 87% lower Unify customer story, CandorIQ, 2026
Time spent on manual prospecting and enrichment tasks 95% less Unify customer story, CandorIQ, 2026
Bounces prevented in outbound enrollments by Unify Managed Deliverability More than 10% Unify customer story, Justworks
Return on investment in first 5 months 6.8X Unify customer story, Justworks

Methodology and limitations. The data decay figure above comes from ZoomInfo's own published research (updated June 30, 2026), which attributes the aggregate 22.5% number to HubSpot benchmark modeling. Field-level decay for direct dials and job titles specifically can run higher than the aggregate. The Unify numbers in this article are not a blended "platform benchmark." There is no single unified Unify dataset behind them: each is attributed to the specific customer case study it came from, CandorIQ or Justworks, and neither should be read as a guarantee for a different ICP, provider mix, or geography. The confidence-threshold and cadence numbers elsewhere in this guide are practitioner starting points, not measured outcomes. Validate them against your own bounce and connect data before locking them in, and tighten them in regulated industries or GDPR and CCPA contexts, where a bad record carries compliance risk on top of a wasted send.

What Does a Waterfall Enrichment Workflow Actually Include?

A complete waterfall enrichment workflow has five functional layers: a provider stack, cascade trigger logic, conflict resolution rules, cost controls, and a monitoring system. Most teams that struggle with enrichment have the first layer roughly configured but have built little or nothing for the other four. That gap is where match rates stall, data quality degrades, and provider spend creeps up without anyone noticing.

A waterfall enrichment workflow is not just a vendor contract or a stack of API integrations. It's an architecture where each layer depends on the ones before it. Cascade logic only works if the provider stack is ordered correctly. Conflict resolution only produces reliable output if confidence thresholds are calibrated. Cost controls only fire when field-state checks are wired into the pre-call logic. This guide walks through each layer in the order you actually need to build them.

If you already understand why waterfall enrichment beats single-source enrichment and just want the implementation steps, keep reading. If you want the foundational "why" first, read What Is Waterfall Enrichment? Why It Beats Single-Source B2B Data before coming back here.

How Do You Choose Your Provider Stack and Priority Order?

The right provider stack depends on the data types you need, your target geography, and the size of the companies you sell into. There's no universal stack. The mistake most teams make is choosing providers on brand recognition or list price alone, instead of coverage benchmarks for their specific ICP.

Evaluate providers by data type, not by brand

Work email, direct-dial phone, firmographic data, technographic data, and intent signals are not equally covered by any single vendor. A provider strong on North American enterprise emails may have weak direct-dial coverage and almost no international contact data. Rather than naming specific vendors here (vendor strengths shift and a name-drop ages badly), use the criteria below to score any candidate provider for each data type in your stack.

Data type Strong primary source Strong secondary source Strong tertiary source
Work email High match rate on your exact ICP at the lowest per-record cost Fills gaps the primary misses, often with regional or vertical depth Premium or verification-heavy fallback, called sparingly
Direct-dial phone Broad general coverage with reasonable connect rates Live-verified numbers confirmed active recently, for higher-value accounts Compliance-heavy or region-specific verification for the hardest records
Firmographic data Broad company database with frequent refresh cycles Specialized depth the primary lacks, such as funding or org structure Narrow, high-precision source for edge segments
Technographic data General tech-stack detection across your whole TAM Specialized detection for niche or emerging tools Rarely needed as a third layer
Intent signals First-party signals you already control, like website or product usage Third-party aggregated intent for broader market coverage High-cost, high-precision signals reserved for named accounts

Run a sample of 500 to 1,000 known target accounts through each candidate provider's API and measure match rate, field completeness, and email deliverability before you commit. Most providers offer trial API access for exactly this purpose, and skipping this step is the single most common reason a stack underperforms once it's live. If you want a repeatable scorecard for this test rather than building one from scratch, see How to Compare B2B Enrichment Providers: A Scorecard.

Put your most complete, cost-effective provider first

The first provider in your waterfall handles the highest volume of queries, so it should be the one with the best match rate for your ICP at the lowest per-record cost. High-cost, high-precision providers belong further down the cascade, called only for records the cheaper primary source couldn't cover. This ordering decision drives most of the cost efficiency in a waterfall architecture, well before you touch any cost-control settings.

How Do You Set Up Cascade Logic and Confidence Thresholds?

Cascade logic determines when your workflow stops querying one provider and moves to the next. The most common mistake is treating cascade triggers as binary: either the field is empty (move on) or it has a value (stop). Binary logic leaves quality on the table, because a provider can return an email that's syntactically valid but flagged as low-confidence, or a phone number formatted correctly but never verified against a live directory.

What is a confidence threshold, and why does it matter?

A confidence threshold is the minimum quality score a returned value must clear before the waterfall stops querying for that field on that record. Instead of asking "did the provider return a value?", your logic asks "did the provider return a value above X% confidence?" A reasonable starting point for email is 85% or above on the provider's own validation score. For phone numbers, the meaningful line is usually whether the provider performs live verification at all, not just a confidence score. For firmographic fields like headcount or revenue, a source under 12 months old is typically an acceptable bar. These are starting points to calibrate against your own bounce and connect data, not fixed rules.

Calibrate thresholds against your own outcomes

Run your first enrichment batch without strict thresholds, then measure real deliverability on the output: bounce rates, connection rates, reply rates. Work backward from that data to find the confidence range above which delivered data was actually accurate. Set your production threshold at the low end of that reliable range, and revisit it quarterly as provider quality shifts.

How Do You Handle Conflicting Data Between Providers?

Data conflicts occur when two providers return different values for the same field on the same record. A well-designed field-level waterfall minimizes conflicts by design, since each field stops at the first provider that clears the confidence threshold. But conflicts still show up during initial stack evaluation, during re-enrichment of existing records, and whenever you add a new provider mid-cycle.

Recency is the default tiebreaker

When two providers disagree and confidence scores alone don't settle it, recency is the most defensible rule. B2B contact data decays at roughly 22.5% per year according to ZoomInfo's data decay research, so a more recently sourced or verified value is statistically more likely to be correct. Most providers include a sourced-at or last-verified timestamp in their API response; build your conflict logic to compare those timestamps whenever confidence scores tie.

Layer in source reliability for persistent conflicts

Recency resolves most conflicts, but some providers are structurally more accurate for a specific field regardless of when their data was pulled. Live-verified mobile data is more reliable than a number scraped from a web directory, even if the scraped number happens to be newer. A simple reliability model assigns each provider a field-level score based on your own observed deliverability, then multiplies that score by a recency weight. The provider with the higher composite score wins the conflict. This two-factor approach catches the case where a highly reliable provider has slightly older data than a lower-reliability one, but is still the right call.

How Do You Optimize Cost in a Waterfall Enrichment Workflow?

Cost control in waterfall enrichment comes down to one principle: never pay a provider to return data you already have. It sounds obvious, but most teams violate it constantly because their workflow doesn't check existing field values before issuing an API call.

Check field state before every call

Before calling any provider for a record, check the current state of every field for that record in your CRM or warehouse. If a field already has a value sourced within your acceptable staleness window, often 90 to 180 days depending on the field type, skip the call for that field entirely. This one check is what separates a field-level cascade from a wasteful record-level one.

Route around known coverage gaps

If you know from historical data that a given provider has weak coverage for a segment, small companies, a specific region, don't query it for those records. Route them directly to a provider with better coverage for that segment instead. This eliminates a whole tier of calls for a subset of your records rather than running everything through the full cascade from the top.

Tier your re-enrichment cadence by account priority

Not every record deserves the same enrichment investment. A named, high-value account your reps are actively working is worth re-enriching on a tight cycle with your full stack. A long-tail account with no engagement doesn't warrant more than a quarterly refresh with your cheapest primary source. Tiering cadence by account priority is one of the highest-leverage cost controls available, and it's a configuration decision, not an engineering one. For how this fits into broader CRM hygiene, see CRM Data Hygiene for RevOps: Waterfall Enrichment, Sync, and Deduplication.

How Do You Monitor and Optimize a Waterfall Workflow Over Time?

A waterfall enrichment workflow is not a set-it-and-forget-it system. Provider coverage changes, your ICP shifts, and new vendors enter the market. Without active monitoring, a carefully configured workflow degrades quietly, and you usually don't notice until deliverability or connect rates have already slipped.

Track three metrics, not one

Match rate by provider and by field tells you whether each provider is doing its job in the cascade. Email deliverability rate on enriched contacts tells you whether your confidence thresholds are actually holding up in production. Cost per enriched record tells you whether your pre-call checks and coverage routing are working as intended. If all three are healthy, the workflow is running well. If one degrades, it points to exactly which layer needs attention, rather than leaving you guessing across the whole system.

Review the provider stack every quarter

Compare each provider's match rate, deliverability contribution, and cost against the prior quarter. Providers with declining match rates or worsening deliverability should move down the cascade or get replaced. New providers worth evaluating should run against a sample of your ICP before entering production. This cadence keeps a waterfall calibrated as the underlying data landscape shifts.

How Unify covers this

Unify combines proprietary databases with a waterfall across 11+ email and phone vendors, plus 40+ signal and intent sources for niche coverage like local businesses or ecommerce, searchable against 1.1B+ contacts and 65M+ companies from a single chat interface, according to Unify's B2B Company & Contact Data product page. Instead of building and maintaining the five layers above yourself, agents adapt the waterfall to the industry, geography, and nuances of each search, and match rate and deliverability data update in the same interface reps already use to prospect and sequence.

Worked Example: How CandorIQ Rebuilt Its Enrichment Stack

CandorIQ, a compensation and headcount management software company, had a founding SDR, Zach Dettlinger, tasked with building outbound infrastructure from nothing. He inherited a stack stitched together from four separate tools: one for list building and sequencing, one for one-off contact lookups, a web-intent tool that required manually cleaning leads before he could act on them, and a chat tool for drafting email copy that required re-explaining business context every time.

Within the first two weeks of evaluating alternatives, he tried a signals-only tool and found it still piped data back into his sequencing tool to actually send anything, solving one problem while leaving the stack just as fragmented. Once CandorIQ's leadership introduced Unify, onboarding to a single agentic outbound engine, covering prospecting, enrichment, and multi-channel sequencing, took 12 days from first evaluation to live use.

The measurable result, per Unify's published CandorIQ case study: $1.8M in pipeline attributed to Unify, a 95% reduction in time spent on manual tasks, a 3.4% reply rate and climbing, and an 87% lower bounce rate once enrichment and deliverability ran through one waterfall instead of several disconnected tools. The bounce-rate improvement is the part most relevant to this guide: it's a direct downstream signal that field-level cascades and confidence thresholds, not just having "a" data source, were doing their job.

Sign up for Unify to see how a waterfall built into your prospecting workflow compares to the stack you're running today.

Should You Build or Buy Your Waterfall Enrichment Infrastructure?

Building this architecture from scratch is a real option, not just a rhetorical setup for "buy instead." Whether it's the right one depends on how differentiated your enrichment logic actually needs to be. Use these as decision points rather than a hard rule:

  • If your enrichment needs are standard (email, phone, firmographic data against a common ICP), prioritize buying: this is commodity infrastructure, and the value you'd add by building it yourself is low relative to the maintenance cost.
  • If you have unique, proprietary conflict-resolution logic tied to your specific business model, building may be worth evaluating, but budget for ongoing maintenance, not just the initial build.
  • If you're a lean team (one or two people own GTM data), prioritize buying: per Unify's build-versus-buy research, the costs teams plan for up front are typically a small share of what a homegrown system actually costs once scope creep, bug fixes, and single-person key risk show up later.
  • If you're already paying for three or more disconnected point tools to cover prospecting, enrichment, and sequencing separately, prioritize consolidating onto one platform before optimizing any single layer in isolation.
  • If your provider mix needs to change every time a vendor updates its API, prioritize a platform that owns those integrations, so your team isn't the one absorbing every vendor-side change.
  • If compliance requirements (GDPR, CCPA) require precise, field-level attribution for every enriched value, prioritize whichever option, built or bought, gives you that attribution by default rather than as an afterthought.

What Are the Most Common Waterfall Enrichment Setup Mistakes?

Even teams that understand the architecture tend to make a recurring set of implementation mistakes. Knowing them before you build saves weeks of debugging later.

  • Using record-level cascades instead of field-level cascades. A record-level cascade re-queries every field from the next provider even if only one field was missing, wasting calls on data you already have. Field-level cascades only call for the specific field that's missing.
  • Not attributing enriched data back to its source provider. Without provider attribution and a sourced-at timestamp on every field, you can't diagnose which vendor is causing bounces, can't measure real match rate per provider, and can't fulfill a GDPR or CCPA deletion request across every source you queried.
  • Setting confidence thresholds too low to save on cascade calls. Lowering thresholds to cut API costs looks like a savings until the downstream cost shows up as bounces, domain reputation damage, and wasted rep time on bad phone numbers. The apparent savings at the enrichment layer usually creates a larger cost in deliverability and pipeline.
  • Never re-enriching existing records. Treating enrichment as a one-time onboarding step ignores that B2B contact data decays roughly 22.5% per year. A record enriched 18 months ago has roughly a 32% chance of having a materially wrong field today, based on that decay rate compounding over a year and a half.
  • Mixing tools without a single source of truth for enriched data. When enrichment, prospecting, and sequencing live in separate tools that don't share attribution, nobody can tell which layer is actually responsible when quality drops.

Decision Framework: What Should You Prioritize First?

Use this to decide where to start if you're building or fixing a waterfall enrichment workflow this quarter:

  • If your bounce rate is already above 4-5%, prioritize confidence thresholds and provider attribution before anything else: you have a live deliverability problem, not a coverage problem.
  • If your match rate is fine but costs feel high, prioritize pre-call field checks and coverage-aware routing: you're likely paying for data you already have.
  • If you're a PLG motion on a self-serve CRM with a small team, prioritize a platform that bundles enrichment, prospecting, and sequencing rather than assembling point tools.
  • If you're sales-led with a dedicated RevOps function, prioritize field-level attribution and a documented conflict-resolution model, since you have the headcount to maintain it and the scale where conflicts compound.
  • If you sell into the EU or handle any GDPR-covered contacts, prioritize provider attribution and a documented legitimate-interest basis before scaling volume.
  • If you've never run a coverage audit on your current providers, do that before changing anything else: you may not have a cascade problem at all, just the wrong primary provider for your ICP.

Role and Segment Variants

The core five-layer framework holds across contexts, but where you invest first shifts by motion, size, and region:

  • PLG motion: Weight product-usage and website-intent signals into your triggering logic ahead of firmographic depth. Speed from signal to enriched contact matters more than exhaustive coverage, since the signal itself already qualified the account.
  • Sales-led motion: Weight firmographic and technographic depth higher, since account-based targeting depends on precision at the company level before a rep ever touches a contact record. Slower cascades with a full three-tier stack are worth the extra latency here.
  • SMB-focused teams: A single strong primary provider plus tight cost controls usually beats a three-tier cascade. The added complexity of a full waterfall rarely pays for itself at SMB volumes and price points.
  • Enterprise-focused teams: Invest in the full three-tier cascade and source reliability scoring. The cost of a bad record at a strategic, named account is high enough to justify the added complexity and re-enrichment frequency.
  • US teams: CCPA deletion requests are the main compliance driver, so provider attribution matters most for fulfilling "right to know" and deletion requests quickly across every source queried.
  • EU and UK teams: Legitimate interest is the usual legal basis for B2B enrichment, which requires a documented balancing test and an easy opt-out built into your sequencing layer, not just your enrichment layer. Budget extra review time before scaling volume in-region.

Edge Cases and Disambiguation

  • Job-seeker data versus buyer data. A contact updating their profile because they're job-hunting is not the same as a buying signal. Filter enrichment triggers tied to profile activity against actual role tenure and company-side signals, not profile edits alone.
  • Technographic false positives. A tool mentioned in an old job posting or case study is not proof a company currently uses it. Weight technographic signals by recency and corroborate with a second source before treating them as a targeting trigger.
  • Opens-only versus genuine engagement. A high open rate on enriched contacts can reflect deliverability working, not interest. Don't confuse a healthy waterfall (low bounce, high open) with message-market fit; they're measuring different things.
  • Stale "current employer" fields. Data decay means a contact's listed employer can lag their actual one by months. Cross-check job-change signals against enrichment refresh dates before enrolling a contact in an account-based sequence tied to their old employer.
  • Consent versus legitimate interest. Enriching a record is not the same as having permission to contact someone under every regional rule. Confirm your legal basis (legitimate interest in the EU, opt-out compliance under CCPA) as a separate step from the enrichment cascade itself.

Stop Rules and Red Flags

Signal Next action Wait time
Bounce rate rises above 4-5% on enriched contacts Pause sends from that segment, raise the confidence threshold, re-verify the batch Immediate
A single provider's match rate drops sharply quarter over quarter Move it down the cascade or pull it pending a coverage re-audit Before next send cycle
Deletion or opt-out request received Purge the record across every provider cache and your CRM, not just the sending tool Immediate and permanent
A record hasn't been re-enriched in over 180 days and is in an active sequence Force a re-enrichment pass before continuing outreach Before next touch
Two providers disagree on a field with no clear recency or reliability winner Flag for manual review rather than auto-resolving on a coin-flip rule Within 24 hours

Frequently Asked Questions About Waterfall Enrichment

What is a waterfall enrichment workflow?

A waterfall enrichment workflow queries B2B data providers in a defined priority order, moving to the next provider only when the previous one returns a blank or low-confidence value for a given field. It's the opposite of parallel multi-source enrichment, which queries every provider at once and pays for every response whether you needed it or not. Waterfall logic reduces spend because expensive or specialized providers are only called for the records your cheaper primary source couldn't cover.

How do you decide the order of providers in a waterfall?

Order providers by match rate for your specific ICP relative to cost per record, not by brand recognition. Your highest-coverage, lowest-cost provider for your exact ICP goes first. Higher-cost or specialized providers, such as those with live-verified mobile numbers or region-specific compliance coverage, belong second or third, reserved for records the primary couldn't cover. Run a sample of your own records through each candidate provider's API before locking in the order, since advertised coverage rarely matches your actual ICP coverage.

What is a confidence threshold and why does it matter?

A confidence threshold is the minimum quality score a returned value must meet before your cascade stops calling providers and accepts the value. Without a threshold, cascade logic is binary, which lets through syntactically valid but low-quality emails that bounce and unverified phone numbers that waste rep time. Most teams start with a stricter threshold for email and phone, since a bad value there carries a real cost, and a looser one for firmographic fields like headcount, where a slightly stale value is a smaller problem.

How do you handle conflicts when two providers return different values?

Default to recency: B2B contact data decays at roughly 22.5% per year according to ZoomInfo's data decay research, so a more recently sourced or verified value is statistically more likely correct. For conflicts that keep recurring on the same field, layer on source reliability scoring: track each provider's field-level accuracy against your own deliverability outcomes, then weight that against how recently the data was sourced. The highest composite score wins.

How often should you re-enrich existing CRM records?

Tier re-enrichment by account priority rather than one schedule for the whole database. Named, high-value accounts your reps are actively working benefit from re-enrichment roughly every 30 days with your full stack. Engaged accounts in active sequences can run every 60 to 90 days. Long-tail accounts with no recent engagement usually only need a quarterly refresh with your cheapest primary source.

What are the most common mistakes when setting up a waterfall?

Four mistakes account for most failed implementations: record-level cascades instead of field-level cascades, which re-queries every provider for fields you already have; not attributing enriched data back to its source provider, which makes it impossible to diagnose bounces or fulfill deletion requests; setting confidence thresholds too low just to reduce cascade calls; and treating enrichment as a one-time onboarding step instead of a recurring, tiered cycle.

Should you build a waterfall enrichment workflow in-house or use a platform?

It depends on whether enrichment is a differentiated capability for your team or commodity infrastructure. A custom-built waterfall requires ongoing engineering attention as provider APIs change, and per Unify's build-versus-buy research, the costs teams plan for up front are typically a small share of what a homegrown system actually costs once maintenance and scope creep show up later. A platform like Unify handles provider integrations, field-state checks, coverage routing, conflict resolution, and monitoring natively, which is usually the better trade unless your logic is genuinely unique to your business.

Is waterfall enrichment compliant with GDPR and CCPA?

Waterfall enrichment is a technical pattern, not a legal basis, so compliance depends on how you use it. In the EU and UK, enriching a contact for B2B outbound generally relies on legitimate interest, which requires a documented balancing test and an easy opt-out. In the US, CCPA gives California residents the right to know what was collected and to request deletion, which is exactly why provider attribution (knowing where every field came from) matters: it's what makes a deletion request actually executable across every source you queried.

Glossary

  • Waterfall enrichment: Querying B2B data providers in priority order, moving to the next one only when the previous returns a blank or low-confidence value.
  • Cascade logic: The rules that decide when a workflow stops querying one provider and moves to the next.
  • Confidence threshold: The minimum quality score a returned value must clear before the cascade accepts it and stops querying further.
  • Field-level cascade: A cascade that evaluates and queries for each field independently, rather than moving the whole record to the next provider at once.
  • Record-level cascade: A less efficient cascade that re-queries an entire record from the next provider even if only one field was missing.
  • Source reliability scoring: Rating each provider's accuracy for a given field based on observed deliverability, used to resolve conflicts alongside recency.
  • Match rate: The percentage of records a provider or workflow successfully returns a usable value for.
  • Data decay: The rate at which contact and company data goes stale as people change jobs, roles, and contact details.
  • ICP (Ideal Customer Profile): The firmographic and behavioral profile of the accounts most likely to buy, used to benchmark provider coverage against your actual targets rather than generic averages.

Common Mistakes to Avoid

  • Choosing providers by brand reputation instead of a coverage audit against your own ICP.
  • Running record-level cascades that re-query fields you already have good data for.
  • Lowering confidence thresholds to save on API costs without measuring the downstream bounce impact.
  • Treating enrichment as a one-time setup step instead of a tiered, recurring cycle.
  • Skipping provider attribution, which breaks both deliverability diagnostics and compliance requests.

Sources

About the Author
Austin Hughes is Co-Founder and CEO of Unify, outbound AI for sellers where AI agents and reps work side by side, from finding the buyers already in market to reaching them with the right message. Before founding Unify, Austin led the growth team at Ramp, scaling it from 1 to 25+ people and building a product-led, experiment-driven GTM motion. Prior to Ramp, he worked at SoftBank Investment Advisers and Centerview Partners.