Sales intelligence data collection is the continuous process of sourcing B2B contact and company records from many independent inputs, verifying them against live signals, and resolving them into a single usable record. It is not a one-time scrape. The quality of every outbound campaign you run is decided here, months before anyone writes a subject line.
Most teams evaluate a data provider on the size of its database. That is the least informative number available. A 500-million-record database assembled once and left to age is worth less than a 50-million-record database that re-verifies on access. What matters is where the data comes from, how often it is checked, and what happens when it is wrong.
What is sales intelligence data collection?
It is three distinct jobs that most buyers collapse into one word:
- Sourcing — finding the raw record in the first place, across public filings, crawled pages, job boards, contributory networks, and licensed feeds.
- Verification — confirming that the record is still true right now, not that it was true when collected.
- Enrichment and resolution — merging fragments from different sources into one person or one company, without creating duplicates or stitching two different people together.
A provider can be excellent at one and poor at the others. Vendors with enormous coverage often source aggressively and verify lazily. Vendors with pristine accuracy often verify well but cover a narrow slice of the market. Understanding which trade-off you are buying is the whole exercise. For the broader picture of how this data gets used once collected, see our guide to what sales intelligence is.
Where does B2B contact data actually come from?
No single source covers the B2B universe. Every serious provider blends several, and the mix determines where their coverage is strong and where it quietly fails.
| Source | What it yields | Weakness |
|---|---|---|
| Public web crawling | Company details, team pages, tech stack signals | Surface-level; misses anything not published |
| Company registries and filings | Legal entity, size, structure, ownership | Slow to update; no contact-level detail |
| Job postings | Hiring intent, internal tooling, team growth | Inferred, not stated |
| Technographic detection | Scripts, tags, and vendors in use | Only sees client-side and public technology |
| Contributory networks | Contact-level data pooled by participating users | Coverage skews to the network’s own user base |
| Licensed and partner feeds | Depth in specific regions or verticals | Cost; provenance depends on the partner |
| First-party signals | Your own site visits, product usage, replies | Only covers accounts that already found you |
| Human research | Accuracy on high-value or obscure accounts | Does not scale |
The practical consequence: ask a vendor which of these they lean on hardest, then test their coverage on the segment you actually sell into. A provider strong in North American SaaS may be close to useless for European manufacturing.
How is B2B data verified before it reaches you?
Verification is where providers differ most, and where the difference is least visible at purchase time. Four methods are in common use:
- SMTP validation — asking the receiving mail server whether a mailbox exists. Fast and cheap, but neutralised by catch-all domains that accept everything.
- Multi-source cross-referencing — treating a record as confirmed only when independent sources agree. Slower, considerably more reliable.
- Recency and confidence scoring — attaching a “last confirmed” timestamp and a confidence value, so stale records can be filtered rather than silently served.
- Human spot-checking — sampling records manually to catch systematic errors that automation reproduces at scale.
The critical question is not whether a provider verifies, but when. Verification performed at database-build time and stamped onto a record is close to meaningless six months later. Verification performed at the moment you access the record reflects reality. This distinction is the single largest driver of bounce rate, which we cover in detail in how to reduce email bounce rate.
What is waterfall enrichment, and does it work?
Waterfall enrichment queries data providers in sequence rather than relying on one. If the first and cheapest source cannot fill a field, the request falls through to the next, and only then to a premium source. You pay for the expensive lookup only when the cheap one misses.
The coverage gain is substantial and well documented:
| Field | Single source | Waterfall |
|---|---|---|
| Business email | 55–65% match rate | 88–95% fill rate |
| Mobile phone | 40–55% match rate | 70–85% fill rate |
| Account-level match | ~50% in vendor benchmarks | ~94% in vendor benchmarks |
The trade-offs are real, though. Sequential lookups add latency, so waterfall is a poor fit for real-time in-session enrichment. Costs are harder to forecast because they depend on hit rates. And merging records from providers who disagree with each other creates identity-resolution problems: two sources returning different job titles for the same person is common, and something has to decide which wins.
Treat the published fill rates as ceilings measured on favourable data, not as guarantees for your list.
How are buying signals collected?
Contact data tells you who to reach. Signal data tells you when. They are collected by completely different mechanisms, and the mechanism determines how much you should trust the output.
First-party signals
Website visits, pricing page views, content downloads, product usage. The highest-fidelity intent data available, because you observed it directly. Its limitation is reach: it only ever covers accounts that have already found you.
Content co-op data
Consortia of B2B publishers and trade sites pool their first-party behavioural data. Consumption is anonymised at the individual level and aggregated to the account using reverse-IP resolution. The system establishes a historical baseline per account per topic, then flags a “surge” when activity exceeds that baseline by a statistically significant margin.
Because it compares against a baseline rather than counting keywords, co-op data produces fewer false positives than bidstream.
Bidstream data
Captured from programmatic advertising auctions. Every ad request carries the URL of the page being loaded; aggregate millions of these and you get a picture of what is being read across the open web. It is broad and cheap, but keyword-driven and prone to false positives — someone reading an article is not necessarily evaluating a purchase.
Hiring signals
Job postings are collected at scale and parsed with language models to infer both what a company runs and where it is investing. Hiring is among the more reliable leading indicators in B2B, for a simple reason: companies hire before they buy.
- SDR and AE hiring precedes sales tooling purchases
- Engineering surges precede infrastructure spend
- Security and compliance hires precede governance tool evaluation
Job-posting analysis also surfaces backend and internal systems that a website crawler physically cannot detect, which makes it a useful complement to technographic scanning.
Technographic detection
Crawling for scripts, tags, and DNS records reveals the client-side stack. Additions and removals are the signal worth watching — a company that just dropped a vendor has budget and an unmet need simultaneously.
Why does B2B data decay faster than teams expect?
Because B2B records are tied to employment, and employment is unstable. Roughly 15–20% of professionals change jobs each year, and the widely cited HubSpot decay figure is 2.1% per month, compounding to about 22.5% annually. Recent studies of live databases report faster rates still.
Email addresses decay fastest of any field, precisely because they are the field most tightly coupled to a job. A direct dial follows close behind.
This reframes what you are buying. A database is not an asset that holds its value; it is inventory with a shelf life. Refresh cadence is a product feature, not an operational detail — and it is the feature least likely to appear on a pricing page.
What makes B2B data collection compliant?
Compliance is a collection-side property. By the time a record reaches your CRM, its lawful basis has already been established or already been skipped.
Under GDPR, cold B2B outreach generally relies on legitimate interest (Article 6(1)(f)) rather than consent. That basis is conditional: it requires a documented balancing test, and it holds only when your message is genuinely relevant to the recipient’s role. Selling supply-chain software to a Head of Supply Chain inside your ICP passes comfortably. Blasting the same message to every marketing manager you could find does not.
In the United States, the picture changed in January 2023 when the CCPA’s B2B exemption expired. Under CPRA, California residents’ contact data is covered regardless of whether it is used in a business context.
What auditors examine is provenance: where the record originated, on what basis it was collected, and whether deletion requests propagate. Scraped, purchased, or “enriched from unknown sources” data is the fastest route to a complaint. If a provider cannot tell you where a record came from, that is the answer.
How do you evaluate a provider’s data collection?
Run this before signing, not after:
- Test on your ICP, not their sample. Take 100 accounts you already know well and measure match rate against reality.
- Ask when verification happens. At build time, or at access time? The answer predicts your bounce rate.
- Separate coverage from accuracy. Fill rate tells you how often a field is populated. It says nothing about whether the value is correct.
- Request the refresh cadence per field. Firmographics can age gracefully. Emails and direct dials cannot.
- Ask for provenance and deletion handling in writing.
- Check what happens when data is wrong. Credit returns on hard bounces signal that a vendor is willing to stand behind its own verification.
A vendor bake-off on 100 known records takes an afternoon and tells you more than any coverage claim in a deck.
How ZenBee handles collection and verification
ZenBee maintains 700M+ verified contacts across 35M+ companies, combined with live buying and hiring signals.
Verification happens at access, not at build. When you reveal a contact, the address is checked against multiple live sources at that moment. If it cannot be verified, it is not returned — you see “Could not find a business email address for this person” rather than a plausible guess.
Exports are filtered, not just flagged. Only verified records reach your CSV or CRM, with skipped contacts reported explicitly rather than passed through.
Credits return on hard bounces, and verification carries no separate charge.
Collection quality compounds downstream. It determines whether lookalike company search surfaces genuine matches, whether AI scoring in your funnel has anything reliable to learn from, and whether prospecting workflows run on fact or on decay.
Frequently asked questions
Where do B2B data providers get email addresses?
From a blend of public web crawling, company registries, job postings, contributory networks where users share data in exchange for access, licensed partner feeds, and pattern inference validated against live mail servers. Any provider relying on a single source will have visible coverage gaps.
Is buying B2B contact data legal?
In most jurisdictions, yes, provided the data was lawfully collected and your use has a valid basis. Under GDPR that is usually legitimate interest, which requires a documented balancing test and role-relevant messaging. Legality depends far more on provenance and use than on the act of purchase.
What is the difference between intent data and hiring signals?
Intent data measures research behaviour — what an account is reading. Hiring signals measure operational commitment — where an account is spending headcount. Hiring is generally the stronger signal because it reflects a budget decision already made.
How often should B2B data be refreshed?
Contact-level fields should be re-verified before every send, and anything older than 60 days should be treated as unverified. Firmographic data tolerates a quarterly cycle.
What is a good match rate?
Single-source providers typically return 55–65% on business email and 40–55% on mobile. Multi-source waterfall approaches reach 88–95% and 70–85% respectively. Measure it on your own ICP — vendor-published rates reflect favourable test sets.
Does more data mean better data?
No. Database size measures sourcing effort, not accuracy. A smaller database verified at the point of access will outperform a larger one verified at build time and left to decay.
The takeaway
Data collection is not a background process that happens before the interesting work starts. It sets the ceiling on everything downstream: which accounts you can find, whether your email arrives, whether your scoring models learn anything true, and whether your outreach is defensible when someone asks where the record came from.
Judge providers on verification timing, refresh cadence, and provenance — not on database size.
Book a demo to test ZenBee’s coverage against your own target accounts, or review pricing. More on this subject in our B2B Data & Sales Intelligence topic hub.