Pivix Logo
Back to glossary

Lead Database

A lead database is the central, structured store of captured prospect records, including contact details, source, qualification scores, and engagement history.

Key takeaways

  • The identity rule decides whether a repeat email becomes one record or two.
  • Provenance on every record answers which capture created it and when.
  • Strict entry validation raises data quality and costs some genuine conversions.
  • Archive stale records instead of deleting them, so history survives for re-scoring.
  • Consent collected for one purpose does not automatically cover another use.

In depth

A lead database is defined by three things: which fields exist, what counts as the same person, and what is allowed to write to it. The identity rule does most of the work, because it decides whether an email that arrives twice creates one record with two events or two competing records. Most systems match on normalised email first, then on company domain and name, and keep a provenance field on every record so you can always say which capture created it and when.

Quality falls in two ways at once. Records decay because people change jobs and mailboxes are retired, and the database dilutes because every new intake adds records without removing any. Adding validation at entry raises quality but costs conversion, since a strict email check rejects some real people alongside the fake ones. Deleting aggressively keeps the base clean but destroys history that later scoring would have used. Most teams settle on archiving rather than deleting, and re-checking rather than trusting age alone.

Day to day the database is worked through saved segments rather than by browsing records. Useful segments combine a fit attribute, a recency window and a status, so a rep opens a list rather than a table. Records that arrive from a scorecard quiz land with the answers attached, which means the segment can be built on what the person said about their situation, not only on what channel they came from. Suppression lists sit beside the segments and stop already-contacted records reappearing.

Size is not an asset. A database of a hundred thousand records where only a few thousand match the current ideal customer profile is a storage cost with a compliance risk attached, not a pipeline. Consent also has a shelf life: a record collected for one purpose under one privacy notice cannot always be used for another, and that limit is legal rather than technical. Scores age too, so a fit rating from two years ago describes a company that may no longer exist in that shape.

Example in practice

A revops analyst audited a 50,000-record lead database and found 22% were duplicates or invalid. After merging records and tagging each remaining lead with its quiz tier and source, the sales team's email reply rate rose from 4% to 11% because reps could finally target by fit and recency.

How to measure it

Two ratios describe database health. Deliverable share is the number of records whose email still accepts mail divided by all active records; a falling figure means decay is outrunning intake. Addressable share is the number matching your current ideal customer profile divided by the total, and it is the honest measure of size. A database that grows while addressable share falls is accumulating cost.

Track duplicate creation rate as a monthly count of new records that later merge into an existing one; a rising number points at an intake path that skips the match rule. Add field completeness for the two or three attributes your segments actually use, measured only on records created in the last quarter, since older records will always look worse and mask a present-day problem.

Common mistakes

Teams commonly treat deduplication as a one-off cleanup project rather than a rule at the point of entry. The base gets merged, everyone celebrates, and duplicates reappear within a quarter because three forms still write records without checking. Set the match rule in the intake path, not in a spreadsheet afterwards, and log every merge so you can tell a genuine second person from a matching error.

The other failure is storing free text where a controlled list belongs. Job title typed by the visitor produces dozens of spellings of the same role, and no segment can be built on it reliably. Ask for role from a short list, store the raw text separately if you want it, and keep the field that segmentation depends on constrained. The same applies to company size and industry, which are worth constraining even at the cost of precision.

Frequently asked questions

What data should a lead database store?

Beyond contact details, capture source, qualification score, quiz tier, and engagement history. That context lets teams segment and prioritize rather than treating every record the same.

Is a lead database the same as a CRM?

Not quite. A lead database focuses on storing and organizing raw lead records, while a CRM manages the full relationship and sales process built on top of those records.

How often should a lead database be cleaned?

Continuously for duplicates and on a fixed cycle for decay. Duplicate prevention belongs in the intake path and runs every time a record is written. Decay checks are better done quarterly against records older than six months, verifying deliverability and archiving what fails. An annual clean-up project is too infrequent to stop the problem compounding.

Should I buy a lead list to fill my database?

Bought lists carry three problems: you did not collect consent, the records were sold to others too, and there is no engagement history to score against. If you use one, keep it in a separate segment with its own source tag so it never contaminates the metrics from leads you earned. Judge it on reply rate, not on record count.

How many fields should a lead record have?

Only those that change a decision. A field earns its place if it appears in a segment, a routing rule or a scoring model; otherwise it is a box someone has to fill and later nobody trusts. Most B2B teams need contact details, company size, role, source, qualification score and last activity, plus whatever the quiz answers add automatically.

What should happen to leads that never responded?

Move them to a dormant status rather than deleting them, and stop them appearing in active segments. Re-engage on a schedule with a low-cost touch, and archive after a defined number of failed attempts. Keep the record, because a contact who ignored you last year may respond when their circumstances change, and the history explains why the score is what it is.

Does a lead database need to be separate from the CRM?

For most small teams, no; the CRM's lead object is enough. A separate store starts to earn its cost when intake volume is high, when several capture sources need normalising before they reach sales, or when marketing needs to hold records the CRM should not see. In that case the CRM stays the system of record for anything a rep is working.

How do I keep unsubscribes and consent status accurate?

Store consent as a field on the record with a timestamp and the source of the permission, not as membership of a mailing list. Lists get rebuilt and copied; a field travels with the record. Check it at send time rather than at segment build time, since a person may withdraw permission between the two, and keep a suppression list that no segment can override.

Related terms

Turn glossary theory into qualified leads

Build a scorecard quiz funnel that qualifies and captures leads in minutes — no code required.

Start for free
  • No credit card
  • Free plan
  • Launch in minutes