Pivix Logo
Back to glossary

Lookalike Modeling

Lookalike modeling is a technique that analyzes the traits of a known group of valuable customers and finds new prospects who share similar characteristics.

Key takeaways

  • The percentage you pick is a cutoff on a ranked list, not a similarity threshold.
  • Only the matched portion of an uploaded seed actually trains the model.
  • A mixed seed blends distinct patterns and produces an audience resembling no one specific.
  • Value-weighted seeds tell the model which customer patterns are worth more.
  • Lookalikes reproduce existing base skew and cannot discover untried segments.

In depth

Under the hood the platform builds a feature vector for every user it knows about, fits a model that separates your seed members from everyone else, and then ranks the remaining population by that score. The percentage you select is a cutoff on that ranked list rather than a similarity threshold: a one percent audience is simply the top slice of the country's users by model score. Only the portion of your seed the platform can match to real accounts contributes, so upload size and match rate are two different numbers.

Seed size, homogeneity and recency pull against each other. A seed too small leaves the model fitting noise; a seed padded with every buyer you ever had mixes several distinct patterns and produces an audience that resembles no one in particular. Widening the percentage cutoff buys reach at the cost of precision, and each step out dilutes toward the general market. Value-weighted seeds, where each member carries a revenue figure, usually beat flat lists because they tell the model which patterns are worth more.

Practically, teams build a small ladder of seeds: closed-won accounts, high-lifetime-value customers, and recent high-intent leads, then test narrow and broad cutoffs against each. Existing customers are excluded from delivery so budget goes to acquisition, and extra demographic layers are kept light because heavy manual targeting fights the model it sits on. A scorecard quiz is a convenient seed source, since only leads above a scoring threshold are exported, which keeps the seed intentionally narrow rather than accepting every form fill.

A lookalike can only find people who resemble whoever is already in your seed, so it reinforces the shape of your current base and will not discover a segment you have never sold to. Any accidental skew in that base, a channel that over-delivered one region or one job family, is reproduced and amplified at scale. The model is also bounded by what the platform observes about its own users, which is opaque and shifts without notice, and in narrow B2B markets the qualifying population may simply be too small to rank meaningfully.

Example in practice

Imagine a B2B analytics company that exports its top 500 closed-won accounts from the prior year and uploads them as a seed to LinkedIn to build a 1% lookalike audience. They pair this with a Pivix qualification quiz so only leads scoring above 70 feed the seed, keeping it clean. Cost per qualified lead might then fall from $140 to around $86 over two months as the lookalike audience consistently matches their proven ICP.

How to measure it

Judge a lookalike on downstream quality, not on click metrics. Compare cost per qualified lead and win rate against a broad-targeting control running the same creative, because a lookalike that only improves click-through is selecting for curiosity. Track the share of leads from the audience that match your fit criteria; if that share is no higher than from untargeted traffic, the seed is not carrying the signal you assumed.

Watch decay over the life of the audience. Model scores are recomputed as the platform learns, but your seed ages, so refresh it on a regular cadence and compare performance before and after. Frequency and reach also tell you when a narrow cutoff has exhausted itself: rising frequency with flat conversion means the audience is saturated and widening the percentage is cheaper than raising bids.

Common mistakes

The most frequent error is seeding on the cheapest available list, usually all newsletter subscribers or every form fill. The model then learns what a content downloader looks like and the resulting audience converts to content rather than to revenue. Seed on the outcome you actually want more of, closed-won accounts or customers past a retention milestone, even when that shrinks the list. A smaller seed built on the right outcome beats a large one built on availability.

The second error is stacking heavy manual targeting on top of a lookalike, adding job titles, interests and narrow age bands. Each layer removes people the model ranked highly for reasons you cannot see, and the audience shrinks until delivery becomes expensive. Use exclusions freely, since removing current customers and recent converters is unambiguous, but add positive filters one at a time and only when a test shows the unfiltered audience underperforms.

Frequently asked questions

How large should a seed audience be for lookalike modeling?

Large enough for the model to find a stable pattern, which usually means at least a few thousand matched records rather than a few hundred. What matters more is that the members share a genuine pattern. If you must choose, a tight seed at the lower end of the size range outperforms a large one assembled by combining unrelated groups just to clear a minimum.

Is lookalike modeling the same as predictive scoring?

No. Lookalike modeling ranks an outside population by resemblance to a seed you supply, and runs inside an ad platform on its data. Predictive scoring ranks people already in your database by their likelihood of a specific future action, using your data. One finds new prospects to reach; the other decides who among known contacts deserves attention first.

How often should a lookalike audience be refreshed?

Refresh the seed whenever a meaningful share of it has aged out of relevance, which for most teams means monthly or quarterly. Platforms recompute scores continuously, but they cannot correct for a seed describing last year's customers. Refresh sooner after any change that alters who buys from you, such as a new pricing tier, a new market or a repositioning.

What percentage lookalike should I start with?

Start narrow, typically the smallest available slice, and widen only when delivery stalls or costs rise from saturation. The narrow tier tests whether the seed carries signal at all; if it fails there, wider tiers will not rescue it. Once the narrow audience performs and exhausts, stepping out one tier at a time shows exactly where precision starts falling away.

Why is my lookalike audience performing worse than broad targeting?

Usually the seed is wrong or the audience is over-layered. A seed of general sign-ups teaches the model to find more sign-ups, not buyers, and stacked demographic filters strip out the people it ranked highest. Modern broad targeting is also strong, so a lookalike must beat it on qualified pipeline, not impressions. Test both with identical creative before concluding.

Can lookalike modeling work without an ad platform?

Yes. The same logic applies to your own database: take your best customers as a seed, identify the attributes that separate them, and score existing contacts or an enrichment list by resemblance. This is slower and needs data of your own, but the scoring is transparent and you can inspect which attributes drive it, which platform lookalikes never show you.

Related terms

Turn glossary theory into qualified leads

Build a scorecard quiz funnel that qualifies and captures leads in minutes — no code required.

Start for free
  • No credit card
  • Free plan
  • Launch in minutes