Pivix Logo
Back to glossary

Predictive Audience

A predictive audience is a segment built by machine-learning models that forecast which users are most likely to take a specific action, such as buying, upgrading, or churning.

Key takeaways

  • Define the label with an explicit time window before choosing any features.
  • Features must be observable before the outcome window opens, or the model leaks.
  • Rare outcomes make accuracy meaningless; compare lift over the base rate instead.
  • The threshold trades precision against reach and is a business decision.
  • A high probability may mark someone who would convert without any spend.

In depth

Every predictive audience has two halves that must be defined before any modelling happens: a label and a feature set. The label is the outcome event bounded by a time window, such as upgrading within fourteen days, and the features are only those signals observable before that window opened. The model is fitted on a historical period, then applied to current users to produce a probability for each one. A threshold on that probability turns a continuous ranking into a membership list, and moving the threshold is a business decision, not a technical one.

Three things govern quality: how many positive examples exist, how fresh the features are when scoring runs, and how the label is defined. Because the target event is usually rare, overall accuracy is meaningless and lift over the base rate is the only useful comparison. Tightening the threshold raises precision and shrinks reach; loosening it does the reverse, and the right point depends on what a wasted contact costs versus a missed one. Acting on scores also changes future outcomes, so training data slowly absorbs your own interventions.

In practice the score is used to allocate scarce attention: which trials get a human touch, which accounts get a save offer, which contacts are suppressed from paid retargeting because they were converting anyway. Bands work better than a single cutoff, with a different treatment per band. Quiz answers make unusually good features because they are declared, structured and cheap to collect, and the score a scorecard already computes gives a readable layer to compare against the model when the two disagree.

Predictive audiences need history, so they fail on cold starts: a new product, a new market or a rarely occurring high-value event leaves too few positive examples to learn from. They also predict what correlated in the past, so any deliberate change in pricing, positioning or channel mix quietly invalidates the model while it keeps producing confident scores. The deeper limit is that likelihood is not incrementality. A high score often marks someone who would have converted unaided, and spending on them buys nothing.

Example in practice

A subscription fitness app used its Pivix onboarding quiz, which captured goals, current activity level, and budget, as labeled features for a predictive model. The model flagged a "likely to upgrade within 14 days" audience of about 2,300 free users, and the lifecycle team sent them a targeted annual-plan offer. That cohort converted at 11% versus 3% for an untargeted blast, roughly tripling upgrade revenue from the campaign.

How to measure it

Start with lift by decile: sort scored users into ten groups and compare the actual outcome rate in each. A working model shows a steep drop from top decile to bottom; a flat curve means the features carry no signal. Compare the top decile's rate against the overall base rate to express the model's value in one number, and recheck it on data from after the model was deployed rather than on the training period.

Then measure incrementality, which lift cannot show. Randomly withhold treatment from part of the predicted audience and compare outcomes between treated and untreated members of the same band. The difference is what the campaign actually caused. Track calibration too: if the model says twenty percent and the observed rate is five, the ranking may still be useful but any forecast built on those probabilities will be wrong.

Common mistakes

The classic failure is leakage: a feature that only exists because the outcome already happened. Including a field the sales team fills in after a deal closes, or a usage metric recorded post-upgrade, produces a model that looks near-perfect in validation and collapses in production. Guard against it by asking, for every feature, whether its value would have been known at scoring time. Anything that would not have been visible then gets removed, however predictive it appears.

The second failure is treating a score as a reason. Teams see a churn-risk audience, send a discount, watch churn fall, and conclude the model worked, when the same discount sent to random accounts might have done as much. Hold out a portion of every predicted audience and leave it untreated. Without that control you cannot separate a model that finds the right people from a treatment that works on everyone.

Frequently asked questions

What data do predictive audiences need to work well?

They need a clear outcome to predict and enough historical examples of that outcome, plus features like behavior, firmographics, and engagement signals. Clean, labeled data, such as scored quiz responses, dramatically improves accuracy.

How much data does a predictive audience need?

It depends on positive examples, not total records. A database of a million contacts with forty upgrades cannot train a reliable upgrade model, while a smaller list with thousands of conversions can. As a rule of thumb, you want enough positive cases that they remain plentiful after splitting into training and validation sets. Below that, simple rules usually outperform a model.

How is a predictive audience different from a lookalike audience?

A predictive audience ranks people already known to you by their likelihood of a specific future action, using your own data and label. A lookalike ranks strangers inside an ad platform by resemblance to a seed you supply. One prioritises contacts you can already reach; the other finds new people to reach. They also fail differently, since only the predictive model can leak future information.

Can I trust a predictive audience without a data science team?

You can use one, provided you check three things yourself: what outcome it predicts, over what window, and whether it beats the base rate on recent data. Packaged predictive audiences in marketing platforms hide the model but usually expose these facts. If the vendor cannot state the label and window plainly, treat the segment as a heuristic rather than a prediction.

How often should a predictive model be retrained?

Retrain on a schedule matched to how fast your inputs change, and immediately after any deliberate shift in pricing, product or channel mix. Scores can be recomputed daily while the model itself is refitted far less often. The signal to retrain early is drift: the distribution of scores or of key features moves noticeably compared with the training period.

What should you do with a low-probability segment?

Do not automatically abandon it. Low scores can mean genuinely poor fit, or simply that you have little data on those people. Split the two: contacts with rich histories and low scores are safe to deprioritise, while sparse, newly acquired contacts deserve a cheap qualifying step first. A short quiz or a single profiling question often converts an unscoreable contact into a scoreable one.

Which signals are most useful as model features?

Recent behaviour usually outranks static attributes: actions taken in the last few sessions, frequency and recency of engagement, and progress through onboarding steps. Declared data such as role, company size and quiz answers adds context that behaviour alone misses. Avoid features that depend on your own outreach, since they mostly record who your team already decided to contact.

Related terms

Turn glossary theory into qualified leads

Build a scorecard quiz funnel that qualifies and captures leads in minutes — no code required.

Start for free
  • No credit card
  • Free plan
  • Launch in minutes