A lookalike audience, called a similar audience on some platforms, is a targeting method in which an advertising platform identifies people who resemble a group the advertiser supplies. The advertiser provides a seed audience, typically a customer list or a set of people who took a valuable action, and the platform's model finds others in its user base sharing similar characteristics and behaviour patterns.
The appeal is that it addresses the hardest problem in acquisition, which is finding people who do not yet know a business exists but would value it. Manual targeting relies on the advertiser guessing which demographics, interests, and behaviours predict a good customer, and those guesses are frequently wrong. A model trained on actual customers can identify predictive signals no marketer would have specified, including combinations that make no intuitive sense but hold empirically.
Seed quality determines everything about the output, and this is where most implementations fail. A model trained on all past purchasers will find people resembling the average customer, including the low-value and quickly churning ones. A model trained on the most valuable customers, defined by lifetime value, repeat purchase, or margin, finds people resembling those instead. The difference in downstream performance between these two seeds is frequently larger than the difference between any two targeting strategies.
Seed size involves a genuine trade-off. Platforms typically require a minimum, often around a thousand matched individuals, and models generally improve with more data up to a point. But a larger seed usually means a looser definition of a good customer, so the choice is between a precise model built on fewer examples and a broader one built on more. Testing several seeds, defined by different value thresholds, is more informative than reasoning about which should work.
The similarity percentage controls the trade-off between reach and precision. A tightly defined audience matching the top one percent of similarity is smaller and typically performs better per impression, while broader definitions reach more people at lower average quality. The right setting depends on whether the constraint is efficiency or volume, and running several tiers as separate campaigns, rather than blending them, is what makes the difference measurable.
Refreshing matters more than most advertisers expect. Seeds built from a customer list become stale as the customer base changes, and models built on a business's customers from two years ago will find people resembling a business the company has since moved on from. Updating seed audiences on a regular schedule, and rebuilding models when positioning or product mix changes materially, keeps the targeting aligned with the current business.
Data protection obligations apply directly, since uploading a customer list to an advertising platform is a processing activity requiring a lawful basis, appropriate disclosure, and respect for any objection or withdrawal. Platforms provide hashing to avoid transmitting identifiers in plain form, but hashing does not remove the obligation, and audiences built from data collected for an unrelated purpose are a common compliance gap.
Because the method depends entirely on knowing which customers are actually valuable, it fails quietly in organizations that cannot distinguish them. In practice the seed definition draws on the customer value analysis maintained through data analytics, the campaign execution sits with marketing services, and the resulting audience strategy is reviewed within growth management whenever the business changes which segment it is trying to grow.