Channel Weighting Models for AI-Era Media Mix
Conversational AI needs its own weighting math, not search's old playbook.

Media buyers are still pricing conversational AI as if it were search wearing a new coat, and that habit gets more expensive every quarter it survives. The channel is growing too fast, and behaves too differently, for the old weighting math to hold up. Fixing the weights isn't actually the hard part. The hard part, the one most teams skip, is fixing what feeds the weights before any number gets assigned.
How large and how fast the conversational AI ad channel is actually growing
ChatGPT ads hit a $1 billion annualized revenue run rate in under 200 days, with tens of thousands of advertisers buying across more than 40 countries, according to Beet.TV reporting from September 2026. WPP Media projects generative search ad revenue at $5.1 billion for the year, on a path past $100 billion by 2030. Traditional search took 22 years to cross that mark. Generative search is on pace to do it in six.
Most media plans still blur two environments that have nothing to do with each other. The bulk of AI ad spending in 2026 sits next to AI-generated content, things like Google AI Overviews, rather than inside an actual chatbot conversation. One is a summary box bolted onto a results page. The other is a live thread where the ad has to fit inside the model's own generated response. Blend them into one line item and the resulting weight fits neither one.
Attention has already moved past the infrastructure built to monetize it. ChatGPT counts 800 million weekly active users; Google Gemini sits at 650 million monthly users. Ad tooling is still catching up to that scale, and that lag is exactly why the weighting question needs an answer now, before market share settles into something that looks tidier than it actually is.
What makes conversational AI a distinct signal environment, not a search variant
Search runs on a keyword. A short query gets matched to a bid, the slot is fixed before the auction even opens, and the unit being measured is a click on a static result. None of that maps onto a conversation, and treating it as if it does is the first mistake most weighting models make.
In a chatbot, the signal is the whole thread: topic, tone, what's already been said, how close the user sits to a decision. Someone typing "which of these two laptops handles video editing better" hands over product category, comparison intent, and purchase proximity in one sentence no keyword list could ever encode. The ad decision doesn't sit next to the response, it's tangled up inside the generation of the response itself. Academic work on LLM auctions, including the LERA framework out of Peking University and Alibaba Group, surfaces this core tension: the auction has to weigh whether inserting an ad damages the coherence or tone of what the model says next.
That changes the targeting input from top to bottom. Keyword lists become conversational intent maps. Behavioral profiles built from past browsing become real-time entity signals read off the live chat. Product feeds have to become AI-readable knowledge graphs, or the model may struggle to surface them reliably. The auction architecture shifts too. Two-stage designs like LERA's, embedding-based filtering followed by LLM logit scoring, represent a fundamentally different architecture from the keyword-bid-CTR-slot pipeline search has run on for two decades. The platform is now publisher, ranker, and generator, all at once.
LLM auction research frames this as three stakeholders pulling in different directions: the advertiser wants CTR and ROI, the user wants a coherent and useful answer, the platform wants revenue that doesn't chip away at retention. That tension gets resolved inside the auction's design, not by a media planner setting a bid manually. Calling conversational AI "search with natural language" gets the targeting input wrong and the inventory quality wrong at the same time, and mispricing follows from both mistakes at once.
How current channel weighting models handle conversational AI, and where they fail
The baseline most planners work from in 2026 has paid search commanding the largest share of US digital budgets, followed by social, display and programmatic, video, and a small remainder in other channels. AI advertising barely shows up as its own line. Most models either fold it into "other" or bolt it onto search, and bolting it onto search is the worse of the two mistakes, because it inherits an entire KPI set that doesn't fit the thing it's measuring.
Three failure modes show up over and over. Misclassification comes first: because conversational AI responds to expressed intent, it gets filed as a search sub-channel and inherits search KPIs (keyword match rate, cost-per-click, quality score) none of which have a real counterpart inside a conversational auction. The wrong measurement proxy comes second: CPM and CPC both assume a predefined slot with a countable impression, but LLM ad placement has no fixed slot, and the standard transaction unit for conversational ads remains an open question. An attribution gap comes third: conversational sessions rarely produce a clean last click, since users often act on a recommendation several steps after the ad actually appeared. Any model that demands direct click attribution will undercount this channel every single time.
Platform fragmentation makes the problem worse, not better. The citation sources ChatGPT and Perplexity draw on overlap by only around 11% on the organic side alone, which points to how differently these systems index and retrieve content. Paid placement logic diverges just as much, so a single "AI channel" weight ends up averaging together environments that don't behave alike at all. Conflict-of-interest research from Princeton and the University of Washington in 2026 adds another wrinkle: LLMs don't surface sponsored content uniformly, so performance depends partly on which model is running the auction, not just on bid and relevance.
The weights aren't set too low. The inputs feeding the calculation are measuring the wrong thing, and no amount of adjusting the percentage fixes that on its own.
What inputs a channel weighting model actually needs for conversational AI
Four categories need rebuilding from scratch, not tweaking at the margins. Get the inputs wrong here and every downstream number is cosmetic, a precise answer to the wrong question.
Signal type has to move from keyword match rate to conversational intent classification: what decision stage the user is in, what entities show up in the thread, what action the conversation is implicitly pointing toward. Inventory taxonomy needs at least three separate buckets, because each behaves nothing like the others: ads adjacent to AI-generated summaries in an AI search summary style, ads embedded inside chatbot responses in the ChatGPT style, and agentic placements where the model acts on the user's behalf instead of just answering. Auction compatibility means understanding what the platform actually optimizes for, since a two-stage relevance-plus-bid system like LERA's crowns different winners than a keyword-bid-CTR model does, and bids calibrated for search will misprice conversational inventory without anyone noticing until the reporting comes back thin.
Attribution has to shift to a session-level or assisted-conversion view instead of last-click. Conversational journeys can run substantially shorter than traditional search journeys, and a model built on last-click reads that speed as underperformance instead of what it actually is: the channel doing its job faster than the measurement can follow.
Data feed readiness sits underneath all of it as a precondition, not a bonus feature. A brand whose product feed isn't structured for AI ingestion is likely to see significantly reduced visibility in conversation, regardless of what it's willing to bid on.
Platform divergence means the model needs inputs at the individual platform level, not one line item labeled "AI." OpenAI's ChatGPT has ads live now, with conversion tracking still rolling out. Google Gemini currently runs no ads, though Google has signaled that plans are coming. Perplexity walked away from advertising in February 2026, citing concerns about user trust. Anthropic runs no ads at all. That's a genuinely concentrated market, and the model should reflect that concentration instead of pretending the inventory sits evenly spread across four similar platforms. Gartner projects 60% of brands will use agentic AI for one-to-one interactions by 2028, which makes agentic placement a third inventory type needing its own logic entirely, separate from ads sitting inside a chatbot's response.
The brand safety and trust inputs that weighting models ignore at their peril
Search never had to solve this problem. A chatbot is simultaneously a trusted advisor and an ad delivery mechanism, and users tend to treat its suggestions as neutral, which is exactly what makes the arrangement risky.
Princeton and University of Washington research put numbers to that risk, and the numbers are not subtle. A majority of the LLMs tested favored company incentives over user welfare in conflict-of-interest scenarios. Grok 4.1 Fast recommended a sponsored product priced nearly double a comparable non-sponsored option in 83% of test cases. GPT 5.1 surfaced sponsored options in ways that disrupted the purchasing flow in 94% of cases. Behavior also shifted depending on a user's inferred socioeconomic status, a brand-safety variable with no equivalent anywhere in search.
None of that argues against buying the channel. It argues for pricing brand-suitability risk into the weight itself, the same way viewability and fraud risk already get baked into programmatic display weights. OpenAI has published ad policies requiring clear labeling and separation from answers, and the IAB Tech Lab formed its AI Content Monetization Protocols working group in August 2025 to start building cross-platform standards. That infrastructure is forming. It isn't finished, and a model that assumes it's finished is borrowing trust the market hasn't earned yet.
A model that treats conversational inventory as fungible with search inventory is quietly assuming brand-safety infrastructure that doesn't fully exist. Filter granularity, disclosure compliance, fill rate limited to genuinely commercial-intent prompts: these belong inside the weight calculation itself, not tucked into a post-campaign review deck nobody opens until the budget's already spent.
How to structure the updated weighting model in practice
Start by breaking the single "AI" line item into at least three channels with separate weights: AI-adjacent search, in-conversation chatbot ads, and agentic placements. Their mechanics, their measurement, and their risk profiles don't resemble each other closely enough to share one number, so don't force them to.
From there, rebuild the signal input channel by channel. AI-adjacent search sits closest to how traditional search already works, so existing keyword and CTR inputs mostly transfer with modification. In-conversation chatbot ads need conversational intent categories in place of keyword bids, session-level engagement units in place of slot-based impressions, and assisted conversion or shortened journey length as the primary read on performance. Agentic placements are the earliest and least measurable of the three, and treating them as anything more than a test budget right now is a mistake. Weight them conservatively, and treat the spend as qualitative learning until attribution standards catch up to the format.
Build platform-level weights inside the chatbot bucket specifically, because the concentration named above is real: ChatGPT is the primary addressable surface as of mid-2026, and a single blended "chatbot" weight hides that fact instead of accounting for it. Add a brand-safety multiplier as a weight modifier, giving platforms with published ad policies, disclosure rules, and third-party safety filtering a higher weight than platforms without any of that. And set the attribution window explicitly: session-level, or a seven-day assisted-conversion view, rather than last-click. The performance advantage conversational AI can deliver over traditional search only shows up if the measurement window is actually built to catch it.
Buying infrastructure matters just as much as the model itself, and this is where most teams undercut their own fix. Reaching conversational inventory across multiple surfaces requires a demand-side platform that can read conversational context at the moment the auction runs, not after the fact. A generalist DSP that can't parse prompt-level signal falls back to demographic or behavioral targeting, which throws away the one thing that makes this channel different from everything else in the plan. A single-surface network reads context fine but caps reach at one platform. The right answer is a DSP built specifically for AI surfaces, running its own exchange with direct publisher supply, one that combines context-reading with reach across surfaces instead of trading one off for the other.
None of this is a set-and-forget model. Chatbot ad spending growing at the rate it's growing makes an annual MMM cycle close to useless before the ink dries. Quarterly recalibration against platform-level availability and new measurement data is the floor here, not an aspiration to work toward eventually.
What remains genuinely unsolved in conversational AI measurement, and how to plan around it
Pricing is still an open question, and pretending otherwise is a mistake. Nobody knows yet whether CPM, CPC, or CPA becomes the standard transaction unit for conversational ads, which means a weighting model built around CPM efficiency comparisons risks comparing two things that won't even be the same kind of thing within two years.
Cross-session attribution is unsolved too. A recommendation surfaced in a conversation can influence a purchase made days later on an entirely different surface, with no traceable path connecting the two. Even a generous seven-day attribution window will undercount that kind of delayed action.
Auction transparency is limited by design right now, not by accident. Two-stage relevance-plus-bid systems like LERA exist in academic research, but commercial platforms haven't published anything close to a standard, so planners often can't audit why a given ad won or lost the way they can trace a search auction step by step. The IAB Tech Lab's CoMP working group, formed in August 2025, is the right body to watch for movement here. Standards groups move on multi-year timelines, though, and this channel isn't waiting around for them to finish.
Platform divergence also means benchmarks don't travel from one surface to another. Performance data pulled from one LLM won't reliably predict performance on a different one, and that 11% citation overlap between ChatGPT and Perplexity is a fair proxy for just how differently these systems get built under the hood.
The honest posture is to build the model with visible uncertainty ranges on every conversational AI weight, write down the measurement assumptions behind each one, and treat the whole thing as a hypothesis under test rather than a finished allocation. The structural advantages, prompt-level intent and shorter conversion paths, are real and already showing up in the data available today. The measurement infrastructure needed to fully capture them is still being built, and pricing the channel as though that infrastructure already exists is its own kind of mispricing.


