Your next-best-action is optimizing applause, not prescriptions
In March 2026, IQVIA unveiled a unified agentic-AI platform, reporting more than 150 intelligent agents deployed and 19 of the top 20 pharma companies already incorporating its agents into their workflows. ZS followed in May with agentic launches of its own. Within twelve months, the agentic announcement has become the category’s standard press release, and next-best-action (NBA), the engine that decides which doctor a rep should see next and with what, is the flagship use case.
The engines are getting faster. The question that decides whether they pay is quieter: what is the engine rewarded for while it learns? I have led a Next-Best-Action program whose lift was read against a control cohort, so I ask from the builder’s side of the table. Across the category, the honest answer is that the engine mostly learns from what it can observe: the accepted suggestion, the opened email, the completed call — applause. And applause is not a prescription.
What does an agentic next-best-action engine actually optimize?
The signal it can see: engagement. An NBA engine recommends the next interaction, then needs feedback to learn. The outcome it exists for, a change in prescribing, is slow, noisy and confounded. The click is instant. So the learning loop closes on what arrives fast: suggestion acceptance, dismissals, opens, call completions.
The vendors describe the loop this way themselves. Aktana, one of the category’s pioneers, writes that “high engagement not only leads to adherence of ‘next best actions,’ but also provides critical feedback to refine strategy,” and advises measuring “the acceptance, dismissal, and engagement rates” of suggestions. IQVIA’s 2025 white paper on commercial AI expects capabilities to include “triggering actions such as content delivery, account targeting, and sequencing with minimal human input”: that is where IQVIA says the capability is heading. No one has audited it yet.
The design choice is understandable: learning systems need fast feedback. It is also where the trouble starts: the fast signal and the true signal are not the same signal.
Why is engagement the wrong reward?
Because a learning system does not pursue your intention; it pursues its reward. The anthropologist Marilyn Strathern compressed the general law into one sentence: “When a measure becomes a target, it ceases to be a good measure.” Engagement was a reasonable measure of commercial attention — right up until the engine was told to maximize it.
AI research has a name for what happens next. Amodei and colleagues catalogued “reward hacking” in 2016: agents that game their objective instead of solving the task, and from the agent’s point of view this “is not a bug, but simply how the environment works.” The canonical demonstration is OpenAI’s boat-race agent: the game scored points for hitting targets along the route, so the boat found an isolated lagoon, circled it knocking over respawning targets (repeatedly catching fire, crashing into other boats) and scored about 20% higher than human players without ever finishing the race. The boat was not broken. The reward was.
Map that onto commercial decisioning:
| The reward the engine sees | What it learns to produce | What the business wanted |
|---|---|---|
| Suggestion acceptance | Suggestions reps find easy to accept | Interactions that change prescribing |
| Email opens and clicks | Better subject lines and send times | HCP behavior change |
| Call completion | Visits to accessible, friendly HCPs | Lift with the right HCPs |
| Channel engagement score | More touches on responsive channels | Incremental outcome per touch |
None of the left-column signals is worthless; they are diagnostics. They corrupt only when promoted to the objective.
The adherence trap: the engine learns the rep, not the market
The subtlest version of the failure closes the loop inside the sales force. A rep accepts the suggestions that are convenient (the familiar doctor, the comfortable route, the call that was going to happen anyway) and dismisses the awkward ones. The engine reads acceptance as success and adapts.
Aktana states this plainly: its learning module “prioritizes suggestions that are more likely to fit the reps preferences.” Over time the system converges on the field force’s habits, and the adoption dashboard describing that convergence points up and to the right.
I have watched this pattern from inside an NBA program. Adherence is the most flattering metric in the stack, because both sides of the loop are rewarded for agreeing with each other: the rep gets suggestions that are easy to act on, and the engine gets acceptance it can book as success. Nobody inside that loop is paid to notice that the prescriptions would have been written anyway.
Is this a pharma quirk?
No. Rewarding the observable proxy is the default failure of engagement optimization wherever it runs. Kleinberg, Mullainathan and Raghavan formalized it for recommender platforms: a platform that “simply wants to maximize user utility, but only observes user engagement” can provably make users worse off; they characterize exactly when “increasing engagement fails to increase user utility.”
Milli and colleagues then measured it in the field: Twitter’s engagement-based ranking “amplifies emotionally charged, out-group hostile content,” and “users do not prefer the political tweets selected by the algorithm.”
Marketing science has known its own version for a decade. When eBay finally ran the experiment on its paid search, brand-keyword ads, long credited by attribution models, showed “no measurable short-term benefits”; the clicks came from people who would have arrived anyway. Ascarza showed the same shape in churn programs: the customers a model flags as highest-risk “are not necessarily the best targets”: the ones worth treating are the ones the treatment actually moves. The same law shows up in every industry that has run the experiment: the measurable proxy drifts from the outcome the moment you optimize it.
How do you reward the outcome instead?
Point the objective at incrementality, and make the counterfactual part of operations rather than a post-mortem. Concretely: hold comparable physicians or territories out of an action, read the difference in outcomes, and let that difference, not acceptance, be the score the system answers to. Engagement metrics keep their place in the stack, as diagnostics.
That is the design we held on a Next-Best-Action and marketing-mix program I led at a top-5 pharma in the CIS. The engine’s actions were judged against a control cohort: +7% incremental Rx, read visited-versus-not-visited, indexed 100 → 107; the role was mine: AI lead and builder. That number holds for that cohort and that design, nothing wider. What made it defensible was a choice made before scaling: what the system would be rewarded for. The measurement discipline follows from that choice.
- Scope
- Next-Best-Action + marketing-mix, top-5 CIS pharma
- Baseline
- indexed 100, visited-vs-not-visited
- Role
- AI lead & builder
- Result
- +7% incremental Rx (100 → 107)
An engine rewarded on applause will learn applause. If you want prescriptions, pay it in prescriptions — measured against the doctors it never touched.
Doesn’t “twice as effective” settle it?
It is a vendor’s sentence, not a measurement you can inspect. ZS writes that “dynamic, targeted calls are twice as effective as other types of calls, leading to a 5%-10% lift in top-line brand sales” — with no methodology attached. Effective at what, measured against which counterfactual? The claim may even be true. The point is that a buyer cannot know, and an engine cannot learn the right thing, from a number with no control behind it.
The category’s own numbers recommend the skepticism. Gartner names the relabeling “agent washing” and estimates that only about 130 of the thousands of self-described agentic-AI vendors are real, from the same press release that expects over 40% of agentic-AI projects to be canceled by 2027. And ZS itself now sells test-and-control analytics for “moving beyond surface-level engagement metrics to understand true sales impact”: the vendors can see the gap too. When the pitch and the fine print disagree, read the fine print.
Where to start: fix the reward before you scale the engine
If an agentic NBA is on your roadmap or already running, the reward function is the first artifact worth auditing. Four moves, cheapest first:
- Write down what the engine is rewarded on today. The actual feedback signal in the loop, however unflattering. If the answer is acceptance, opens and completions, you are training for applause.
- Build the holdout into operations. A standing control of comparable HCPs, territories or accounts the engine leaves alone turns incrementality from a quarterly study into a live signal the system can learn against.
- Re-point the objective. Reward outcome lift against that control; engagement goes back to being telemetry. Expect adoption charts to look worse; the applause was subsidizing them.
- Interrogate every vendor lift claim for its counterfactual. One question does most of the work: compared to whom? A vendor who can answer has a measurement; a vendor who cannot has a brochure.
The engines get faster every quarter, which is exactly the reason to fix the objective now: a system that learns the wrong thing quickly gets very good at it.
Choosing what a commercial engine is paid for is operating-model work — the same discipline that held on the Next-Best-Action program measured against control. If your NBA dashboard is rich in adoption and silent on incrementality, let’s look at the reward function.