AI token prices are falling but enterprise AI bills are rising. Here's what that means for CEP strategy and agentic martech in Southeast Asia.
The headline sounds reassuring: AI inference costs have dropped from $20.00 per million tokens in late 2022 to $0.07 by mid-2025, per Stanford’s AI Index. So why are marketing teams opening their cloud invoices and quietly losing their minds?
The Jevons Paradox Is Running Your AI Budget
Barr Moses at Monte Carlo puts a precise name to what’s happening: consumption is outpacing unit cost reduction. When something gets cheaper, organisations use dramatically more of it — a dynamic economists have called the Jevons Paradox since the 19th century. In martech terms, this plays out as follows: your team deploys one AI agent to personalise email subject lines, another to score inbound leads, a third to generate campaign briefs, and a fourth to summarise CRM notes before sales calls. Each agent call is cheap. The aggregate bill is not.
For Southeast Asian brands running always-on engagement across Shopee, LINE, and their own apps simultaneously, the compounding effect is severe. A regional retailer with 8 million active customers, triggering even modest AI inference per session, can cross six-figure monthly AI spend before the use cases have been formally signed off. The cost isn’t in the model — it’s in the architecture around it.
Agentic Martech Is a Governance Problem Before It’s a Technology Problem
Tealium’s Zack Wenthe makes an uncomfortable observation: every software vendor is now building an AI agent, which means brands are accumulating agents the same way they accumulated SaaS tools in the 2010s — opportunistically, without a coherent data layer underneath. The result is predictable: agents operating on inconsistent customer data, making contradictory decisions across channels, and creating engagement experiences that feel fractured rather than orchestrated.
This is the walled garden problem reframed. If your personalisation agent lives inside your ESP, your next-best-action agent inside your CDP, and your content agent inside your CMS, you haven’t built a CEP — you’ve built three separate point solutions that occasionally talk to each other. The customer on the other end experiences the seams.
Wenthe’s argument is that the architecture of the data layer determines the quality of agentic output. Agents are only as coherent as the customer profile they’re drawing from. For markets like Indonesia or Vietnam, where a single customer might engage across a brand’s Tokopedia storefront, WhatsApp Business account, and mobile app in the same purchase journey, a fragmented agent layer doesn’t just cost more — it actively destroys the continuity that drives conversion.
Uncertainty Quantification: The Missing Layer in AI-Driven Engagement
Here’s where things get strategically interesting. Most martech AI systems return point predictions: this customer will churn, that segment will respond to discount messaging, this send time will maximise open rate. Tom Narock’s work on Bayesian Neural Networks in Towards Data Science makes the case that point predictions are often the wrong output to optimise for — because they discard the uncertainty that should be informing the decision.
A Bayesian approach doesn’t just tell you what a model predicts — it tells you how confident the model is in that prediction. For CEP design, this matters enormously. A next-best-action engine that knows it’s uncertain about a customer’s intent should behave differently than one that’s highly confident. Concretely: uncertain predictions should trigger lower-cost engagement touchpoints (a push notification rather than a high-production email), or route the customer to a human-assisted channel rather than a fully automated journey.
This isn’t theoretical. Teams building engagement frameworks on top of probabilistic outputs can build explicit confidence thresholds into their orchestration logic — agents only act autonomously above a certain confidence score, and escalate or pause below it. That architecture both reduces the cost of incorrect AI-driven actions and provides a governance mechanism for agent behaviour that stakeholders can actually understand and audit.
Building a CEP That Governs Agent Sprawl
The practical implication for marketing and data teams is that the CEP layer needs to do something it wasn’t originally designed to do: manage agents, not just channels. That means three specific things.
First, centralise the customer profile before you deploy agents. Agents drawing from inconsistent data sources will produce inconsistent decisions. A unified, real-time customer graph — covering identity resolution across devices and platforms, behavioural signals, and transaction history — is the prerequisite, not the nice-to-have. In Southeast Asia’s multi-platform reality, this often means building identity bridges between first-party data and platform IDs (LINE UIDs, Grab customer IDs, Lazada seller data) before a single agent goes live.
Second, instrument every agent call. Monte Carlo’s point about rising AI costs being invisible until they’re painful is a data observability problem. Treat agent token consumption as a monitored metric alongside open rates and conversion rates. Set spend thresholds per journey, per segment, per campaign. This isn’t about being cheap — it’s about making AI spend a strategic variable you control rather than a line item that surprises you at month-end.
Third, build confidence thresholds into your orchestration logic. Don’t let agents act uniformly across the uncertainty spectrum. Design journeys that respond differently when model confidence is low — slower cadences, simpler content, human review queues. This both protects the customer experience and gives you a natural audit trail for AI-driven decisions that can be explained to compliance teams across Southeast Asia’s varied regulatory environments.
The brands that will extract durable value from agentic martech aren’t the ones that deploy the most agents. They’re the ones that build the infrastructure to govern them — and know the difference between a cheap token and an expensive mistake.
Is your CEP architecture designed to orchestrate agents, or just channels? Because in 12 months, that distinction will be the one that separates the campaigns that feel seamless from the ones that feel like they were assembled by committee.
At grzzly, we spend a lot of time helping brands in Southeast Asia untangle exactly this: building customer engagement frameworks that can absorb agentic AI without fragmenting the customer experience or blowing the infrastructure budget. If your stack is accumulating agents faster than your data layer can support them, that’s a conversation worth having. Let’s talk
Sources
- https://montecarlo.ai/blog-token-prices-are-falling-so-why-is-your-ai-bill-going-up
- https://tealium.com/blog/artificial-intelligence/the-future-of-agentic-martech-is-choice-not-another-walled-garden/
- https://towardsdatascience.com/beyond-point-predictions-a-practical-introduction-to-bayesian-neural-networks/
Written by
Brooding GrizzlyDesigning CEP frameworks that move beyond batch-and-blast into real-time, context-aware engagement — across channels, devices, and the messiness of actual human behaviour.