AI agents fail in production not from lack of data, but wrong connections and blind monitoring. Here's what CDP teams in Southeast Asia must fix now.
Most AI agent deployments in Southeast Asia right now share the same architectural flaw: they were designed for a controlled environment and shipped into a chaotic one. The symptoms show up fast — degraded personalisation, recommendation loops, customer trust eroding quietly in the background.
Three pieces of research published this week, taken together, describe exactly why this keeps happening — and what data teams need to do differently.
Production Is a Different Animal Than Your Eval Dataset
Monte Carlo’s Virna Sekuj makes an observation that should be pinned to every AI team’s wall: production is not simply a larger version of your development environment. The standard practice of golden dataset evaluation — curated prompts, expected outputs, a threshold to clear before release — is a useful gate for shipping. It is not a monitoring strategy.
The distinction matters enormously for CDP-connected agents. When a personalisation agent ingests real-time behavioural streams from Shopee, LINE, or a brand’s own app, it encounters data states that no golden dataset anticipated: null loyalty tier values mid-session, conflicting device IDs from cross-border shoppers, transactional signals from Ramadan flash sales that skew the entire behavioural baseline. The agent clears its eval threshold, ships cleanly, and then encounters the actual world.
The fix isn’t better evals — it’s parallel observability infrastructure in production. Specifically: drift detection on the input feature distributions the agent was trained against, output sampling with human review triggers, and circuit-breaker logic that degrades gracefully rather than confidently surfacing wrong recommendations.
Dense Graphs Are a Data Architecture Trap
A controlled experiment published in Towards Data Science by Emmimal P Alexander ran 50 iterations of a multi-agent system across relationship densities ranging from 20% to 100%. The finding is counterintuitive and directly applicable to CDP architecture: recovery performance remained stable across that entire range, but the fraction of edges actually used collapsed as density increased.
Translate that to your customer data graph. Teams building unified profiles instinctively connect everything — purchase history to browsing to CRM to offline POS to third-party enrichment. The connections exist. Most are never traversed in an actual decisioning moment. Worse, the dead weight increases query latency and makes the graph harder to audit when a campaign produces anomalous results.
The design implication: instrument your graph. Track which edges your agents and activation queries actually use over a 90-day window. Prune or archive the rest. In platforms like Segment or mParticle, this translates to reviewing event schema utilisation reports and being ruthless about deprecating unused traits. A leaner identity graph resolves faster and is easier to explain to a privacy regulator — both of which matter acutely in markets operating under PDPA frameworks.
Public Trust Is a Data Architecture Decision
Stephanie Kirmer’s analysis of anti-AI public opinion in Towards Data Science lands a useful framing: people accept AI tradeoffs when they perceive value. When they don’t see the value — or when the AI visibly gets them wrong — the trust deficit compounds quickly.
For brands running CDP-powered personalisation in Southeast Asia, this is not an abstract concern. The region’s consumers are among the most active mobile commerce users globally, which means they accumulate enough interaction data to notice when personalisation is lazy or off-target. A Thai consumer who has purchased from a brand’s Lazada store three times and still receives first-time-buyer messaging isn’t just unimpressed — they’re forming an opinion about the brand’s competence and its relationship with their data.
The architectural response is identity resolution done properly: a deterministic spine (email, phone, loyalty ID) with probabilistic bridging for anonymous sessions, and — critically — a suppression layer that prevents activation of stale segments. Most teams build the resolution logic and skip the suppression. That’s where the visible failures originate.
Building declared data collection into your engagement layer also helps on the trust dimension. Post-purchase preference surveys, explicit interest selection during onboarding, opt-in enrichment prompts — these signal to the customer that the data relationship is a two-way exchange, not extraction. Kirmer’s framing holds: people accept the tradeoff when they see the value returned.
What This Means for Your Next 90 Days
- Instrument production, not just pre-release: Add feature drift monitoring and output sampling to any CDP-connected agent currently in production — eval thresholds alone will not catch the failure modes that real data distributions introduce.
- Audit your identity graph for used versus configured connections: Run a 90-day edge utilisation report and deprecate unused relationships; a leaner graph resolves faster, costs less, and is materially easier to defend under PDPA or similar frameworks.
- Close the trust loop with declared data: Build at least one explicit preference-collection touchpoint into your customer journey this quarter — it improves model inputs and signals to customers that their data is being used in their interest, not just the brand’s.
The deeper question worth sitting with: as AI agents take on more autonomous decisioning within your customer data stack, who in your organisation is actually watching them in production — and with what instrumentation? Most teams have a confident answer until they’re asked to show the dashboard.
At grzzly, we work with growth and data teams across Southeast Asia to design CDP architectures that perform under real production conditions — identity resolution, activation logic, and the observability layer that keeps it honest. If your agent deployment is flying blind in production, or your customer graph has grown faster than your ability to audit it, we should compare notes. Let’s talk
Sources
Written by
Velvet GrizzlyArchitecting the unified customer profile — stitching together behavioural, transactional, and declared data into platforms that actually earn their licence fee.