Indonesia Singapore ไทย Pilipinas Việt Nam Malaysia မြန်မာ ລາວ
← Back to Blog

Topic Clusters, llms.txt, and the New Rules of AI Search

Build topical authority through tightly structured content clusters before AI search engines decide your brand isn't worth citing.

By Sneaky Grizzly →
Editorial illustration of a figure navigating a web of interconnected content nodes and AI search signals
Illustrated by Mikael Venne

Topic clusters and llms.txt are reshaping how LLMs discover brands. Here's what Southeast Asian marketers need to act on now.

Somewhere between Google’s last core update and the moment your CMO started asking why ChatGPT recommends your competitor, the rules of search visibility quietly rewrote themselves. Most teams are still optimising for a search engine that now shares the stage with generative systems that neither rank pages nor care about your keyword density.

Two developments — one structural, one technical — explain most of what’s shifting. Understanding them together is the fastest way to stop optimising for a world that no longer exists.

Topic Clusters Are Now Your Entity Reputation System

Semrush’s breakdown of topic cluster strategy makes a point that deserves more weight than it typically gets: clusters aren’t just an internal linking tactic anymore. They are the primary signal through which LLMs establish whether a brand is genuinely authoritative on a subject — or merely present.

Here’s the distinction that matters. A traditional keyword strategy gets you found when someone searches. A topic cluster strategy gets you cited when an AI synthesises an answer. Generative engines pull from sources that demonstrate consistent, interconnected depth on a subject — not isolated pages that rank for individual queries.

For Southeast Asian brands, this has a compounding effect. If your content exists primarily in fragmented campaign microsites (a pattern common across Shopee and Lazada brand stores, LINE campaigns, and seasonal landing pages), you likely have width but no depth. An LLM crawling your footprint sees isolated signals, not a coherent knowledge graph. Building a cluster around, say, skincare ingredients or SME financing in the Thai market — with a strong pillar page and tightly linked supporting content — creates the kind of topical density that generative engines reward with mentions.

Implementation note: start with one cluster, not five. Map your pillar topic, identify eight to twelve subtopics with genuine search demand, and link them with deliberate semantic consistency. Measure success not just by organic traffic, but by whether your brand starts appearing in AI-generated answers for category queries.

The llms.txt Experiment — Useful Signal or False Comfort?

Common Crawl’s analysis of 584,107 llms.txt files is one of the more quietly damning research findings of the year. The format, designed to give websites a way to communicate instructions to LLM crawlers, has been widely misimplemented. Most files were generated from templates. Many contained no links. And a significant share included crawler rules — blocking instructions, access controls — that llms.txt simply cannot enforce.

That last point is the one to sit with. A meaningful portion of the web has essentially put up a sign asking AI crawlers to behave in ways the format was never designed to mandate. It’s the digital equivalent of a ‘no photography’ notice in a public park.

This doesn’t make llms.txt useless. It means the format is a declaration of intent, not a control mechanism. Used correctly — as a structured, link-rich document that guides LLM crawlers toward your most authoritative content — it can improve how your site is understood and indexed by AI systems. Used as a copy-paste compliance exercise, it’s noise.

For teams managing multilingual sites across SEA markets, the structural challenge is real: if your llms.txt points to your English-language pillar content only, you’re potentially starving AI systems of your Bahasa, Thai, or Vietnamese content clusters entirely. Localisation of llms.txt is an underappreciated implementation gap.


Entity Authority Is the New Domain Authority

The through-line connecting both developments is entity authority — how clearly and consistently your brand is understood by AI systems as a credible, well-defined entity within a specific knowledge domain.

Domain authority was a proxy metric for trust in a link-based search world. Entity authority operates differently: it’s built through semantic consistency (does your content say the same things about your brand across formats and platforms?), topical depth (do you own a subject area, or just touch it?), and structured data that helps machines understand what your brand is, not just what it publishes.

Brands in Southeast Asia face a specific challenge here. Operating across markets with different languages, platform ecosystems, and consumer behaviours makes semantic consistency genuinely hard to maintain. A brand that positions itself as a sustainability leader in Singapore but leads with price in Vietnam sends inconsistent entity signals. AI systems aren’t confused by this — they simply default to whichever signal is stronger, which is rarely the one you’d choose.

The practical implication: treat your brand’s knowledge graph as a product. Audit how your brand entity is described across your own content, third-party mentions, and structured data. Where there’s inconsistency, there’s entity erosion — and that erosion shows up as absence in AI-generated answers.

Infrastructure Is Now an AI Visibility Decision

This may seem like a digression, but it isn’t: the infrastructure your content sits on affects how reliably AI crawlers can access and process it. InMotion Cloud’s launch of fixed-price Managed Private Cloud on bare metal is aimed squarely at the predictability problem — erratic crawl access due to shared hosting throttling, egress surprises, and inconsistent uptime all create friction for the automated systems that determine your AI search visibility.

For larger SEA brands managing high-traffic content operations, the calculus is shifting. AI crawlers are not polite. They hit content at scale, and hosting environments that weren’t built to handle that load will quietly degrade your indexability. Fixed-egress, dedicated infrastructure removes one variable from a discoverability equation that already has too many.

Key takeaways:

  • Build one deep topic cluster before expanding — topical density beats keyword breadth in generative search citation patterns.
  • Audit your llms.txt for implementation quality, not just existence — include multilingual content and avoid pasting in crawler rules the format can’t enforce.
  • Treat entity consistency across markets as a GEO prerequisite — AI systems penalise brand ambiguity with invisibility, not demotion.

The deeper question for 2026 is whether the brands investing in topical authority and entity hygiene today are building a durable GEO moat — or simply staying ahead of a curve that will flatten as these practices become table stakes. The window where this work creates genuine competitive separation is probably measured in months, not years.


At grzzly, we work with growth and marketing teams across Southeast Asia to build content architectures and entity strategies that hold up inside generative search — not just traditional rankings. If your brand is starting to ask why it’s invisible in AI answers, that’s exactly the conversation we’re set up to have. Let’s talk

Sneaky Grizzly

Written by

Sneaky Grizzly

Tracking the quiet revolution inside LLM-powered search — where brand mentions, structured semantics, and entity authority rewrite the rules of discoverability before most marketers notice.

Enjoyed this?
Let's talk.

Start a conversation