Structured Data for AI Crawlers: JSON-LD, Entities and llms.txt
Most advice about AI visibility is strategic. This one is not. This is the plumbing — the layer that determines whether a model can parse your business accurately before any question of content quality arises.
It is also the part most Dubai businesses have never touched, which makes it the cheapest available advantage.
**Start with access, because it is binary**
Before anything else, confirm the crawlers can reach you. The ones that matter most are GPTBot, PerplexityBot, ClaudeBot, Google-Extended and CCBot.
Check two places, not one. Your robots.txt is visible and easy to audit. Your CDN or web application firewall is neither, and bot-protection rules there routinely block AI crawlers independently of anything in that file. We find this more often than any other single issue, and it is almost never deliberate — it is inherited configuration nobody revisited.
Blocking these crawlers is a legitimate choice for some businesses. Discovering eighteen months later that you made it by accident is not.
**JSON-LD: state facts unambiguously**
Schema markup is how you assert facts in a form that needs no interpretation. Prose can be misread; structured data cannot.
The types that carry most of the weight:
*Organization* — your legal name, trading name, address, contact details, licence identifiers, parent company. This is the anchor everything else attaches to.
*LocalBusiness* or its sector-specific subtypes — for anything with a physical presence and service area.
*Service* and *Product* — what you actually sell, described as discrete entities rather than as paragraphs on a page.
*Person* — named professionals with credentials, qualifications and roles. In regulated sectors this is frequently the single strongest signal you can publish.
*FAQPage* — question and answer pairs, which map almost directly onto how an assistant assembles a response.
The test that matters: could a model answer a factual question about your business from your markup alone, without reading a word of your prose? For most sites the honest answer is no.
**Entity consistency: the unglamorous one**
Models cross-reference. Your business name, address, licence details and category should read identically everywhere something might check — your site, your Google Business Profile, the relevant regulator's register, LinkedIn, industry directories, trade press.
In Dubai this takes real work. Free zone entities, DED licences and trading names routinely produce several legitimate versions of one company's name. Commercially that is normal. Computationally it means a model has multiple candidate entities and no confident way to merge them, and the usual resolution is to describe someone whose identity is unambiguous instead.
Pick a canonical form. Use it everywhere. Reconcile the rest. This is administrative rather than technical and it is frequently the highest-impact item on the list.
**llms.txt: a convention, not a standard**
llms.txt is a proposed file at your domain root that points AI systems to the content you consider authoritative, in a form that is cheap to parse. A companion llms-full.txt carries the expanded content.
Be clear-eyed about its status. It is a community convention that has gained traction, not a ratified standard, and adoption across AI providers is uneven and changing. It costs very little to publish and may help; it is not the foundation of anything.
Anyone selling llms.txt deployment as the core of an AI visibility programme is selling you the cheapest part of the work.
**Content structure: writing for retrieval**
When a model retrieves from your site, it works with passages rather than whole pages. It needs a chunk of text that stands alone — a specific claim, a number, a direct answer — liftable into a response without the surrounding context.
Practically that means: real headings that describe what follows. Direct answers before elaboration rather than after. Self-contained sections. Specific figures, dates, qualifications and ranges rather than adjectives. And honest hedging where hedging is warranted, because models are measurably more willing to cite a source that acknowledges limits than one making absolute claims.
The tension is real and worth naming: this is close to the opposite of how persuasive marketing copy is structured. Resolve it across pages rather than within one — reference content built to be cited, conversion pages built to convince.
**The order to do it in**
1. Crawler access. Binary, fast, and everything downstream is wasted without it. 2. Organization and Service schema. The foundation other markup attaches to. 3. Entity consistency. Tedious, unglamorous, disproportionately effective. 4. Person schema, if you are in a sector where credentials matter. 5. Content restructuring on your highest-value pages only. 6. llms.txt, last, because it is cheap and speculative.
**What this does not do**
Structured data makes you parseable. It does not make you authoritative. A model that can read your business perfectly still needs a reason to prefer you, and that comes from corroboration in sources it trusts — which is a slower, harder, more valuable piece of work.
Do the plumbing first because it is cheap and gating. Do not mistake it for the whole job.
هل أنت مستعد لتنمية عملك في دبي؟
احصل على تدقيق مجاني واكتشف كيف يمكننا مساعدتك على التفوق في سوقك المحلي.
احصل على تدقيق مجاني