We tested 15 well-known Indian D2C brand homepages for AI-crawler accessibility, structured data, and llms.txt adoption. The real numbers, brand by brand.
We ran 15 well-known Indian D2C and consumer brand homepages through our AI Visibility Checker — the tool that shows whether AI crawlers like GPTBot, ClaudeBot, and PerplexityBot can actually fetch and parse a page. One brand's firewall blocked the checker itself before robots.txt was even read. Of the 14 that returned a score, the average was 86/100, but nearly a third had zero structured data on their homepage, half were missing basic section headings, and not one of the 15 had published an llms.txt file. Here's the full brand-by-brand breakdown, what it actually measures, and what it means if you're wondering whether your own site would score any better.

The AI Visibility Checker: one URL in, then a breakdown of all 10 AI crawlers it checks and exactly what the 100-point score is made of. This is the exact tool used to test all 15 brands below.
What This Tool Actually Measures (Read This Before the Numbers)
Before the data: it matters what "AI visibility" means here, because it's easy to assume this is about whether ChatGPT recommends a brand by name. It isn't, and that distinction is the most important thing in this entire study.
The AI Visibility Checker makes four requests to a domain — the homepage, /robots.txt, /llms.txt, and /sitemap.xml — identifying itself honestly as ScultToolsBot/1.0. It then scores the result out of 100 based on whether the ten major AI crawlers (OpenAI's GPTBot, OAI-SearchBot, and ChatGPT-User; Anthropic's ClaudeBot and anthropic-ai; Perplexity's PerplexityBot; Google's Google-Extended; Common Crawl's CCBot; ByteDance's Bytespider; and Meta's meta-externalagent) are permitted by robots.txt, whether the homepage carries JSON-LD structured data, whether basic on-page fundamentals are present (title, meta description, h1, h2 headings, Open Graph tags, canonical link, alt text, minimum word count), whether llms.txt exists, and whether a sitemap is published.
That is a technical crawler-accessibility and machine-readability score — it tells you whether AI systems can read and parse a page cleanly. It does not, and cannot, tell you whether ChatGPT will recommend that brand when someone asks "what's a good moisturizer" or "best D2C mattress brand in India." Recommendation depends on training data, real-world reputation, third-party mentions, and review consensus — none of which a single homepage crawl can measure. The tool's own documentation is explicit about this: a perfect score "means AI engines can read and parse you — not that they will cite you." Read what GEO actually is for the fuller picture of what does influence citation and recommendation.
So the honest framing of this study is: is a defined set of major Indian consumer brands even technically reachable and parseable by the crawlers that feed today's AI answer engines? That's a narrower question than "does AI love this brand," but it's a real one, it's completely measurable without guessing, and — as the results below show — the answer is far from universally yes.
Why We Ran This
Every brand assumes its site is fine for crawlers because it ranks fine on Google. But Google's crawler and the newer AI crawlers are not the same user-agent, and a robots.txt file written five years ago — or a WAF rule added last month to block "bot traffic" — can quietly wall off GPTBot or ClaudeBot while Googlebot sails through untouched. Nobody checks this until it becomes a problem, because there's no dashboard for "am I invisible to ChatGPT's browsing crawler." We wanted a real answer for a real, checkable set of brands rather than a hypothetical one, so we picked a spread of well-known Indian D2C and consumer companies across beauty, fashion, food, home, and other categories and ran every one of them through the checker exactly as any visitor would use it.
Method
- Sample: 15 established Indian D2C/consumer brands, chosen for name recognition and spread across categories — beauty and personal care, fashion and apparel, food and grocery delivery, home and furniture, baby and kids, and eyewear.
- Tool used: the AI Visibility Checker, run once per brand's primary
.com/.indomain, homepage only. - What was recorded: overall score out of 100, how many of the 10 named AI crawlers robots.txt allows, which specific crawlers (if any) are blocked, which JSON-LD
@typevalues are present, and which on-page/llms.txt/sitemap checks failed. - Date: results captured directly from the live tool in a single session; robots.txt rules and homepage markup can change at any time, so treat this as a snapshot, not a permanent verdict on any brand.
- What we did not measure: brand recognition, sentiment, or recommendation frequency inside any AI chat product. See the section above for why.
The Headline Numbers
Of the 15 domains tested, 14 returned a score and one could not be checked at all — its firewall returned an HTTP 403 to the checker itself, before robots.txt was ever evaluated. That's covered in its own section below, because it's arguably the most important finding in the whole study.
Across the 14 that scored:
- Average score: 86/100.
- Highest score: 100/100 (one brand — full marks on every check).
- Lowest score: 66/100.
- 12 of 14 (86%) allow all 10 AI crawlers to fetch their homepage without restriction.
- 2 of 14 (14%) block at least one named AI crawler in robots.txt.
- 4 of 14 (29%) have zero JSON-LD structured data on their homepage.
- 7 of 14 (50%) are missing h2 section headings on the homepage.
- 0 of 14 (0%) have published an
llms.txtfile. - 14 of 14 (100%) have a working
sitemap.xml— the one check nobody failed.
Full Brand-by-Brand Results
| Brand | Score | Crawlers allowed | Schema types found | Biggest gap |
|---|---|---|---|---|
| Sugar Cosmetics | 100/100 | 10/10 | 3 (Organization, WebSite, BreadcrumbList) | None — full marks |
| Mamaearth | 98/100 | 10/10 | 2 (Organization, WebSite) | Missing Open Graph tags |
| boAt | 98/100 | 10/10 | 4 (Organization, WebSite, WebPage, BreadcrumbList) | Missing h2 headings |
| Mokobara | 98/100 | 10/10 | 3 (Organization, WebSite, BreadcrumbList) | Missing h1 heading |
| The Man Company | 96/100 | 10/10 | 2 (Organization, WebSite) | Missing OG tags + partial alt text (73%) |
| FirstCry | 96/100 | 9/10 (blocks CCBot) | 2 (Organization, WebSite) | Only named-crawler block in the "high scorer" group |
| Nykaa | 93/100 | 10/10 | 2 (Organization, WebSite) | Thin homepage text (52 words) + missing OG tags |
| Lenskart | 93/100 | 10/10 | 2 (WebSite, Organization) | Thin homepage text (82 words), no llms.txt |
| Zivame | 82/100 | 10/10 | 2 (WebSite, MobileApplication) | Missing meta description, h1, h2, OG tags, canonical |
| Purplle | 76/100 | 10/10 | 0 | No structured data at all; missing h1/h2 |
| Country Delight | 71/100 | 10/10 | 0 | No structured data; 0% image alt-text coverage |
| WOW Skin Science | 69/100 | 4/10 (blocks 6) | 2 (Organization, WebSite) | Blocks GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, meta-externalagent |
| Licious | 69/100 | 10/10 | 0 | No structured data; thin homepage text (33 words) |
| Bewakoof | 66/100 | 10/10 | 0 | No structured data; missing h1, h2, canonical; thin text (37 words) |
| Wakefit | Not scored | — | — | Firewall returned HTTP 403 to the checker itself |
Crawler Access: Two Very Different Blocking Patterns
Twelve of the fourteen scored brands allow every one of the 10 crawlers the tool checks, which is the right default for any brand that wants to be discoverable through AI-mediated search. The two exceptions were instructive precisely because they blocked in such different ways.
FirstCry blocks exactly one crawler: CCBot, Common Crawl's crawler, via an explicit Disallow: / rule in a group naming it directly. Every other crawler — including GPTBot, ClaudeBot, and PerplexityBot — is allowed. This looks like a narrow, likely deliberate decision rather than an accident, and it cost only 4 of the possible 40 crawler-access points.
WOW Skin Science blocks six of ten: GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, and meta-externalagent are all disallowed, while OAI-SearchBot, ChatGPT-User, anthropic-ai, and PerplexityBot are still allowed. That's a much stranger pattern than an all-or-nothing AI block — it blocks the crawlers that mainly train future models (GPTBot, Google-Extended, Bytespider, meta-externalagent, CCBot) while leaving the crawlers that power live, in-the-moment answers (ChatGPT-User's live browsing, OAI-SearchBot's search index, PerplexityBot, and legacy anthropic-ai) untouched. Whether that's an intentional "opt out of training, stay visible in live answers" strategy or simply an old robots.txt block-list that was never revisited for the newer, more targeted crawlers, the effect is the same: this is the only brand in the sample actively shaping which parts of the AI ecosystem see it, rather than either fully allowing or fully blocking.
Common Crawl's CCBot is the single most-blocked crawler in this sample — the only one to appear in both blocklists. CCBot's dataset is a foundational training source for a wide range of language models beyond just the two hyperscalers, which makes it a plausible target for a brand more focused on limiting model-training exposure than on limiting live-answer visibility. GPTBot, ClaudeBot, Google-Extended, Bytespider, and meta-externalagent were each blocked once. PerplexityBot, OAI-SearchBot, ChatGPT-User, and anthropic-ai were never blocked by a single brand in the sample — for these 14 sites, at least, the crawlers behind live AI search and chat answers had a clean run every time.
The Silent Failure: When Your Own Firewall Blocks the Checker
Wakefit is the finding that matters most operationally, and it isn't a low score — it's no score at all. When the checker tried to fetch the homepage, the site's own infrastructure answered with an HTTP 403, refusing the request outright. Robots.txt was never even reached, because whatever bot-protection layer sits in front of the site rejected the connection before that file could be read.
The tool's own guidance on this is worth repeating verbatim, because it's the single most useful caveat in the entire report: "blocking ScultToolsBot does not mean AI bots are blocked. Plenty of firewalls reject unfamiliar crawlers while letting GPTBot or ClaudeBot through." In other words, this result is genuinely ambiguous — it could mean Wakefit's WAF is configured with an explicit allowlist that happens to include the named AI crawlers and simply rejects anything unrecognized (in which case its real AI-crawler access could be fine), or it could mean the same aggressive bot-blocking that caught our checker is also silently catching GPTBot, ClaudeBot, and PerplexityBot, in which case the brand could be substantially less visible to AI systems than a robots.txt read alone would ever reveal — because robots.txt is a voluntary courtesy crawlers choose to honor, while a WAF block happens whether or not the crawler asked politely.
This is exactly the scenario the checker's own documentation warns site owners to rule out before trusting a score at all: if your WAF blocks unfamiliar user-agents, you have to explicitly allowlist the real AI crawler user-agents to know whether you're actually reachable, because otherwise a clean-looking robots.txt is meaningless — the request never gets that far. For any brand running aggressive bot-mitigation (Cloudflare's bot-fight mode, a WAF with a strict default-deny posture, or similar), this is worth checking directly rather than assuming robots.txt tells the whole story.
Structured Data: Nearly a Third of Big Brands Have Nothing
Four of the fourteen scored brands — Bewakoof, Licious, Country Delight, and Purplle — have no JSON-LD structured data at all on their homepage. No Organization, no WebSite, nothing. For AI systems that lean on structured data to unambiguously identify what a business is, what it sells, and how it's named and branded, these four homepages offer nothing beyond whatever a language model can infer from raw HTML text and images — which is strictly harder and less reliable than being told directly.
This is a 20-point gap that's genuinely simple to close: a single <script type="application/ld+json"> block with Organization and WebSite types would have moved any of these four brands into the 85-90+ range on its own, without touching a single line of visible page copy. It's also, notably, unrelated to brand size or resourcing — all four of these are large, well-funded, well-known consumer brands, which suggests structured data is simply an overlooked line item on the technical-SEO checklist rather than something correlated with company maturity.
At the other end, boAt led the sample with four distinct schema types (Organization, WebSite, WebPage, BreadcrumbList), and Sugar Cosmetics, Mokobara, and FirstCry all combined multiple schema types with full crawler access to post the highest scores in the study.
The Missing H2 Problem — and Other Recurring On-Page Gaps
Half the scored brands — seven of fourteen — are missing h2 section headings on their homepage, the single most common on-page issue in the whole study. This one is easy to miss because a homepage can look complete to a human scrolling through hero banners, product carousels, and promotional blocks, while structurally containing no semantic heading hierarchy an AI system (or, for that matter, a screen reader) can use to understand the page's outline.
Several other gaps showed up repeatedly:
- Thin visible text: four brands — Nykaa (52 words), Bewakoof (37 words), Licious (33 words), and Lenskart (82 words) — had homepages the checker flagged as containing too little extractable text, typically because the page is dominated by images and interactive components rather than readable copy.
- Missing Open Graph tags: at least four brands (Mamaearth, Nykaa, The Man Company, and Zivame) were missing
og:title,og:description, orog:image— tags that matter for how a page previews and gets described when shared or referenced, not just for social platforms. - Incomplete image alt text: WOW Skin Science (33% of 24 images), The Man Company (73% of 157 images), Zivame (73% of 199 images), and Country Delight (0% of 27 images) all had meaningful gaps in alt-text coverage.
llms.txt: zero adoption. Not one of the 14 scored brands has published anllms.txtfile. This is the most unambiguous finding in the entire study, and the least surprising —llms.txtis a genuinely new, still-emerging convention, so near-total non-adoption among mainstream consumer brands (as opposed to developer-tool or AI-native companies, where it's caught on faster) tracks with how new the idea still is. It's also the cheapest point of the bunch to pick up: a plain markdown file listing key pages with one-line summaries, no code change required.- Sitemap.xml: the one thing nobody got wrong. All 14 scored brands had a working sitemap, which lines up with how long XML sitemaps have been a baseline SEO expectation — this is the one part of "AI readiness" that mainstream brands had already solved for a previous era of search.
What This Means for Any Indian Brand
A few patterns hold up across this sample that are worth generalizing carefully, without overstating a 15-brand snapshot into a universal law:
- Crawler access is mostly a solved problem, but not universally. Most large Indian consumer brands aren't accidentally blocking AI crawlers wholesale — but 2 of 14 do block at least one, and a third brand's own infrastructure blocked the check before robots.txt even mattered. "We probably haven't blocked anything" is not the same as "we've verified we haven't blocked anything."
- Structured data is the most common real gap, and the cheapest to fix. Close to a third of brands tested had none at all, despite it being a static, one-time addition rather than an ongoing content commitment.
- A technically excellent homepage and a content-thin one can both belong to a major, well-known brand. Company size and brand recognition did not reliably predict a high score — some of the most recognizable names in the sample scored in the 60s and 70s, while less universally famous brands in the same category scored near-perfect.
llms.txtis a genuine first-mover opportunity, not table stakes yet. With adoption at zero across this sample, any brand that publishes one today is doing something essentially none of its direct competitors have done.- Robots.txt is necessary but not sufficient evidence of AI-crawler access. A WAF, bot-mitigation product, or CDN-level rule can block a crawler that robots.txt would have happily allowed — and the only way to know is to test the real user-agent against the real infrastructure, not just read the text file.
What to Fix First if You're Starting from a Low Score
If a scan of your own site would land closer to the 66-76 range than the 90s, the fixes in rough order of effort-to-impact are: add a basic Organization/WebSite JSON-LD block if you have none (biggest single point gain for zero ongoing cost — see the structured data guide if you're building it from scratch); add semantic h2 headings to the homepage if the page is currently a stack of unlabeled visual blocks; fill in missing Open Graph tags and image alt text; publish a plain-text llms.txt; and — separately from anything robots.txt-related — check your WAF or bot-mitigation configuration directly to confirm it isn't rejecting the named AI crawler user-agents outright, the way Wakefit's did in this study. None of this replaces the deeper work of actually earning citations and mentions, which is a different and harder problem covered in how to get your brand mentioned by ChatGPT — but a brand that fails the technical basics never gets the chance to compete on the harder problem in the first place.
Frequently Asked Questions
What does the AI Visibility Checker actually measure?
Technical AI-crawler accessibility and machine-readability: whether the 10 major AI crawlers can fetch your homepage per robots.txt, whether structured data and on-page basics are present, and whether llms.txt and a sitemap exist. It does not measure whether AI chat products recommend or cite your brand by name — that depends on far more than a single homepage crawl.
Did any of the 15 brands block AI crawlers entirely?
No brand in the sample blocked all 10 crawlers. WOW Skin Science blocked 6 of 10 (GPTBot, ClaudeBot, Google-Extended, CCBot, Bytespider, and meta-externalagent, while still allowing OAI-SearchBot, ChatGPT-User, anthropic-ai, and PerplexityBot), and FirstCry blocked only CCBot. The other 12 scored brands allowed all 10.
Why did Wakefit's website show no score at all?
Its firewall returned an HTTP 403 to the checker's request before robots.txt was even read — the site rejected the crawl outright at the infrastructure level. This means its real AI-crawler accessibility is unknown from this test alone: the same rule could be blocking legitimate AI crawlers too, or it could be allowlisting them specifically while rejecting unfamiliar bots like the checker. It has to be verified directly against the WAF configuration.
How common is missing structured data among big Indian brands?
In this sample, 4 of 14 scored brands (29%) had no JSON-LD structured data at all on their homepage — Bewakoof, Licious, Country Delight, and Purplle. It was the single most impactful, cheapest-to-fix gap in the study.
Is llms.txt common among Indian consumer brands yet?
Not at all — zero of the 14 scored brands had published one. It remains a genuine, low-effort first-mover opportunity rather than an expected baseline.
My site passes robots.txt for all 10 crawlers — does that guarantee AI crawlers can reach me?
Not necessarily. Robots.txt is a voluntary rule crawlers choose to honor; a web application firewall or bot-mitigation layer can reject a crawler's request at the network level regardless of what robots.txt says, exactly as happened to the checker itself when testing Wakefit. If you run aggressive bot protection, confirm the named AI crawler user-agents are explicitly allowlisted, not just permitted by robots.txt.
How can I check my own brand's AI visibility?
Run your homepage through the AI Visibility Checker directly — it takes one URL, no account or email required, and returns the same crawler-access, structured-data, and on-page breakdown used throughout this study.
Want a full technical AI-readiness audit, not just a homepage snapshot? Talk to Scult's AI and automation team.



