Week one is a log-file audit, not a content plan
Most AI visibility proposals start with content. Ours starts with a question that takes a day to answer and decides whether the rest of the work is worth doing: in the last 30 days, did an AI crawler request a page on your site, and what status code did it get back?
The answer is regularly ugly. We see sites where GPTBot and PerplexityBot have been served a 403 on every request for months, sites where the whole /blog/ directory is disallowed by a robots.txt somebody copied from a template in 2019, and sites where the crawler got a 200 with an empty HTML shell because the content only exists after JavaScript runs.
None of those are content problems. You can rewrite every page on a blocked site and change nothing. So we read the server logs first, in week one, before anyone opens a document.
- Pull 30–90 days of raw server or CDN logs and filter by AI user agent. Analytics won't show you this; bots don't run JavaScript.
- Record status codes per bot. A 403 or 429 pattern is a block. A 200 that returns a near-empty body is a rendering problem, which is worse because it looks fine.
- Diff robots.txt against intent. Not what you meant to allow — what the file actually says today.
- Check the CDN and WAF rules separately. Robots.txt can permit a crawler that Cloudflare then blocks before it ever reaches your server.
- Test render without JavaScript. If your content isn't in the initial HTML, assume a meaningful share of crawlers never see it.
Who's actually crawling you, and what blocking each one costs
The single most useful distinction in this whole field, and the one most 'AI SEO' proposals blur: some of these crawlers build training data, and some of them fetch pages to answer a question right now. Blocking the second group removes you from citations. Blocking the first group is a licensing and copyright decision with a much smaller visibility cost.
We map every one of them against your logs and give you a decision per bot, with the trade-off written down.
| Agent | Who runs it | What it does | Cost of blocking it |
|---|---|---|---|
GPTBot | OpenAI | Crawls broadly, primarily for model training data. | Mostly a training and licensing decision, not a citation one. |
OAI-SearchBot | OpenAI | Indexes pages so ChatGPT can surface and link them in search results. | High. Block it and you remove yourself from ChatGPT's linked results. |
ChatGPT-User | OpenAI | Fetches a page live because a user's question needed it. | High. This is the real-time citation path. |
PerplexityBot | Perplexity | Indexes pages for Perplexity's answers. | High. Perplexity cites sources prominently and sends real clicks. |
ClaudeBot | Anthropic | Crawls for Anthropic's models and assistant retrieval. | Moderate to high, depending on how your buyers use it. |
Google-Extended | A robots token controlling use of your content in Gemini and Vertex AI grounding. | Low for search. It does not affect Google Search indexing or AI Overviews inclusion. | |
Bingbot | Microsoft | Builds the Bing index that Copilot and several assistants read from. | Very high. This is the least glamorous and most consequential row here. |
Applebot-Extended | Apple | Opt-out token for Apple's generative training, separate from Applebot search crawling. | Low for visibility, relevant if training use concerns you. |
CCBot | Common Crawl | Builds the open dataset a great many models are trained on. | Training exposure only, and it's a one-way door once you're out of the archive. |
Cloudflare is usually the culprit
When a site is invisible to assistants and the robots.txt looks fine, the answer is almost always sitting in front of the origin. India's small-and-mid-market web runs heavily on Cloudflare's free and Pro tiers, and its bot controls are aggressive by design — that's what they're for.
Cloudflare has also shipped one-click AI crawler blocking, and since mid-2025 has been blocking AI crawlers by default for newly onboarded domains. If nobody at your company has looked at that setting, the default may well be making your decision for you.
- Bot Fight Mode challenges clients that don't behave like browsers. Legitimate crawlers frequently fail that test and receive a challenge page instead of your content.
- Managed AI-crawler rules may be on without anyone enabling them deliberately, particularly on domains added recently.
- Rate limiting at aggressive thresholds looks like a partial block: some pages get through, deep pages never do, and the pattern is invisible without logs.
- Firewall rules by country or ASN, written years ago to stop scrapers, that also catch the datacentre IPs these crawlers use — or a security level left on 'I'm Under Attack' after an incident and never turned back down.
What we change, and what we ask you to decide
We write the robots.txt, propose the CDN rule set, and hand you a one-page decision sheet: allow, allow-with-rate-limit, or block, per agent, with the visibility trade-off for each. Implementation is ours if we have access, your team's if you'd rather. Then we re-pull logs in 14 days to confirm the crawler actually came back with a 200 — because 'we changed the rule' and 'the bot returned' are two different claims and only one of them is evidence.
Bing and IndexNow, because half this surface runs on Bing
Indian companies spend years on Google and never open Bing Webmaster Tools. That was a reasonable trade when Bing was three percent of search. It stopped being reasonable when Microsoft grounded Copilot in Bing's index, and it's worth noting DuckDuckGo has long drawn on Bing results too.
We check whether you're indexed there at all — the answer is often 'partially, and nobody noticed' — and fix it. Bing's tooling makes this cheap: you can import your Google Search Console verification and property in a few clicks rather than re-verifying from scratch.
- Verify and import the property, then read Bing's own index coverage report. It flags different problems than Google's, and some of them are real.
- Submit sitemaps to Bing explicitly. Discovery is slower there and it will not find you as forgivingly as Google does.
- Wire up IndexNow. It's supported by Bing, Yandex, Seznam and Naver, and Cloudflare offers a one-toggle integration. New and updated URLs get pushed instead of waiting to be found.
- Fix Bing-specific crawl blocks.
bingbotis sometimes rate-limited or blocked by rules written to stop scrapers, entirely by accident. - Check your pages actually render for Bingbot, which historically handles JavaScript less generously than Googlebot.
What's included, what it costs, and what the guarantee covers
AI search optimisation starts at ₹75,000 a month, ex-GST, month-to-month after the first quarter, thirty days' notice, and you keep every asset. Smaller sites start at ₹40,000 a month. Most clients pair it with answer engine optimisation — this service makes sure you're reachable, that one makes sure you're quotable. Full numbers on pricing.
| Included | Not included |
|---|---|
| Log-file audit of AI crawler access, with status codes per agent | Server, CDN and hosting costs, or migrating you off a host that won't give logs |
| Robots.txt rewrite and a per-bot allow/block decision sheet | The commercial decision itself — you own that, we implement it |
| CDN and WAF bot-rule review, with implementation if we have access | Security architecture and DDoS response beyond crawler access |
| Bing Webmaster Tools setup, sitemap submission and IndexNow wiring | Bing Ads or any paid placement on Microsoft properties |
| Rendering checks and a fix specification where content needs JavaScript | Front-end engineering to implement server-side rendering |
| Frozen prompt panel, monthly runs, share-of-voice reporting with competitors | Enterprise AI-monitoring tool licences if you want one alongside ours |
| Content restructuring for citation — only if you add AEO | Wholesale content rewriting inside this retainer |
The guarantee, stated exactly
We freeze two numbers on day one: your mention count across the agreed prompt panel, and your trailing-90-day qualified leads from organic search. Miss either at day 90 and we keep working free until we beat it. It isn't a refund and we won't market it as one — it's unpaid work until the numbers clear. Three clients a month, because you cannot carry that risk at volume.
We will never guarantee that a named assistant mentions you for a named prompt. Model outputs aren't stable and aren't controllable by anyone outside those companies. That promise is the 2026 version of guaranteeing position one — same trick, newer paint, and fewer founders have learned to spot it yet.