Infrastructure feature

Watch AI crawlers read your site

Before an engine can cite you, its crawler has to fetch you. Siftly shows every AI crawler hit on your site as it happens: which bot, which page, whether the visit was for indexing or for training, and where it came from. It is the earliest signal you have that new content will be citable.

Crawler trafficLive
WhenPlatformTypePathCountry
12:04GPTBotTraining/fit-guide
11:58OAI-SearchBotIndexing/blog/stability-vs-neutral
11:41PerplexityBotIndexing/fit-guide
11:20Google-ExtendedTraining/blog/marathon-shoe-picks
10:57PerplexityBotIndexing/sizing

the live log

Crawl comes before citation

Every bot hit, as it lands. GPTBot, OAI-SearchBot, PerplexityBot, Google-Extended and the rest, each with the time, the path, and the country the request came from. When you publish, you can watch the crawlers arrive instead of guessing whether they did.

Two visits that mean different things. An indexing fetch means the page can be cited in an answer now. A training fetch means it may inform the model later. Siftly labels each hit, so you can tell a page that is answer-ready from one that is only feeding a future model.

PagesExport
PageCitationsIndexingAI SessionsCCRClicks
/fit-guide1832,4107.6%1,204
/blog/stability-vs-neutral1211,8426.6%918
/blog/marathon-shoe-picks961,3577.1%742
/blog/trail-under-150446127.2%287

per page

Turn crawl data into a content decision

Which of your pages AI actually reads. Crawl counts sit next to citations for the same page, so you can separate the two failures that look identical from the outside: a page nothing crawls, and a page that is crawled constantly and never cited.

Bots on one side, people on the other. Connect analytics and Search Console and the same table carries AI sessions, human sessions, clicks, and impressions. That is the full picture of a page: who fetched it, who cited it, and who arrived because of it.

PagesExport
PageCitationsIndexingAI SessionsCCRClicks
/fit-guide1832,4107.6%1,204
/blog/stability-vs-neutral1211,8426.6%918
/blog/marathon-shoe-picks961,3577.1%742
/blog/trail-under-150446127.2%287

AI Crawler Traffic

Why AI Crawler Traffic Matters

AI crawler traffic is the record of requests that AI companies' bots make to your site: which bot fetched which page, when, and for what purpose. It matters because citation has a precondition. An engine cannot cite a page it has never fetched, so a page with no crawler hits has a distribution problem, not a quality problem, and no amount of rewriting will fix it.
Crawler traffic is also the fastest feedback loop in AI search. Visibility moves over weeks. A crawl either happened today or it did not.
The two silent failures: a page nothing crawls is invisible, and a page crawled constantly but never cited is being read and rejected. They look the same in a visibility report and need opposite fixes.

AI Crawler Traffic

The Bots Siftly Tracks

CrawlerRun byWhat the visit is for
GPTBotOpenAITraining: content may inform a future model
OAI-SearchBotOpenAIIndexing: content can be surfaced and cited in answers
PerplexityBotPerplexityIndexing: content can be cited in answers
Google-ExtendedGoogleTraining: governs use in Gemini and grounded answers
ClaudeBotAnthropicTraining and retrieval, depending on the product
BingbotMicrosoftIndexing: feeds Copilot as well as search
Which of these can reach you is a robots.txt decision. Run the free crawler audit to see what your current rules allow before you change anything.

How it works

Indexing Versus Training

An indexing fetch is a citation opportunity

The bot is building a retrievable index. A page it fetched can appear as a source in an answer within days, so this is the visit to want on anything you have just published.

A training fetch is a long bet

The bot is gathering data for a future model. There is no near-term citation, but it shapes what the model knows about your brand once it ships.

The two are governed separately

Allowing one does not allow the other. Many sites block training and unintentionally block the indexing bot beside it, which removes them from answers entirely.

Decide per bot, not per company

OpenAI runs both kinds. Blocking GPTBot while allowing OAI-SearchBot keeps you out of training and in answers, which is the setting most brands actually want.

AI Crawler Traffic

What To Do With The Data

No hits

Access problem: check robots.txt, the firewall, and the CDN

Hits, no citations

Content problem: the page is read and passed over

Training only

Indexing bot is blocked or has not reached the page

Hits after publish

The page is answer-ready, so start measuring visibility

Read crawl next to citations for the same page and the diagnosis is usually immediate. That is why both live in the same table rather than in two dashboards.
Before you block anything: blocking a training crawler is a reasonable brand decision. Blocking an indexing crawler removes you from the answers your buyers are reading. Know which one a rule affects before you ship it.

Getting started

Time to value, not time to configure

Hour 1

Point the log at your site

Connect your site through the integration or the log endpoint. Crawler hits start appearing immediately; there is no waiting period before the first data.

Week 1

Read the baseline

See which bots visit, how often, and which sections of the site they favour. Pages nothing has fetched are the first thing to fix.

Week 2

Publish and watch

Ship a page, then watch for the indexing fetch that makes it citable. If it never comes, the problem is access, not content.

Questions

FAQ

It is the record of AI companies' bots fetching pages on your site: which bot, which page, when, and whether the visit was for indexing or for training. It is the earliest evidence that a page can be cited, because an engine cannot cite what it has never fetched.
The named bots run by the major AI companies, including GPTBot and OAI-SearchBot from OpenAI, PerplexityBot, Google-Extended, ClaudeBot, and Bingbot. Each hit records the bot, the path, the time, and the country.
An indexing crawl builds a retrievable index, so the page can appear as a source in an answer soon after. A training crawl gathers data that may inform a future model, with no near-term citation. They are governed by separate rules, which is why blocking one does not block the other.
It means access is fine and the content is losing on merit. The engine fetched the page and chose something else as its source. That is a content and structure problem, and the citation data tells you which page it chose instead.
That depends on the bot. Blocking a training crawler is a defensible brand decision. Blocking an indexing crawler takes you out of the answers your buyers read, which is usually the opposite of the intent. Check what your current rules actually do before changing them.

Your buyers are asking AI.
Be the answer.