Infrastructure feature
Watch AI crawlers read your site
Before an engine can cite you, its crawler has to fetch you. Siftly shows every AI crawler hit on your site as it happens: which bot, which page, whether the visit was for indexing or for training, and where it came from. It is the earliest signal you have that new content will be citable.
the live log
Crawl comes before citation
Every bot hit, as it lands. GPTBot, OAI-SearchBot, PerplexityBot, Google-Extended and the rest, each with the time, the path, and the country the request came from. When you publish, you can watch the crawlers arrive instead of guessing whether they did.
Two visits that mean different things. An indexing fetch means the page can be cited in an answer now. A training fetch means it may inform the model later. Siftly labels each hit, so you can tell a page that is answer-ready from one that is only feeding a future model.
per page
Turn crawl data into a content decision
Which of your pages AI actually reads. Crawl counts sit next to citations for the same page, so you can separate the two failures that look identical from the outside: a page nothing crawls, and a page that is crawled constantly and never cited.
Bots on one side, people on the other. Connect analytics and Search Console and the same table carries AI sessions, human sessions, clicks, and impressions. That is the full picture of a page: who fetched it, who cited it, and who arrived because of it.
AI Crawler Traffic
Why AI Crawler Traffic Matters
AI Crawler Traffic
The Bots Siftly Tracks
| Crawler | Run by | What the visit is for |
|---|---|---|
| GPTBot | OpenAI | Training: content may inform a future model |
| OAI-SearchBot | OpenAI | Indexing: content can be surfaced and cited in answers |
| PerplexityBot | Perplexity | Indexing: content can be cited in answers |
| Google-Extended | Training: governs use in Gemini and grounded answers | |
| ClaudeBot | Anthropic | Training and retrieval, depending on the product |
| Bingbot | Microsoft | Indexing: feeds Copilot as well as search |
How it works
Indexing Versus Training
An indexing fetch is a citation opportunity
The bot is building a retrievable index. A page it fetched can appear as a source in an answer within days, so this is the visit to want on anything you have just published.
A training fetch is a long bet
The bot is gathering data for a future model. There is no near-term citation, but it shapes what the model knows about your brand once it ships.
The two are governed separately
Allowing one does not allow the other. Many sites block training and unintentionally block the indexing bot beside it, which removes them from answers entirely.
Decide per bot, not per company
OpenAI runs both kinds. Blocking GPTBot while allowing OAI-SearchBot keeps you out of training and in answers, which is the setting most brands actually want.
AI Crawler Traffic
What To Do With The Data
No hits
Access problem: check robots.txt, the firewall, and the CDN
Hits, no citations
Content problem: the page is read and passed over
Training only
Indexing bot is blocked or has not reached the page
Hits after publish
The page is answer-ready, so start measuring visibility
Getting started
Time to value, not time to configure
Hour 1
Point the log at your site
Connect your site through the integration or the log endpoint. Crawler hits start appearing immediately; there is no waiting period before the first data.
Week 1
Read the baseline
See which bots visit, how often, and which sections of the site they favour. Pages nothing has fetched are the first thing to fix.
Week 2
Publish and watch
Ship a page, then watch for the indexing fetch that makes it citable. If it never comes, the problem is access, not content.
Questions