Tip 1 of 31: Let AI Read Your Website
One Client Generated £45,000 in Placement Fees in a Single Month. Here Is Where We Started.
One of our clients generated £45,000 in placement fees attributed to their website in a single month. Their attribution data also indicated that AI-assisted search and referral sources contributed to website visits.
This is the first post in a series breaking down the practical work completed on the website. Each tip is specific and actionable. Start here.
Tip 1: Let Useful Crawlers Access Your Website
This sounds obvious. It is not.
In 22 years of building recruitment websites, we have tested hundreds of sites where crawler access has been restricted by robots.txt, firewall rules, content delivery networks or bot-protection settings.
Some restrictions are intentional and valuable. Others block search or retrieval agents that the website owner would prefer to admit.
If you use Cloudflare, review its AI-bot settings rather than assuming that every AI-related agent is allowed or blocked. Cloudflare currently provides controls for blocking AI bots and analysing AI crawler activity.
Do not simply disable all bot protection. Decide which agents serve a useful purpose, verify them where possible and allow only the access that supports your objectives.
Excellent content cannot be directly retrieved by a particular crawler or user agent when that agent is prevented from accessing it. Allowing access, however, does not guarantee indexing, citations, recommendations or commercial results.
The Quick Check: Can Search and Retrieval Agents Access Your Pages?
One useful starting point is checking whether your website allows the agents used by the services you want to support.
For example, OpenAI documents three separate agents:
OAI-SearchBot, which supports search discoveryChatGPT-User, which may visit a page in response to a user requestGPTBot, which is associated with potential model training
These controls are independent. A website can permit search discovery while making a different decision about model-training access.
In Plain English
Most websites have a file called robots.txt at the root of the domain.
Think of it as a set of instructions telling compliant crawlers which areas they may access. It is primarily a crawler-management mechanism, not a security system.
Ask your website developer to check whether relevant public pages can be accessed by the search and retrieval agents you have chosen to support.
They should also inspect firewall and content delivery network settings. A crawler can be permitted by robots.txt but still be blocked before it reaches the website.
Search, Retrieval and Training Are Different
Do not treat every AI-related user agent as though it does the same job.
There are three broad categories:
- Search crawlers, which discover or index public content for a provider’s search functionality
- User-requested agents, which retrieve a page when a person asks a service to access it
- Training crawlers, which collect public content that may contribute to model development
For example, Anthropic documents Claude-SearchBot, Claude-User and ClaudeBot for these different purposes.
Perplexity similarly documents PerplexityBot for search discovery and Perplexity-User for user-requested retrieval. It advises website owners using firewalls to verify and configure access using its published network information.
The exact agent names and provider policies can change. Your developer should use each provider’s current official documentation rather than copying an old generic list into robots.txt.
Agents to Review
Depending on which services you want to support, ask your developer to review access for:
OAI-SearchBotChatGPT-UserClaude-SearchBotClaude-UserPerplexityBotPerplexity-User
This is not a universal allowlist.
Your organisation may choose to permit some agents and block others. Training crawlers should be considered separately from search and user-requested retrieval.
Google Search access should also be reviewed independently. Do not assume that allowing or blocking a Google AI-related product token has the same effect as controlling Googlebot.
Check More Than Robots.txt
A complete review should include:
- The live
robots.txtfile - Cloudflare or other content delivery network settings
- Web application firewall rules
- Bot-management settings
- Server access logs
- Rate-limiting rules
- CAPTCHA or challenge pages
- HTTP status codes returned to crawlers
- Rules affecting individual subdomains
- Published verification details for legitimate agents
Do not trust a request solely because its user-agent string contains the name of a recognised company. User-agent strings can be copied, so use provider-supplied IP information or verification methods where available.
What Comes Next
Allowing appropriate access is only the first step.
Crawler access does not make a page useful, authoritative or relevant. It simply removes one possible technical barrier to direct discovery or retrieval.
The next tips in this series cover indexing controls, snippet settings, sitemaps, canonical URLs and the structure of the content itself.
Frequently Asked Questions
Does blocking AI-related crawlers affect my visibility?
It can. Blocking a search crawler may prevent that provider from directly discovering or indexing your pages through that crawler. Blocking a user-requested agent may prevent it from fetching a page when someone asks for it. The effect depends on the particular agent and system, and allowing access does not guarantee that your website will be cited or recommended.
Will allowing these crawlers create security or performance problems?
Legitimate crawlers normally request public pages, but all automated traffic consumes some server resources and user-agent names can be impersonated. Use appropriate firewall rules, rate limits, logs and the verification information published by each provider rather than allowing every request that claims to be an AI crawler.
How do I know whether my website is blocking useful crawlers?
Ask your developer to review your robots.txt file, server logs, firewall rules, content delivery network and bot-management settings. A robots.txt file can permit access while a firewall still blocks the request, so both crawler instructions and server-level access must be checked.
Is allowing a crawler the same as being cited by an AI-assisted system?
No. Access only makes direct discovery or retrieval technically possible for that agent. Whether a page is selected, cited or recommended can depend on the query, the system’s sources, content relevance, quality, indexing and other factors. There are no guarantees.
Do I need to allow model-training crawlers as well as search crawlers?
No. Search crawling, user-requested retrieval and model training are separate activities. You can allow search or retrieval agents while blocking a provider’s training crawler where the provider supports separate controls. The right choice depends on your organisation’s commercial, legal and content-use policies.
Darren Revell, Co-Founder, RecruiterWEB
Co-Founder, RecruiterWEB
Darren Revell began working in recruitment technology in 2004 when he founded Recruitwise Technology. He later became a founder of RecruiterWEB, which acquired the Recruitwise Technology brand, platform and customer base in 2016. Darren remains Co-Founder and Co-Owner of RecruiterWEB.
Darren came to Rectech after eleven years working in recruitment. He started as a trainee recruiter in 1993 and progressed through the ranks to recruiter, billing manager, billing director, and eventually recruitment company owner. During that career, he delivered permanent hires, contract hires, client campaign advertising, team moves, retained search, master vendor services, and RPO.
In 2004, he switched focus to recruitment technology and began building websites and job boards specifically for recruitment agencies. RecruiterWEB has since built websites for 667+ agencies and executive search firms in the UK and internationally. The platform runs on custom code built explicitly for recruitment, with built-in job board functionality, ATS and job poster integration, Google for Jobs structured data, and GDPR-compliant candidate registration included as standard on every plan.
Darren writes on recruitment website design, SEO and AI visibility for recruitment agencies, candidate data protection, and the commercial impact of digital investment on recruitment businesses.
Specialist Areas
- Recruitment website design and technology
- SEO and AI visibility for recruitment agencies
- ATS and job poster integration (Bullhorn, Vincere, idibu, and others)
- GDPR and candidate data protection
- Branding for recruitment agencies and executive search firms
Connect
LinkedIn: linkedin.com/in/
Phone: 01223 655278
Darren has also appeared as a guest on the RecTalk podcast, covering his background in recruitment and the founding of RecruiterWEB.


