Skip to Main Content Enable / Disable keyboard navigation (ENTER key) Accessible Menu Accessibility panel Reset accessibility Sitemap Accessibility statement
Help Center
Welcome to our knowledge base, we are here to help.
Bytheweb 3 dots icon

Robots.txt AI Policy Manager

ByTheWeb GEO includes a Robots.txt AI Policy Manager for controlling access by selected AI and search crawlers through your WordPress virtual robots.txt file.

By default, ByTheWeb GEO does not block any of the supported crawlers.

You decide individually which user-agents, if any, should receive a site-wide Disallow: / rule.

The feature is available when ByTheWeb GEO is managing the site’s overlapping SEO and robots.txt functionality. If Yoast SEO or Rank Math SEO is active, ByTheWeb GEO leaves this functionality to the external SEO provider.

For a product overview, see:

Robots.txt AI Policy Manager Feature Overview

No ByTheWeb AI API Key or AI credits are required.

What Is robots.txt?

robots.txt is a crawler policy file normally available at the root of a website:

https://example.com/robots.txt

It can contain instructions for specific crawler user-agents.

ByTheWeb GEO works with the virtual robots.txt output provided through WordPress.

This is different from placing a physical robots.txt file directly in the website root.

What Does Robots.txt AI Policy Manager Do?

When ByTheWeb GEO is responsible for the site’s robots output, the feature can:

  • Keep all supported crawlers allowed by default.
  • Let you selectively block individual supported user-agents.
  • Add the ByTheWeb GEO XML Sitemap URL to the virtual robots.txt output.
  • Provide direct access to the current /robots.txt URL.
  • Warn when a physical robots.txt file may prevent WordPress virtual rules from being used.
  • Remind you that cache, security, firewall, or server-level rules can affect the final output.

Nothing is blocked automatically simply because ByTheWeb GEO is installed.

Where Do I Find the Robots.txt AI Policy Manager?

In your WordPress Dashboard, go to:

ByTheWeb > SEO Settings > Robots.txt

The page is labeled:

Robots.txt AI Policy Manager

Starting with ByTheWeb GEO 1.3.2, the section also includes an Enable Robots.txt AI Controls option. When enabled, ByTheWeb GEO can manage the supported AI and search crawler rules in the WordPress virtual robots.txt. When disabled, ByTheWeb GEO does not add its crawler-policy rules to the robots.txt output.

When ByTheWeb GEO is the active manager for these settings, you will see the supported crawler groups and a:

Block in robots.txt

checkbox for each crawler.

When Does ByTheWeb GEO Manage robots.txt?

Starting with ByTheWeb GEO 1.3.2, ByTheWeb GEO manages its robots.txt functionality when Enable Robots.txt AI Controls is enabled and a supported external SEO provider is not managing the overlapping SEO layer.

This means the ByTheWeb GEO Robots.txt AI Policy Manager is active when ByTheWeb GEO is responsible for this functionality.

When Yoast SEO or Rank Math SEO is active as the supported external SEO provider, ByTheWeb GEO does not add its own crawler policy to robots.txt.

This prevents competing robots.txt management.

For the broader provider behavior, see:

Requirements & Compatibility

What Happens If Yoast SEO Is Active?

When Yoast SEO is the active supported SEO provider, ByTheWeb GEO leaves the overlapping robots.txt functionality to Yoast.

The ByTheWeb GEO SEO Settings screen displays a notice that the external SEO provider is active.

The overlapping ByTheWeb GEO controls are disabled rather than creating separate competing output.

ByTheWeb GEO does not add its own selected crawler-blocking rules or its own Sitemap reference through the Robots module in this state.

What Happens If Rank Math SEO Is Active?

The same provider-aware behavior applies to Rank Math SEO.

When Rank Math is the active supported external SEO provider:

  • ByTheWeb GEO does not manage its own robots.txt policy output.
  • The overlapping ByTheWeb GEO controls are disabled.
  • ByTheWeb GEO does not append its own crawler rules through the Robots module.
  • Rank Math remains responsible for the overlapping SEO functionality.

You do not need to deactivate ByTheWeb GEO itself.

Its non-overlapping GEO functionality can continue to be used.

What Is the Default Crawler Policy?

The default policy is open.

ByTheWeb GEO displays the following guidance in the settings:

All bots are allowed by default. Enable “Block in robots.txt” only for bots you explicitly want to block.

If you do not select any crawler, ByTheWeb GEO does not create Disallow: / rules for the supported user-agents.

Which Crawlers Can I Control?

The current Robots.txt AI Policy Manager organizes supported user-agents into the following groups.

OpenAI / ChatGPT

  • GPTBot
  • OAI-SearchBot
  • ChatGPT-User

Anthropic / Claude

  • ClaudeBot
  • Claude-SearchBot
  • Claude-User

Perplexity

  • PerplexityBot
  • Perplexity-User

Google

  • Googlebot
  • Google-Extended
  • GoogleOther

Microsoft / Bing

  • Bingbot

Apple

  • Applebot
  • Applebot-Extended

Amazon

  • Amazonbot
  • Amzn-SearchBot
  • Amzn-User

Meta

  • meta-externalagent
  • meta-externalfetcher

Each crawler is controlled independently.

How Do I Block a Crawler?

Go to:

ByTheWeb > SEO Settings > Robots.txt

Find the crawler you want to restrict.

Enable:

Block in robots.txt

for that crawler.

Then save the SEO Settings.

ByTheWeb GEO stores the selected crawler policy and adds the corresponding rule to the WordPress virtual robots.txt output while ByTheWeb GEO is responsible for robots management.

What Rule Is Added When I Block a Crawler?

For each selected crawler, ByTheWeb GEO creates a rule in this format:

User-agent: ExampleBot

Disallow: /

For example, if a supported crawler named ExampleBot were selected, the rule would instruct that user-agent not to crawl paths under the site root.

The actual user-agent name comes from the crawler selected in the Robots.txt AI Policy Manager.

Does Blocking a Crawler Apply to the Whole Website?

The generated ByTheWeb GEO rule uses:

Disallow: /

for the selected user-agent.

The / applies from the site root, so this is a site-wide crawler directive for that selected user-agent.

The current Robots.txt AI Policy Manager does not create a different path-specific rule for individual pages through these crawler checkboxes.

Can I Block One Crawler Without Blocking the Others?

Yes.

Each supported user-agent has its own Block in robots.txt control.

For example, selecting one crawler does not automatically select all other crawlers in the same company group.

Only the user-agents you explicitly select are included in the ByTheWeb GEO blocking rules.

Does ByTheWeb GEO Automatically Add My Sitemap to robots.txt?

Yes, when ByTheWeb GEO is managing the robots.txt output and the site is publicly indexable.

The plugin uses its ByTheWeb GEO Sitemap URL:

https://example.com/sitemap.xml

and adds it to the virtual robots.txt output if that exact Sitemap URL is not already present.

The generated section is identified as the ByTheWeb GEO Sitemap.

For Sitemap configuration, see:

XML Sitemap

What Happens If No Crawlers Are Blocked?

ByTheWeb GEO can still add its Sitemap reference when it is the active robots manager and the site is public.

No crawler-specific Disallow: / rules are added if no supported bots have been selected.

This means installing the feature does not automatically restrict crawler access.

Does Robots.txt AI Policy Manager Replace Page-Level Index or Noindex Settings?

No.

These are different controls.

The Robots.txt AI Policy Manager works at the crawler user-agent level.

Page-level robots settings determine the indexing directive for a particular Post, Page, or supported Taxonomy.

For example, making a page noindex is not the same as blocking a crawler through robots.txt.

For page-level controls, see:

Canonical URL & Robots

What Is the Difference Between robots.txt and noindex?

robots.txt controls crawler access instructions.

A page-level noindex directive controls whether supported search engines should index a particular page.

Because they serve different purposes, you should not treat a crawler block in robots.txt as a replacement for a page’s index/noindex configuration.

ByTheWeb GEO provides these controls separately.

What Happens If a Physical robots.txt File Exists?

ByTheWeb GEO checks whether a physical robots.txt file exists in the website root.

If one is detected, the settings screen displays:

Physical robots.txt file detected

with a warning that WordPress virtual robots.txt rules may not be applied until the physical file is removed or updated manually.

This is important because ByTheWeb GEO adds its rules through WordPress’s virtual robots.txt system.

A physical file can take precedence over that virtual output.

Do I Need to Create a Physical robots.txt File?

No.

The Robots.txt AI Policy Manager is designed to work with the WordPress virtual robots.txt output.

When that virtual output is being used correctly, you do not need to manually create a physical file just to use the ByTheWeb GEO crawler controls.

Can Cache or Security Tools Affect robots.txt?

Yes.

The ByTheWeb GEO settings specifically warn that the final robots.txt response can also be affected by:

  • Cache systems
  • Security plugins
  • Firewalls
  • Server-level rules
  • A physical robots.txt file

As a result, the settings saved in WordPress and the file ultimately returned at /robots.txt may not always be identical if another system is intercepting or replacing the response.

Always check the public robots.txt URL after making changes.

How Do I View My Current robots.txt File?

When ByTheWeb GEO is managing the feature, the Robots.txt settings provide the website’s current robots URL together with:

Copy URL

and:

Open

controls.

You can also open it directly:

https://example.com/robots.txt

Use the public URL to verify the actual output being returned by the website.

What Should I Look for in the robots.txt Output?

If ByTheWeb GEO is the active robots manager and the site is public, you can check for the ByTheWeb GEO Sitemap reference.

If you selected crawler blocks, you can also check that the relevant user-agents appear with:

Disallow: /

Only bots that were explicitly selected should appear in the ByTheWeb GEO crawler-policy section.

What Happens When WordPress Search Engine Visibility Is Disabled?

WordPress has a site-wide setting under:

Settings > Reading > Search Engine Visibility

If the site is configured to:

Discourage search engines from indexing this site

ByTheWeb GEO displays a site-wide indexing warning in its SEO Settings.

For the Robots module specifically, ByTheWeb GEO does not append its own Sitemap or selected crawler-policy rules through its virtual robots filter while WordPress reports the site as non-public.

The WordPress site-wide visibility setting therefore needs to be considered separately from the individual crawler controls.

Does Blocking Google-Extended Block Googlebot?

No.

They are separate user-agents in the ByTheWeb GEO interface.

Googlebot and Google-Extended have independent controls.

The current plugin description for Google-Extended specifically distinguishes it from standard Google Search crawling.

Selecting one does not automatically select the other.

Does Blocking Applebot-Extended Block Applebot?

No.

They are also separate user-agents.

Applebot and Applebot-Extended each have their own checkbox.

Only the selected user-agent receives the ByTheWeb GEO blocking rule.

Are Search Crawlers and AI Crawlers the Same Thing?

Not necessarily.

The Robots.txt AI Policy Manager intentionally lists different user-agents separately because they can represent different crawler or fetcher functions.

For example, the interface contains separate entries for:

  • GPTBot
  • OAI-SearchBot
  • ChatGPT-User

and separate entries for:

  • ClaudeBot
  • Claude-SearchBot
  • Claude-User

This lets you make the crawler decision per user-agent instead of using one global switch for an entire provider.

Should I Block AI Crawlers?

ByTheWeb GEO does not automatically decide this for you.

The feature starts with all supported bots allowed.

You should enable Block in robots.txt only for user-agents you intentionally want to restrict.

The plugin also provides descriptions beside the supported bots to help explain what each user-agent represents and, where relevant, warns that blocking certain search-oriented crawlers may affect discovery through the corresponding service.

Can Blocking Search Crawlers Affect Search Visibility?

It can.

For example, the current ByTheWeb GEO interface warns that blocking Googlebot can prevent Google from crawling the site and that blocking Bingbot can prevent Bing from crawling it.

The interface also warns that blocking certain AI search crawlers may reduce visibility in the related search or answer experience.

Review the description shown beside a crawler before blocking it.

Is robots.txt a Security Feature?

No.

robots.txt is a crawler policy mechanism.

It should not be used as a method for protecting:

  • Private information
  • Passwords
  • Confidential content
  • Restricted administration areas
  • Sensitive files

The ByTheWeb GEO Robots.txt AI Policy Manager controls the policy sent to crawlers. It is not an access-control or authentication system.

Does robots.txt Guarantee That a Bot Will Not Access My Website?

ByTheWeb GEO creates the configured robots.txt directives and serves them through the applicable WordPress virtual output.

The feature itself does not technically authenticate or block incoming HTTP requests from a crawler.

It should therefore be understood as a crawler policy rather than as a firewall or security block.

Does the Feature Use AI Credits?

No.

Robots.txt AI Policy Manager is a local ByTheWeb GEO configuration feature.

It does not require:

  • A ByTheWeb AI API Key
  • A ByTheWeb Cloud account
  • AI credits

Changing or viewing crawler policies does not make an AI-generation request.

Does This Feature Send My robots.txt Settings to an AI Model?

No.

The Robots.txt AI Policy Manager itself generates the applicable WordPress robots.txt rules locally.

It does not need to send the selected crawler settings to an external AI model in order to produce those rules.

What Happens If ByTheWeb GEO Is Deactivated?

ByTheWeb GEO provides its crawler rules dynamically through WordPress while the plugin and Robots module are active.

If ByTheWeb GEO is deactivated, it no longer adds its Robots.txt AI Policy Manager output to the WordPress virtual robots.txt response.

Other WordPress, server, SEO plugin, security, or physical robots.txt behavior can still affect the file independently.

Why Are the Robots Controls Disabled?

If the individual crawler controls are unavailable, first confirm that Enable Robots.txt AI Controls is enabled.

If the controls are visible but disabled, first check whether Yoast SEO or Rank Math SEO is active.

When a supported external SEO provider is active, ByTheWeb GEO intentionally disables the overlapping SEO controls and leaves that functionality to the external provider.

A notice at the top of the SEO Settings identifies the active provider.

For more information, see:

Requirements & Compatibility

Why Are My Changes Not Visible in robots.txt?

Check the following:

  1. Confirm that Yoast SEO or Rank Math SEO is not currently managing the overlapping robots functionality.
  2. Confirm that you saved the ByTheWeb GEO SEO Settings.
  3. Open the public /robots.txt URL rather than relying only on the settings screen.
  4. Check whether a physical robots.txt file exists.
  5. Check whether a cache, security plugin, firewall, CDN, hosting configuration, or server-level rule is affecting the response.
  6. Check the WordPress Settings > Reading > Search Engine Visibility setting.

If the output still does not match the saved configuration, see:

Troubleshooting

What If the Robots Module Is Not Loaded?

If ByTheWeb GEO cannot load the supported crawler definitions, the settings screen can display:

Robots module is not loaded

and:

No crawler list is available. Please make sure the ByTheWeb GEO Robots module file is loaded correctly.

This is not the normal operating state.

If you see this message, use the Troubleshooting guide or contact support.

Who Can Access These Settings?

The ByTheWeb GEO SEO Settings area uses WordPress capability-based access rather than relying only on a specific role name.

The main GEO settings interfaces require the applicable edit_others_posts capability.

For a broader explanation of ByTheWeb permissions, see:

User Access Control & Permissions

Recommended Setup

For most new installations, start by reviewing the default state before making any crawler restrictions.

  1. Open ByTheWeb > SEO Settings > Robots.txt.
  2. Confirm that ByTheWeb GEO is the active manager for the overlapping SEO functionality.
  3. Leave crawlers unblocked unless you have intentionally decided to restrict a particular user-agent.
  4. Read the description shown beside a crawler before changing its policy.
  5. Enable Block in robots.txt only for the specific user-agents you want to restrict.
  6. Save the settings.
  7. Open the public /robots.txt URL.
  8. Confirm that the ByTheWeb GEO Sitemap reference is present when applicable.
  9. Confirm that only the selected user-agents have a Disallow: / rule.
  10. If the public output differs from the settings, check for a physical file, cache, security, firewall, CDN, or server-level rule.

Summary

The Robots.txt AI Policy Manager gives you crawler-specific control over the WordPress virtual robots.txt output when ByTheWeb GEO is responsible for the site’s overlapping SEO and robots functionality.

Key points:

  • All supported crawlers are allowed by default.
  • No bot is blocked automatically.
  • Each supported user-agent can be controlled separately.
  • A selected crawler receives a site-wide Disallow: / directive.
  • ByTheWeb GEO can add its /sitemap.xml URL to the virtual robots output.
  • Yoast SEO and Rank Math SEO take ownership of the overlapping functionality when active.
  • A physical robots.txt file can prevent WordPress virtual rules from being applied.
  • Cache, security, firewall, CDN, hosting, and server-level rules can affect the final response.
  • The feature does not use AI credits.
  • robots.txt is a crawler policy, not a security system.

Where Should I Go Next?

For the product overview:

Robots.txt AI Policy Manager Feature Overview

For the XML Sitemap added to the applicable robots.txt output:

XML Sitemap

For page-level index/noindex and canonical controls:

Canonical URL & Robots

For AI-oriented content discovery files:

llms.txt & llms-full.txt

For Yoast SEO and Rank Math compatibility:

Requirements & Compatibility

For permissions:

User Access Control & Permissions

For robots.txt problems:

Troubleshooting