MCP connector for Claude

AI Search Kit: built to be cited.

AI Search Kit helps any business publish what answer engines read: crawler rules that let AI search agents in, structured data that states who the business is, an llms.txt that points to the pages that matter, and the same name, address and phone everywhere. It audits, it writes the files, and it shows its work. Every rule it applies links to the public spec or vendor page it comes from, with the date we read it. Nothing here can promise that an AI will recommend a business; it makes a business easier to find, read and cite.

Connector URL

https://mcp.modernmustardseed.com/ai-search-kit

Streamable HTTP. No sign-in. Every tool is read-only. Rules as of 2026-10-11; schema.org release 30.1.

Connect it

  1. In Claude, open Settings, then Connectors.
  2. Choose Add custom connector, paste https://mcp.modernmustardseed.com/ai-search-kit, and save. No account or key is needed.
  3. In a chat, ask about a business website, its schema, its llms.txt or its listings. Claude calls the tools when the question needs them.

In Claude Code: claude mcp add --transport http ai-search-kit https://mcp.modernmustardseed.com/ai-search-kit

Try these

The tools

audit_ai_readiness

Audit AI search readiness

Reads one public page plus its site's robots.txt, llms.txt and sitemap. Reports, per AI crawler (OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Googlebot, Google-Extended, Applebot, Applebot-Extended), whether robots.txt lets it read the page under RFC 9309 matching; noindex and nosnippet directives; JSON-LD validity against schema.org and Google's LocalBusiness rules; whether the name, phone and address in the markup also appear on the visible page; llms.txt format; canonical tags; and the sitemap. Every finding cites the spec or vendor page behind it, with the date it was read.

Inputs: url (required); detail: "summary" (default, omits passing findings) or "full".

generate_local_business_schema

Generate LocalBusiness schema

Turns business facts into schema.org JSON-LD using the most specific LocalBusiness subtype (130 to choose from, such as RoofingContractor, Dentist or Brewery), with address, geo, hours in Google's documented format, areaServed, sameAs, founders and services. Validates its own output against the schema.org vocabulary and returns a ready-to-paste script tag plus notes on anything left out.

Inputs: name (required); business_type or schema_type; url, telephone, email, logo, images, street_address, address_locality, address_region, postal_code, address_country, hide_street_address, latitude, longitude, area_served, hours, open_24_7, seasonal_closures, same_as, price_range, payment_accepted, currencies_accepted, founders, founding_date, services, serves_cuisine, menu_url, accepts_reservations (all optional).

generate_llms_txt

Generate llms.txt

Writes an llms.txt file in the llmstxt.org format, either from facts you supply (site name, summary, sections of links) or from a live site: its home page, sitemap and up to 30 pages' own titles and descriptions, honoring the site's robots.txt and skipping noindex pages. Checks the result against the format before returning it.

Inputs: url or site_name (one is required); summary, details, sections, optional_links, max_pages (all optional).

check_entity_consistency

Check entity consistency

Compares business name, phone, street, city, region and postal code across up to 8 sources: URLs it reads (JSON-LD first, then tap-to-call links and page text) and facts you supply for profiles that block automated readers. Normalizes formats so "123 North Main Street" matches "123 N Main St", then reports each field as consistent, variant or mismatch, with the exact fixes.

Inputs: profiles (required, 1 to 8, each a url and or supplied facts); business (optional canonical record to compare against).

validate_structured_data

Validate structured data

Validates pasted JSON-LD against the schema.org vocabulary (every type, property, domain, range and enumeration value) and Google's documented LocalBusiness and Organization rules. No network access.

Inputs: json_ld (required): the JSON-LD text, with or without its script tag.

The AI crawlers we check

robots.txt is read the way RFC 9309 says a crawler must: tokens match without regard to case, every group naming an agent is merged, the * group applies only when no group names the agent, and the longest matching rule wins, with Allow winning a tie. A missing robots.txt (4xx) allows everything; a server error (5xx) or no answer means compliant crawlers assume everything is blocked. Blocking a search or user-fetch agent is reported as a failure. Opting out of a training agent or control token is reported as information: it is the site owner's choice, and the vendors document it as separate from their search products.

TokenOperatorRoleWhat the vendor says it does
OAI-SearchBotOpenAISearchSurfaces websites in search results in ChatGPT search features. Source
ChatGPT-UserOpenAIUser fetchVisits a page when a ChatGPT or Custom GPT user action asks for it. Not used for automatic crawling. Source
GPTBotOpenAITrainingCrawls content that may be used to train OpenAI generative AI foundation models. Source
Claude-SearchBotAnthropicSearchNavigates the web to improve search result quality for Claude users. Source
Claude-UserAnthropicUser fetchAccesses a website when a Claude user asks a question that needs it. Source
ClaudeBotAnthropicTrainingCollects web content to help enhance the utility and safety of Anthropic generative AI models (model training). Source
PerplexityBotPerplexitySearchSurfaces and links websites in Perplexity search results. Not used to crawl content for AI foundation models. Source
Perplexity-UserPerplexityUser fetchVisits a page when a Perplexity user asks a question, and may link it in the answer. Source
GooglebotGoogleSearchCrawls for Google Search. Google: a page must be indexed and eligible to show in Search with a snippet to appear as a supporting link in AI Overviews and AI Mode. Source
Google-ExtendedGoogleControl tokenControls whether content Google crawls may be used for training future Gemini models and for grounding. Not a separate crawler. Source
ApplebotAppleSearchCrawls for Apple search features in Spotlight, Siri and Safari. Its data may also help train Apple foundation models. Source
Applebot-ExtendedAppleControl tokenControls whether Applebot data may train Apple generative AI models. Does not crawl. Source

Where every rule comes from

Tool output repeats the citation and the date beside each finding, so an answer can always be checked. Structured data is validated against the full schema.org vocabulary (every type, property, domain, range and enumeration value), then against the properties Google documents for local businesses and organizations. Google states its AI features need no special schema or AI text file, and our output says so where it matters: structured data and llms.txt help machines read a business accurately; they are not a switch that makes an engine cite it.

Data sources and licenses

Limits

What it will not fetch

Every request goes through the same guard as our Site Check connector: public web servers on ports 80 and 443 only; private, loopback, link-local, carrier-grade NAT, multicast and reserved addresses (IPv4 and IPv6) refused; names like localhost and .internal refused; every redirect checked again, at most five followed; every response capped in size and time.

Troubleshooting

“The page could not be read”
The site is down, the name does not resolve, or a firewall refuses automated readers. Open it in a browser to confirm, then try again after ten minutes.
llms.txt shows as missing but the site has one
The file must be at the root, /llms.txt, and answer with plain text. Sites that answer every unknown path with their home page are reported as returning HTML.
A profile comes back not_readable
That site did not give its content to an automated reader. Add the name, phone and address shown on it to the same source and run the check again.
The schema type is LocalBusiness instead of something specific
No subtype matched the words in business_type. Pass schema_type with the exact schema.org name, such as Plumber or AutoRepair.

For reviewers

No account, key or setup is needed. Add the connector URL above, then:

  1. audit_ai_readiness with url warbyparker.com returns crawler verdicts for 12 agents, structured data, llms.txt, canonical and sitemap findings. Add detail: "full" to include passing findings.
  2. generate_local_business_schema with name Juniper Ridge Roofing, business_type roofing company, telephone (406) 555-0142, address_locality Kalispell, address_region MT and hours [{"days":["Monday-Friday"],"opens":"7 AM","closes":"5 PM"}] returns RoofingContractor JSON-LD with zero validation errors.
  3. generate_llms_txt with url allbirds.com builds a file from the live site; with site_name Example Co and one section of links it builds one from facts.
  4. check_entity_consistency with profiles [{"url":"allbirds.com"}, {"label":"Google","name":"Allbirds","phone":"(406) 555-0100"}] reads the site and reports the phone mismatch against the supplied test number.
  5. validate_structured_data with json_ld {"@context":"https://schema.org","@type":"localBusiness","Name":"X"} returns errors with case-sensitive hints.
  6. Refusals: http://169.254.169.254, http://localhost and http://10.0.0.1 each return a specific error and are never fetched.

Support and security

AI Search Kit by Modern Mustard Seed is built and run by Modern Mustard Seed LLC. Questions, problems and security reports go to sarah@modernmustardseed.com. A machine-readable contact is at /.well-known/security.txt. How we handle data is in the privacy policy.