MCP connector for Claude
AI Search Kit: built to be cited.
AI Search Kit helps any business publish what answer engines read: crawler rules that let AI search agents in, structured data that states who the business is, an llms.txt that points to the pages that matter, and the same name, address and phone everywhere. It audits, it writes the files, and it shows its work. Every rule it applies links to the public spec or vendor page it comes from, with the date we read it. Nothing here can promise that an AI will recommend a business; it makes a business easier to find, read and cite.
Connector URL
https://mcp.modernmustardseed.com/ai-search-kit
Streamable HTTP. No sign-in. Every tool is read-only. Rules as of 2026-10-11; schema.org release 30.1.
Connect it
- In Claude, open Settings, then Connectors.
- Choose Add custom connector, paste
https://mcp.modernmustardseed.com/ai-search-kit, and save. No account or key is needed. - In a chat, ask about a business website, its schema, its llms.txt or its listings. Claude calls the tools when the question needs them.
In Claude Code: claude mcp add --transport http ai-search-kit https://mcp.modernmustardseed.com/ai-search-kit
Try these
- “Audit warbyparker.com for AI search readiness: which AI crawlers can read it, whether its structured data is valid, and what to fix first.”
- “Write LocalBusiness schema for Juniper Ridge Roofing, a roofing contractor at 125 Example Ave, Kalispell, MT 59901, phone (406) 555-0142, website juniperridgeroofing.example, open Monday to Friday 7 AM to 5 PM, serving Kalispell, Whitefish and Columbia Falls.”
- “Generate an llms.txt for allbirds.com from its live pages.”
The tools
audit_ai_readiness
Audit AI search readiness
Reads one public page plus its site's robots.txt, llms.txt and sitemap. Reports, per AI crawler (OAI-SearchBot, ChatGPT-User, GPTBot, Claude-SearchBot, Claude-User, ClaudeBot, PerplexityBot, Perplexity-User, Googlebot, Google-Extended, Applebot, Applebot-Extended), whether robots.txt lets it read the page under RFC 9309 matching; noindex and nosnippet directives; JSON-LD validity against schema.org and Google's LocalBusiness rules; whether the name, phone and address in the markup also appear on the visible page; llms.txt format; canonical tags; and the sitemap. Every finding cites the spec or vendor page behind it, with the date it was read.
Inputs: url (required); detail: "summary" (default, omits passing findings) or "full".
generate_local_business_schema
Generate LocalBusiness schema
Turns business facts into schema.org JSON-LD using the most specific LocalBusiness subtype (130 to choose from, such as RoofingContractor, Dentist or Brewery), with address, geo, hours in Google's documented format, areaServed, sameAs, founders and services. Validates its own output against the schema.org vocabulary and returns a ready-to-paste script tag plus notes on anything left out.
Inputs: name (required); business_type or schema_type; url, telephone, email, logo, images, street_address, address_locality, address_region, postal_code, address_country, hide_street_address, latitude, longitude, area_served, hours, open_24_7, seasonal_closures, same_as, price_range, payment_accepted, currencies_accepted, founders, founding_date, services, serves_cuisine, menu_url, accepts_reservations (all optional).
generate_llms_txt
Generate llms.txt
Writes an llms.txt file in the llmstxt.org format, either from facts you supply (site name, summary, sections of links) or from a live site: its home page, sitemap and up to 30 pages' own titles and descriptions, honoring the site's robots.txt and skipping noindex pages. Checks the result against the format before returning it.
Inputs: url or site_name (one is required); summary, details, sections, optional_links, max_pages (all optional).
check_entity_consistency
Check entity consistency
Compares business name, phone, street, city, region and postal code across up to 8 sources: URLs it reads (JSON-LD first, then tap-to-call links and page text) and facts you supply for profiles that block automated readers. Normalizes formats so "123 North Main Street" matches "123 N Main St", then reports each field as consistent, variant or mismatch, with the exact fixes.
Inputs: profiles (required, 1 to 8, each a url and or supplied facts); business (optional canonical record to compare against).
validate_structured_data
Validate structured data
Validates pasted JSON-LD against the schema.org vocabulary (every type, property, domain, range and enumeration value) and Google's documented LocalBusiness and Organization rules. No network access.
Inputs: json_ld (required): the JSON-LD text, with or without its script tag.
The AI crawlers we check
robots.txt is read the way RFC 9309 says a crawler must: tokens match without regard to case, every group naming an agent is merged, the * group applies only when no group names the agent, and the longest matching rule wins, with Allow winning a tie. A missing robots.txt (4xx) allows everything; a server error (5xx) or no answer means compliant crawlers assume everything is blocked. Blocking a search or user-fetch agent is reported as a failure. Opting out of a training agent or control token is reported as information: it is the site owner's choice, and the vendors document it as separate from their search products.
| Token | Operator | Role | What the vendor says it does |
|---|---|---|---|
| OAI-SearchBot | OpenAI | Search | Surfaces websites in search results in ChatGPT search features. Source |
| ChatGPT-User | OpenAI | User fetch | Visits a page when a ChatGPT or Custom GPT user action asks for it. Not used for automatic crawling. Source |
| GPTBot | OpenAI | Training | Crawls content that may be used to train OpenAI generative AI foundation models. Source |
| Claude-SearchBot | Anthropic | Search | Navigates the web to improve search result quality for Claude users. Source |
| Claude-User | Anthropic | User fetch | Accesses a website when a Claude user asks a question that needs it. Source |
| ClaudeBot | Anthropic | Training | Collects web content to help enhance the utility and safety of Anthropic generative AI models (model training). Source |
| PerplexityBot | Perplexity | Search | Surfaces and links websites in Perplexity search results. Not used to crawl content for AI foundation models. Source |
| Perplexity-User | Perplexity | User fetch | Visits a page when a Perplexity user asks a question, and may link it in the answer. Source |
| Googlebot | Search | Crawls for Google Search. Google: a page must be indexed and eligible to show in Search with a snippet to appear as a supporting link in AI Overviews and AI Mode. Source | |
| Google-Extended | Control token | Controls whether content Google crawls may be used for training future Gemini models and for grounding. Not a separate crawler. Source | |
| Applebot | Apple | Search | Crawls for Apple search features in Spotlight, Siri and Safari. Its data may also help train Apple foundation models. Source |
| Applebot-Extended | Apple | Control token | Controls whether Applebot data may train Apple generative AI models. Does not crawl. Source |
Where every rule comes from
Tool output repeats the citation and the date beside each finding, so an answer can always be checked. Structured data is validated against the full schema.org vocabulary (every type, property, domain, range and enumeration value), then against the properties Google documents for local businesses and organizations. Google states its AI features need no special schema or AI text file, and our output says so where it matters: structured data and llms.txt help machines read a business accurately; they are not a switch that makes an engine cite it.
- RFC 9309: Robots Exclusion Protocol (IETF, read 2026-10-11)
- Overview of OpenAI crawlers (OpenAI, read 2026-10-11)
- Does Anthropic crawl data from the web, and how can site owners block the crawler? (Anthropic, read 2026-10-11)
- Perplexity Crawlers (Perplexity, read 2026-10-11)
- Google common crawlers (Google-Extended) (Google, read 2026-10-11)
- AI features and your website (Google, read 2026-10-11)
- About Applebot (Apple, read 2026-10-11)
- The /llms.txt file (llmstxt.org, read 2026-10-11)
- Local business (LocalBusiness) structured data (Google, read 2026-10-11)
- Organization (Organization) structured data (Google, read 2026-10-11)
- General structured data guidelines (Google, read 2026-10-11)
- How to specify a canonical URL with rel="canonical" and other methods (Google, read 2026-10-11)
- Sitemaps XML format (sitemaps.org, read 2026-10-11)
- schema.org vocabulary, release 30.1 (schema.org, read 2026-10-11)
- JSON-LD 1.1 (W3C, read 2026-10-11)
- Guidelines for representing your business on Google (Google, read 2026-10-11)
- Actions: property -input and -output annotations (schema.org, read 2026-10-11)
- sameAs (schema.org, read 2026-10-11)
Data sources and licenses
- Websites you name. The audit, llms.txt and consistency tools read public pages, robots.txt, llms.txt and sitemaps of the addresses in the request, as a browser would, identified by the user agent
ModernMustardSeedAISearchKit/1.0. Building an llms.txt from a site honors that site's robots.txt for our user agent. - schema.org vocabulary, release 30.1. Vendored into the server from schema.org's published release file and used under the Creative Commons Attribution-ShareAlike 3.0 license (schema.org terms).
- Vendor and standards documents. The pages listed above are linked and summarized in our own words, never copied in bulk.
- No other API, no model and no database of ours is called. Nothing is posted or published anywhere.
Limits
- Audits are saved for 10 minutes and site-built llms.txt drafts for 30 minutes. Repeating the same request in that time returns the saved result, marked cached.
- Fresh audits across all users: 60 an hour and 600 a day, one per site every 10 minutes.
- Site-built llms.txt: 20 an hour and 200 a day, one per site every 30 minutes, reading at most 30 pages (20 by default), four at a time.
- Profile reads for the consistency check: 200 an hour and 1500 a day, 6 per site every 10 minutes, up to 8 sources per call.
- Over a limit, the tool says which one and the time to try again. Limits are shared by everyone, never set per person or per network.
- Many profile sites (Google Maps, Facebook, Yelp) build their pages with scripts or refuse automated readers. The consistency check says so per source and takes the facts shown there when you supply them.
What it will not fetch
Every request goes through the same guard as our Site Check connector: public web servers on ports 80 and 443 only; private, loopback, link-local, carrier-grade NAT, multicast and reserved addresses (IPv4 and IPv6) refused; names like localhost and .internal refused; every redirect checked again, at most five followed; every response capped in size and time.
Troubleshooting
- “The page could not be read”
- The site is down, the name does not resolve, or a firewall refuses automated readers. Open it in a browser to confirm, then try again after ten minutes.
- llms.txt shows as missing but the site has one
- The file must be at the root,
/llms.txt, and answer with plain text. Sites that answer every unknown path with their home page are reported as returning HTML. - A profile comes back not_readable
- That site did not give its content to an automated reader. Add the name, phone and address shown on it to the same source and run the check again.
- The schema type is LocalBusiness instead of something specific
- No subtype matched the words in business_type. Pass schema_type with the exact schema.org name, such as
PlumberorAutoRepair.
For reviewers
No account, key or setup is needed. Add the connector URL above, then:
audit_ai_readinesswith urlwarbyparker.comreturns crawler verdicts for 12 agents, structured data, llms.txt, canonical and sitemap findings. Adddetail: "full"to include passing findings.generate_local_business_schemawith nameJuniper Ridge Roofing, business_typeroofing company, telephone(406) 555-0142, address_localityKalispell, address_regionMTand hours[{"days":["Monday-Friday"],"opens":"7 AM","closes":"5 PM"}]returns RoofingContractor JSON-LD with zero validation errors.generate_llms_txtwith urlallbirds.combuilds a file from the live site; with site_nameExample Coand one section of links it builds one from facts.check_entity_consistencywith profiles[{"url":"allbirds.com"}, {"label":"Google","name":"Allbirds","phone":"(406) 555-0100"}]reads the site and reports the phone mismatch against the supplied test number.validate_structured_datawith json_ld{"@context":"https://schema.org","@type":"localBusiness","Name":"X"}returns errors with case-sensitive hints.- Refusals:
http://169.254.169.254,http://localhostandhttp://10.0.0.1each return a specific error and are never fetched.
Support and security
AI Search Kit by Modern Mustard Seed is built and run by Modern Mustard Seed LLC. Questions, problems and security reports go to sarah@modernmustardseed.com. A machine-readable contact is at /.well-known/security.txt. How we handle data is in the privacy policy.