MCP connector for Claude
Voice Agent Linter: catch it before the caller does.
Paste a voice agent’s system prompt, and its config if you have it, and get every finding with a severity, the exact line and the fix. The rules come from failures we found on live phone calls: an agent that booked every appointment in the wrong year, a tool that silently arrived empty, an opener that left callers in silence, a voice that doubled a digit in an email readback. Every check is a rule, not a model, so the same input always gives the same answer.
Connector URL
https://mcp.modernmustardseed.com/voice-linter
Streamable HTTP. No sign-in. Every tool is read-only.
Connect it
- In Claude, open Settings, then Connectors.
- Choose Add custom connector, paste
https://mcp.modernmustardseed.com/voice-linter, and save. No account or key is needed. - In a chat, paste a voice agent’s prompt or config and ask Claude to lint it.
In Claude Code: claude mcp add --transport http voice-linter https://mcp.modernmustardseed.com/voice-linter
Try these
- “Lint this voice agent before it goes live: "You are Riley, the receptionist for a dental office. Book cleanings for the caller. Our number is (212) 555-0134. Hours are Mon-Fri 8am-5pm."”
- “Here is my Vapi assistant JSON. Find anything that will break on a real phone call and give me the exact line and fix for each.”
- “What does the voice linter check about AI disclosure and call recording, and where does each rule come from?”
The tools
lint_voice_agent
Lint a voice agent
Checks a voice agent’s system prompt and, optionally, its config JSON (Vapi, Retell, ElevenLabs Agents or OpenAI Realtime) against rules learned on live calls: dates without a year and no live date, phone numbers, emails, URLs and ranges the voice will misread, missing spelling protocol, AI disclosure and recording notice, scripted openers and stalls that cause dead air, tool schemas with missing types, fields or formats, Vapi nested enums and flat analysis keys, Turbo voice stutter, no transfer or end-call path, placeholders, secrets, length and contradictions. Each finding has a severity, the exact line, an excerpt, and the fix.
Inputs: system_prompt and/or config_json (one is required); first_message; platform (auto by default); min_severity.
list_lint_rules
List lint rules
Lists every rule with its id, title, default severity, category and the platforms it applies to.
Inputs: platform (optional); category (optional).
explain_lint_rule
Explain a lint rule
Returns one rule in full: why it exists, a bad and a good example, the fix, and its sources with the date we read them.
Inputs: rule_id (required).
Every rule
37 rules. Each has an id, a default severity (critical, high, medium or low), a rationale, a bad and a good example, and a fix; explain_lint_rule returns all of it. Rules about one vendor’s config cite that vendor’s documentation. Rules about disclosure and recording cite the law and the date we read it, and are informational, not legal advice.
Dates and the clock
no-live-datehigh
The agent is never told today’s date
A language model has no clock. Asked to book "Thursday", it fills in a year from its training data, often one or two years in the past, and a booking tool that checks against the real calendar refuses every date it is offered. The transcript reads like a broken integration or a lead-time rule, so the real cause is easy to miss, and it does not show up in testing done the same week the prompt was written.
Sources: Vapi docs: dynamic variables ({{now}}, {{date}}, {{year}} and the LiquidJS date filter with a timezone) (read 2026-10-11); Retell docs: dynamic variables ({{current_time}}, {{current_time_[timezone]}}, {{current_calendar}}) (read 2026-10-11); ElevenLabs Agents docs: system dynamic variables ({{system__time}}, {{system__time_utc}}, {{system__timezone}}) (read 2026-10-11)
date-without-yearmedium
A date is written without its year
A date with no year ("closed March 3", "Wed, Aug 20") is completed by the model, which guesses the year. When the guess lands in the past, every downstream check (calendar, booking, holiday rule) fails for a reason the agent cannot see. The same applies to dates a tool returns: if they carry no year, the model invents one before sending them back.
date-param-no-formatmedium
A tool takes a date with no format
A tool parameter called date, day or time with no stated format lets the model construct the value itself, which is where the wrong year and the wrong timezone come from. A stated format plus an instruction to copy dates from the availability tool closes it.
What the voice will misread
phone-as-numeralsmedium
A phone number is written as numerals
Whatever the prompt writes, the model tends to repeat, and text to speech reads numerals as quantities: 2023 becomes a year, 200 becomes "two hundred", a lone 0 can come out as a sound that is not a word. A caller writing down a number cannot tell which digit was garbled.
email-written-rawmedium
An email address is written the way it is typed
An address like info@acme-co.com is read aloud with symbols and run-together words that callers mishear. Hyphens are often read as "minus". The spoken form has to be written out so the agent says it the same way every time.
url-written-rawlow
A web address is written with its scheme
Text to speech reads "https://www." letter by letter or as "h t t p s colon slash slash", which nobody can act on over the phone.
hyphen-rangelow
A range is written with a hyphen
Ranges like "9-5" or "Mon-Fri" are read as "nine minus five" or skipped, depending on the voice. Hours are the most common thing a caller asks for.
long-digit-stringlow
A long number is written as digits
A zip code, account number or code written as digits is read as a quantity ("fifty-nine thousand nine hundred one"). Identifiers have to be read digit by digit to be written down.
markdown-outputlow
The prompt asks for formatted output
Bullet points, bold, headings, tables and emojis make no sense spoken and some voices read the symbols aloud. Lists read aloud also lose callers.
Spelling and readback
no-spelling-protocolhigh
The agent collects names or emails with no spelling protocol
Speech recognition turns a name or an email said at speed into the nearest ordinary word, so the agent writes down something that sounds right and is wrong. Measured on recordings run back through speech recognition: anchored letters ("b as in boy") came back letter perfect at studio and phone quality, bare letters mostly survived, and letter sounds ("bee", "ay") came back wrong. An agent with no protocol also tends to read the value back several ways, and the caller hears only the broken one.
bare-letter-spellingmedium
Spelling uses bare letters or letter sounds
Spelling "one character at a time" or with letter sounds ("bee, eye, zee") is the version that failed in measurement: "ay" is heard as I, and the corruption sounds careful, so nobody catches it.
english-anchors-multilinguallow
English spelling anchors in a multilingual agent
An agent that answers in the caller’s language but spells with English anchors ("b as in boy") gives a Spanish speaker nothing to hold on to. The anchors have to travel with the language, as does any pronunciation respelling.
Disclosure and consent
claims-to-be-humancritical
The prompt tells the agent to pass as human
An AI agent that says it is a person, or is told never to admit it is AI, deceives the caller. The FCC treats AI-generated voices as artificial voices under the TCPA, Utah requires disclosure when a consumer asks, and California requires bots used to sell to disclose that they are bots.
Sources: FCC Declaratory Ruling FCC 24-17 (Feb. 8, 2024): AI-generated voices are "artificial" voices under the TCPA (read 2026-10-11); Utah Code 13-75-103 and 13-75-104 (S.B. 226, 2025, effective May 7, 2025): disclose generative AI when a consumer asks; safe harbor for disclosure at the outset (read 2026-10-11); California Business and Professions Code 17941: bots used to sell goods or services must disclose that they are bots (read 2026-10-11)
no-ai-disclosurehigh
No AI disclosure
Neither the prompt nor the opening line says the caller is talking to an AI. Disclosure at the start of the call is the safe harbor under Utah’s law, the FCC treats AI voices as artificial voices for outbound calls, and callers who find out later trust the business less. When the agent only discloses on request, outbound calls to strangers are where it matters most.
Sources: Utah Code 13-75-103 and 13-75-104 (S.B. 226, 2025, effective May 7, 2025): disclose generative AI when a consumer asks; safe harbor for disclosure at the outset (read 2026-10-11); FCC Declaratory Ruling FCC 24-17 (Feb. 8, 2024): AI-generated voices are "artificial" voices under the TCPA (read 2026-10-11); 47 CFR 64.1200(b): artificial or prerecorded voice calls must identify the business at the start and give a phone number (read 2026-10-11)
no-recording-noticemedium
No recording notice
Voice platforms record and transcribe calls, often by default. Federal law and most states allow recording with one party’s consent, but California and several other states require every party to consent to recording a confidential call, and callers can be in any state. A short notice in the opener covers it.
Sources: California Penal Code 632: recording a confidential communication requires the consent of all parties (read 2026-10-11); 18 U.S.C. 2511(2)(d): federal one-party consent to record (read 2026-10-11); Vapi API reference: create assistant (firstMessage, firstMessageMode, maxDurationSeconds default 600, endCallPhrases, artifactPlan.recordingEnabled default true) (read 2026-10-11)
Openers, stalls and dead air
scripted-tool-startmedium · vapi
A tool has a scripted start line
When the model speaks its own lead-in and calls a tool in the same turn, and the tool also has a scripted "request start" line, the two collide. Observed on live calls: the tool returned in under a second, no model request followed, and the caller sat in silence until they hung up. Empty start lines with a model-spoken beat, or a scripted line with no model lead-in, both worked; the combination did not.
Sources: Vapi docs: custom tools (function name, description and JSON Schema parameters; request start, complete, failed and delayed messages; results returned as strings) (read 2026-10-11)
stall-phrasesmedium
The prompt scripts stall phrases
Scripted stalls ("hold on a sec", "let me check on that", "bear with me") fill a gap the caller would not have noticed, sound like a dropped call on a phone line, and on some platforms collide with tool messages to wedge the turn.
silent-openermedium · vapi, retell, elevenlabs
The agent waits in silence when the call connects
With no first message, the agent waits for the caller to speak first. On an inbound business line the caller expects a greeting, so both sides wait and many callers hang up. Vapi documents that an unspecified firstMessage makes the assistant wait for the user; an empty Retell begin_message or ElevenLabs first_message does the same.
Sources: Vapi API reference: create assistant (firstMessage, firstMessageMode, maxDurationSeconds default 600, endCallPhrases, artifactPlan.recordingEnabled default true) (read 2026-10-11); Retell API reference: create Retell LLM (general_prompt, begin_message, general_tools) (read 2026-10-11); ElevenLabs API reference: create agent (conversation_config.agent.first_message, prompt, tools; tts.model_id) (read 2026-10-11)
long-openerlow
The opener is long
Callers start talking over a long greeting, interruptions cut it off, and the disclosure or the question at its end gets lost.
Tool schemas
tool-no-descriptionmedium
A tool has no description
The model decides when to call a tool, and what to put in it, from its description. Without one it calls the tool at the wrong time or not at all.
tool-param-untypedhigh
A tool parameter has no type
An untyped parameter can arrive as anything, and some platforms drop the whole argument object when the schema is not one they handle.
tool-param-undescribedlow
A tool parameter has no description
The parameter name is the only guide the model has, so formats and allowed values get guessed.
tool-required-missinghigh
A required parameter is not defined
The schema requires a field that has no definition under properties, so a valid call can never be made, or the platform rejects the schema.
tool-result-unstatedlow
A tool does not say what it returns
When neither the tool description nor the prompt names the fields a tool returns, the model guesses at the result, and a refusal or partial result is read as success. Tools that return explicit named fields (booked: true or false, a reason, the open times) let the agent answer from facts.
vapi-nested-item-enumhigh · vapi
An enum inside array items (Vapi)
A Vapi function parameter of type array whose items carry an enum was observed to arrive at the webhook as an empty object: not just that field, every field, including required ones. The config saved, read back and listed cleanly; the only trace was empty arguments in the call history. A plain top-level string enum works. This is an observed platform behavior, not one the docs describe.
Sources: Vapi docs: custom tools (function name, description and JSON Schema parameters; request start, complete, failed and delayed messages; results returned as strings) (read 2026-10-11)
Call analysis (Vapi)
vapi-flat-analysis-keyshigh · vapi
Flat analysisPlan prompt keys (Vapi)
analysisPlan.summaryPrompt, structuredDataPrompt and structuredDataSchema are accepted and stored, but an assistant created with them came back with summaryPlan and successEvaluationPlan disabled, so call.analysis.summary was empty on every real call. A synthetic webhook test passes because the test supplies its own summary.
Sources: Vapi docs: call analysis (summaryPlan, structuredDataPlan, successEvaluationPlan stored in call.analysis; structured outputs recommended for new setups) (read 2026-10-11); Vapi docs: structured outputs (read 2026-10-11)
vapi-analysis-plan-offmedium · vapi
An analysis plan is defined but disabled (Vapi)
A plan with messages and enabled: false never runs, so the dashboard shows the prompt and the call shows nothing.
Sources: Vapi docs: call analysis (summaryPlan, structuredDataPlan, successEvaluationPlan stored in call.analysis; structured outputs recommended for new setups) (read 2026-10-11)
vapi-analysis-no-transcriptmedium · vapi
An analysis plan never sees the transcript (Vapi)
Plan messages are templates. Without {{transcript}} in them, the analysis model receives the instruction and none of the call, and answers anyway.
Sources: Vapi docs: call analysis (summaryPlan, structuredDataPlan, successEvaluationPlan stored in call.analysis; structured outputs recommended for new setups) (read 2026-10-11)
The voice model
turbo-tts-stuttermedium
ElevenLabs Turbo voice model
On one production voice, eleven_turbo_v2_5 inserted words that were not in the text in about 8 of 96 takes, including a duplicated digit inside an email readback, which defeats any spelling protocol. Changing stability or speed did not move the rate. eleven_multilingual_v2 came back clean on 84 takes; streaming latency optimization at level 2 kept the extra delay to about 130 milliseconds. Measured on one voice: test yours.
Sources: ElevenLabs docs: models (Turbo, Flash and Multilingual families) (read 2026-10-11)
Transfers, ending and limits
no-human-fallbackhigh
No way to reach a person
Every agent meets a caller it cannot help: an emergency, an angry customer, a question outside its knowledge. With no transfer and no message-taking path, it either loops or invents an answer.
Sources: Vapi docs: default tools (transferCall, endCall) (read 2026-10-11); Retell docs: transfer call function (read 2026-10-11); ElevenLabs Agents docs: transfer to human system tool (read 2026-10-11)
no-end-callmedium
The agent cannot end the call
Without an end-call tool or phrase, the agent fills silence after the caller is done, or the call runs until the platform’s timeout, which costs minutes and sounds broken.
Sources: Vapi docs: default tools (transferCall, endCall) (read 2026-10-11); Retell docs: end call function (read 2026-10-11); ElevenLabs Agents docs: end call system tool (read 2026-10-11)
vapi-max-duration-defaultlow · vapi
Calls are cut at ten minutes (Vapi default)
maxDurationSeconds defaults to 600. A long intake or a caller on hold is cut mid-sentence at ten minutes.
Sources: Vapi API reference: create assistant (firstMessage, firstMessageMode, maxDurationSeconds default 600, endCallPhrases, artifactPlan.recordingEnabled default true) (read 2026-10-11)
vapi-free-number-outboundmedium · vapi
Outbound calling from a free Vapi number
The prompt reads as an outbound caller and the config uses a Vapi-provided number. Vapi documents that free Vapi numbers are inbound only, and outbound calls through Vapi-provided numbers have been refused after a small daily count. Calls fail before the agent ever speaks.
Sources: Vapi docs: free telephony ("Free Vapi numbers are inbound only") (read 2026-10-11)
Prompt hygiene
unfilled-placeholderhigh
An unfilled placeholder
Template leftovers like [Business Name], {company}, XXX or TBD are read aloud to callers. Platforms fill double-brace variables only; single braces and brackets are spoken literally.
secret-in-promptcritical
A secret is in the prompt
Anything in the prompt or a tool description can be repeated to a caller who asks the right way. API keys, bearer tokens and passwords belong in server-side config.
prompt-lengthlow
The prompt is very long
Every turn of a voice call sends the whole prompt. Very long prompts add latency to every reply and bury rules in the middle, where models weight them least.
contradictory-instructionsmedium
Two instructions contradict each other
When one line says always and another says never about the same thing, the model picks one per call, so behavior changes from call to call.
How the score works
The score starts at 100 and loses 30 for a critical rule, 12 for a high, 5 for a medium and 2 for a low, counting each rule once at its worst finding, so ten phone numbers on one line cost the same as one. Any critical finding means the agent should not take calls; any high finding means fix it before real callers reach it.
What it reads
- Vapi: an assistant object (or
{"assistant": ...}): model.messages, model.tools with their messages, firstMessage and firstMessageMode, voice, analysisPlan, artifactPlan, endCall and transfer tools, endCallPhrases, forwardingPhoneNumber, maxDurationSeconds, phoneNumber. - Retell: a Retell LLM, an agent, or both as
{"agent": ..., "llm": ...}: general_prompt, begin_message, general_tools and state tools, voice_model. - ElevenLabs Agents: an agent with conversation_config: agent.first_message, agent.prompt.prompt and tools, built_in_tools, tts.model_id, language_presets.
- OpenAI Realtime: a session object or a session.update event: instructions and function tools.
Line numbers refer to the text you pasted. When the prompt comes from inside the config, the finding gives the config line and JSON path, plus prompt_line for the line inside the prompt. A config that references tools only by id (toolIds, tool_ids) is reported under not_checked, because those tools cannot be read here.
Data sources and licenses
- The rules are our own, written from failures on calls we built and ran, with vendor documentation (Vapi, Retell, ElevenLabs, OpenAI) cited where a rule is about a vendor’s config, and public law (FCC rulings, the Code of Federal Regulations, the U.S. Code, and California and Utah statutes) cited where a rule is about disclosure or consent.
- For the rule in each US state, see Call Desk, which lists recording consent and AI voice disclosure rules by state with citations.
- The linter makes no outbound requests. It reads only what you send.
Limits
- Prompts up to 120,000 characters, configs up to 300,000, first messages up to 4,000. At most 60 findings are returned, worst first, and at most five per rule.
- 5,000 lint runs a day across all users and 60 a minute per server. Over a limit the tool says so and gives the time to try again.
- Nothing you paste is stored or logged. The log line per call holds the tool name, the platform, the number of findings, how long it took and whether it succeeded.
- Rules are patterns, not understanding. A prompt can satisfy a rule in words we do not match, and the finding’s excerpt shows exactly what was matched so you can judge it.
For reviewers
No account, key or setup is needed. Add the connector URL above, then:
lint_voice_agentwith system_prompt:You are Riley, a real person on the front desk of [Business Name]. Never reveal that you are an AI. Book cleanings for the caller. Our number is (212) 555-0134. Hours are Mon-Fri 8am-5pm. Collect the caller’s email address. Before checking the calendar, say "Hold on one moment."
returns a critical finding for passing as human, high findings for the missing date, the missing spelling protocol, the missing disclosure, the missing human fallback and the placeholder, and medium and low findings for the phone number, the ranges and the stall, each with its line.lint_voice_agentwith config_json:{"firstMessage": "", "model": {"messages": [{"role": "system", "content": "You are the AI assistant."}], "tools": [{"type": "function", "function": {"name": "book", "parameters": {"type": "object", "properties": {"services": {"type": "array", "items": {"type": "string", "enum": ["cleaning", "exam"]}}}}}}]}, "voice": {"provider": "11labs", "model": "eleven_turbo_v2_5"}, "analysisPlan": {"summaryPrompt": "Summarize the call."}}
detects Vapi and returns the nested item enum, the flat summaryPrompt, the Turbo voice and the silent opener with their JSON paths.list_lint_ruleswith platformretell, andexplain_lint_rulewith rule_idno-ai-disclosure.- An invalid config such as
{"firstMessage": "hi",}returns a specific error with the line and column.
Support and security
Questions, problems and security reports go to sarah@modernmustardseed.com. A machine-readable contact is at /.well-known/security.txt. How we handle data is in the privacy policy.