MCP connector for Claude

Voice Agent Linter: catch it before the caller does.

Paste a voice agent’s system prompt, and its config if you have it, and get every finding with a severity, the exact line and the fix. The rules come from failures we found on live phone calls: an agent that booked every appointment in the wrong year, a tool that silently arrived empty, an opener that left callers in silence, a voice that doubled a digit in an email readback. Every check is a rule, not a model, so the same input always gives the same answer.

Connector URL

https://mcp.modernmustardseed.com/voice-linter

Streamable HTTP. No sign-in. Every tool is read-only.

Connect it

  1. In Claude, open Settings, then Connectors.
  2. Choose Add custom connector, paste https://mcp.modernmustardseed.com/voice-linter, and save. No account or key is needed.
  3. In a chat, paste a voice agent’s prompt or config and ask Claude to lint it.

In Claude Code: claude mcp add --transport http voice-linter https://mcp.modernmustardseed.com/voice-linter

Try these

The tools

lint_voice_agent

Lint a voice agent

Checks a voice agent’s system prompt and, optionally, its config JSON (Vapi, Retell, ElevenLabs Agents or OpenAI Realtime) against rules learned on live calls: dates without a year and no live date, phone numbers, emails, URLs and ranges the voice will misread, missing spelling protocol, AI disclosure and recording notice, scripted openers and stalls that cause dead air, tool schemas with missing types, fields or formats, Vapi nested enums and flat analysis keys, Turbo voice stutter, no transfer or end-call path, placeholders, secrets, length and contradictions. Each finding has a severity, the exact line, an excerpt, and the fix.

Inputs: system_prompt and/or config_json (one is required); first_message; platform (auto by default); min_severity.

list_lint_rules

List lint rules

Lists every rule with its id, title, default severity, category and the platforms it applies to.

Inputs: platform (optional); category (optional).

explain_lint_rule

Explain a lint rule

Returns one rule in full: why it exists, a bad and a good example, the fix, and its sources with the date we read them.

Inputs: rule_id (required).

Every rule

37 rules. Each has an id, a default severity (critical, high, medium or low), a rationale, a bad and a good example, and a fix; explain_lint_rule returns all of it. Rules about one vendor’s config cite that vendor’s documentation. Rules about disclosure and recording cite the law and the date we read it, and are informational, not legal advice.

Dates and the clock

What the voice will misread

  • phone-as-numerals

    medium

    A phone number is written as numerals

    Whatever the prompt writes, the model tends to repeat, and text to speech reads numerals as quantities: 2023 becomes a year, 200 becomes "two hundred", a lone 0 can come out as a sound that is not a word. A caller writing down a number cannot tell which digit was garbled.

  • email-written-raw

    medium

    An email address is written the way it is typed

    An address like info@acme-co.com is read aloud with symbols and run-together words that callers mishear. Hyphens are often read as "minus". The spoken form has to be written out so the agent says it the same way every time.

  • url-written-raw

    low

    A web address is written with its scheme

    Text to speech reads "https://www." letter by letter or as "h t t p s colon slash slash", which nobody can act on over the phone.

  • hyphen-range

    low

    A range is written with a hyphen

    Ranges like "9-5" or "Mon-Fri" are read as "nine minus five" or skipped, depending on the voice. Hours are the most common thing a caller asks for.

  • long-digit-string

    low

    A long number is written as digits

    A zip code, account number or code written as digits is read as a quantity ("fifty-nine thousand nine hundred one"). Identifiers have to be read digit by digit to be written down.

  • markdown-output

    low

    The prompt asks for formatted output

    Bullet points, bold, headings, tables and emojis make no sense spoken and some voices read the symbols aloud. Lists read aloud also lose callers.

Spelling and readback

  • no-spelling-protocol

    high

    The agent collects names or emails with no spelling protocol

    Speech recognition turns a name or an email said at speed into the nearest ordinary word, so the agent writes down something that sounds right and is wrong. Measured on recordings run back through speech recognition: anchored letters ("b as in boy") came back letter perfect at studio and phone quality, bare letters mostly survived, and letter sounds ("bee", "ay") came back wrong. An agent with no protocol also tends to read the value back several ways, and the caller hears only the broken one.

  • bare-letter-spelling

    medium

    Spelling uses bare letters or letter sounds

    Spelling "one character at a time" or with letter sounds ("bee, eye, zee") is the version that failed in measurement: "ay" is heard as I, and the corruption sounds careful, so nobody catches it.

  • english-anchors-multilingual

    low

    English spelling anchors in a multilingual agent

    An agent that answers in the caller’s language but spells with English anchors ("b as in boy") gives a Spanish speaker nothing to hold on to. The anchors have to travel with the language, as does any pronunciation respelling.

Disclosure and consent

Openers, stalls and dead air

Tool schemas

  • tool-no-description

    medium

    A tool has no description

    The model decides when to call a tool, and what to put in it, from its description. Without one it calls the tool at the wrong time or not at all.

  • tool-param-untyped

    high

    A tool parameter has no type

    An untyped parameter can arrive as anything, and some platforms drop the whole argument object when the schema is not one they handle.

  • tool-param-undescribed

    low

    A tool parameter has no description

    The parameter name is the only guide the model has, so formats and allowed values get guessed.

  • tool-required-missing

    high

    A required parameter is not defined

    The schema requires a field that has no definition under properties, so a valid call can never be made, or the platform rejects the schema.

  • tool-result-unstated

    low

    A tool does not say what it returns

    When neither the tool description nor the prompt names the fields a tool returns, the model guesses at the result, and a refusal or partial result is read as success. Tools that return explicit named fields (booked: true or false, a reason, the open times) let the agent answer from facts.

  • vapi-nested-item-enum

    high · vapi

    An enum inside array items (Vapi)

    A Vapi function parameter of type array whose items carry an enum was observed to arrive at the webhook as an empty object: not just that field, every field, including required ones. The config saved, read back and listed cleanly; the only trace was empty arguments in the call history. A plain top-level string enum works. This is an observed platform behavior, not one the docs describe.

    Sources: Vapi docs: custom tools (function name, description and JSON Schema parameters; request start, complete, failed and delayed messages; results returned as strings) (read 2026-10-11)

Call analysis (Vapi)

The voice model

  • turbo-tts-stutter

    medium

    ElevenLabs Turbo voice model

    On one production voice, eleven_turbo_v2_5 inserted words that were not in the text in about 8 of 96 takes, including a duplicated digit inside an email readback, which defeats any spelling protocol. Changing stability or speed did not move the rate. eleven_multilingual_v2 came back clean on 84 takes; streaming latency optimization at level 2 kept the extra delay to about 130 milliseconds. Measured on one voice: test yours.

    Sources: ElevenLabs docs: models (Turbo, Flash and Multilingual families) (read 2026-10-11)

Transfers, ending and limits

Prompt hygiene

  • unfilled-placeholder

    high

    An unfilled placeholder

    Template leftovers like [Business Name], {company}, XXX or TBD are read aloud to callers. Platforms fill double-brace variables only; single braces and brackets are spoken literally.

  • secret-in-prompt

    critical

    A secret is in the prompt

    Anything in the prompt or a tool description can be repeated to a caller who asks the right way. API keys, bearer tokens and passwords belong in server-side config.

  • prompt-length

    low

    The prompt is very long

    Every turn of a voice call sends the whole prompt. Very long prompts add latency to every reply and bury rules in the middle, where models weight them least.

  • contradictory-instructions

    medium

    Two instructions contradict each other

    When one line says always and another says never about the same thing, the model picks one per call, so behavior changes from call to call.

How the score works

The score starts at 100 and loses 30 for a critical rule, 12 for a high, 5 for a medium and 2 for a low, counting each rule once at its worst finding, so ten phone numbers on one line cost the same as one. Any critical finding means the agent should not take calls; any high finding means fix it before real callers reach it.

What it reads

Line numbers refer to the text you pasted. When the prompt comes from inside the config, the finding gives the config line and JSON path, plus prompt_line for the line inside the prompt. A config that references tools only by id (toolIds, tool_ids) is reported under not_checked, because those tools cannot be read here.

Data sources and licenses

Limits

For reviewers

No account, key or setup is needed. Add the connector URL above, then:

  1. lint_voice_agent with system_prompt:
    You are Riley, a real person on the front desk of [Business Name]. Never reveal that you are an AI. Book cleanings for the caller. Our number is (212) 555-0134. Hours are Mon-Fri 8am-5pm. Collect the caller’s email address. Before checking the calendar, say "Hold on one moment."
    returns a critical finding for passing as human, high findings for the missing date, the missing spelling protocol, the missing disclosure, the missing human fallback and the placeholder, and medium and low findings for the phone number, the ranges and the stall, each with its line.
  2. lint_voice_agent with config_json:
    {"firstMessage": "", "model": {"messages": [{"role": "system", "content": "You are the AI assistant."}], "tools": [{"type": "function", "function": {"name": "book", "parameters": {"type": "object", "properties": {"services": {"type": "array", "items": {"type": "string", "enum": ["cleaning", "exam"]}}}}}}]}, "voice": {"provider": "11labs", "model": "eleven_turbo_v2_5"}, "analysisPlan": {"summaryPrompt": "Summarize the call."}}
    detects Vapi and returns the nested item enum, the flat summaryPrompt, the Turbo voice and the silent opener with their JSON paths.
  3. list_lint_rules with platform retell, and explain_lint_rule with rule_id no-ai-disclosure.
  4. An invalid config such as {"firstMessage": "hi",} returns a specific error with the line and column.

Support and security

Questions, problems and security reports go to sarah@modernmustardseed.com. A machine-readable contact is at /.well-known/security.txt. How we handle data is in the privacy policy.