Skip to article
Journal strattlabs.com · Markdown
← Back to Journal

live 8 min read

Your MCP server passes conformance. Will a connector directory accept it?

Protocol tests and connector review answer different questions. A practical guide to tool titles, safety annotations, descriptions and the evidence a directory needs.

The official MCP conformance suite is useful, public and worth running. It checks the protocol behaviors included in its scenarios. A connector directory asks additional questions about the product that speaks that protocol.

A valid tool definition can still be misleading. A successful request can still produce an unusable result. A server can pass a test suite and fail a directory's review without either process being wrong.

This is the English edition of our Czech analysis, originally published on 6 September 2026. The field observations below are from that earlier review, not a new September 17 scan. Links to current documentation were checked for this edition.

Start by naming what passed

A statement such as “MCP compliant” is incomplete without the protocol revision, test-suite version, scenarios run and results.

The official conformance repository distinguishes the evolving suite from frozen requirements for a particular protocol revision. It documents --requirements for selecting the latter. A passing result under one revision does not establish compatibility with another.

Record an exact package version or commit and the requirements selected. Preserve the machine-readable output alongside the claim. If a test was skipped or could not run, retain that distinction.

Our earlier Czech article discussed a snapshot of the suite and its release channels. Package versions and scenario counts are historical observations, not installation instructions for every future reader. Check the repository's current instructions before running the suite.

That establishes a reproducible protocol result. It still leaves the application's promises to be reviewed.

A directory has its own acceptance criteria

Claude's submission documentation requires tool titles and applicable safety annotations, alongside security, authentication and documentation requirements. Its review checklist and directory policy also covers tool naming, descriptions, functional behavior and ongoing maintenance.

Those requirements belong to that directory. They should not be presented as universal rules imposed on every MCP implementation or as a guarantee of how every other host makes permission decisions.

Directory membership is also distinct from whether someone can configure and use a remote connector. Decide whether your goal is protocol interoperability, a specific directory listing, a particular customer workflow, or all three.

For each goal, keep the relevant acceptance evidence separate.

Where the gaps appear

Three areas deserve an early review: names and titles, safety annotations, and descriptions.

Names and titles. A machine identifier and a readable title serve different audiences. A protocol-valid identifier does not automatically meet a host's display or directory requirements. Claude's published policy limits tool names to 64 characters. Check the host you are targeting instead of assuming your SDK's accepted input is the directory's accepted input.

Safety annotations. An annotation communicates a claim about behavior. It does not make that claim true. Calling a tool read-only cannot prevent its implementation from creating a job, sending a message or changing a record.

Descriptions. A tool description needs to explain the capability, inputs and meaningful limits. Instructions that attempt to control the host's broader behavior are a different kind of content. Read the description as something an agent will ingest alongside other tools, not as a sales paragraph viewed in isolation.

These distinctions are visible before anyone runs a production workflow.

Why hints need behavioral evidence

The MCP model treats tool annotations as hints. Consumers cannot assume that metadata from an arbitrary server truthfully describes its implementation. That is a trust boundary, not a defect in the protocol.

A directory reviewer, meanwhile, has to assess whether the metadata matches the product being submitted. The practical work includes tracing the tool's side effects, checking its actual authorization scope and exercising representative inputs.

For example, a tool advertised as a status lookup may call an upstream API that creates a record as part of the lookup. A generic “request” tool may support both observation and mutation. A helper that fetches URLs may expose an outbound network boundary even if it never writes business data.

None of those cases is resolved by adding a boolean field to the schema. Narrow tools, truthful descriptions and tests of actual behavior make the claim reviewable.

Our SSRF analysis examines the last case in detail.

A finding visible in one catalogue response

During the work behind the Czech article, we read the public tools/list response of a connector from a large vendor. Two descriptions directed the model's calling order: one told it always to call another tool first, and the other repeated that prerequisite. The definitions also lacked titles and safety annotations.

The observation was checked again on 4 September 2026. We did not call the tools. The vendor is intentionally unnamed, and this edition does not claim the definitions are still unchanged today.

The finding was a review concern, not proof that the connector had been rejected. There is a further distinction: a real data dependency between tools may need to be documented, while an instruction that tries to govern the host's behavior warrants a different kind of scrutiny. Explain required inputs and their provenance precisely; do not assume that every use of “before” proves an injection attempt.

This is why a static report needs the definition, the relevant criterion and a reasoned interpretation. Keyword matching alone is insufficient.

What our sample showed, and what it did not

The earlier investigation began with 398 randomly selected registry records that declared themselves accessible without authentication. Registry metadata was not a reliable access guarantee: the article reports that 43.2% actually required authorization. We could read 194 catalogues anonymously; other records did not become readable observations.

Those 194 endpoints represented 128 distinct operators, deduplicated by registrable domain. One operator accounted for 58 endpoints. Repeating a template across many deployments must not be mistaken for independent evidence from many businesses.

Our static checks found at least one finding on every readable endpoint. Findings classified as failures occurred on 164 endpoints, representing 102 operators; the remaining endpoints had at least one warning. These were our classifications against published criteria, not actual directory rejection decisions.

The readable sample had a median of four tools. Its accessibility and size bias it toward smaller deployments. It does not establish a failure rate for large commercial connectors or the whole MCP ecosystem.

Within the 128-operator sample, 69 operators had tool titles and 65 had safety hints according to the checks recorded in the original report. Only 17 operators were observed on revision 2026-07-28, so results for that subset are particularly limited.

The original report also compared several checks with different eligible populations. Their denominators must stay attached to each result. They should not be merged into a single universal pass rate.

One result is useful without that extrapolation: only 9 of the 128 operators passed all eight static tool-definition checks in that review. This measures visible definition quality under that checklist. It cannot tell us whether an annotation was truthful or whether an agent could complete the intended task.

Review in three passes

Pass 1: inspect the declared interface

Read your own catalogue as an external reviewer would. Check readable titles, target-host naming requirements, applicable annotations and descriptions that explain the tool's real scope.

For a tool with several unrelated responsibilities, decide whether separate operations would make its behavior clearer. For free-form queries or generic API access, identify the upstream interface and constraints. Do not conceal a large capability behind a reassuringly narrow name.

Record missing information as missing. Static inspection cannot turn an assertion into verified behavior.

Pass 2: exercise behavior in a controlled environment

Use test accounts and representative data. Check successful calls, invalid inputs, authorization boundaries, retries, timeouts and failures from upstream systems.

Pay particular attention to operations with side effects. Establish what happens when a request is repeated or a connection is interrupted after the underlying operation succeeds.

The result should be evidence you can hand to a reviewer: input, expected behavior, observed behavior and enough context to reproduce it. A green dashboard without that context is hard to evaluate.

Pass 3: test the actual user journey

A server can pass the first two passes and still be difficult for an agent to select or use.

Give the target host a realistic task. Observe whether it finds the right tool, supplies sensible arguments, handles the response and reaches the requested outcome. Keep the host, model and configuration attached to the observation, because those affect the result.

This is a separate test from directly calling an endpoint. Both are useful; neither substitutes for the other.

Five changes a team can make this week

  1. Save a dated copy of your tool catalogue and the exact conformance report.
  2. Compare definitions against the current requirements of the directory you actually intend to join.
  3. Verify each safety annotation against implementation behavior and side effects.
  4. Rewrite descriptions that dictate model behavior into precise descriptions of capabilities, inputs and constraints.
  5. Run a small set of real user journeys in a test environment and preserve the failures as well as the successes.

You do not need a paid audit to start that work. Good teams can do it themselves.

At Stratt Labs, our free Fit Check starts with the interface that can be inspected without invoking business tools. When access prevents a check, the report states the limitation. A clean static report means those static checks passed; it is not a promise of directory acceptance.

A protocol result, a directory review and a completed user journey answer three different questions. Make each answer explicit, and the next engineering decision becomes much easier.