Skip to content

feat(mcp): add experimental version-only stdio server - #4822

Merged
mnriem merged 6 commits into
github:mainfrom
mnriem:mnriem-experimental-version-mcp-server
Oct 2, 2026
Merged

mnriem merged 6 commits into
github:mainfrom
mnriem:mnriem-experimental-version-mcp-server

Conversation

@mnriem

@mnriem mnriem commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

  • add an experimental root specify mcp command using the official Python MCP SDK
  • expose the compact specify_list_commands, specify_describe_command, and specify_run_command tool surface over stdio only
  • support only the stable dotted version command and execute specify version --json through an isolated child process using the running CLI's Python environment
  • propagate the direct version success payload and structured CLI failures without banners, Rich output, fallback success values, or protocol contamination
  • document the experimental version-only/stdio-only limitations and refresh the committed dependency-audit snapshot

Scope

The supported inventory is deliberately explicit and contains only version. Unsupported commands return an unavailable_command tool error. This change does not expose HTTP/SSE transports, project discovery, artifact commands, mutations, installation/update/removal, workflows, confirmations, access tiers, or generalized timeout/cancellation infrastructure.

The official mcp>=2.2.0,<3.0.0 SDK is a production dependency. Although the SDK brings HTTP/SSE-related transitive packages, specify mcp registers and documents only stdio transport.

Validation

  • uv sync --extra test — passed
  • .venv/bin/python -m pytest tests/specify_cli/test_command_mcp.py tests/specify_cli/mcp_server -q — 22 passed
  • .venv/bin/python -m pytest tests/specify_cli -q — 2,932 passed, 1 skipped
  • .venv/bin/python -m pytest — 9,533 passed, 19 skipped, 62 warnings; 9,552 collected in 11m 06s
  • collection comparison — 9,552 after versus 9,530 before (+22; no decrease)
  • uvx ruff@0.15.0 check src tests — passed
  • uvx --from pip-audit==2.10.0 pip-audit --disable-pip --require-hashes -r .github/security-audit-requirements.txt --progress-spinner off — no known vulnerabilities
  • real stdio subprocess integration — initialization, tool discovery, direct version execution, unsupported-command error, and empty server stderr all passed

AI assistance

Implemented with GitHub Copilot using GPT-5.6 Sol in autonomous mode; code, tests, documentation, dependency updates, review, and validation were AI-assisted.

Expose the stable version JSON command through an stdio-only MCP server with explicit discovery, subprocess isolation, structured errors, focused tests, and reference documentation.

Assisted-by: GitHub Copilot (model: GPT-5.6 Sol, autonomous)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 19:27

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

Pydantic needs a direct dependency declaration, and two output-validation branches lack regression coverage.

Review effort: Balanced
Findings: 1 Medium severity · 1 Low severity

Open (2)
What changed in this PR

Adds an experimental, stdio-only MCP server exposing the stable version command.

Changes:

  • Adds MCP command catalog, server, and subprocess adapter.
  • Adds unit and real-stdio integration tests.
  • Documents MCP limitations and updates dependencies/audit data.
File Description
src/​specify_cli/​__init__.py Registers specify mcp.
src/​specify_cli/​command_mcp.py Defines the root MCP command.
src/​specify_cli/​mcp_server/​__init__.py Exports server entry points.
src/​specify_cli/​mcp_server/​_worker.py Runs CLI commands in a child process.
src/​specify_cli/​mcp_server/​catalog.py Defines the version-only inventory and schemas.
src/​specify_cli/​mcp_server/​executor.py Executes and validates version output.
src/​specify_cli/​mcp_server/​server.py Registers MCP tools and stdio transport.
tests/​specify_cli/​test_command_init.py Updates root registration expectations.
tests/​specify_cli/​test_command_mcp.py Tests MCP command registration and startup.
tests/​specify_cli/​mcp_server/​__init__.py Initializes the MCP test package.
tests/​specify_cli/​mcp_server/​test_catalog.py Tests inventory behavior.
tests/​specify_cli/​mcp_server/​test_executor.py Tests subprocess adaptation and failures.
tests/​specify_cli/​mcp_server/​test_server.py Tests tool discovery and dispatch.
tests/​specify_cli/​mcp_server/​test_stdio.py Tests the real stdio protocol flow.
pyproject.toml Adds the MCP SDK dependency.
.github/​security-audit-requirements.txt Refreshes audited dependency pins.
docs/​toc.yml Adds MCP documentation navigation.
docs/​reference/​overview.md Introduces the MCP reference.
docs/​reference/​mcp.md Documents tools, transport, and limitations.
docs/​reference/​core.md Links the MCP command from core commands.

💡 Configure MCP servers for context-aware, tailored reviews. Learn more in the docs.

Comment thread pyproject.toml
Comment thread tests/specify_cli/mcp_server/test_executor.py
Declare Pydantic as a direct runtime dependency and cover schema-invalid success and failure JSON payloads in the subprocess adapter tests.

Assisted-by: GitHub Copilot (model: GPT-5.6 Sol, autonomous)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 19:51
@mnriem

mnriem commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the review findings in 1755d88f:

  • declared pydantic>=2.13.0,<3.0.0 directly in both runtime dependency declarations and regenerated the audit snapshot
  • added regression cases for schema-invalid success and failure JSON payloads

Validation: focused MCP tests 18 passed; full suite 9,535 passed, 19 skipped with 9,554 collected; Ruff passed; dependency audit found no known vulnerabilities.

AI disclosure: Posted on behalf of @mnriem by GitHub Copilot using GPT-5.6 Sol in autonomous mode; the code, tests, validation, and this review-round summary were AI-assisted.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

Success payload validation remains coercive, and the explicit invalid-UTF-8 failure path lacks regression coverage.

Review effort: Balanced
Findings: None

Resolved since last review (2)
Previously missed (2)

In code that hasn't changed since last review

Medium severity Enable strict validation for child payloads

src/​specify_cli/​mcp_server/​executor.py:92

Validate the child payload in strict mode. Pydantic's default lax validation accepts values such as 1 or "true" for a boolean feature, so malformed CLI output is silently normalized and returned as a success instead of producing invalid_success_payload; that also breaks the promised direct-payload contract. Pass strict=True here and add a coercible-type case to the invalid-payload matrix.

Low severity Cover non-UTF-8 bytes and invalid_utf8 failures

tests/​specify_cli/​mcp_server/​test_executor.py:89

Add a case with non-UTF-8 bytes so the explicit invalid_utf8 adapter failure remains covered. The current negative matrix reaches empty, malformed JSON, schema-invalid, and mixed-output branches, but deleting the UnicodeDecodeError normalization would still leave it green, contrary to the repository's deterministic-behavior coverage requirement.

Reject coercible machine-output types and cover invalid UTF-8 subprocess output as a sanitized adapter failure.

Assisted-by: GitHub Copilot (model: GPT-5.6 Sol, autonomous)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 20:09
@mnriem

mnriem commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the latest review findings in 1f8cde98:

  • enabled strict Pydantic validation for both successful version payloads and structured CLI failure payloads, preventing coercion of malformed machine output
  • added regression coverage for coercible boolean values and non-UTF-8 subprocess output normalized to invalid_utf8

Validation: focused MCP tests 20 passed; full suite 9,537 passed, 19 skipped with 9,556 collected; Ruff passed.

AI disclosure: Posted on behalf of @mnriem by GitHub Copilot using GPT-5.6 Sol in autonomous mode; the code, tests, validation, and this review-round summary were AI-assisted.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟡 Changes recommended

The child module can be shadowed by an untrusted package in the MCP host’s working directory.

Review effort: Balanced
Findings: 1 High severity

Open (1)

Comment thread src/specify_cli/mcp_server/executor.py
Launch the child CLI with Python safe-path mode so a project-local package cannot shadow the installed MCP worker, with a real cwd-shadow regression test.

Assisted-by: GitHub Copilot (model: GPT-5.6 Sol, autonomous)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 20:32
@mnriem

mnriem commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the worker-shadowing finding in 5ae9c349:

  • launch the child CLI with Python safe-path mode (-P) so the MCP host working directory is not prepended to module lookup
  • added a real regression that changes into a project containing a shadow specify_cli.mcp_server._worker package and verifies the installed worker still runs

Validation: focused MCP tests 21 passed; full suite 9,538 passed, 19 skipped with 9,557 collected; Ruff passed.

AI disclosure: Posted on behalf of @mnriem by GitHub Copilot using GPT-5.6 Sol in autonomous mode; the code, tests, validation, and this review-round summary were AI-assisted.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

MCP failures are exposed as prefixed text rather than the documented structured error payload.

Review effort: Balanced
Findings: None

Resolved since last review (1)
Previously missed (1)

In code that hasn't changed since last review

Medium severity Preserve structured CLI errors in MCP responses

src/​specify_cli/​mcp_server/​server.py:25

This conversion drops the structured CLI error at the MCP boundary. MCP SDK 2.2 turns a raised ToolError into is_error=True with only str(exc) in text content (prefixed with Error executing tool ...), leaving structured_content empty. Clients therefore cannot read error.code, error.message, and error.details as the structured failure promised by the PR/docs without scraping JSON from free-form text. Return an error CallToolResult carrying both readable text and the payload in structured_content, and assert that wire shape in the stdio test.

Return explicit error CallToolResult values so MCP clients receive readable content and the unchanged structured CLI error payload, with in-memory and real stdio coverage.

Assisted-by: GitHub Copilot (model: GPT-5.6 Sol, autonomous)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 20:58
@mnriem

mnriem commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the structured MCP error finding in 56ff5b30:

  • return explicit error CallToolResult values with readable text, isError: true, and the unchanged structured error.code, error.message, and error.details payload
  • updated in-memory coverage and the real stdio integration test to assert the wire-level structuredContent error shape

Validation: focused structured-error tests 6 passed; full suite 9,538 passed, 19 skipped with 9,557 collected; Ruff passed.

AI disclosure: Posted on behalf of @mnriem by GitHub Copilot using GPT-5.6 Sol in autonomous mode; the code, tests, validation, and this review-round summary were AI-assisted.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🔵 Needs a closer look

The unbounded stdio integration test can stall CI indefinitely when the server stops responding.

Review effort: Balanced
Findings: None

Previously missed (1)

In code that hasn't changed since last review

Medium severity Add read deadline to prevent subprocess test hangs

tests/​specify_cli/​mcp_server/​test_stdio.py:70

Bound this real subprocess protocol test. The session currently has no read deadline, so a regression that starts the child but stops it replying will hang the test job instead of producing a failure. An outer timeout also ensures cancellation unwinds the contexts and terminates the child.

Add per-read and whole-test deadlines so a non-responsive MCP subprocess fails deterministically while context cleanup terminates the child.

Assisted-by: GitHub Copilot (model: GPT-5.6 Sol, autonomous)

Co-authored-by: Copilot App <223556219+Copilot@users.noreply.github.com>
Copilot AI balanced review requested due to automatic review settings October 2, 2026 21:35
@mnriem

mnriem commented Oct 2, 2026

Copy link
Copy Markdown
Collaborator Author

Addressed the unbounded stdio integration test in d13c0350:

  • configured a 10-second MCP session read deadline
  • wrapped the complete subprocess handshake in a 30-second asyncio.timeout, so cancellation unwinds both async contexts and terminates the child

Validation: the bounded stdio integration test passed; Ruff passed. A full local run completed 9,512 non-PowerShell tests successfully, but 26 pre-existing PowerShell cases failed because the host pwsh runtime now crashes before script execution with System.IO.FileLoadException: The given assembly name was invalid; a direct pwsh -NoProfile health check reproduces the same runtime failure. The preceding full run on this branch passed all 9,557 collected tests before this test-only timeout change.

AI disclosure: Posted on behalf of @mnriem by GitHub Copilot using GPT-5.6 Sol in autonomous mode; the test change, validation, investigation, and this review-round summary were AI-assisted.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot review overview

🟢 Approval recommended

The constrained implementation is documented and covered by positive, negative, isolation, and real-protocol tests.

Review effort: Balanced
Findings: None

@mnriem
mnriem merged commit e1fa857 into github:main Oct 2, 2026
15 checks passed
@mnriem
mnriem deleted the mnriem-experimental-version-mcp-server branch October 2, 2026 21:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants