Skip to content

chore: OAuth pre-flight harness for the hosted MCP endpoint - #187

Open
karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:chore/oauth-preflight
Open

karaposu wants to merge 1 commit into
brightdata:mainfrom
karaposu:chore/oauth-preflight

Conversation

@karaposu

Copy link
Copy Markdown

Why

ChatGPT connects to a remote MCP server only through OAuth 2.1 (PKCE with S256, dynamic client registration, the resource indicator). mcp.brightdata.com advertises all of that, but nobody had walked the flow end to end, and a plugin submission fails at its first step if any part of it does not work. This adds scripts that check each step and record what they saw, with no secret in any output.

What it contains

scripts/oauth-preflight/, not wired into npm test.

  • 01-discovery.mjs: the 401 challenge on /mcp, the protected-resource metadata, and the authorization-server metadata. 13 of 14 assertions pass today.
  • 02-registration.mjs: dynamic client registration with a throwaway client and a delete check. Gated: it refuses to run unless PREFLIGHT_NOTIFIED is set, because it creates records on the production authorization server. notification-draft.md is the message to send to whoever owns that server first.
  • 03-authorize-and-token.mjs: a pre-login probe that the authorize endpoint validates the two callback shapes ChatGPT uses, then a local listener for the code, token exchange, decode, and refresh. Needs one interactive login.
  • 05-usage.mjs and 06-failures.mjs: a 2×2 matrix of token type against endpoint, and the failure paths.
  • mock/: a local mock of the whole flow so the scripts can be rehearsed offline (10 of 10).
  • common.mjs: the secrets boundary. Tokens are read from the environment, sent only as Authorization: Bearer, never placed in a URL, and every recorded output is scrubbed by exact value.

Two things discovery found on the authorization server, not in this repository

  1. The authorization-server metadata is published at two URLs and the copies disagree. The issuer copy advertises resource_parameter_supported: true and an agent_auth block that the copy under the resource host lacks. Clients that read the issuer copy work today; the other copy is stale.
  2. A request with an invalid bearer token gets a bare 401 with no WWW-Authenticate header, while a request with no token gets the full challenge. RFC 6750 §3 requires the header on every 401. An expired session takes the unhelpful path.

Not run

Registration, login and token exchange have not been executed against production. The gate is deliberately closed until the owner of the authorization server has been told; the scripts are ready to run in one sitting once that is done.

Why this exists
---------------

OpenAI's ChatGPT and Codex plugin directory accepts exactly one authentication
method for an MCP server: OAuth 2.1 per the MCP authorization specification.
Their documentation is explicit that ChatGPT cannot present an API key. The
hosted endpoint at mcp.brightdata.com is documented everywhere as
`?token=<API_TOKEN>`, so the documented path is not the path a submission can
use.

The OAuth machinery does exist and its metadata is correct: S256 PKCE, dynamic
client registration and refresh tokens are all advertised. What nobody had done
is run the flow. Every link between "the metadata looks right" and "a client
holds a working token" was an assumption, and the authorization server is a
separate system from this repository, so anything broken there is a slow fix.

These scripts walk the flow in the order ChatGPT walks it and record what the
server actually answers. They drive the MCP SDK's own OAuth primitives rather
than hand-written HTTP, so a failure implicates the server rather than the test.

Nothing here ships. It is not in package.json's files list, so it stays out of
the published package, and it imports nothing that server.js depends on.

What has been verified
----------------------

Against the live endpoint, read-only, creating nothing:

  - Discovery: 13 of 14 assertions pass. The unauthenticated 401 carries the
    challenge naming the metadata document; protected-resource metadata is
    correct at both paths; the authorization server advertises S256, both
    grants, and a registration endpoint.
  - Tool exposure over the documented API-token path: 5 tools by default, 74
    with ?pro=1. A submission names one fixed URL, so which of those an OAuth
    session gets decides what URL to submit.
  - Failure paths, three of four cases.

Against a local mock that stands in for the authorization server (mock/), the
whole flow passes 10 of 10 rehearsal checks, including registration, the
reversibility probe, ChatGPT's redirect URL being accepted, code capture,
audience binding, refresh, all four cells of the tool-count matrix, and the
absence of secret values in generated output.

That rehearsal already earned its place. It found that the MCP SDK's
registerClient() drops registration_client_uri and registration_access_token
from the registration response -- the only two fields that reveal whether a
registration can be deleted afterwards. Read from the SDK's parsed result, the
answer is always "not deletable", so a production run would have reported that
cleanup requires an internal engineer regardless of the truth. The script now
reads them from the raw response body.

Two defects found, neither previously known
-------------------------------------------

  1. The authorization-server metadata is published in two places and the
     copies disagree. The issuer's copy carries resource_parameter_supported
     and an agent_auth block; the copy on the MCP host carries neither.
     Standards-following clients read the issuer's, so nothing is broken today,
     but one copy is stale and will keep drifting.

  2. A 401 for a request carrying an invalid token omits WWW-Authenticate,
     while a request with no credentials at all receives the full challenge.
     RFC 6750 section 3 requires the header when a request "does not contain an
     access token that enables access to the protected resource", and
     recommends an error="invalid_token" attribute in exactly that case, while
     saying a credential-less request should NOT carry an error code. The two
     cases are handled backwards. This matters because an expired token takes
     the invalid-token path, and an expired token is the ordinary case in
     production: it is how a client learns it should log in again.

What is NOT verified
--------------------

  - Whether an *expired* token behaves like an invalid one. Defect 2 rests on
    an invalid-token probe; the expired case needs a real token aged past its
    lifetime, and is recorded as unverified rather than inferred.
  - Whether Bright Data's registration endpoint accepts the per-connector
    redirect URL ChatGPT is forced to use, since issuer identification is not
    advertised. This is the single question that decides whether a connector
    can be created at all, and it needs a real registration.
  - Whether an OAuth session sees 5 tools or 74, and whether the mode flag can
    be carried on a submitted URL without breaking audience verification.
  - Whether the access token is a JWT, what its lifetime is, whether refresh
    works, and whether its audience matches the resource indicator.
  - ChatGPT itself. It cannot be tested until the plugin exists as a draft in
    an OpenAI organisation with Apps Management: Write. A green run here proves
    a standards-conformant client succeeds, not that ChatGPT will.

What must be done by a person
-----------------------------

  1. Tell whoever owns the authorization server before running
     02-registration.mjs. It creates real client registrations on production,
     and registration traffic plus a burst of authorize requests is the shape
     of thing that looks like probing to whoever monitors it. A draft message
     is in notification-draft.md; it already carries both defects above. The
     script refuses to run until PREFLIGHT_NOTIFIED names the date.

  2. Accept, or decline, the cleanup position. Phase A registers one disposable
     client and tries to delete it. If the server implements creation but not
     the management half of the spec, every test registration is permanent from
     the outside and only an administrator can remove it. The script stops
     there and requires PREFLIGHT_OWNER_CONFIRMED=yes to continue.

  3. Sign in once, in a private window with no existing session. Signed-out
     matters twice: it exercises the first-time path a reviewer's fresh account
     will face, including a consent screen an existing account would skip, and
     it keeps an identity out of the screenshots the submission needs. The
     authorization code is captured by a local listener, so nobody handles it.

  4. Note what the login demands -- password only, or MFA, SMS, email
     confirmation, a bot check -- because OpenAI requires reviewer credentials
     that complete login without any of those, and somebody has to provision an
     account against that list.

Run artefacts are gitignored: evidence, tokens, registered client ids and
registration management tokens never leave the working tree.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant