Skip to content

Publish and validate per-model thinking levels (#1087) - #1089

Open
AnthonyRonning wants to merge 2 commits into
masterfrom
codex-thinking-levels-maple
Open

AnthonyRonning wants to merge 2 commits into
masterfrom
codex-thinking-levels-maple

Conversation

@AnthonyRonning

@AnthonyRonning AnthonyRonning commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Implements the backend, SDK and Research parts of #1087: every model's thinking levels are published in the catalog, validated per resolved model on both inference APIs, and selectable from the Research composer. Additive only: existing requests, responses and error bodies are unchanged.

What changes

Catalog (services/opensecret/src/model_config.rs)

  • One per-model table, ModelReasoning, verified live against both providers on 2026-10-07, drives everything below. It is published as a reasoning object on each model in /v1/models and /v1/models/catalog, in the shape OpenRouter already publishes for these model families: { mandatory, default_enabled, supported_efforts (highest first), default_effort }. capabilities.reasoning is unchanged.
  • max_completion_tokens is published per model and alias (vLLM has no output cap of its own; input and output share the window).
  • Alias entries publish only the efforts every model they can resolve to accepts (paid Quick and Powerful: max, high, low, mandatory; free Quick: high, medium, low, default medium), plus context_window and max_completion_tokens.

Chat Completions (web/model_request_policy.rs, web/openai.rs)

  • reasoning_effort is validated after alias and route resolution against the model that will run. An explicit model rejects an unsupported value before any provider is called, in OpenAI's shape (400, error.type: invalid_request_error, error.code: unsupported_value, error.param: reasoning_effort, message listing the accepted values). The legacy status/message keys stay.
  • On auto:quick / auto:powerful the value is moved to the nearest effort the resolved model accepts (up, then down), so a health fallback never fails a running task.
  • chat_template_kwargs.enable_thinking: false / thinking: false are dropped for models whose reasoning is mandatory (GLM): on every deployed build they make the reasoning come back inside content.
  • developer messages are sent as system to Kimi K3, whose renderer rejects the role (Pi sends it by default).
  • Provider payloads are canonicalized to one reasoning field: reasoning_content is dropped when reasoning is present and promoted to reasoning if a route ever sends only the old name.
  • Absent or null effort, and non-reasoning models, are forwarded untouched. The request log gains an allowlisted reasoning_effort field.

Responses (web/responses/handlers.rs)

  • reasoning: { effort } is accepted (OpenAI's shape; summary is accepted and ignored), validated against the tier's model before any write, and forwarded as the model turn's reasoning_effort after clamping to the model that actually runs. none on Gemma turns enable_thinking off; otherwise Maple's defaults stand.

SDKs (unpublished; in-tree consumers build from source)

  • Rust: Model/CatalogModel/CatalogAlias gain reasoning and max_completion_tokens; ChatMessage gains reasoning (plus reasoning_text() that falls back to the deprecated reasoning_content); ChatCompletionRequest gains reasoning_effort.
  • TypeScript: ReasoningEffort, ModelReasoning, ModelListItem; ModelCatalogItem and ModelAlias gain the new fields; ResponsesCreateRequest gains reasoning.

Research (apps/maple-research/frontend)

  • A thinking-level control in the composer, next to the model selector, built from the selected model's catalog reasoning: Default, Off where reasoning is not mandatory, then the supported efforts. Hidden for models that never reason. The choice persists in localStorage and only ever sends a level the current model publishes.

Docs: proxy README section on reasoning effort, catalog metadata and the error shape; backend development contract.

Decisions taken (from the issue)

  1. Reject on explicit models, clamp on aliases.
  2. Kimi K3 none is published (it works on the deployed build and OpenRouter lists Kimi as non-mandatory).
  3. Each model's own default is published as default_effort.
  4. Reasoning text is returned in reasoning on every route. Continuum's deprecated reasoning_content copy (byte-identical to reasoning, never sent by Tinfoil routes) is folded into reasoning at the response boundary, for the non-streaming message and every stream delta, so the field set is the same on all eight models and the two can never be concatenated. Nothing in-tree read the copy; the SDKs keep reasoning_content readable and replayable for other servers.
  5. Per-route engine metadata is not added here; the table carries its verification date instead.

Validation

  • Backend: cargo fmt, cargo clippy --all-targets --all-features -D warnings, cargo test --all-features (762 passed). New tests cover every model's accepted and rejected values, alias clamping, GLM switch stripping, the Kimi role mapping, catalog//v1/models parity, alias intersection, the Responses request shape and the error body.
  • Rust SDK: fmt, clippy, doc, cargo test --lib (112 passed). TypeScript SDK: tsc, format, build, unit tests. Proxy: compiles against the changed SDK (pass-through, no code change).
  • Research: typecheck, lint, prettier, bun test (new chatThinkingLevel tests).
  • Live, through the Maple proxy against local OpenSecret with Free and Pro fixtures (one Pro account per GLM provider route): see the comment below for the matrix.
  • Live, in the Research web app with the Pro fixture: the control appears for reasoning models and hides for Llama, options match each model's catalog entry, and a message sent at Low reaches the backend with reasoning_effort: low.

Closes nothing by itself; the Agent integration in #1087 is a follow-up.

🤖 Generated with Claude Code

@AnthonyRonning

Copy link
Copy Markdown
Contributor Author

Live verification after the change (local maple-dev-env, Maple proxy → OpenSecret → providers)

194 requests through the proxy with the Free fixture, the Pro fixture (GLM route: Tinfoil) and a second Pro account on the Continuum GLM route: 156 returned 200, 38 were rejected with 400 unsupported_value, and no response carried reasoning inside content (before the change every GLM thinking-off control did).

  • Explicit models reject exactly the values outside their catalog entry, before any provider call, with OpenAI's body (error.type: invalid_request_error, error.code: unsupported_value, error.param: reasoning_effort, message listing the accepted values), streaming and non-streaming alike. GLM none/minimal/medium/xhigh, Kimi minimal/medium/xhigh, DeepSeek minimal/medium, gpt-oss none/minimal/xhigh/max; Gemma accepts everything; Llama ignores everything; explicit null is unset.
  • Aliases clamp on both GLM routes: auto:powerful with none/minimal runs at low (31–33 reasoning tokens on the marbles prompt), medium at high (36), xhigh at max (106); auto:quick (Flash) likewise; free auto:quick (gpt-oss) accepts every value.
  • GLM thinking-off switches (chat_template_kwargs.enable_thinking: false, thinking: false) are dropped: reasoning comes back in reasoning on Tinfoil and on Continuum (where the duplicate reasoning_content still passes through unchanged).
  • Kimi K3 accepts a Pi-shaped request (developer role, store: false, max_completion_tokens, reasoning_effort: low), which was a 400 before.
  • Research web app (Pro and Free fixtures, headless Chrome): the thinking control appears next to the model selector with each model's published options and hides for Llama; a message sent at Low reaches the backend as reasoning_effort=low on the Responses model turn.

Raw results (results-verify-*.jsonl), the harness and the pre-change matrix live in ~/workspaces/_notes/thinking-levels-2026-10-07/ on the dev VM; REPORT.md there is the full Report A/B write-up.

Every model's reasoning controls now come from one table in
model_config.rs, verified live against both providers on 2026-10-07
(issue #1087). It is published as an OpenRouter-shaped `reasoning`
object (mandatory, default_enabled, supported_efforts, default_effort)
plus `max_completion_tokens` on `/v1/models` and the catalog, with
aliases publishing the intersection of their candidates.

Chat Completions validates `reasoning_effort` after alias and route
resolution: an explicit model rejects an unsupported value with
OpenAI's `unsupported_value` error (type, param and the accepted list
in the message) before any provider is called; `auto:` aliases move
the value to the nearest effort the resolved model accepts. Thinking-off
template switches are dropped for models whose reasoning is mandatory,
because the deployed GLM builds leak reasoning into `content` on every
off control. Kimi K3 receives `developer` messages as `system`.

Responses accepts `reasoning: { effort }`, validates it before any
write and forwards it to the model turn. The Rust and TypeScript SDKs
gain the catalog types, `ChatMessage.reasoning` and `reasoning_effort`,
and the Research composer gets a thinking-level control built from the
selected model's catalog entry. Existing requests, responses and error
bodies are unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Maple development preview: https://34705363.maple-ca8.pages.dev

Commit: b8953df591f6e9989ba8f36552c0704249c14064

Uses development API, billing, flags and PCR configuration. Cloudflare Access applies.

The Continuum proxy copies vLLM's `reasoning` into the deprecated
`reasoning_content` on every message and delta; Tinfoil routes never
send it, so the same model answered with a different field set depending
on the route an account landed on. Fold the copy into `reasoning` where
provider payloads are already canonicalized (the public model id), for
the non-streaming message and every stream chunk, and promote the old
name to `reasoning` should a route ever send only that. Nothing in-tree
read the copy; the Rust SDK keeps `reasoning_content` readable and
replayable for other OpenAI-compatible servers.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@AnthonyRonning

Copy link
Copy Markdown
Contributor Author

Follow-up: one reasoning field on every route

Second commit folds Continuum's deprecated reasoning_content copy into reasoning at the response boundary (non-streaming message and every stream delta), so GLM answers carry the same field set whichever provider route an account lands on. A route that ever sends only the old name gets it promoted to reasoning.

Verified live through the proxy after the change, 16 GLM-5.3 / GLM-5.3-Flash cells on the Continuum-route account plus GLM-5.3 and Kimi K3 on a Tinfoil-route account, streaming and non-streaming: every response has reasoning, none has reasoning_content, token accounting unchanged. Backend gate: fmt, clippy, 786 tests.

This branch was successfully deployed

1 active deployment
pages-pr-1089 — b8953df5 Deployed Oct 7, 2026 by github-actions[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant