Repository navigation
Publish and validate per-model thinking levels (#1087) - #1089
AnthonyRonning wants to merge 2 commits into
Conversation
Live verification after the change (local maple-dev-env, Maple proxy → OpenSecret → providers)194 requests through the proxy with the Free fixture, the Pro fixture (GLM route: Tinfoil) and a second Pro account on the Continuum GLM route: 156 returned 200, 38 were rejected with
Raw results ( |
Every model's reasoning controls now come from one table in model_config.rs, verified live against both providers on 2026-10-07 (issue #1087). It is published as an OpenRouter-shaped `reasoning` object (mandatory, default_enabled, supported_efforts, default_effort) plus `max_completion_tokens` on `/v1/models` and the catalog, with aliases publishing the intersection of their candidates. Chat Completions validates `reasoning_effort` after alias and route resolution: an explicit model rejects an unsupported value with OpenAI's `unsupported_value` error (type, param and the accepted list in the message) before any provider is called; `auto:` aliases move the value to the nearest effort the resolved model accepts. Thinking-off template switches are dropped for models whose reasoning is mandatory, because the deployed GLM builds leak reasoning into `content` on every off control. Kimi K3 receives `developer` messages as `system`. Responses accepts `reasoning: { effort }`, validates it before any write and forwards it to the model turn. The Rust and TypeScript SDKs gain the catalog types, `ChatMessage.reasoning` and `reasoning_effort`, and the Research composer gets a thinking-level control built from the selected model's catalog entry. Existing requests, responses and error bodies are unchanged. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
b4413c0 to
9923f4e
Compare
|
Maple development preview: https://34705363.maple-ca8.pages.dev Commit: Uses development API, billing, flags and PCR configuration. Cloudflare Access applies. |
The Continuum proxy copies vLLM's `reasoning` into the deprecated `reasoning_content` on every message and delta; Tinfoil routes never send it, so the same model answered with a different field set depending on the route an account landed on. Fold the copy into `reasoning` where provider payloads are already canonicalized (the public model id), for the non-streaming message and every stream chunk, and promote the old name to `reasoning` should a route ever send only that. Nothing in-tree read the copy; the Rust SDK keeps `reasoning_content` readable and replayable for other OpenAI-compatible servers. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Follow-up: one reasoning field on every routeSecond commit folds Continuum's deprecated Verified live through the proxy after the change, 16 GLM-5.3 / GLM-5.3-Flash cells on the Continuum-route account plus GLM-5.3 and Kimi K3 on a Tinfoil-route account, streaming and non-streaming: every response has |
Implements the backend, SDK and Research parts of #1087: every model's thinking levels are published in the catalog, validated per resolved model on both inference APIs, and selectable from the Research composer. Additive only: existing requests, responses and error bodies are unchanged.
What changes
Catalog (
services/opensecret/src/model_config.rs)ModelReasoning, verified live against both providers on 2026-10-07, drives everything below. It is published as areasoningobject on each model in/v1/modelsand/v1/models/catalog, in the shape OpenRouter already publishes for these model families:{ mandatory, default_enabled, supported_efforts (highest first), default_effort }.capabilities.reasoningis unchanged.max_completion_tokensis published per model and alias (vLLM has no output cap of its own; input and output share the window).max,high,low, mandatory; free Quick:high,medium,low, default medium), pluscontext_windowandmax_completion_tokens.Chat Completions (
web/model_request_policy.rs,web/openai.rs)reasoning_effortis validated after alias and route resolution against the model that will run. An explicit model rejects an unsupported value before any provider is called, in OpenAI's shape (400,error.type: invalid_request_error,error.code: unsupported_value,error.param: reasoning_effort, message listing the accepted values). The legacystatus/messagekeys stay.auto:quick/auto:powerfulthe value is moved to the nearest effort the resolved model accepts (up, then down), so a health fallback never fails a running task.chat_template_kwargs.enable_thinking: false/thinking: falseare dropped for models whose reasoning is mandatory (GLM): on every deployed build they make the reasoning come back insidecontent.developermessages are sent assystemto Kimi K3, whose renderer rejects the role (Pi sends it by default).reasoning_contentis dropped whenreasoningis present and promoted toreasoningif a route ever sends only the old name.nulleffort, and non-reasoning models, are forwarded untouched. The request log gains an allowlistedreasoning_effortfield.Responses (
web/responses/handlers.rs)reasoning: { effort }is accepted (OpenAI's shape;summaryis accepted and ignored), validated against the tier's model before any write, and forwarded as the model turn'sreasoning_effortafter clamping to the model that actually runs.noneon Gemma turnsenable_thinkingoff; otherwise Maple's defaults stand.SDKs (unpublished; in-tree consumers build from source)
Model/CatalogModel/CatalogAliasgainreasoningandmax_completion_tokens;ChatMessagegainsreasoning(plusreasoning_text()that falls back to the deprecatedreasoning_content);ChatCompletionRequestgainsreasoning_effort.ReasoningEffort,ModelReasoning,ModelListItem;ModelCatalogItemandModelAliasgain the new fields;ResponsesCreateRequestgainsreasoning.Research (
apps/maple-research/frontend)reasoning: Default, Off where reasoning is not mandatory, then the supported efforts. Hidden for models that never reason. The choice persists inlocalStorageand only ever sends a level the current model publishes.Docs: proxy README section on reasoning effort, catalog metadata and the error shape; backend development contract.
Decisions taken (from the issue)
noneis published (it works on the deployed build and OpenRouter lists Kimi as non-mandatory).default_effort.reasoningon every route. Continuum's deprecatedreasoning_contentcopy (byte-identical toreasoning, never sent by Tinfoil routes) is folded intoreasoningat the response boundary, for the non-streaming message and every stream delta, so the field set is the same on all eight models and the two can never be concatenated. Nothing in-tree read the copy; the SDKs keepreasoning_contentreadable and replayable for other servers.Validation
cargo fmt,cargo clippy --all-targets --all-features -D warnings,cargo test --all-features(762 passed). New tests cover every model's accepted and rejected values, alias clamping, GLM switch stripping, the Kimi role mapping, catalog//v1/modelsparity, alias intersection, the Responses request shape and the error body.cargo test --lib(112 passed). TypeScript SDK:tsc, format, build, unit tests. Proxy: compiles against the changed SDK (pass-through, no code change).bun test(newchatThinkingLeveltests).reasoning_effort: low.Closes nothing by itself; the Agent integration in #1087 is a follow-up.
🤖 Generated with Claude Code