Skip to content

fix: support reasoning models through opt-in Responses API - #104

Draft
DevJoaoLopes wants to merge 2 commits into
CopilotKit:mainfrom
DevJoaoLopes:fix/issue-58-responses
Draft

DevJoaoLopes wants to merge 2 commits into
CopilotKit:mainfrom
DevJoaoLopes:fix/issue-58-responses

Conversation

@DevJoaoLopes

Copy link
Copy Markdown

Summary

Addresses #58 with an explicit Dot model API selector, configurable output budget,
and optional native reasoning effort. Chat Completions and the existing 2200
budget remain the defaults; an empty effort setting preserves the provider default.

Draft / blocked on SDK release: depends on
CopilotKit/CopilotKit#7672.
Published @copilotkit/runtime@1.77.0 does not contain the encrypted-reasoning and
HITL continuation fixes. This PR deliberately retains the compatibility tests:
the current public lock produces 7 failures / 337 passes, comprising three
SDK compatibility gates and four review-resume cases. No tests are skipped or
weakened to hide that dependency.

After a corrected SDK release is published, update the public dependency/lock
resolution and verify a clean npm ci before marking this PR ready.

Changes

  • Add validated server settings DOT_MODEL_API, DOT_MAX_OUTPUT_TOKENS, and
    optional DOT_REASONING_EFFORT.
  • Select Chat Completions or Responses through the existing TanStack-compatible
    adapter, with native parameters for each API.
  • Responses uses store: false and includes encrypted reasoning for stateless
    continuation. The configured Intelligence deployment still persists Threads.
  • Keep effort opt-in: use reasoning_effort for Chat Completions or
    reasoning: { effort } for Responses. There is no automatic API/model fallback
    or implicit reasoning downgrade.
  • Apply selection in the shared Dot executor, covering chat, page chat, scheduled
    turns, Slack turns, and voice compute/receipts. The legacy research call and
    Realtime speech endpoint are unchanged.
  • Align CopilotKit packages at 1.77.0 and direct AG-UI dependencies at 1.0.1;
    manifests/lock contain public dependencies only, no local-package paths.
  • Forward settings through Compose and document API support, reasoning levels,
    combined output budgets and the pending SDK release.

Example configuration after the corrected dependency is available:

DOT_MODEL_API=responses
DOT_MAX_OUTPUT_TOKENS=8192
DOT_REASONING_EFFORT=medium

8192 and medium are explicit example choices, not new application defaults.
Support for APIs and effort levels depends on the selected model/provider.

Verification

Local corrected SDK package

  • 344 tests passed across 46 files with the reviewed local runtime package.
  • Formatter, lint, typecheck and client/server production build passed.
  • Wire tests use the real adapter and Dot executor, with only network responses
    mocked. They cover tool execution, canonical review, approve/decline,
    serialization/reopened history, opaque signatures and tool item/call IDs.
  • Error coverage includes HTTP 400, failed/incomplete Responses, truncated stream,
    network failure, pause and timeout for both APIs, without a false clean finish.
  • Docker Compose was evaluated using isolated fixture credentials to check
    defaults, explicit settings and blank effort handling.

The local runtime package declares 1.77.0 but includes the unmerged SDK patch;
it is not the published 1.77.0. Validation package SHA-256:

dc88855b0a5f3a8f45a4bc555786e603060a472deab684470cc1c5583f21dbe9

Live configuration / UI

On macOS/Node 24, an isolated workspace used gpt-6.1-sol, Responses,
medium, output budget 8192, and the local corrected runtime:

  • Final flow: 9 model requests returned HTTP 200, no HTTP 400.
  • Actual page-tool call and reply; review approval by keyboard saved one page;
    declined review saved no extra page; both decisions resumed the model.
  • A fresh browser context reopened the conversation and continued successfully.
    The outgoing provider request contained an encrypted reasoning item from history.
  • 34 reasoning tokens were observed in that flow. Simple list/draft requests
    can report zero reasoning even with an explicit effort; the calculation task
    supplied positive provider evidence.
  • Owner pause aborted an active model request with AbortError; a new turn
    succeeded after unpausing. The chat was inspected at a 375px viewport.

Only synthetic page content was used. No credentials, encrypted values, local
databases, screenshots or private planning documents are included in the PR.
Connected Slack and speech providers were not part of this live verification.

Prepared with AI assistance.

@jerelvelarde jerelvelarde left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This opt-in model API selector has useful template value: default Chat Completions and the 2200 output budget remain intact, and explicit reasoning effort is passed in the native shape for the selected API without endpoint/billing fallback. The settings reach the shared Dot executor used by web/page chat, Slack, scheduled turns and delegated voice compute; legacy research and Realtime speech stay unchanged. Setup/Compose documentation accurately distinguishes provider response storage from persisted Intelligence Threads. No additional application security blocker found in this review.

Retain this draft and its existing corrected-SDK-release hold. At 038fb79 with the public locked @copilotkit/runtime 1.77.0, credential-free wire tests produce 89 passes and 7 failures: three SDK encrypted-reasoning/history compatibility gates and four serialized approve/decline review-resume cases. The installed converter drops reasoning history and fails to forward encrypted reasoning through the BuiltInAgent boundary; review resume also fails AG-UI reasoning-message verification. This reproduces the stated dependency limitation, rather than establishing readiness from the locally corrected package. CopilotKit/CopilotKit#7672 remains open and unmerged.

Before readying this PR, update the public dependency/lock to a released corrected SDK, verify clean npm ci, rerun the complete compatibility/HITL matrix and repository gates, then recheck current-main integration. Focused adapter tests verify native URLs/options, defaults, budgets, explicit effort, tool execution, error/incomplete/truncated streams, pause and timeout without live providers. The prospective merge with main 1b2425d is conflict-free. Live UI/provider claims in the description were not independently reproduced here; full repository gates are recorded separately by the coordinating reviewer. Official stateless reasoning guidance: https://developers.openai.com/api/docs/guides/reasoning .

@DevJoaoLopes

Copy link
Copy Markdown
Author

@jerelvelarde — update following your review.

I am retaining this PR as Draft and keeping the corrected-SDK-release hold.

SDK prerequisite update

CopilotKit/CopilotKit#7672 now includes follow-up commit d13d771:

  • Normalize empty REASONING_START IDs with getNonEmptyString before UUID fallback.
  • Add two public runAgent() regressions for explicit/automatic closure, signature targeting, message materialization and history replay when both input IDs are empty.
  • Document the converter's metadata, lifecycle, alias and encrypted-update helpers.

That follow-up passed 2,833 runtime tests, typecheck, build, formatting and lint locally. The SDK branch was subsequently merged with current main; its current PR head is 61e49f4.

CodeRabbit is passing, but the upstream Actions runs for that merge head are still action_required and have not executed their jobs. A maintainer with write access needs to approve the fork workflows; these are not being reported as green CI. The SDK PR remains open/unmerged.

Current-main integration for this PR

I merged OpenDots main 625452e in 86a9eeb and resolved the import conflict in tests/tanstack-agent.test.ts, retaining both the Responses coverage and the new MCP connection-tool test. The shared executor keeps the MCP approval/permission controls alongside the API selector. Upstream dark-mode, page-deletion and scheduled-turn changes are retained.

Validation of this merged tree:

  • Clean npm ci with the current public locked runtime: 385 passes / 7 failures across 51 files. The same three encrypted-reasoning/history gates and four approve/decline resume cases remain red; no skip or weakened assertion was introduced.
  • With the previously reviewed local corrected runtime 1.77.0 package: 392/392 tests passed, plus formatter, lint, typecheck and production build. This is local-package evidence, not proof that the published runtime is corrected.
  • Public manifest/lock content is unchanged by the local-package overlay and contains no private package paths.

I have not repeated the earlier live-provider/UI validation after this main merge. Before readying the PR, I will still update the public dependency/lock to a released SDK containing #7672, verify a clean install, rerun the complete compatibility/HITL matrix and repository gates, and recheck current-main integration as requested.

Could a maintainer approve the pending fork workflows on SDK PR #7672 so its upstream checks can execute?

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants