merge(main): merge v2.7.0 release from develop - #4044
Merged
Merged
Conversation
* feat: support AIDP knowledge file management Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: Codex * fix: unblock AIDP PR quality checks Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: Codex * fix: reduce Sonar duplication and stabilize web install Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: Codex * test: improve AIDP file operation coverage * fix: support unicode filenames in AIDP mock download * fix: align AIDP document deletion with local KB * fix: stream AIDP document downloads * fix: preserve AIDP permission test compatibility * fix: simplify AIDP document download streaming * fix(aidp): align document removal request with API * chore: remove unrelated assistant ui dependency change * test(aidp): align mock document id type
* Feature: 知识库界面优化 * Fix: 修复知识库门禁问题 --------- Co-authored-by: hzw <hzw@qq.com>
* perf(frontend): enable Turbopack for local development Enable Turbopack in the custom development server and stabilize its configuration dependencies. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * perf(frontend): upgrade Next.js for faster Turbopack Upgrade Next.js and React to 16.3.5 and 19.2.3, migrate lint and proxy configuration, and resolve React 19 type compatibility.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * fix(frontend): silence Next HMR upgrade logging Allow Next.js to handle its own HMR WebSocket upgrades without emitting a misleading proxy log.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * fix(frontend): migrate Ant Design 6 component APIs Replace deprecated modal, drawer, and alert props with their Ant Design 6 equivalents. Mount hidden memory forms before use and make memory configuration card spacing explicit. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * build(web): align image runtime with frontend dependencies Use Node 22 for the web image build and runtime stages so freshly resolved dependencies meet their engine requirements. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5
…troller, and break adapter on terminal human_run (#3948) * fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests Root cause: the worker thread writes human_interaction events synchronously via SQLAlchemy in ask_user, while model_output_thinking/parse observer messages flow through the async consumer and are flushed only on a batched threshold (32 chunks or 250ms). When the worker suspends before that flush fires, human_interaction gains a lower event_seq number than the already-buffered observer chunks, causing the SSE replay stream to show them in the wrong order. Fix: replace the plain async-for consumer loop with a manual asyncio.wait iterator using a 50ms timeout. Once the worker finishes producing model output (i.e. right before ask_user), the loop times out and flushes any buffered observer chunks to the DB first, guaranteeing they precede the subsequent human_interaction row. Empty queue idle periods are essentially zero-cost; overall DB write frequency stays on par with the original. * fix(hitl): guarantee observer chunks precede human_interaction in DB event order When the agent invokes ask_user, two independent write paths caused the human_interaction row to be persisted BEFORE model_output_thinking / parse chunks, breaking the SSE replay ordering: the worker thread writes HITL events synchronously via SQLAlchemy, while observer messages flow through the async consumer which only flushes on a batched threshold. Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort (port.add_chunk / port.take_chunks). The async consumer pushes every processed chunk there; the worker thread calls flush_chunks_until_idle() before dispatching any HITL event — it polls the shared buffer and waits for the async loop to drain the observer queue (20ms idle window, 500ms max wait), then persists every chunk in its own transaction. This guarantees chunk event_seq < human_interaction event_seq regardless of async scheduling latency. Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel now falls back from httpx2 to httpx at import time. * test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer and the flush_chunks_until_idle poll loop that guarantees observer chunks precede human_interaction events in DB order. All five HITL entry points (dispatch / boundary / receipt / finish / _wait_until_ready) are verified to invoke the idle flush before opening their transaction. Also add two tests for the openai_llm httpx2 → httpx ImportError fallback introduced to support openai >= 2.50 where the httpx2 shim was removed: one covers the fallback path, one confirms httpx2 still wins when present. * fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety Fix 4 real issues flagged by github-code-review: 1. Non-resettable hard_deadline in flush_chunks_until_idle — previously reset on every drain, meaning a model that kept producing chunks could stall the worker forever. deadline is now computed once at entry and the sleep call clips to hard_deadline - now. 2. _emit_in_flight Event bridges the async emit path and the worker's idle poll. Without this, buffer-empty = 'persisted' was confused with buffer-empty = 'taken for emit but still in run_blocking queue'. The worker now checks both 'buffer empty for settle_ms' AND 'no emit in flight' before deciding the async side is truly idle. 3. peek_chunks() replaces the take-put-back pattern in _flush_if_due. Previously the async loop drained the buffer, decided it was not yet due, then put everything back. That transiently-empty window (16 us normally, arbitrarily long under GIL/GC/preemption) was enough for the worker's 20 ms poll to mis-fire. We now peek (read count, no drain) and only take_chunks when we actually intend to persist. 4. emit_chunks wrapped in try/except that puts drained chunks back into the shared buffer before re-raising, and finish() wraps its flush call in try/except: pass. Guarantees (a) no chunk loss on DB failure and (b) the terminal human_run row is always written even if the flush step fails. Tests added: - 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit, try/except path, and every HITL entry-point's flush-before-transaction. - 1 async execute_attempt integration test in new test_application_execute_attempt.py drives the full consumer loop through _flush_if_due (peek → take → begin/end_emit) and the final flush, verifying that every patch line added in application.py is hit. * perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected When isRunning=true the EventSource already pushes human_interaction and human_execution events in real time — the 1.5s polling loop duplicated that work, hitting the DB and re-rendering the frontend for every tick. Disable polling entirely while the SSE stream is alive, and drop to 5s intervals only when the stream is closed (e.g. page load before the first run, or after a run finishes) so we can still discover WAITING_HUMAN requests that were created while the client was disconnected. Add isRunning to the useEffect dependency array so the polling cadence resets immediately when the SSE connection state changes. * perf(hitl): replace polling with SSE subscription and move snapshot off the write path Frontend — /conversation polling → /{run_id}/events SSE: - Replace the 5s conversation snapshot polling with a native EventSource subscription to the backend's /{run_id}/events SSE stream. Discovery is now one-shot: conversationId change and the agent stream pause (isRunning true→false), the exact moment a HITL run is most likely to exist. The SSE stream then keeps run state live with native auto-reconnect. - Add dual guards inside refresh() to absorb the thundering herd from adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all concurrent callers and returns cached state; (2) 3s minimum interval so bursts after the in-flight resolves do not immediately re-hit DB. - Use a runRef mirror so refresh() stays stable and downstream effects do not re-run on every snapshot. - Detect terminal status inside SSE onmessage and proactively es.close() to prevent EventSource from reconnecting forever against COMPLETED runs. Backend — snapshot off the write path: - Add repository.read_only() context manager: plain SELECT without WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads must not contend with worker writes on the same row lock. - Add service.light_snapshot() using read_only. Retain snapshot() as a writer-path API for any future lock-held callers. - Route conversation_snapshot, run snapshot endpoint, and both snapshot calls inside stream_run() through light_snapshot. - Move expiration to the writer path: decide() still calls _expire inline before processing each request, and expire_waiting() remains the periodic scheduler sweep. Impact: conversation snapshot calls drop from 12+/min (polling) or 10+/s (burst from adapter + SSE) to at most one every 3s. Each call is now two plain SELECTs instead of a lock-held transaction with a possible write from _expire. Read and write paths are fully decoupled. * fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning stayed true — the stop button remained visible and new messages went into the queue buffer instead of being sent normally. The root cause is that isRunning is driven entirely by the ChatModelRun generator lifetime, which only returns when the backend SSE HTTP connection closes (reader.read() -> done=true). The backend stream_run loop can hang on heartbeat even after the run is terminal when the SSE was opened during WAITING_HUMAN with attempt_active=true: the break condition requires both cursor >= event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false and empty rows). If continueHitl fires mid-flight with a stale after_event, the cursor never catches up, so the SSE stays alive forever and the generator never returns. Stop depending on the backend closing first. Inside the adapter's SSE chunk loop, detect a terminal human_run event (status in COMPLETED, FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner for-loop, and let the outer while-loop exit via the same flag on the next iteration. Assistant-UI sees the generator return and flips isRunning false immediately. Only affects HITL streams — the normal non-HITL agent path never emits human_run events so this branch is never taken. * test(hitl): update mock from snapshot to light_snapshot after read-path refactor test_human_interaction_app.py still mocked service.snapshot after commit 288ae4e moved conversation_snapshot and the run snapshot endpoint to service.light_snapshot (read-only path, no lock, no _expire). The fixture return_value and the two assert_called_once_with/assert_not_called assertions all referenced the old method name, causing CI to fail because MagicMock.snapshot was never called. * test(hitl): raise diff coverage above the 90% merge gate Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.
* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear) * Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key * Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation * Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names) * Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it * Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots) * Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)" This reverts commit 2376aa5. * Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves. --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* fix(evaluation): support multi-agent task filtering Default the evaluation task list to all tenant agents, add searchable multi-select filtering with one JSON agent_ids parameter, and preserve the active filter after task creation. Co-authored-by: Codex <noreply@tool-provider.com> Generated-by: gpt-5-codex * refactor(evaluation): simplify agent id validation Use Pydantic strict JSON validation for the agent_ids query parameter while preserving the all-agents default and stable deduplication. Co-authored-by: Codex <noreply@tool-provider.com> Generated-by: gpt-5-codex
* fix(newchat): reject attachments larger than 10 MB Validate attachments before upload in the newchat page and show a localized size error so oversized files never enter the upload queue. Cover the exact 10 MB boundary. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(i18n): clarify local A2A development URLs 区分本地直启、本地容器和 k8s 启动时使用的 A2A 基础 URL。 Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol * fix(i18n): sync English local A2A URL hint Clarify the local, local-container, and k8s A2A base URLs in English. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol * fix(newchat): limit attachments to 50 files Track pending newchat attachments per runtime and reject the 51st file with a localized error message. Releasing slots on removal or successful upload keeps subsequent selections available. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 --------- Co-authored-by: Codex <noreply@openai.com>
…permissions to create folders and files. (#3964)
* fix(agent-version): return not found for absent version Reject direct lookups for absent version metadata so the existing API mapping returns HTTP 404 rather than 200 null.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * refactor(agent): remove legacy draft creation endpoint Remove the unused draft allocation API and its frontend client.\nMigrate creation coverage to the existing update endpoint.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * test(agents): expose stable variable name locator Add an explicit test contract for the agent variable-name field so browser flows do not depend on localized placeholder text. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-version): scope version lookup to tenant Filter version metadata by tenant ID to prevent cross-tenant lookups and cover the predicate with a regression test. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex
* [codex] fix(agent): silently retry transient model errors (#3965) * fix(agent): retry transient model failures silently * fix(model): support OpenAI httpx2 timeout client * test(model): add deterministic OpenAI-compatible mock * fix(agent): keep stream runtime within line budget * fix(agent): enforce explicit CodeAgent termination * fix(test): restore CodeAgent CI compatibility * test(agent): keep develop isolation checks compatible * fix(agent): hide protocol repair generation stream * fix(ci): satisfy CodeAgent quality gate * test(model): cover retry classification branches
…ncurrency (#3977) * Fix StopAsyncIteration leak in execute_attempt finally block Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors. Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate. * Fix HITL form not appearing until page refresh Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload. Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh. * Fix SSE chunk/HITL event ordering race under real server load Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order. Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order. * fix(hitl): recover pending form when SSE delivery stalls silently Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones. Changes: - Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls - Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it - Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop * fix(hitl): stop parked human-input waits from consuming scheduler slots Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed. Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments. Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots. Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43. * test(scheduler): fix flaky waiting-concurrency assertion on fast event loops Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently. Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.
* feat:新增需求设计开发阶段测试用例及脚本生成skill * feat:新增a2a测试mock服务 * feat:一键部署mock服务 * test:ad test
* fix: standardize tenant resource limit errors * fix: roll back rejected tenant user registrations * fix: stabilize tenant limit test imports * test: import HTTPStatus in northbound tests * test: cover tenant resource limit error paths * config: make tenant resource limits configurable
* refactor(resource): unify repository resource cards Use the shared resource grid and card shell for Agent, MCP, and Skill repository listings while preserving their existing search, filtering, pagination, and actions. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(resource): share resource create cards Route Agent, MCP, and Skill create-card entry points through the shared CreateResourceCard and use the shared grid for Agent and Skill mine lists. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(resource): share mine resource card shells Apply the shared ResourceCard shell to Agent, MCP, and Skill mine-resource cards while retaining their existing content and actions. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(resource): align public repository card actions Use the shared card's stacked footer layout for Agent, MCP, and Skill repository listings so metadata and actions remain consistent across resource types. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(agent-space): split tab content components Move repository, my-agent, and review-center tab content behind same-level components and expand the card grids to four desktop columns.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * refactor(agent-space): move my agent view Place the complete my-agent view beside the route page and preserve its resource-card grid behavior.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * fix(agent-space): resolve moved my agent imports Update relative component imports after moving the my-agent view and keep layout regression coverage aligned.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * refactor(agent-space): remove legacy tab views Delete duplicate repository and review tab implementations after moving them into dedicated components.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * refactor(agent-space): localize detail dialogs Move repository copy and detail dialogs into the repository tab and version detail into the my-agent tab.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * refactor(agent-space): localize tab state and queries Move list, filter, pagination, detail, and mutation state into each tab view. Keep tab counts on the page and preserve filters across tab switches while disabling inactive queries. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(agent-space): use shared cards and pagination Render repository listings with ResourceCard and delegate server pagination to ResourceCardGrid. Rename the view to space.tsx and remove the unused agent card adapter and pagination controls. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(agent-space): streamline repository card actions Open repository details from the card, move copying into the action menu, show versions as title badges, and place authors in the card footer. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * feat(agent-space): adapt repository cards to the viewport Use Ant Design breakpoints and measured card-region height to select the server page size. Constrain ResourceCardGrid rows to the available viewport space. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): clamp card descriptions to available height Reduce repository-card description lines as card rows shrink and make the shared card description overflow-safe. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): fit repository cards within viewport Remove duplicate tab spacing and reserve the responsive page bottom padding when sizing the card grid. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): match repository card minimum height Use the shared card minimum height when calculating responsive rows so cards and pagination remain within the viewport. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(agent-space): reposition repository card actions Move copying to the card footer and show download counts beside repository names. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): defer four-column repository layout Keep repository cards in three columns through xl and use four columns only at the xxl breakpoint. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): keep lg repository cards in two columns Limit the repository grid to two columns through lg, then use three columns at xl and four at xxl. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(resource-card): align small-screen grid breakpoint Start the two-column resource grid at 576px so CSS matches Ant Design responsive state and pagination capacity. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): emphasize repository copy action Use a bordered default button for copying repository agents from the card footer. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): strengthen repository copy button Increase the copy action size and use the primary button treatment for a clear visual affordance. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): simplify my agent toolbar Remove redundant import and create-agent toolbar actions, align tag filtering with search, and reuse the shared resource pagination component. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol * fix(agent-space): reorganize my agent card actions Move evaluation into the card overflow menu and keep editing in the lower-right action area. Reposition version, lifecycle status, repository status, and creation date for clearer scanning. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol * fix(agent-space): consolidate repository status badge Show the listed version status beside the agent name and remove the redundant Hub badge. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol * fix(agent-space): use blue badge for listed agents Match the listed repository status badge to the former Hub blue styling. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol * fix(agent-space): clarify current version label Change the agent card version label to Current Version and add a primary color indicator for the active version. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol * fix(agent-space): align lifecycle badge with more menu Place the published or draft badge beside the card menu so both share the top-right header alignment. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): place lifecycle badge below more menu Stack the published or draft badge below the top-right menu to match the card header layout. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): unify repository version display Show the repository card version with the current-version label and typography used by my agent cards. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): unify my agent action button style Use the repository card text-button treatment for editable and read-only my agent card actions. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(agent-space): fill my agent grid to viewport Measure the available viewport height and distribute my agent cards across responsive grid rows. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(mcp-space): split tab content into components Move MCP repository, mine, and review content into focused components and keep shared actions in a controller. Match agent-space page spacing and tab styling while preserving existing interactions. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(mcp-space): keep three tab components Consolidate the MCP tab content into space, my-mcp, and review-center files while keeping the page focused on tab selection and preserving shared interactions. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * refactor(skill-space): split three tab views Move repository, mine, and review state and actions into their corresponding components. Keep the route focused on tab selection and align its spacing and tabs with the other spaces. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * fix(mcp-space): refine repository card actions Move downloads beside the more menu, show the publisher in the metadata row, and make installation the sole footer action. Open the detail modal when the card itself is clicked. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5-codex * fix(mcp-space): align repository card footer Match the agent repository card layout by placing the publisher and install action in the same inline footer row. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5-codex * fix(agent-space): align repository download action Place the repository download count in the top-right action area so it remains adjacent to the conditional more menu.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * fix(agent-space): align repository search height Use the default Ant Design input height so the repository search field matches the adjacent tag filter button. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-space): match repository copy button styling Use the same transparent text-button treatment as My Agent card actions for repository copy controls, and add a regression assertion. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5-codex * fix(mcp-space): align search filter control heights Match the MCP search input and status select height to Ant Design's default tag filter button height. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol Co-authored-by: Cursor <cursoragent@cursor.com> * fix(mcp-space): match repository action button style Use the same borderless text-button treatment as the Agent repository card while keeping installed repository entries disabled. Co-authored-by: Cursor <noreply@cursor.com> Generated-by: gpt-5.6-sol * refactor(mcp-space): align responsive repository cards Align Agent and MCP repository layouts, and simplify My MCP card actions for direct editing and compact controls. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-space): align repository status badges Place the repository listing status beside the lifecycle state on My Agent cards. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(resource): top-align card header actions Align multi-line header action groups from their top edge so adjacent controls stay level. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-space): center repository header actions Keep the download indicator vertically aligned with the adjacent more menu in repository cards.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * fix(mcp-space): align mine card header actions Keep the more menu beside the health-check action while the review badge stays below it.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * fix(mcp-space): group mine card header actions Place health check and more actions on the first header row, with the review badge beneath them.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * fix(mcp-space): right-align mine review badge Align the review status badge below the more action at the card header's right edge.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * refactor(mcp-space): unify card grid pagination Use ResourceCardGrid for repository and mine MCP cards so pagination follows the shared responsive layout.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * fix(resource): always show card pagination Keep the shared pagination control visible for single-page card lists.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * fix(resource): reserve space for card pagination Include the always-visible pagination area in responsive card grid height calculations.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * docs(skill-repository): document card design Document the approved Skill repository card layout and implementation plan. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * refactor(skill-repository): align cards with resource grid Use responsive ResourceCardGrid pagination and streamline Skill repository and My Skills card actions. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(skill-repository): refine card metadata layout Square Skill search controls, move filtering beside search, reserve tag space, and show repository authors. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(skill-repository): enforce card metadata layout Apply square controls without relying on CSS ordering and display the repository submitter when author is absent. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(skill-repository): match agent search controls Use the Agent repository search radius and default filter button styling in both Skill tabs. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(skill-repository): use default component radius Remove the Skill page radius override so Ant Design controls match the Agent repository. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(skill-repository): render listing tags directly Render repository tags as direct ResourceCard content while preserving empty tag space when no tags exist. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(skill-repository): show managed Skill tags Populate My Skill cards from structured tag assignments with a legacy tag fallback. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(skill-repository): use default description lines Remove the two-line override from Mine Skill cards so descriptions use the shared ResourceCard three-line default.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * refactor(resource-cards): align default action styles Use default small text-button styling in Agent and MCP cards, keep the Agent page full height, and rename the Chinese MCP page title.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * test(resource-cards): remove temporary layout tests Remove the Agent and Skill layout test files created during the refactor-card work.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * test(resource-cards): remove remaining temporary tests Remove the remaining Agent and MCP layout test files requested for the refactor-card branch.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * docs(skill-repository): remove temporary plans Remove the Skill repository design and implementation plans created for this branch.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 * test(skill-repository): isolate managed tag lookup Mock tag assignment lookups in Skill repository service unit tests so they do not connect to PostgreSQL.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5 --------- Co-authored-by: Codex <noreply@openai.com> Co-authored-by: Cursor <cursoragent@cursor.com>
* 🐛 Fix(evaluation): run trials in runtime service (#3954) * fix(evaluation): run trials in runtime service Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * test(evaluation): stub config thread manager Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * test(evaluation): stub runtime jwt helper * test(evaluation): cover trial proxy error paths * Fix/override delete (#3958) * Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear) * Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key * Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation * Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names) * Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it * Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots) * Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)" This reverts commit 2376aa5. * Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves. --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> * cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959) * Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948) * fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests Root cause: the worker thread writes human_interaction events synchronously via SQLAlchemy in ask_user, while model_output_thinking/parse observer messages flow through the async consumer and are flushed only on a batched threshold (32 chunks or 250ms). When the worker suspends before that flush fires, human_interaction gains a lower event_seq number than the already-buffered observer chunks, causing the SSE replay stream to show them in the wrong order. Fix: replace the plain async-for consumer loop with a manual asyncio.wait iterator using a 50ms timeout. Once the worker finishes producing model output (i.e. right before ask_user), the loop times out and flushes any buffered observer chunks to the DB first, guaranteeing they precede the subsequent human_interaction row. Empty queue idle periods are essentially zero-cost; overall DB write frequency stays on par with the original. * fix(hitl): guarantee observer chunks precede human_interaction in DB event order When the agent invokes ask_user, two independent write paths caused the human_interaction row to be persisted BEFORE model_output_thinking / parse chunks, breaking the SSE replay ordering: the worker thread writes HITL events synchronously via SQLAlchemy, while observer messages flow through the async consumer which only flushes on a batched threshold. Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort (port.add_chunk / port.take_chunks). The async consumer pushes every processed chunk there; the worker thread calls flush_chunks_until_idle() before dispatching any HITL event — it polls the shared buffer and waits for the async loop to drain the observer queue (20ms idle window, 500ms max wait), then persists every chunk in its own transaction. This guarantees chunk event_seq < human_interaction event_seq regardless of async scheduling latency. Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel now falls back from httpx2 to httpx at import time. * test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer and the flush_chunks_until_idle poll loop that guarantees observer chunks precede human_interaction events in DB order. All five HITL entry points (dispatch / boundary / receipt / finish / _wait_until_ready) are verified to invoke the idle flush before opening their transaction. Also add two tests for the openai_llm httpx2 → httpx ImportError fallback introduced to support openai >= 2.50 where the httpx2 shim was removed: one covers the fallback path, one confirms httpx2 still wins when present. * fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety Fix 4 real issues flagged by github-code-review: 1. Non-resettable hard_deadline in flush_chunks_until_idle — previously reset on every drain, meaning a model that kept producing chunks could stall the worker forever. deadline is now computed once at entry and the sleep call clips to hard_deadline - now. 2. _emit_in_flight Event bridges the async emit path and the worker's idle poll. Without this, buffer-empty = 'persisted' was confused with buffer-empty = 'taken for emit but still in run_blocking queue'. The worker now checks both 'buffer empty for settle_ms' AND 'no emit in flight' before deciding the async side is truly idle. 3. peek_chunks() replaces the take-put-back pattern in _flush_if_due. Previously the async loop drained the buffer, decided it was not yet due, then put everything back. That transiently-empty window (16 us normally, arbitrarily long under GIL/GC/preemption) was enough for the worker's 20 ms poll to mis-fire. We now peek (read count, no drain) and only take_chunks when we actually intend to persist. 4. emit_chunks wrapped in try/except that puts drained chunks back into the shared buffer before re-raising, and finish() wraps its flush call in try/except: pass. Guarantees (a) no chunk loss on DB failure and (b) the terminal human_run row is always written even if the flush step fails. Tests added: - 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit, try/except path, and every HITL entry-point's flush-before-transaction. - 1 async execute_attempt integration test in new test_application_execute_attempt.py drives the full consumer loop through _flush_if_due (peek → take → begin/end_emit) and the final flush, verifying that every patch line added in application.py is hit. * perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected When isRunning=true the EventSource already pushes human_interaction and human_execution events in real time — the 1.5s polling loop duplicated that work, hitting the DB and re-rendering the frontend for every tick. Disable polling entirely while the SSE stream is alive, and drop to 5s intervals only when the stream is closed (e.g. page load before the first run, or after a run finishes) so we can still discover WAITING_HUMAN requests that were created while the client was disconnected. Add isRunning to the useEffect dependency array so the polling cadence resets immediately when the SSE connection state changes. * perf(hitl): replace polling with SSE subscription and move snapshot off the write path Frontend — /conversation polling → /{run_id}/events SSE: - Replace the 5s conversation snapshot polling with a native EventSource subscription to the backend's /{run_id}/events SSE stream. Discovery is now one-shot: conversationId change and the agent stream pause (isRunning true→false), the exact moment a HITL run is most likely to exist. The SSE stream then keeps run state live with native auto-reconnect. - Add dual guards inside refresh() to absorb the thundering herd from adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all concurrent callers and returns cached state; (2) 3s minimum interval so bursts after the in-flight resolves do not immediately re-hit DB. - Use a runRef mirror so refresh() stays stable and downstream effects do not re-run on every snapshot. - Detect terminal status inside SSE onmessage and proactively es.close() to prevent EventSource from reconnecting forever against COMPLETED runs. Backend — snapshot off the write path: - Add repository.read_only() context manager: plain SELECT without WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads must not contend with worker writes on the same row lock. - Add service.light_snapshot() using read_only. Retain snapshot() as a writer-path API for any future lock-held callers. - Route conversation_snapshot, run snapshot endpoint, and both snapshot calls inside stream_run() through light_snapshot. - Move expiration to the writer path: decide() still calls _expire inline before processing each request, and expire_waiting() remains the periodic scheduler sweep. Impact: conversation snapshot calls drop from 12+/min (polling) or 10+/s (burst from adapter + SSE) to at most one every 3s. Each call is now two plain SELECTs instead of a lock-held transaction with a possible write from _expire. Read and write paths are fully decoupled. * fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning stayed true — the stop button remained visible and new messages went into the queue buffer instead of being sent normally. The root cause is that isRunning is driven entirely by the ChatModelRun generator lifetime, which only returns when the backend SSE HTTP connection closes (reader.read() -> done=true). The backend stream_run loop can hang on heartbeat even after the run is terminal when the SSE was opened during WAITING_HUMAN with attempt_active=true: the break condition requires both cursor >= event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false and empty rows). If continueHitl fires mid-flight with a stale after_event, the cursor never catches up, so the SSE stays alive forever and the generator never returns. Stop depending on the backend closing first. Inside the adapter's SSE chunk loop, detect a terminal human_run event (status in COMPLETED, FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner for-loop, and let the outer while-loop exit via the same flag on the next iteration. Assistant-UI sees the generator return and flips isRunning false immediately. Only affects HITL streams — the normal non-HITL agent path never emits human_run events so this branch is never taken. * test(hitl): update mock from snapshot to light_snapshot after read-path refactor test_human_interaction_app.py still mocked service.snapshot after commit 288ae4e moved conversation_snapshot and the run snapshot endpoint to service.light_snapshot (read-only path, no lock, no _expire). The fixture return_value and the two assert_called_once_with/assert_not_called assertions all referenced the old method name, causing CI to fail because MagicMock.snapshot was never called. * test(hitl): raise diff coverage above the 90% merge gate Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass. * fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare import httpx2 causes ImportError in CI and on systems with recent openai versions. This restores the try/except fallback introduced in PR #3948 commit 9521b93 and later accidentally reverted by commit d086da2. * Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm" This reverts commit 8e06ce3. * 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963) * Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962) * Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in. * chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> * [codex] fix(agent): silently retry transient model errors (#3965) * fix(agent): retry transient model failures silently * fix(model): support OpenAI httpx2 timeout client * test(model): add deterministic OpenAI-compatible mock * fix(agent): keep stream runtime within line budget * Fix: AIDP knowledge base bug fix (#3967) * Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection * refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change * [fix] enforce explicit CodeAgent termination and silent recovery (#3969) * fix(agent): enforce explicit CodeAgent termination * fix(test): restore CodeAgent CI compatibility * cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981) * Fix StopAsyncIteration leak in execute_attempt finally block Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors. Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate. * Fix HITL form not appearing until page refresh Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload. Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh. * Fix SSE chunk/HITL event ordering race under real server load Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order. Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order. * fix(hitl): recover pending form when SSE delivery stalls silently Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones. Changes: - Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls - Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it - Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop * fix(hitl): stop parked human-input waits from consuming scheduler slots Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed. Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments. Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots. Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43. * test(scheduler): fix flaky waiting-concurrency assertion on fast event loops Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently. Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window. * Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980) The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which also carried two unrelated upstream changes that this release branch never had: - #3909 Knowledge base interface optimization (AIDP UI refactor) - #3930 support AIDP knowledge file deletion and download Both are removed here so the release line keeps only the bug fix. - Restore the AIDP frontend components to their pre-refactor layout and drop the helper modules only the refactor used: AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils. - Drop the #3930 document remove/download endpoints from services/api.ts and the AIDP translations that only those screens referenced. - Keep the fix itself unchanged: knowledge-base scoped Channels and KnowledgeFiles/History paths, all-status document listing, keyword search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the upload-triggered polling. Verified: - pytest test/ext_components/aidp -q -> 583 passed - frontend `npm run type-check` (tsc --noEmit) -> no errors * fix(agent): accept reasoning-prefixed code actions (#3990) * refactor: remove human interaction features and related configurations (#3988) * refactor: remove human interaction features and related configurations * refactor: remove human interaction features and related configurations --------- Co-authored-by: cj2026-bit <647646783@qq.com> Co-authored-by: lijiayang619 <1170349871@qq.com> Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> Co-authored-by: bernard1234 <840646206@qq.com> Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com> Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com> Co-authored-by: gs-aion <gs597153711@qq.com> Co-authored-by: Dallas98 <40557804+Dallas98@users.noreply.github.com> Co-authored-by: chase <byzhangxin11@126.com>
…st suite (#4002) - quota: reject negative GB/MB inputs and warning>=critical threshold pairs (previously persisted as 200 with semantically inverted config) - api-key: mask access_key in /api-keys and /user/tokens list responses; only create/refresh may return the full secret once - tag: translate the DB trigger's 'Tag assignment limit exceeded' (without the 'Resource ' prefix) into a structured 409, and flush replacement deletes before inserts so a full tag replacement no longer trips the assignment-capacity trigger - model: propagate ValueError from create/batch_create so duplicate display names map to 409 and malformed batch entries to 422 instead of 500; include model_type in /model/llm_list items; tolerate quick-config entries without model_repo in get_model_name_from_config - test: align model service tests with the ValueError passthrough contract Co-authored-by: chase <byzhangxin11@126.com>
…ry (#4004) The guard rejected private/reserved IP literals and DNS names resolving into non-routable ranges before fetching {base_url}/models. Internal and IP-based deployments (including private-network OpenAI-compatible endpoints) were blocked from batch model discovery. Remove the _reject_non_public_ip / _validate_provider_base_url / _assert_resolved_ips_public chain and the related tests; keep the ssl_verify TLS toggle, which is independent of the guard. Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
…e authorization (#3798) * ✨ Feature: Pass user context to MCP tools and A2A agents for tool-side authorization * test: update NexentAgent exact-call assertions for user_context kwarg * test: cover user_context injection branches for codecov patch coverage * docs: remove internal tool user context design * fix: support legacy agent run info context * fix: hide injected MCP context from tool signatures * fix: hide injected MCP context from action traces * Revert "fix: hide injected MCP context from action traces" This reverts commit de25e57. * fix: hide injected context from model tool metadata
* bugfix:支持思考挡位配置 * bugfix: support model reasoning effort configuration * test: improve reasoning configuration coverage * test: complete reasoning coverage * fix: resolve CI quality gate failures * test: complete coverage for gate branches * bugfix: support configurable model reasoning effort * test: complete reasoning coverage * test: cover reasoning override branch * bugfix: support model reasoning effort settings * bugfix: complete reasoning capability matching * test: align reasoning capability assertions * test: close reasoning coverage gaps --------- Co-authored-by: hzw <hzw@qq.com>
…3996) * fix(file): reject invalid upload destinations at the API boundary An unknown destination value previously reached the service layer, where it raised a bare Exception and surfaced as HTTP 500 with a generic "File upload error." message. Constrain the form field to the documented "local"/"minio" values so FastAPI rejects unknown destinations with 422 before any business logic runs. Co-authored-by: ZCode <noreply@zcode.ai> Generated-by: new-provider/glm-5.3 * fix(data-process): call module-level get_task_details in the details route The route called service.get_task_details, a method that has never existed on DataProcessService (regression introduced by the coding-style refactor in 4a70a99), so every request to GET /tasks/{task_id}/details failed with AttributeError and returned 500. Call the data_process.utils implementation that the celery support introduced instead. Co-authored-by: ZCode <noreply@zcode.ai> Generated-by: new-provider/glm-5.3 * fix(file): stop the upload conflict scan from matching its own batch The filename conflict scan in upload_files_impl calls list_files after the current batch's lifecycle rows have been created, and list_files has merged durable lifecycle rows into the file list since the lifecycle ledger landed. The scan therefore counted the upload's own row as an existing document and renamed every first upload of a name to <name>_1. Exclude the current batch's file_ids when collecting existing names. ES entries carry no file_id and are still treated as pre-existing documents, so genuine conflicts keep resolving to _1 and within-batch duplicates keep getting sequential suffixes. Covered by a unit test whose list_files fixture exposes only the batch's own row: the uploaded name is kept unchanged and no filename rewrite is recorded. The shared upload mock harness is extracted so the new test does not duplicate the existing conflict-resolution fixtures. * fix(evaluation): allow tenant administrators to delete evaluation runs DELETE /agent-evaluations/{id} only allowed the run's creator (160208 / HTTP 403). When the creator account was deleted, the run could not be removed until the 30-day cleanup, and runs with NULL created_by could not be deleted by anyone. After the creator check, also allow callers whose tenant role is in consts.const.CAN_EDIT_ALL_USER_ROLES (SU/ADMIN/SPEED/ASSET_OWNER). The role is resolved inside the service via database.user_tenant_db (unknown callers fall back to USER, fail-closed), keeping the endpoint signature unchanged. Error copy updated in error_message.py and the zh/en locales; code 160208 stays the same. Covered by three new unit tests (admin may delete others' runs, DEV cannot, unknown role fails closed) on top of the existing creator-only test. * test(evaluation): stub database.user_tenant_db in the pure-logic harness CI only failed on test_evaluation_pure_logic: the harness reloads services.agent_evaluation_service, whose new 'from database.user_tenant_db import get_user_tenant_by_user_id' pulled the real module outside the harness's stub set; the real module then executed 'from database.db_models import TenantGroupInfo' against this file's two-class db_models stub and raised ImportError at setup for all 27 tests. Register database.user_tenant_db in the harness (same pattern as the other database submodules) so the reload stays fully stubbed. --------- Co-authored-by: ZCode <noreply@zcode.ai>
* bugfix: remove automatic model renaming * fix: restore reasoning compatibility for test gate * test: cover reasoning wire adapters * fix: validate provider hosts before reasoning mapping --------- Co-authored-by: hzw <hzw@qq.com>
Non-array skill_tags values persisted in the skill_t JSON column caused the "t.tags || []).forEach is not a function" error on the agents page when the skill selection dialog computed its tag filter set. - Backend _to_dict now normalizes skill_tags into a validated string list - Frontend fetchSkills reuses normalizeTags for the same defense in depth - Add malformed-tags coverage to test_skill_db.py
…3978) * fix(resource-manage): paginate knowledge base list in quota settings modal The quota settings modal rendered the knowledge base breakdown table without pagination, which froze the UI on deployments with many KBs. Use a 10-row small pagination with a total count hint (zh/en locales). * fix(resource-manage): harden platform quota panel on slow and large deployments Three related issues in the SU platform quota overview: - The panel never listened to QUOTA_USAGE_CHANGED_EVENT, so quota changes made elsewhere did not refresh it until reopen. - With data still loading, the header rendered misleading \"0 B / Unlimited\" placeholders; on a 786-tenant deployment the overview aggregation is slow enough that users saw quota values apparently reset to zero. Show a Card loading state until real data arrives. - toQuotaInput could floor a nonzero finite quota below 1 MB down to 0; clamp the MB view to at least 1 MB. Also add destroyOnHidden to the system settings modal so it always remounts with fresh data. * fix(resource-manage): stop bare 0 leak in quota fair share line The fair share reference rounded hard_limit_gb / kb_count with Math.round, so a 2 GB limit across many KBs produced 0. Because the render condition used the value directly, React printed the bare number 0 between the usage tag and the per-KB breakdown heading. Keep the raw ratio and format it to 2 decimals when it is not an integer (e.g. 0.17 GB/KB), and guard the JSX with != null so a null value never renders.
…elopers (#4008) - Add RBAC checks to all /model/* endpoints via permissions.depends.require - Restrict cross-tenant /manage/* endpoints to the SU role - Gate model page buttons by model:update / model:delete permissions - Remove the /models menu entry from the DEV role via a new migration - Keep model:read so agent editors can pick admin-configured models Co-authored-by: chase <byzhangxin11@126.com>
…options (#4013) * bugfix: preserve AIDP permissions when submitting collapsed advanced options * test: add AIDP regression evidence screenshot * chore: remove local regression screenshot artifact --------- Co-authored-by: hzw <hzw@qq.com>
* feat(logging): 新增模型调用日志分类与路由能力 1. 新增model_call日志分类,将SDK模型层日志与系统日志分离 2. 配置运行时服务日志同时包含runtime和model_call分类 3. 为core_agent、openai_llm等模型相关日志添加专用日志器 4. 新增.gitignore规则忽略.serena目录 5. 移除run_agent.py中多余的DEBUG日志级别设置 6. 完善.env.example日志配置注释 7. 新增完整的模型日志路由测试用例 * feat(logging): route context_evidence logger into model_call file - Add "context_evidence" to MODEL_CALL_LOGGERS so the run-level "Agent loop context evidence" record lands in the model_call category file (console + file_model_call, propagate=False) - Extend whitelist coverage and routing regression tests Co-authored-by: ZCode <noreply@zcode.ai> Generated-by: glm-5.3-flash * refactor(test): deduplicate model_call routing regression tests - Merge test_dictconfig_routes_model_records_to_dedicated_file and test_configure_logging_routes_model_records_to_dedicated_file into one parametrized test (ids: dictconfig / configure_logging) - Shared log+assert body extracted to _log_and_assert_routing helper; per-path setup moved to module-level _apply_* functions Co-authored-by: ZCode <noreply@zcode.ai> Generated-by: glm-5.3-flash * feat(logging): land model input/output body records in model_call file - Pin the MODEL_CALL_LOGGERS whitelist to DEBUG in both the dictConfig "loggers" section and configure_logging bind/unbind, and widen the file_model_call handler to DEBUG, so MODEL INPUT PARAMETERS logged at DEBUG on model_call.core_agent is no longer dropped by root INFO - Add the symmetric MODEL OUTPUT body record in core_agent - Merge the core_agent module logger into the model_call.core_agent namespace so every core-agent record routes to the model_call file instead of runtime; verified end-to-end (model records land only in model_call, runtime records unaffected, 41 logging tests pass) Co-authored-by: ZCode <noreply@zcode.ai> Generated-by: glm-5.3-flash * test: 增加核心代理日志记录相关的测试用例并添加调试日志 在 core_agent 的运行逻辑中添加调试日志以记录新任务,同时为日志记录功能新增多项测试,包括验证模型输出是否写入 model_call 日志、日志记录器名称是否保持在指定命名空间、日志参数是否正确记录到文件和面板等,确保日志功能符合预期且敏感数据被正确脱敏。 * fix(core_agent): 修复模型输出日志记录时未读取最新值的问题 调整日志记录的位置,确保在记录模型输出前已完成 model_output 的赋值,避免记录到空值或旧值 * refactor(logging): 将模型调用日志级别调整为跟随根日志级别 调整模型相关日志的级别设置,不再单独固定为DEBUG,而是跟随根日志级别。同时修改SDK中的模型输入、输出和任务日志从DEBUG级别改为INFO,确保在默认根日志级别为INFO时这些关键模型日志能正常输出。更新测试用例以匹配新的日志级别配置,同时修改日志配置文件中的相关注释和设置,使日志行为更符合预期,避免关键模型日志在默认配置下被丢弃。 --------- Co-authored-by: ZCode <noreply@zcode.ai>
* feat: redesign model configuration page (v0 design) - inline default-model slot grid replaces the DefaultModelDialog (priority badges, per-slot connectivity dots and hints, configured counter, one-click verify), model library section keeps the existing table, filters, add/edit dialogs and all data flows
* feat: redesign model library as custom list (v0 design) - replaces the antd Table with lightweight rows (name + default badge, type badge, provider, connectivity dot with spinner, icon actions with tooltips), custom pagination, and a simplified toolbar (search + type + provider). Moves filtering/paging state into ModelLibraryList; drops the context/max-output columns and status filter per the design
* feat: batch edit / delete by connection group (v0 design) - new ModelManagerDialog groups the library by connection (source + API key + base URL); batch edit rotates key/URL for a whole group via partial updates, batch delete supports per-model checkboxes with group select-all. Extracts shared default-slot cleanup so deleted models always vacate their slots
* feat: one-click verify checks default-slot models only (v0 design) - the button lives in the default-config section, so it now delegates to the selection-based verifier instead of probing the whole library; non-default models stay checkable via the per-row action
* feat: rebuild the model config redesign on the project's shadcn/ui primitives - the first pass used antd components whose visual language diverges from the v0 design. Adds radix-based Select/Card/Label ui components and rewrites ModelSlotSelect, ModelLibraryList and ModelManagerDialog on them (quiet borders, status dots inside selects, custom pagination, connection-group cards), matching the v0 aesthetic
* fix: drop the per-slot status line under the select - the v0 design conveys connectivity via the dot inside the select only, and the raw status value (available) leaked untranslated text
* style: widen the gap between the status dot and model name in slot selects - the radix trigger renders selected content with its own tight spacing, wrapping dot+name in one flex span keeps dropdown and trigger consistent
* style: align section titles with the page header - add px-2 to the redesign content to match CARD_HEADER.PADDING (8px), so 默认配置 lines up with 模型设置
* feat: full-height layout for the models page - the default-config card keeps its natural height, the library section fills the remaining height with the list as an internal scroll area (toolbar and pagination stay fixed); rows keep adaptive height and 8 items per page as the default
* fix: legend shows all three status dots with labels (available/unavailable/not-detected) instead of a single green dot labeled as connected
* style: legend keeps a single green dot, reworded to 'green means connected'
* style: widen the model library row columns (model 256px / type 160px / source 128px / actions 128px) so the flex status column no longer leaves a wide gap before the actions
* style: cap the models page content column at 1440px - the redesign was authored for a ~1150px column; stretched to the shared 1920px container the library rows showed wide empty gaps. Scoped to this page only
* style: prettier
* style: widen models page content column to 1600px
* style: widen models page content column to 1760px
* feat: v0-style add-model dialog (single / batch tabs) and reorder buttons - single tab: provider / type / name / display name / base url / key; batch tab: provider + key + fetch list with checkboxes (all selected by default, per-row type badges). No per-row capacity editing or pre-submit connectivity gate (v0 logic) - models land not_detected and are checked/edited from the library. Add-model button moves to the first position; sync-ModelEngine opens the batch tab. The old ModelAddDialogV2 stays for edit mode.
* fix: batch add no longer pre-selects all fetched models - user picks what to import
* feat: search filter for the batch-add fetched model list - filters display only, selections survive searching; toggle-all applies to visible rows
* feat: per-row advanced settings in batch add - gear button on each fetched model opens a dialog with capacity fields (context/input/output/reserve tokens) and inference params (temperature, top_p, enable_thinking, custom params via ModelAdvancedSettings). Overrides apply on submit; a dot indicator marks rows with overrides
* fix: per-row advanced settings now loads real inference specs so temperature/top_p/thinking/custom-params actually render, matching the old dialog's field set (capacity + inference, no display name)
* feat: batch add auto-fills capacity from suggestCapacity on fetch, per-row + batch connectivity check (optional, no submit gate) - suggestions are local catalog lookups applied automatically; user-modified overrides win over suggestions; blue dot only marks user edits
* fix: advanced settings pre-fills catalog suggestions - key={settingsRowId} forces remount so useState initializes from the correct row's suggestions instead of sticking with the first mount's empty state
* fix: remove duplicated capacity fields in per-row settings - ModelAdvancedSettings in override mode already renders capacity + inference, so the manual capacity grid was redundant; suggestions now flow directly into the component value (snake_case keys)
* fix: hide tokenizer_family from the per-row advanced settings
* feat: advanced settings (capacity + inference params) in the single-add tab - gear button opens the same RowSettingsDialog used by batch add; overrides apply on submit
* feat: batch add provider dropdown includes a custom option - custom requires Base URL (validated on fetch), modelFactory set to OpenAI-API-Compatible
* fix: capacity coverage warning now references the redesigned UI (edit button instead of removed manage button)
* fix: simplify capacity warning text to avoid JSX parser edge case with special characters
* feat: require connectivity check before batch add + remove capacity coverage banner - submit blocked until all selected models pass the probe (same gate as the old dialog); banner referenced removed UI and the gate makes it redundant
* fix: remove dangling JSX fragments from banner deletion
* fix: show human-readable type labels (not raw ids like llm/vlm) in batch delete rows
* fix: embedding models group with their parent connection - strip /embeddings suffix when computing the group key so BAAI/bge-m3 joins the SiliconFlow group instead of forming its own
* fix: show '已配置' instead of misleading '未设置' when the list API doesn't return the key value
* fix: remove redundant API key display from connection group cards - every model in the list necessarily has a key configured, so showing a label for it is noise; only the URL line remains
* fix: also strip trailing slashes in connection grouping - v1/ and v1 were different keys after /embeddings removal
* style: rename 选整条 to 全选
* feat: batch delete models are collapsed by default with expand/collapse toggle (same as batch edit)
* fix: batch add carries the verified connectivity status into the created record - the gate requires all selected models to pass the probe, so connectStatus: available is sent instead of defaulting to not_detected
* feat: per-model display name in batch add settings dialog - optional field at the top, empty means auto-generate (model name + random suffix)
* feat: v0-style model edit dialog - single-page form with basic info (display name, readonly model name, type/provider badges, API key, Base URL) and expandable advanced config (capacity + inference params). Replaces the old tab-based ModelAddDialogV2 for edit mode; connectivity check intentionally dropped (use library row check / one-click verify)
* fix: hide tokenizer_family from the edit dialog advanced config
* style: edit dialog advanced config now uses v0-style fields with hints (如 131072 placeholders, range descriptions) instead of the shared ModelAdvancedSettings component
* feat: add deep thinking toggle and custom params section to the edit dialog advanced config
* style: remove descriptive hint text from edit dialog advanced config fields
* fix(model): probe image-generation (vlm2) models via /images/generations
Image-generation models are not served on the chat-completions endpoint,
so the shared VLM chat probe reported them as "model does not exist".
Split vlm2 out of the chat-probe group: the free provider-catalog check
runs first, then a minimal real generation against
{base_url}/images/generations (2xx = connected, 60s default timeout).
* fix(model): carry capacity/inference params through add and edit flows
The add dialogs merged buildInferenceParamsPayload output (snake_case)
into create params, but addCustomModel's buildCapacityRequestBody only
reads camelCase keys - catalog capacity suggestions were silently dropped
on create (rows landed with NULL capacity). Mirror the explicit mapping
ModelAddDialogV2 does: capacity fields to camelCase, max_output mirrored
into max_tokens, capacitySource "operator" when capacity is sent.
The edit dialog also only pre-filled temperature/top_p/custom params -
its init passed an empty spec list to advancedSettingsValueFromRecord, so
nothing loaded. Build the value directly from the model record (capacity,
inference params, enable_thinking, custom params) so the advanced config
reflects the stored configuration, and re-send capacitySource/maxTokens
on save for parity with the old edit dialog.
frontend-only companion to the vlm2 probe fix: unify the advanced-config
section into a shared ModelAdvancedConfig component (v0 field grid) used
by both the edit dialog and the add-dialog advanced settings; add
per-row type editing in batch add and type editing + in-dialog
connectivity probe in the edit dialog (probe carries inference params so
invalid custom params fail at probe time); align the default badge with
the display-name line in the library list.
* style: prettier modelService after develop merge
* feat(model): reasoning effort / budget controls in the v0 advanced config
Port the reasoning-effort feature (#3953 on develop) onto the v0-style
ModelAdvancedConfig so the models-page dialogs expose the same controls
as the legacy ModelAdvancedSettings: below the deep-thinking switch,
render the effort-level select or the budget-tokens input when the
catalog-declared capability provides one (budget takes precedence over
effort, mirroring the payload builder). Toggling thinking off clears
both keys so they are dropped from the wire payload.
Capability wiring:
- edit dialog: seeded from the model record, refreshed (debounced) via
suggestCapacity when the URL or type changes
- batch add: captured per row from the existing suggestCapacity lookup
on fetch and passed into the per-row settings dialog
- single add: looked up (debounced) while the settings dialog is open
* test(model): cover the vlm2 probe network-error branch
Codecov patch coverage flagged the except path of
_image_generation_connectivity_check as uncovered; add a transport-error
test so the v0 PR diff clears the 90% patch target.
* refactor(model): dedupe type maps and reasoning resolution for Sonar gate
The SonarCloud quality gate failed on new-code duplication (5.3% > 3%)
and cognitive complexity (18 > 15 in ModelAdvancedConfig). Consolidate the
duplicated blocks:
- new modelTypeUi.ts holds the single copy of TYPE_LABEL_KEY_MAP,
TYPE_BADGE_CLASS, STATUS_DOT_CLASS, typeLabel and useTypeOptions
(previously copied across four dialogs) plus a shared ProviderSelect
for the add-dialog tabs
- resolveReasoningControls / clampReasoningBudget extracted from
ModelAdvancedSettings and reused by ModelAdvancedConfig, replacing the
ported inline block (fixes the cognitive-complexity failure and the
nested-ternary warnings)
- normalizeUrlForGrouping drops the backtracking regex Sonar flagged
- clean unused imports (Alert, Badge, DialogFooter, ChevronUpIcon,
useMemo) and add a keyboard stopPropagation to the per-row type editor
wrapper span
* feat(model): probe gate and auto-detection in the single-add form
The batch tab gates submit on a passing connectivity probe and auto-fills
capacity / reasoning capability from suggestCapacity; the single-add form
did neither - models landed with a hardcoded 4096 max_tokens placeholder
and unverified credentials.
- add an in-form 检测连通性 button with inline result; submit is blocked
until the probe passes (same gate as batch) and the verified status is
carried into the created record
- run the catalog lookup (debounced) whenever name + URL are filled:
capacity prefill and reasoning capability flow into the gear dialog as
suggestions (user overrides still win) and a compact "已自动识别" hint
shows what was detected (context window / max output / thinking mode)
- any name / type / URL / key / settings change resets the probe
- extract capacitySuggestionToSettings shared by both tabs
* fix(model): disable browser autofill on the single-add form inputs
* fix(model): reset manual advanced settings when the single-add model name changes
* fix(model): seed enable_thinking for LLMs in the v0 advanced config
The payload builder drops reasoning_effort / reasoning_budget_tokens when
enable_thinking is not strictly true, but the v0 switch renders undefined
as ON - a user who only changed the effort level (never toggling the
switch) saved nothing. Materialize the default (true) for LLMs the same
way ModelAdvancedSettings does.
* refactor(model): express the default-model slots as a compact seed table
buildModelSlots' ten inline slot literals read as 21-46 line duplicated
blocks to Sonar (37.6% duplication on the file, the last chunk keeping
the PR above the 3% gate). Move the per-slot metadata into a
MODEL_SLOT_SEEDS table whose rows differ on every line and expand it in
buildModelSlots; behaviour is unchanged.
* chore: restore tsconfig.json after a local next build rewrote it
* refactor(model): shrink remaining Sonar duplication after the develop merge
Sonar's copy-paste detector normalizes string literals, so the slot seed
objects still matched each other in bulk (25.6% on ModelSlotSelect) and
the #4009 strip/seed effects I mirrored into ModelAdvancedConfig
duplicated ModelAdvancedSettings.
- ModelSlotSelect: split the long i18n defaults into SLOT_LABELS /
SLOT_HINTS records with a compact tuple table for the structural fields
- extract useReasoningFormEffects (capability-gated visibility + legacy
strip + enable_thinking seed) shared by both advanced-config components
* fix(model): gate the probe key fallback on the stored endpoint
The temporary-healthcheck fallback substituted a stored api_key while the
caller still controlled base_url, so any authenticated tenant member could
exfiltrate another model's stored key by probing it against their own
server (the key rides the Authorization header to the caller-chosen URL).
Only borrow the stored key when the probe targets the stored endpoint
(exact base_url match after stripping trailing slashes). Edit-dialog
probes with an untouched URL keep working; a changed URL with an empty
key now honestly fails instead of silently sending the old key to the
new address.
* fix(model): stop auto-suffixing display names in the v0 add dialogs
develop #4009 removed the automatic model renaming from the legacy V2
dialog but our v0 ModelAddDialog still appended a random 5-char suffix to
every batch-imported display name (kimi-k3j72bi). Use the clean model
name as the default display name, matching develop; a collision with an
existing row surfaces as a per-row failure the user can resolve with the
gear's display-name field instead of silently creating suffixed
duplicates.
* feat(model): make the model name editable in the v0 edit dialog
The edit dialog kept the model name read-only; allow editing it like
the type. The connectivity probe uses the edited name, the reasoning
capability lookup re-runs (debounced) when it changes, and saving a
changed name sends model_name plus a not_detected status reset so the
renamed record must be re-verified.
* fix(model): address review findings in the v0 model-config dialogs
Selective fixes for the Copilot review findings (outdated or intentional
items skipped):
- batch add: keep the dialog open when every create fails instead of
closing with only a warning; changing the shared API key or base URL
now clears all per-row probe results (they were measured with the
previous credentials)
- batch delete: require an explicit confirmation before the destructive
batch (states the count and the default-slot vacation)
- batch-add rows: render the selectable row as a div instead of a button
— the row hosts a Select and icon Buttons, and nesting interactive
elements inside <button> is invalid HTML; keyboard toggle preserved
- slot selects: associate the visible label with the select trigger
(htmlFor/id) for screen readers
- remove the dead capacity-coverage fetch chain (state, request-id ref
and type import) left over from the removed warning banner
---------
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* fix: preserve agent trace hierarchy and improve monitoring metadata Trace the full agent lifecycle, sandbox callbacks, child agents, and conversation title generation. Use user emails and agent display names, filter noisy spans, and add observability regression coverage. Include the current Docker build adjustments. * Trace sandbox warmup with agent monitoring metadata * Isolate lazy human interaction service in stop endpoint test --------- Co-authored-by: hhhhsc <name>
* feat: enforce knowledge base resource limits * test: cover knowledge resource limit branches * fix: align knowledge resource limits with latest develop * config: expose knowledge resource limits via environment * test: improve knowledge limit patch coverage
* feat: add agent sharing backend foundation * feat: expose northbound API base URL * feat: guide users to published agent usage * feat: add agent usage guide modal * feat: support scoped runtime rate limits * feat: add authenticated agent share session APIs * feat: isolate agent share run identities * feat: protect agent share run lifecycle * feat: add authenticated agent share page * feat: protect agent share responses from caching * feat: secure all agent share responses * feat: rate limit agent share runs * test: cover isolated agent share runtime sessions * test: cover agent share revocation during runs * fix: sanitize agent share management errors * test: cover agent share session history isolation * test: keep shared sessions out of conversation lists * feat: link northbound guide to API key settings * feat: localize agent share page feedback * feat: retry unavailable A2A guide settings * feat: restore published agent usage guide targets * feat: complete agent share guide lifecycle * test: lock agent publish guide outcomes * feat: guard and restore agent share pages * feat: improve agent guide accessibility * chore: align agent share validation style * fix: accept uuid share ids in token resolution * feat: proxy northbound api and relocate runtime config * test: add agent usage guide component coverage * chore: ignore local openspec skill copies * fix: route agent share APIs to runtime * fix: localize shared agent input * fix: redact agent share tokens from access logs * fix: apply share token redaction to uvicorn access logs * fix: configure runtime access log redaction * fix: preserve agent share login return * fix: support share runs without history * fix: correct northbound usage guide calls * feat: guide published agents through card menu * feat: share agents through chat deep links * fix: open agent deep links in current chat * fix: open agent deep links directly * fix: restore agent repository access and usage guide deep-link parsing * fix: wait for agent usage guide activation * fix: refresh editable agents after publish * fix: stabilize agent guide checks * refactor: remove standalone agent share runtime * fix: satisfy runtime and architecture unit tests * fix: remove unrelated deployment changes * test: cover agent share runtime paths * refactor: focus agent usage guide on frontend * fix: use published agent name in northbound guide * fix: stabilize northbound frontend proxy * remove .agents
…rtain scenarios (#4022) * bugfix: fix infinite loop when agent configuration update fails in certain scenarios * bugfix: fix infinite loop when agent configuration update fails in certain scenarios --------- Co-authored-by: hzw <hzw@qq.com>
* fix(agent): preserve model attempt streaming across semantic retries (cherry picked from commit e38fd50) * feat(agent): make strict output repair opt-in per agent * fix(agent): continue bare output and reinforce action protocol * test(agent): keep context reminder fixture lint clean * Restore legacy code action routing and accept final strict attempt * Gate empty model response retries on strict code action mode * Show a quiet final hint for legacy empty model responses * test(agent): adapt protocol regression tests to develop * fix(agent): align protocol tests and preserve guardrail user boundary
…, bump version (#4024) * merge v2.6.1 hotfix release from hotfix/v2.6.1 (#3971) * 🐛 Fix(evaluation): run trials in runtime service (#3954) * fix(evaluation): run trials in runtime service Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * test(evaluation): stub config thread manager Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * test(evaluation): stub runtime jwt helper * test(evaluation): cover trial proxy error paths * Fix/override delete (#3958) * Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear) * Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key * Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation * Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names) * Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it * Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots) * Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)" This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06. * Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves. --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> * cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959) * Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948) * fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests Root cause: the worker thread writes human_interaction events synchronously via SQLAlchemy in ask_user, while model_output_thinking/parse observer messages flow through the async consumer and are flushed only on a batched threshold (32 chunks or 250ms). When the worker suspends before that flush fires, human_interaction gains a lower event_seq number than the already-buffered observer chunks, causing the SSE replay stream to show them in the wrong order. Fix: replace the plain async-for consumer loop with a manual asyncio.wait iterator using a 50ms timeout. Once the worker finishes producing model output (i.e. right before ask_user), the loop times out and flushes any buffered observer chunks to the DB first, guaranteeing they precede the subsequent human_interaction row. Empty queue idle periods are essentially zero-cost; overall DB write frequency stays on par with the original. * fix(hitl): guarantee observer chunks precede human_interaction in DB event order When the agent invokes ask_user, two independent write paths caused the human_interaction row to be persisted BEFORE model_output_thinking / parse chunks, breaking the SSE replay ordering: the worker thread writes HITL events synchronously via SQLAlchemy, while observer messages flow through the async consumer which only flushes on a batched threshold. Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort (port.add_chunk / port.take_chunks). The async consumer pushes every processed chunk there; the worker thread calls flush_chunks_until_idle() before dispatching any HITL event — it polls the shared buffer and waits for the async loop to drain the observer queue (20ms idle window, 500ms max wait), then persists every chunk in its own transaction. This guarantees chunk event_seq < human_interaction event_seq regardless of async scheduling latency. Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel now falls back from httpx2 to httpx at import time. * test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer and the flush_chunks_until_idle poll loop that guarantees observer chunks precede human_interaction events in DB order. All five HITL entry points (dispatch / boundary / receipt / finish / _wait_until_ready) are verified to invoke the idle flush before opening their transaction. Also add two tests for the openai_llm httpx2 → httpx ImportError fallback introduced to support openai >= 2.50 where the httpx2 shim was removed: one covers the fallback path, one confirms httpx2 still wins when present. * fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety Fix 4 real issues flagged by github-code-review: 1. Non-resettable hard_deadline in flush_chunks_until_idle — previously reset on every drain, meaning a model that kept producing chunks could stall the worker forever. deadline is now computed once at entry and the sleep call clips to hard_deadline - now. 2. _emit_in_flight Event bridges the async emit path and the worker's idle poll. Without this, buffer-empty = 'persisted' was confused with buffer-empty = 'taken for emit but still in run_blocking queue'. The worker now checks both 'buffer empty for settle_ms' AND 'no emit in flight' before deciding the async side is truly idle. 3. peek_chunks() replaces the take-put-back pattern in _flush_if_due. Previously the async loop drained the buffer, decided it was not yet due, then put everything back. That transiently-empty window (16 us normally, arbitrarily long under GIL/GC/preemption) was enough for the worker's 20 ms poll to mis-fire. We now peek (read count, no drain) and only take_chunks when we actually intend to persist. 4. emit_chunks wrapped in try/except that puts drained chunks back into the shared buffer before re-raising, and finish() wraps its flush call in try/except: pass. Guarantees (a) no chunk loss on DB failure and (b) the terminal human_run row is always written even if the flush step fails. Tests added: - 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit, try/except path, and every HITL entry-point's flush-before-transaction. - 1 async execute_attempt integration test in new test_application_execute_attempt.py drives the full consumer loop through _flush_if_due (peek → take → begin/end_emit) and the final flush, verifying that every patch line added in application.py is hit. * perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected When isRunning=true the EventSource already pushes human_interaction and human_execution events in real time — the 1.5s polling loop duplicated that work, hitting the DB and re-rendering the frontend for every tick. Disable polling entirely while the SSE stream is alive, and drop to 5s intervals only when the stream is closed (e.g. page load before the first run, or after a run finishes) so we can still discover WAITING_HUMAN requests that were created while the client was disconnected. Add isRunning to the useEffect dependency array so the polling cadence resets immediately when the SSE connection state changes. * perf(hitl): replace polling with SSE subscription and move snapshot off the write path Frontend — /conversation polling → /{run_id}/events SSE: - Replace the 5s conversation snapshot polling with a native EventSource subscription to the backend's /{run_id}/events SSE stream. Discovery is now one-shot: conversationId change and the agent stream pause (isRunning true→false), the exact moment a HITL run is most likely to exist. The SSE stream then keeps run state live with native auto-reconnect. - Add dual guards inside refresh() to absorb the thundering herd from adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all concurrent callers and returns cached state; (2) 3s minimum interval so bursts after the in-flight resolves do not immediately re-hit DB. - Use a runRef mirror so refresh() stays stable and downstream effects do not re-run on every snapshot. - Detect terminal status inside SSE onmessage and proactively es.close() to prevent EventSource from reconnecting forever against COMPLETED runs. Backend — snapshot off the write path: - Add repository.read_only() context manager: plain SELECT without WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads must not contend with worker writes on the same row lock. - Add service.light_snapshot() using read_only. Retain snapshot() as a writer-path API for any future lock-held callers. - Route conversation_snapshot, run snapshot endpoint, and both snapshot calls inside stream_run() through light_snapshot. - Move expiration to the writer path: decide() still calls _expire inline before processing each request, and expire_waiting() remains the periodic scheduler sweep. Impact: conversation snapshot calls drop from 12+/min (polling) or 10+/s (burst from adapter + SSE) to at most one every 3s. Each call is now two plain SELECTs instead of a lock-held transaction with a possible write from _expire. Read and write paths are fully decoupled. * fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning stayed true — the stop button remained visible and new messages went into the queue buffer instead of being sent normally. The root cause is that isRunning is driven entirely by the ChatModelRun generator lifetime, which only returns when the backend SSE HTTP connection closes (reader.read() -> done=true). The backend stream_run loop can hang on heartbeat even after the run is terminal when the SSE was opened during WAITING_HUMAN with attempt_active=true: the break condition requires both cursor >= event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false and empty rows). If continueHitl fires mid-flight with a stale after_event, the cursor never catches up, so the SSE stays alive forever and the generator never returns. Stop depending on the backend closing first. Inside the adapter's SSE chunk loop, detect a terminal human_run event (status in COMPLETED, FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner for-loop, and let the outer while-loop exit via the same flag on the next iteration. Assistant-UI sees the generator return and flips isRunning false immediately. Only affects HITL streams — the normal non-HITL agent path never emits human_run events so this branch is never taken. * test(hitl): update mock from snapshot to light_snapshot after read-path refactor test_human_interaction_app.py still mocked service.snapshot after commit 288ae4e69 moved conversation_snapshot and the run snapshot endpoint to service.light_snapshot (read-only path, no lock, no _expire). The fixture return_value and the two assert_called_once_with/assert_not_called assertions all referenced the old method name, causing CI to fail because MagicMock.snapshot was never called. * test(hitl): raise diff coverage above the 90% merge gate Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass. * fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare import httpx2 causes ImportError in CI and on systems with recent openai versions. This restores the try/except fallback introduced in PR #3948 commit 9521b934 and later accidentally reverted by commit d086da259. * Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm" This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61. * 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963) * Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962) * Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in. * chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> * [codex] fix(agent): silently retry transient model errors (#3965) * fix(agent): retry transient model failures silently * fix(model): support OpenAI httpx2 timeout client * test(model): add deterministic OpenAI-compatible mock * fix(agent): keep stream runtime within line budget * Fix: AIDP knowledge base bug fix (#3967) * Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection * refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change --------- Co-authored-by: cj2026-bit <647646783@qq.com> Co-authored-by: lijiayang619 <1170349871@qq.com> Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> Co-authored-by: bernard1234 <840646206@qq.com> Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com> Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com> Co-authored-by: gs-aion <gs597153711@qq.com> * [fix] enforce explicit CodeAgent termination and silent recovery (#3969) * fix(agent): enforce explicit CodeAgent termination * fix(test): restore CodeAgent CI compatibility * merge(main): merge v2.6.1 hotfixes from hotfix/v2.6.1 (#3983) * 🐛 Fix(evaluation): run trials in runtime service (#3954) * fix(evaluation): run trials in runtime service Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * test(evaluation): stub config thread manager Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * test(evaluation): stub runtime jwt helper * test(evaluation): cover trial proxy error paths * Fix/override delete (#3958) * Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear) * Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key * Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation * Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names) * Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it * Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots) * Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)" This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06. * Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves. --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> * cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959) * Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948) * fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests Root cause: the worker thread writes human_interaction events synchronously via SQLAlchemy in ask_user, while model_output_thinking/parse observer messages flow through the async consumer and are flushed only on a batched threshold (32 chunks or 250ms). When the worker suspends before that flush fires, human_interaction gains a lower event_seq number than the already-buffered observer chunks, causing the SSE replay stream to show them in the wrong order. Fix: replace the plain async-for consumer loop with a manual asyncio.wait iterator using a 50ms timeout. Once the worker finishes producing model output (i.e. right before ask_user), the loop times out and flushes any buffered observer chunks to the DB first, guaranteeing they precede the subsequent human_interaction row. Empty queue idle periods are essentially zero-cost; overall DB write frequency stays on par with the original. * fix(hitl): guarantee observer chunks precede human_interaction in DB event order When the agent invokes ask_user, two independent write paths caused the human_interaction row to be persisted BEFORE model_output_thinking / parse chunks, breaking the SSE replay ordering: the worker thread writes HITL events synchronously via SQLAlchemy, while observer messages flow through the async consumer which only flushes on a batched threshold. Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort (port.add_chunk / port.take_chunks). The async consumer pushes every processed chunk there; the worker thread calls flush_chunks_until_idle() before dispatching any HITL event — it polls the shared buffer and waits for the async loop to drain the observer queue (20ms idle window, 500ms max wait), then persists every chunk in its own transaction. This guarantees chunk event_seq < human_interaction event_seq regardless of async scheduling latency. Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel now falls back from httpx2 to httpx at import time. * test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer and the flush_chunks_until_idle poll loop that guarantees observer chunks precede human_interaction events in DB order. All five HITL entry points (dispatch / boundary / receipt / finish / _wait_until_ready) are verified to invoke the idle flush before opening their transaction. Also add two tests for the openai_llm httpx2 → httpx ImportError fallback introduced to support openai >= 2.50 where the httpx2 shim was removed: one covers the fallback path, one confirms httpx2 still wins when present. * fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety Fix 4 real issues flagged by github-code-review: 1. Non-resettable hard_deadline in flush_chunks_until_idle — previously reset on every drain, meaning a model that kept producing chunks could stall the worker forever. deadline is now computed once at entry and the sleep call clips to hard_deadline - now. 2. _emit_in_flight Event bridges the async emit path and the worker's idle poll. Without this, buffer-empty = 'persisted' was confused with buffer-empty = 'taken for emit but still in run_blocking queue'. The worker now checks both 'buffer empty for settle_ms' AND 'no emit in flight' before deciding the async side is truly idle. 3. peek_chunks() replaces the take-put-back pattern in _flush_if_due. Previously the async loop drained the buffer, decided it was not yet due, then put everything back. That transiently-empty window (16 us normally, arbitrarily long under GIL/GC/preemption) was enough for the worker's 20 ms poll to mis-fire. We now peek (read count, no drain) and only take_chunks when we actually intend to persist. 4. emit_chunks wrapped in try/except that puts drained chunks back into the shared buffer before re-raising, and finish() wraps its flush call in try/except: pass. Guarantees (a) no chunk loss on DB failure and (b) the terminal human_run row is always written even if the flush step fails. Tests added: - 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit, try/except path, and every HITL entry-point's flush-before-transaction. - 1 async execute_attempt integration test in new test_application_execute_attempt.py drives the full consumer loop through _flush_if_due (peek → take → begin/end_emit) and the final flush, verifying that every patch line added in application.py is hit. * perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected When isRunning=true the EventSource already pushes human_interaction and human_execution events in real time — the 1.5s polling loop duplicated that work, hitting the DB and re-rendering the frontend for every tick. Disable polling entirely while the SSE stream is alive, and drop to 5s intervals only when the stream is closed (e.g. page load before the first run, or after a run finishes) so we can still discover WAITING_HUMAN requests that were created while the client was disconnected. Add isRunning to the useEffect dependency array so the polling cadence resets immediately when the SSE connection state changes. * perf(hitl): replace polling with SSE subscription and move snapshot off the write path Frontend — /conversation polling → /{run_id}/events SSE: - Replace the 5s conversation snapshot polling with a native EventSource subscription to the backend's /{run_id}/events SSE stream. Discovery is now one-shot: conversationId change and the agent stream pause (isRunning true→false), the exact moment a HITL run is most likely to exist. The SSE stream then keeps run state live with native auto-reconnect. - Add dual guards inside refresh() to absorb the thundering herd from adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all concurrent callers and returns cached state; (2) 3s minimum interval so bursts after the in-flight resolves do not immediately re-hit DB. - Use a runRef mirror so refresh() stays stable and downstream effects do not re-run on every snapshot. - Detect terminal status inside SSE onmessage and proactively es.close() to prevent EventSource from reconnecting forever against COMPLETED runs. Backend — snapshot off the write path: - Add repository.read_only() context manager: plain SELECT without WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads must not contend with worker writes on the same row lock. - Add service.light_snapshot() using read_only. Retain snapshot() as a writer-path API for any future lock-held callers. - Route conversation_snapshot, run snapshot endpoint, and both snapshot calls inside stream_run() through light_snapshot. - Move expiration to the writer path: decide() still calls _expire inline before processing each request, and expire_waiting() remains the periodic scheduler sweep. Impact: conversation snapshot calls drop from 12+/min (polling) or 10+/s (burst from adapter + SSE) to at most one every 3s. Each call is now two plain SELECTs instead of a lock-held transaction with a possible write from _expire. Read and write paths are fully decoupled. * fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning stayed true — the stop button remained visible and new messages went into the queue buffer instead of being sent normally. The root cause is that isRunning is driven entirely by the ChatModelRun generator lifetime, which only returns when the backend SSE HTTP connection closes (reader.read() -> done=true). The backend stream_run loop can hang on heartbeat even after the run is terminal when the SSE was opened during WAITING_HUMAN with attempt_active=true: the break condition requires both cursor >= event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false and empty rows). If continueHitl fires mid-flight with a stale after_event, the cursor never catches up, so the SSE stays alive forever and the generator never returns. Stop depending on the backend closing first. Inside the adapter's SSE chunk loop, detect a terminal human_run event (status in COMPLETED, FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner for-loop, and let the outer while-loop exit via the same flag on the next iteration. Assistant-UI sees the generator return and flips isRunning false immediately. Only affects HITL streams — the normal non-HITL agent path never emits human_run events so this branch is never taken. * test(hitl): update mock from snapshot to light_snapshot after read-path refactor test_human_interaction_app.py still mocked service.snapshot after commit 288ae4e69 moved conversation_snapshot and the run snapshot endpoint to service.light_snapshot (read-only path, no lock, no _expire). The fixture return_value and the two assert_called_once_with/assert_not_called assertions all referenced the old method name, causing CI to fail because MagicMock.snapshot was never called. * test(hitl): raise diff coverage above the 90% merge gate Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass. * fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare import httpx2 causes ImportError in CI and on systems with recent openai versions. This restores the try/except fallback introduced in PR #3948 commit 9521b934 and later accidentally reverted by commit d086da259. * Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm" This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61. * 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963) * Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962) * Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in. * chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> * [codex] fix(agent): silently retry transient model errors (#3965) * fix(agent): retry transient model failures silently * fix(model): support OpenAI httpx2 timeout client * test(model): add deterministic OpenAI-compatible mock * fix(agent): keep stream runtime within line budget * Fix: AIDP knowledge base bug fix (#3967) * Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection * refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change * [fix] enforce explicit CodeAgent termination and silent recovery (#3969) * fix(agent): enforce explicit CodeAgent termination * fix(test): restore CodeAgent CI compatibility * cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981) * Fix StopAsyncIteration leak in execute_attempt finally block Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors. Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate. * Fix HITL form not appearing until page refresh Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload. Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh. * Fix SSE chunk/HITL event ordering race under real server load Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order. Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order. * fix(hitl): recover pending form when SSE delivery stalls silently Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones. Changes: - Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls - Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it - Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop * fix(hitl): stop parked human-input waits from consuming scheduler slots Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed. Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments. Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots. Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43. * test(scheduler): fix flaky waiting-concurrency assertion on fast event loops Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently. Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window. * Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980) The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which also carried two unrelated upstream changes that this release branch never had: - #3909 Knowledge base interface optimization (AIDP UI refactor) - #3930 support AIDP knowledge file deletion and download Both are removed here so the release line keeps only the bug fix. - Restore the AIDP frontend components to their pre-refactor layout and drop the helper modules only the refactor used: AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils. - Drop the #3930 document remove/download endpoints from services/api.ts and the AIDP translations that only those screens referenced. - Keep the fix itself unchanged: knowledge-base scoped Channels and KnowledgeFiles/History paths, all-status document listing, keyword search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the upload-triggered polling. Verified: - pytest test/ext_components/aidp -q -> 583 passed - frontend `npm run type-check` (tsc --noEmit) -> no errors --------- Co-authored-by: cj2026-bit <647646783@qq.com> Co-authored-by: lijiayang619 <1170349871@qq.com> Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> Co-authored-by: bernard1234 <840646206@qq.com> Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com> Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com> Co-authored-by: gs-aion <gs597153711@qq.com> Co-authored-by: chase <byzhangxin11@126.com> * merge(main): merge v2.6.1 hotfixes from hotfix/v2.6.1 (#3993) * 🐛 Fix(evaluation): run trials in runtime service (#3954) * fix(evaluation): run trials in runtime service Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * test(evaluation): stub config thread manager Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5 * test(evaluation): stub runtime jwt helper * test(evaluation): cover trial proxy error paths * Fix/override delete (#3958) * Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear) * Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key * Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation * Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names) * Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it * Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots) * Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)" This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06. * Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves. --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> * cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959) * Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948) * fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests Root cause: the worker thread writes human_interaction events synchronously via SQLAlchemy in ask_user, while model_output_thinking/parse observer messages flow through the async consumer and are flushed only on a batched threshold (32 chunks or 250ms). When the worker suspends before that flush fires, human_interaction gains a lower event_seq number than the already-buffered observer chunks, causing the SSE replay stream to show them in the wrong order. Fix: replace the plain async-for consumer loop with a manual asyncio.wait iterator using a 50ms timeout. Once the worker finishes producing model output (i.e. right before ask_user), the loop times out and flushes any buffered observer chunks to the DB first, guaranteeing they precede the subsequent human_interaction row. Empty queue idle periods are essentially zero-cost; overall DB write frequency stays on par with the original. * fix(hitl): guarantee observer chunks precede human_interaction in DB event order When the agent invokes ask_user, two independent write paths caused the human_interaction row to be persisted BEFORE model_output_thinking / parse chunks, breaking the SSE replay ordering: the worker thread writes HITL events synchronously via SQLAlchemy, while observer messages flow through the async consumer which only flushes on a batched threshold. Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort (port.add_chunk / port.take_chunks). The async consumer pushes every processed chunk there; the worker thread calls flush_chunks_until_idle() before dispatching any HITL event — it polls the shared buffer and waits for the async loop to drain the observer queue (20ms idle window, 500ms max wait), then persists every chunk in its own transaction. This guarantees chunk event_seq < human_interaction event_seq regardless of async scheduling latency. Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel now falls back from httpx2 to httpx at import time. * test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer and the flush_chunks_until_idle poll loop that guarantees observer chunks precede human_interaction events in DB order. All five HITL entry points (dispatch / boundary / receipt / finish / _wait_until_ready) are verified to invoke the idle flush before opening their transaction. Also add two tests for the openai_llm httpx2 → httpx ImportError fallback introduced to support openai >= 2.50 where the httpx2 shim was removed: one covers the fallback path, one confirms httpx2 still wins when present. * fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety Fix 4 real issues flagged by github-code-review: 1. Non-resettable hard_deadline in flush_chunks_until_idle — previously reset on every drain, meaning a model that kept producing chunks could stall the worker forever. deadline is now computed once at entry and the sleep call clips to hard_deadline - now. 2. _emit_in_flight Event bridges the async emit path and the worker's idle poll. Without this, buffer-empty = 'persisted' was confused with buffer-empty = 'taken for emit but still in run_blocking queue'. The worker now checks both 'buffer empty for settle_ms' AND 'no emit in flight' before deciding the async side is truly idle. 3. peek_chunks() replaces the take-put-back pattern in _flush_if_due. Previously the async loop drained the buffer, decided it was not yet due, then put everything back. That transiently-empty window (16 us normally, arbitrarily long under GIL/GC/preemption) was enough for the worker's 20 ms poll to mis-fire. We now peek (read count, no drain) and only take_chunks when we actually intend to persist. 4. emit_chunks wrapped in try/except that puts drained chunks back into the shared buffer before re-raising, and finish() wraps its flush call in try/except: pass. Guarantees (a) no chunk loss on DB failure and (b) the terminal human_run row is always written even if the flush step fails. Tests added: - 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit, try/except path, and every HITL entry-point's flush-before-transaction. - 1 async execute_attempt integration test in new test_application_execute_attempt.py drives the full consumer loop through _flush_if_due (peek → take → begin/end_emit) and the final flush, verifying that every patch line added in application.py is hit. * perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected When isRunning=true the EventSource already pushes human_interaction and human_execution events in real time — the 1.5s polling loop duplicated that work, hitting the DB and re-rendering the frontend for every tick. Disable polling entirely while the SSE stream is alive, and drop to 5s intervals only when the stream is closed (e.g. page load before the first run, or after a run finishes) so we can still discover WAITING_HUMAN requests that were created while the client was disconnected. Add isRunning to the useEffect dependency array so the polling cadence resets immediately when the SSE connection state changes. * perf(hitl): replace polling with SSE subscription and move snapshot off the write path Frontend — /conversation polling → /{run_id}/events SSE: - Replace the 5s conversation snapshot polling with a native EventSource subscription to the backend's /{run_id}/events SSE stream. Discovery is now one-shot: conversationId change and the agent stream pause (isRunning true→false), the exact moment a HITL run is most likely to exist. The SSE stream then keeps run state live with native auto-reconnect. - Add dual guards inside refresh() to absorb the thundering herd from adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all concurrent callers and returns cached state; (2) 3s minimum interval so bursts after the in-flight resolves do not immediately re-hit DB. - Use a runRef mirror so refresh() stays stable and downstream effects do not re-run on every snapshot. - Detect terminal status inside SSE onmessage and proactively es.close() to prevent EventSource from reconnecting forever against COMPLETED runs. Backend — snapshot off the write path: - Add repository.read_only() context manager: plain SELECT without WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads must not contend with worker writes on the same row lock. - Add service.light_snapshot() using read_only. Retain snapshot() as a writer-path API for any future lock-held callers. - Route conversation_snapshot, run snapshot endpoint, and both snapshot calls inside stream_run() through light_snapshot. - Move expiration to the writer path: decide() still calls _expire inline before processing each request, and expire_waiting() remains the periodic scheduler sweep. Impact: conversation snapshot calls drop from 12+/min (polling) or 10+/s (burst from adapter + SSE) to at most one every 3s. Each call is now two plain SELECTs instead of a lock-held transaction with a possible write from _expire. Read and write paths are fully decoupled. * fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning stayed true — the stop button remained visible and new messages went into the queue buffer instead of being sent normally. The root cause is that isRunning is driven entirely by the ChatModelRun generator lifetime, which only returns when the backend SSE HTTP connection closes (reader.read() -> done=true). The backend stream_run loop can hang on heartbeat even after the run is terminal when the SSE was opened during WAITING_HUMAN with attempt_active=true: the break condition requires both cursor >= event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false and empty rows). If continueHitl fires mid-flight with a stale after_event, the cursor never catches up, so the SSE stays alive forever and the generator never returns. Stop depending on the backend closing first. Inside the adapter's SSE chunk loop, detect a terminal human_run event (status in COMPLETED, FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner for-loop, and let the outer while-loop exit via the same flag on the next iteration. Assistant-UI sees the generator return and flips isRunning false immediately. Only affects HITL streams — the normal non-HITL agent path never emits human_run events so this branch is never taken. * test(hitl): update mock from snapshot to light_snapshot after read-path refactor test_human_interaction_app.py still mocked service.snapshot after commit 288ae4e69 moved conversation_snapshot and the run snapshot endpoint to service.light_snapshot (read-only path, no lock, no _expire). The fixture return_value and the two assert_called_once_with/assert_not_called assertions all referenced the old method name, causing CI to fail because MagicMock.snapshot was never called. * test(hitl): raise diff coverage above the 90% merge gate Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * style(hitl): unify comment style across HITL changes Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change. * refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass. * fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare import httpx2 causes ImportError in CI and on systems with recent openai versions. This restores the try/except fallback introduced in PR #3948 commit 9521b934 and later accidentally reverted by commit d086da259. * Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm" This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61. * 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963) * Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962) * Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in. * chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> * [codex] fix(agent): silently retry transient model errors (#3965) * fix(agent): retry transient model failures silently * fix(model): support OpenAI httpx2 timeout client * test(model): add deterministic OpenAI-compatible mock * fix(agent): keep stream runtime within line budget * Fix: AIDP knowledge base bug fix (#3967) * Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection * refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change * [fix] enforce explicit CodeAgent termination and silent recovery (#3969) * fix(agent): enforce explicit CodeAgent termination * fix(test): restore CodeAgent CI compatibility * cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981) * Fix StopAsyncIteration leak in execute_attempt finally block Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors. Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate. * Fix HITL form not appearing until page refresh Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload. Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh. * Fix SSE chunk/HITL event ordering race under real server load Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order. Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order. * fix(hitl): recover pending form when SSE delivery stalls silently Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones. Changes: - Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls - Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it - Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop * fix(hitl): stop parked human-input waits from consuming scheduler slots Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed. Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments. Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots. Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43. * test(scheduler): fix flaky waiting-concurrency assertion on fast event loops Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently. Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window. * Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980) The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which also carried two unrelated upstream changes that this release branch never had: - #3909 Knowledge base interface optimization (AIDP UI refactor) - #3930 support AIDP knowledge file deletion and download Both are removed here so the release line keeps only the bug fix. - Restore the AIDP frontend components to their pre-refactor layout and drop the helper modules only the refactor used: AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils. - Drop the #3930 document remove/download endpoints from services/api.ts and the AIDP translations that only those screens referenced. - Keep the fix itself unchanged: knowledge-base scoped Channels and KnowledgeFiles/History paths, all-status document listing, keyword search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the upload-triggered polling. Verified: - pytest test/ext_components/aidp -q -> 583 passed - frontend `npm run type-check` (tsc --noEmit) -> no errors * fix(agent): accept reasoning-prefixed code actions (#3990) * refactor: remove human interaction features and related configurations (#3988) * refactor: remove human interaction features and related configurations * refactor: remove human interaction features and related configurations --------- Co-authored-by: cj2026-bit <647646783@qq.com> Co-authored-by: lijiayang619 <1170349871@qq.com> Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)> Co-authored-by: bernard1234 <840646206@qq.com> Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com> Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com> Co-authored-by: gs-aion <gs597153711@qq.co…
…time (#4016) * Fix: keep ingested AIDP documents visible when the channel history is used The document list switched data sources instead of combining them: while the resolved channel directory was empty the list was served from the knowledge-base scoped listing, and as soon as one upload landed in that directory the all-status history took over. The history only covers the resolved channel directory, so every file ingested into another directory disappeared from the list right after an upload — reported on 2.6.1 for files that were created before the upgrade. Merge the two sources instead: - The knowledge-base scoped listing guarantees membership, the history supplies the live statuses, and a file only the history knows about (still uploading, or failed before ingestion) is kept exactly as reported. - Items are matched through every identity they expose, so a file that one payload describes with a uuid and the other with an ino number stays one row. - Files taken from the listing are marked COMPLETED, which is all that listing ever returns. The listing is read in pages of 100, capped at 20 pages, and a failing page degrades to the files read so far (logged) instead of failing the request. Verified: pytest test/ext_components/aidp -q -> 586 passed, including new coverage for ingested files outside the channel directory, a history-only upload, and a failing listing. * Fix: read AIDP file history across pages so bursts of uploads stay visible The history endpoint is paginated and sorts files that are still being processed to the front. Reading only the first page therefore drops exactly the files the document list exists to show: more simultaneous uploads than fit in one page push the rest out of view. - Page the history (body `page`, one-based): the walk stops at an empty page, at an explicit "no next page" signal, or on a page that adds nothing new. The last one keeps a build that ignores `page` from looping over the same files, and logs that everything beyond page one stays invisible. - Cap the walk at 20 pages and log when the cap truncates, so one list request cannot turn into an unbounded number of upstream calls. - Trust only unambiguous pagination signals: this endpoint family reports `total_count` as the size of the current page elsewhere, so it is not read as a grand total. - The mock now mirrors the real contract: `page` in the body, ten entries per page (tunable through POST /_mock/history-page-size), files that are still being processed listed first, and `total_count`/`next_link` in the response. A status pinned through POST /_mock/doc-status also stays PROCESSING instead of advancing to COMPLETED, so that stage can be inspected locally. Verified: pytest test/ext_components/aidp -q -> 588 passed, including a burst of uploads spanning three history pages; end-to-end against the mock -> 10/10 checks covering page order, page metadata and the last page. * Fix: show AIDP files still processing first and keep listed file metadata Order the AIDP document list so files that are still uploading or extracting sit above the finished ones, which keeps an upload's progress on page 1 instead of behind files that were already ingested. An unrecognised stage counts as open as well, so it neither stops the polling nor drops behind the ingested rows. Merge the channel history with the document listing field by field: the history still supplies the live status, but a payload that omits metadata (name, size, creation time) no longer blanks out the fields the listing already carries. * Fix: keep the creation time AIDP history reports as created_at The history endpoint reports the creation time under the canonical created_at (ISO string), but normalization only read first_upload_time / create_time and overwrote created_at with a failed conversion, which left the column empty for every history row. Both spellings are now accepted, numeric or ISO, for created_at and updated_at. * Fix: read the AIDP document creation time through every spelling reported A blank field means 'not reported' in AIDP payloads, so an empty create_time used to shadow a populated canonical created_at and the history rows lost the column. Blank values are skipped now, the listing's upload stamp and the canonical created_at are both accepted, and update_time closes the chain for files AIDP has not finished registering yet.
* 调整mcp添加mcp和编辑页面,与mcp配置页面保持一致; mcp容器化启动点击后就展示日志,配置页面和仓库页面都修改; 修复mcp服务更新后,无法刷新工具的问题; * mcp仓库增加分页; fix test case * fix test case * fix test case * Update custom.json * 修复本地镜像上传部署日志显示问题 * fix bug * fix bug * add test case * add test case * 删除mcp市场 * 优化前端页面 * Update custom.json * Normalize custom locale file endings * Fix MCP registry endpoint typing * Remove obsolete MCP registry frontend code * 修改本地镜像上传MCP错误提示 * 修复mcp提交审核后不及时显示问题 * 修复api转mcp服务切换租户添加会导致原有工具列表被清空的问题 * 修复API转MCP后工具列表未同步 * 恢复工具手动刷新逻辑 * 避免API转MCP工具因扫描失败被隐藏 * docs(agents): define bundled official agent design Document industry-specific official agent packages, reserved platform templates, startup synchronization, repository visibility, and tenant copy permissions. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs(agents): add bundled agent implementation plan Break the approved industry-specific official agent design into test-first backend, repository, frontend, and deployment tasks. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs(agents): add Chinese design and plan Provide Chinese versions of the approved bundled official agent design and implementation plan. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(agents): add official bundle sync foundation Support multi-profile official agent bundles with safe ZIP loading and idempotent synchronization into the reserved platform repository. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(agents): expose official listings in repository Merge selected official listings into repository reads, keep platform templates read-only through tenant scoping, and trigger synchronization at startup or through the deployment script. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(agents): mark official listings in clients Show official repository entries as read-only templates in the UI and inject selected bundle directories into Docker and Kubernetes deployments. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(deploy): select official agent profiles offline Allow offline package builds to retain only explicitly selected official agent profiles and emit the runtime official-agent configuration. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agents): isolate bundle skill validation Resolve official bundle skill entries with a local model so repository service test doubles and runtime imports cannot corrupt the Pydantic forward reference. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(deploy): sync official agents during installation Run official agent synchronization after Docker and Kubernetes deployments become ready, while retaining startup synchronization as an idempotent fallback. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(deploy): let users select agent categories Fetch the official agent catalog during installation, scan its category directories, and persist the user's multi-category selection for Docker and Kubernetes deployment. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(agents): install official repository dependencies Route official repository imports through tenant-level MCP, Skill, and knowledge-base preparation while preserving ordinary imports. Add ZIP and directory bundle loading and retain the current Agent service with only the required skill-link helpers. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agents): resolve bundles by repository agent name Use the official repository name as the bundle directory key and restore the normal listing content value. Keep bundle lookup aligned with the contract that the bundle folder and root Agent name are identical. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(deploy): separate official agent deployment Move official Agent resource selection and synchronization out of the main Nexent installation flow. Add a standalone Docker/Kubernetes/local deployment entry, keep official repository visibility independent of startup profiles, and preserve tenant-level installation for user copies. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(deploy): mark official agent entry executable Allow the standalone official Agent deployment entry to run directly in Git-based deployment environments. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(deploy): accept host paths for official agents Copy local official Agent resources through the running config container and explicitly synchronize the mounted container directory. Remove the conflicting read-only submount so Windows host paths can be used after Nexent is deployed. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(deploy): preserve container paths in git bash Disable MSYS path conversion for Docker and Kubernetes sync arguments so mounted Linux paths are not rewritten into Windows paths. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(repository): resolve official listings during import Allow repository prechecks and imports to fall back to the reserved official tenant so shared official listings can be copied by business tenants. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(repository): report official import requirements Build official repository precheck items from the mounted bundle so copy dialogs show model, knowledge base, MCP, Skill, and tool availability instead of an empty list. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agents): keep official bundle keys stable Store the Bundle directory key in official repository listings and repair existing official records during synchronization so prechecks and imports reload the correct mounted Bundle. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agents): resolve bundles inside profiles Allow official installation and precheck to locate directory, JSON, and ZIP bundles nested under selected profile directories. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agents): ignore skill archives during bundle scan Keep Skill ZIP payloads from being misclassified as standalone official Agent bundles during profile synchronization. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agents): diagnose official bundle lookup Log the resolved official bundle path and validation target so runtime lookup failures can be distinguished from missing mounted resources. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(deploy): pin official bundle path in container Override host-local OFFICIAL_AGENTS_PATH values from backend/.env with the mounted container path so repository prechecks can resolve official bundles. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(repository): choose tenant models during official copy Show tenant-configured language and embedding models in the repository copy dialog and pass the selected IDs through the official install path. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(repository): preserve official flag for model selection Expose the publisher tenant in repository summaries and use it as a fallback so official copy dialogs always show tenant model selectors. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(repository): preserve publisher on summary listings Restore the publisher tenant on lightweight repository records so official listings reach the copy dialog with the official flag. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agents): apply official skill resolutions Honor the copy dialog choice to reuse an existing tenant Skill or create a renamed Skill when installing an official agent. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): ensure dependencies before reuse Prepare official knowledge bases and other dependencies even when the root agent already exists, so partial installs can be repaired instead of being skipped. Add runtime logs for official repository routing and bundle dependency counts. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agents): honor skill conflict resolutions Evaluate official Skill reuse or rename choices before raising duplicate errors, so dependency preparation can continue to knowledge base creation and Agent import. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): repair reused knowledge references Update existing official Agent tool instances after resolving a tenant knowledge base, so earlier partial installs no longer keep the bundle-local kb key. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): isolate user agent copies Treat official dependencies as tenant-scoped while keeping copied Agents user-scoped. Create a uniquely named visible copy when another user already owns the bundle's root Agent. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): share tenant knowledge bases Make official knowledge bases visible to every group in the tenant while keeping them read-only, and show bundle knowledge base display names instead of logical placeholders during import precheck. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(official-agents): label copies by email Use the current account email in copied official Agent and Skill display names while retaining private technical identifiers for tenant uniqueness. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): hide official publisher label Hide the publisher text below the official badge for official Agent listings while preserving author display for regular listings. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): reuse existing official embeddings Hide embedding model selection when all official knowledge bases already exist in the tenant and let the backend reuse their configured models. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): block official copy without embeddings Show a clear tenant embedding-model requirement and prevent official knowledge-base installation from starting when no usable embedding model is configured. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): validate official kb embedding model Treat an official knowledge base as reusable only when its tenant embedding model still exists and is available, while leaving ordinary repository imports unchanged. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(deploy): remove official profile env setting Keep profile selection in the standalone official-agent deployment flow instead of the main environment template. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(deploy): remove legacy official profile variable Keep official agent profile selection exclusively in the standalone deployment script. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(k8s): remove unused market backend config Remove the obsolete external market backend address from Kubernetes deployment values and rendering. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(docs): remove bundled agent plans from code repo Keep bundled default agent design documents in the centralized nexent-doc archive instead of the business repository. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-import): use migrated agent management service Replace stale services.agent_service imports in the official agent installer with the migrated agent management entry points so repository imports can complete after the agent service refactor. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-import): use migrated knowledge base service Replace the removed vectordatabase_service import in official agent knowledge-base installation with the migrated knowledge-base and model management services. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): refresh mine agents after copy Switch to the mine tab and refresh its query after a repository copy succeeds so the created or reused Agent is immediately visible. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): label copied agents with user email Append the current user's email to official Agent display names for both new installs and idempotent retries. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-import): import model status enum Allow repository import precheck to evaluate tenant model availability without raising a NameError. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): create unique copies on name conflicts Match ordinary repository imports by always creating official Agent copies and generating tenant-scoped unique names while preserving tenant dependency reuse. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(official-agents): add default marketplace tags Populate category tags in official agent bundle metadata so synchronized repository listings and local marketplace submissions have default labels. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): preserve default tags on import Apply official Bundle category tags to the newly imported tenant Agent after resolving tenant-specific tag values. Keep ordinary repository imports unchanged. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): remove forced default tags Keep official Agent tags empty so users choose from the tenant tag definitions when applying for listing. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(tags): allow tenant users to read tag definitions Keep tag library mutations restricted to management roles while allowing authenticated tenant users to load tag choices for repository listing forms. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): render tag limit in hint Pass the tag limit to the localized hint so the UI renders the numeric value instead of the raw interpolation placeholder. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: align regression tests with migrated services Update stale mocks, assertions, and service paths after the management-service migration. Cover official agent knowledge-base reuse, skill resolution, MCP headers, tag reads, and asynchronous tool refresh. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs: add official agent deployment guide Document online, local, offline, Docker, and Kubernetes deployment of official Agent bundles, including profile selection, synchronization, verification, and troubleshooting. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs: document agent export bundle conversion Add the conversion workflow from Nexent-exported Agent JSON or ZIP to a profile-scoped official bundle, including Skill and knowledge-base seed handling. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: mock tenant MCP endpoint in tool tests Patch the tenant-scoped local MCP endpoint used by get_all_mcp_tools instead of the obsolete base URL mock. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs: remove official agent deployment guide Withdraw the deployment guide from the feature branch as requested. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: cover official agent synchronization Add unit coverage for source-agent materialization, repository upsert, and failure isolation during official bundle synchronization. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: cover official repository import flow Exercise official listing fallback lookup, bundle precheck, installation outcomes, and download accounting branches. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: cover official agent bundle loading Exercise profile validation, safe archive handling, skill and knowledge-base loading, and directory bundle synchronization paths. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: stub model parameter filter for official agent tests Keep the isolated consts.model test module compatible with the naming service import chain used in CI. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: cover official agent service edge cases Add coverage for bundle discovery, nested archives, resource loading failures, partial KB repair, and existing-agent installation paths. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: cover tenant MCP routing branches Exercise tenant refresh skips, app construction fallbacks, router authorization and lazy initialization, and MCP startup cleanup paths. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: cover repository import precheck branches Test embedding model availability and official knowledge-base name fallback during repository import prechecks. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test: cover official agent bundle defaults Verify official bundle card fields derive from the root agent and fall back to safe defaults when metadata is absent. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(mcp): avoid duplicate tool cleanup on delete Clean up MCP tools exactly once during deletion and use the shared outer-apis name for OpenAPI services. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(deploy): remove obsolete agent TUI branches Align common.sh with develop after official agent deployment was separated from the main installation flow. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(official-agents): keep local snapshots out of remote Stop tracking the two local official agent snapshots while preserving the files for local deployment and testing. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(agents): remove unused attachment fallback Revert the unrelated MinIO bucket fallback that is not used by official agent installation. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(deploy): remove redundant official agent directory setup Keep official agent directory creation in the standalone deployment flow and restore deploy.sh to develop behavior. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(deploy): remove obsolete official agent env example Restore the environment example after official agent deployment moved to the standalone flow. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(deploy): remove official agent helm wiring Restore Kubernetes deployment templates to develop and keep official agent installation outside the main Nexent deployment flow. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(offline): remove obsolete official agent packaging Restore offline package generation to develop and remove the unused official-agent profile filtering path. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(agent): manage official agent visibility Add tenant-scoped visibility overrides and super-admin management APIs for official agent templates. Provide resource-management UI with visibility switch and global template deletion while preserving tenant copies. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * refactor(agent): remove tenant visibility controls Keep official-agent management global and remove the tenant visibility table, migration, switch, and filtering. Super admins can still delete official templates while tenant copies and dependencies remain untouched. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): restrict official agent management to super admins Use the authenticated global role for official agent management APIs and hide the management control from non-super-admin users. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): confirm official agent deletion in modal Replace the inline confirmation with a descriptive modal before deleting the official template and bundle files. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): resolve official knowledge base conflicts Allow official agent imports to reuse or create a tenant knowledge base when names collide, and pass the selected resolution through the repository copy flow. Keep ordinary repository imports on the existing index-name path and add regression coverage. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): separate knowledge base resolution panel Render official knowledge base reuse or create-new choices as a sibling panel to skill conflicts and model selection. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): align knowledge base conflict styling Match the official knowledge base conflict panel to the skill conflict layout with white cards and separate titles and options. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): show accurate dependency status Prefer repository requirement reason codes over generic activation labels so existing knowledge bases with unavailable embedding models are not shown as unactivated. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): show knowledge base copy name Display the expected copy name beside the create-new knowledge base option, matching the skill conflict experience. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): remove copy name parentheses Display the knowledge base copy name without surrounding parentheses in the repository import dialog. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): treat duplicate skills as resolvable Keep duplicate skills in the dedicated reuse-or-rename panel instead of also classifying them as unavailable dependencies. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): show duplicate knowledge bases as conflicts Keep same-name knowledge bases in the abnormal dependency section with a conflict label, matching duplicate skill behavior while preserving reuse-or-create choices. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): avoid duplicate knowledge base status Exclude knowledge bases requiring conflict resolution from the available dependency section so they appear only in the conflict section. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): align official import arguments Use the authenticated tenant and user context when importing official agents instead of passing unsupported arguments to import_agent_impl. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(deploy): support Unicode official agent profiles Allow official agent profile directory names such as Chinese names while retaining path traversal protection. Add coverage for Unicode profile parsing. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(deploy): support nested official agent profiles Add a profile-root option so Hub layouts such as AgentsHub/行业智能体/金融 can be selected without treating the nested directory as a Git repository URL. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(deploy): limit Hub checkout to selected profiles Skip whole-repository Git LFS checkout and sparsely materialize only the selected profile paths. This prevents unrelated unavailable LFS objects from blocking official agent deployment. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(deploy): materialize sparse official agent checkout Update the worktree explicitly after configuring sparse checkout for no-checkout Hub clones, so nested Unicode profiles are available to the synchronization step. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-sync): avoid logging expected missing agents Use a non-raising agent lookup during official bundle synchronization so first-time provisioning does not emit database error logs. Preserve the existing raising lookup for callers that require strict lookup semantics and add coverage for both paths. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(deploy): preserve official profile directories Copy selected profile contents into an explicit profile directory in the target mount so synchronization can resolve paths such as /mnt/nexent/official-agents/金融. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): hide official take-down action from admins Keep official agent management in the super-admin resource page and prevent tenant administrators from seeing the ordinary repository take-down action on official listings. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): preserve container paths in sync script Clean selected profiles before copying and disable MSYS path conversion so Windows Git Bash does not leave stale nested bundles behind. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): convert Git Bash copy sources Keep container paths protected from MSYS conversion while converting temporary local sources to Docker-readable Windows paths. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): publish Nexent as repository author Use Nexent as the marketplace author for official bundles without changing the source agent snapshot metadata. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): restore official listing badge Render the official badge in the unified agent repository card while preserving the existing version badge and listing behavior. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): hide activation for name conflicts Keep skill and knowledge base name conflicts in the resolution flow instead of offering an activation action. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): prioritize dependency name conflicts Detect skill duplicates by their precheck reason and show knowledge base conflicts before availability errors, so conflicting resources do not expose activation actions. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-repository): show conflict reason without activation Use the conflict-prioritized reason label in the non-action status branch so duplicate knowledge bases are not shown as unavailable. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test(agent-repository): update knowledge resolution coverage Align repository and official-agent tests with knowledge-base resolution payloads and the embedding availability checks introduced by official imports. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * test(agent-repository): cover official management branches Add coverage for repository fallback errors, official listing management, bundle deletion, and invalid official bundle paths. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * chore(official-agents): exclude document writing snapshot Remove the document writing official agent snapshot from version control while preserving the local deployment asset. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs(deployment): add official agent deployment guide Document independent official-agent deployment from Agent Hub, local directories, and offline archives, including profile selection, verification, and troubleshooting. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs(deployment): document official agent deletion Refine the official-agent guide to match the Chinese documentation style and explain super-admin template deletion, preserved tenant copies, and post-deletion verification. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs(deployment): align official agent page style Remove the page navigation block so the official-agent deployment guide follows the existing deployment documentation style. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs(deployment): remove page metadata Keep the official-agent deployment guide consistent with deployment documents that use a plain Markdown heading. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs(config): document official agent variables Add the official-agent resource path and optional profile fallback to the environment template while keeping profile selection in the independent deployment command. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * refactor(mcp): move tenant URL builder to utils Keep MCP configuration constants in consts while moving tenant-scoped URL construction into a utility module. Update all production callers and the focused test to use the new location. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * refactor(agent): unify official identity under system IDs Use one system tenant and user identity for official agent synchronization and repository visibility. Keep deprecated aliases for compatibility while updating frontend detection and tests. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): use system identity for official agents Replace the separate official tenant and user identifiers with the shared system identity and update all related repository, sync, frontend, and test references. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * docs(deployment): reorder official agent sections Place profile redeployment guidance before troubleshooting and keep the section numbering sequential. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * refactor(agent): sync official bundles through local API Replace the backend script entrypoint with a loopback-only synchronization API and make Docker and Kubernetes deployment scripts call it from inside the config container. Keep synchronization business logic in the service layer. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent): avoid MSYS path conversion in sync request Use the service default official-agent directory instead of passing the container path through docker exec curl, which MSYS rewrites on Windows Git Bash. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-sync): prevent user-controlled bundle paths Remove the base_dir query parameter from the internal synchronization API so bundle loading always uses the configured container path. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-sync): resolve profiles from mounted directories Avoid concatenating the selected profile directly into a filesystem path and cover Unicode profile directory resolution. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(agent-sync): scan configured root before filtering profiles Avoid recursively traversing a directory derived from user-selected profile input by scanning the configured bundle root once and filtering descendants by their top-level profile. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * feat(deploy): simplify bundled official agent installation Read official bundles from the repository deploy directory and let operators select profiles interactively, while updating the documentation site navigation and deployment guide. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(official-agents): preserve knowledge base display names Keep official knowledge base declarations when synchronizing repository snapshots so logical folder names remain internal references and user-facing display names are retained. Normalize single-object knowledge_base metadata and cover the regression with bundle and sync tests. Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * feat(deploy): bundle official agents in main image Include official Agent seed resources in the main image without auto-synchronizing them at startup. Let the deployment script fall back to the image-bundled resources when the host asset directory is unavailable. Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * fix(official-agents): refresh repository metadata on resync Refresh official repository card fields and snapshots even when the bundle version is unchanged, so edits to agent.json become visible after repeated deployment. Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * Revert "feat(deploy): bundle official agents in main image" This reverts commit 588a3ce. * docs(installation): integrate official agent deployment Move official Agent deployment instructions into the Chinese and English installation guides and update overview links so the deployment documentation has a single entry point. Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex * feat(official-agents): add bundled official agent assets Track the official Agent bundles and their document, Skill, and knowledge-base resources in the Nexent repository for deployment-time installation. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix(repository): align app merge with develop Keep the official-agent and repository-icon endpoints while matching develop import ordering and repository payload serialization behavior. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * revert(docs): remove obsolete official agent links Remove the official agent deployment navigation and overview links because the referenced documentation page is no longer present. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * merge: adopt latest develop agent repository UI Use the post-PR3899 develop versions of the Agent page, repository listing modal, and repository API tests so subsequent develop changes are not reverted. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * fix * test(official-agents): cover knowledge base declaration branches Cover single-object knowledge base declarations and reject invalid metadata types in official bundle parsing. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex * refactor(deploy): move official agents under deploy root Relocate official Agent bundles to deploy/official-agents and update the deployment script and installation documentation to use the new repository path. Co-authored-by: Codex <noreply@openai.com> Generated-by: gpt-5-codex
* update version style * 删除右下角“联系我们”,优化超级管理员界面
…上限;(3)应用评测:单个评测集大小上限 (#4025) * feat: enforce MCP resource and request timeout limits * test: load MCP error helpers in planning test shim * test: improve MCP timeout coverage * fix: distinguish generic and MCP timeouts * fix: adapt MCP timeout handling to latest managed runtime * test: cover managed MCP request timeouts * config: make MCP limits and timeout configurable * refactor: group MCP environment settings * refactor: address MCP resource limit review feedback * feat: enforce Skill tenant and upload limits * test: update Skill quota database mocks * test: cover Skill limit propagation branches * feat: limit evaluation set Excel upload size * test: provide tenant limit exception mock for skill db * ci: rerun PR checks * fix: limit MCP timeout to connection establishment * fix: resolve Sonar reliability findings
…nd default directly to the Agent Workbench interface. (#4036) * ♻️ Refactor: When launching the Agent Workbench, skip the home page and default directly to the Agent Workbench interface. * ♻️ Refactor: When launching the Agent Workbench, skip the home page and default directly to the Agent Workbench interface.
* fix(frontend): move skill listing action to more menu * fix(frontend): clarify skill listing action label
* update version style * 删除右下角“联系我们”,优化超级管理员界面 * fix(frontend): keep agent conversations visible after returning * fix: show repository status in paged agent list * fix(frontend): move skill listing action to more menu * fix: restrict agent repository review data in paged list * fix(frontend): return from chat when thread reload fails --------- Co-authored-by: Summer-Si <mingmingsu22@gmail.com>
… in the Agent Workbench; fixed the issue where agents were not fully displayed on the agent configuration page. (#4040)
* update version style * 删除右下角“联系我们”,优化超级管理员界面 * fix(frontend): keep agent conversations visible after returning * fix: show repository status in paged agent list * fix(frontend): move skill listing action to more menu * fix: restrict agent repository review data in paged list * fix(frontend): return from chat when thread reload fails * fix(frontend): handle agent name overflow and card tag layout --------- Co-authored-by: Summer-Si <mingmingsu22@gmail.com> Co-authored-by: panyehong <2655992392@qq.com>
#4008 closed a horizontal-privilege hole on /model/manage/* by restricting the endpoints to the SU role. The whitelist did not distinguish a cross-tenant call from one naming the caller's own tenant, so ADMIN users lost the whole Models tab on /resource-manage: manage/list returned 403 and the page rendered an empty table with no error, because the create, update, delete, healthcheck and provider endpoints share the same guard. The page always sends the caller's own tenant_id (UserManageComp falls back to user.tenantId for non-SU), so rejecting ADMIN blocked no cross-tenant access -- it only broke tenant admins managing their own models. Replace the role whitelist with a role + tenant scope check: SU may target any tenant, ADMIN only the tenant its token belongs to, and every other role is rejected as before. ADMIN and SU share the same model:* permission seeds, so the role itself still has to be checked -- permissions alone cannot separate them. The existing ADMIN cross-tenant test kept passing because it used a foreign tenant_id, which is why the regression went uncaught. Add own-tenant coverage for list and update, plus a DEV case proving require(model:read) is not sufficient on its own to reach the manage surface. Co-authored-by: jeffwu <meetjeff.wu@gmail.com> Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
jeffwu-1999
requested review from
Dallas98,
Jasonxia007,
WMC001,
YehongPan and
hhhhsc701
as code owners
September 29, 2026 13:36
…on for unauthorized users (#4042) * 🐛 Fix(evaluation): surface run delete errors and hide button for unauthorized users * ♻️ Refactor(evaluation): extract shared error-toast helper to fix Sonar duplication
* fix(model): let backfill swap auto-picked defaults for larger-context models The default-model backfill runs after EVERY model creation, and a slot it fills is treated as final. Batch adds create models one by one, so the first-created model permanently occupied the slot before better candidates landed -- the "available first, then larger context window" ranking never got to compare across the batch. Observed live: a 5-model batch import left a 256K-context model as the default LLM while two 1M-context models arrived right after it. Distinguish user choices from backfill placeholders via the config row's user_id: the UI save path (set_single_config) stamps the acting user on rows it writes, backfill-inserted rows leave it empty. Backfill now: - never touches a slot whose row carries a user_id (user's explicit choice) - re-evaluates a previously auto-configured slot on every create and swaps in the best candidate (available first, then larger context); the first user save flips the row to user-owned and locks it - repairs dangling rows and fills never-configured slots as before get_single_config_info now also returns the row's user_id for this classification. * fix(model): fill empty default slots only from newly created models The backfill used to pick the best candidate from ALL live models of a slot's type. When a user deliberately cleared a default slot and then added one new model, the backfill resurrected an older, larger-context model they had passed over -- silently overriding the clear. Pass the ids of the models created by the current call into the backfill: - an empty slot (never configured, or cleared by the user) is now filled only from those newly created models; if the call added none of the slot's type, the slot stays empty - dangling rows still repair from the full pool (the previous choice is gone, so the best remaining replacement is appropriate) - auto-configured slots keep re-evaluating among all candidates (the larger-context swap from the previous commit) - user-configured slots remain locked create_model_record returns only a bool, so the new ids are recovered via display-name lookup (_ids_for_created_models); multi_embedding creates include their embedding twin automatically. * fix(model): freeze settled auto-defaults, swap only within an import session The auto-slot swap from the earlier commit never expired: an auto-configured default could be replaced by a better model at any later create, so adding models months after an import could still move the default. Users expect a default that has been sitting in the slot to stay put -- only the batch import that is still in progress should converge on the best model. Gate the swap on occupant freshness: - swap candidates are the current occupant plus the models created in the current call; older models the user passed over are never resurrected through the swap path (previously the swap re-ranked ALL models of the type, so a cleared-then-refilled slot could drift back to an old giant) - the swap only runs while the occupant was created within _AUTO_SLOT_SWAP_WINDOW (5 minutes) of the newest model in the current call -- batch imports create their rows seconds apart, so mid-batch convergence still works; an occupant from an earlier session is frozen - timestamps come from the DB on both sides, so no clock/timezone skew; missing create_time disables the gate (permissive, legacy behaviour) - empty-slot fill (new-only), dangling repair and user-choice locking are unchanged * fix(model): never move occupied default slots; batches finalize once The auto-slot swap was removed: a default slot occupied by ANY live model -- user-picked or system-backfilled -- is now never touched by later creates. Users expect "the slot already has a model" to mean exactly that; swapping in a better model months after an import silently moved defaults and compounded with the empty-slot rules into hard-to-predict behaviour. The previous commit's freshness window tried to reconcile this with mid-batch convergence via heuristics; explicit batch context replaces it. Batch imports now carry their own flow control: - ModelRequest gains an optional skip_default_backfill flag (popped by the app layer before the dict reaches the service/DB layer, same contract as the accept-signal fields). The batch dialog marks every per-row create with it, so no row claims empty slots as it lands. - A new POST /model/backfill_defaults endpoint finalizes the batch: it resolves the created display names to ids and runs the backfill ONCE with the whole batch as candidates, so an empty slot gets the best model of the batch (available first, then larger context) in a single decision. - The user-facing batch dialog calls the finalize after its loop; the manage-tenant path keeps its existing per-row behaviour (its request model ignores the flag and its frontend service does not forward it). Final slot semantics: occupied -> never touched; empty -> best of the current call's new models (or all models for legacy callers); dangling -> repaired from the full pool; user choice -> locked. --------- Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
The v2.6.1 hotfixes landed on main but were never merged back into develop, so both branches evolved the same files independently. The resulting divergence makes the v2.7.0 release PR (develop -> main) conflict on 45 files. This merge restores the missing back-flow. Resolution: - Conflicted files take the develop side. develop already carries the hotfix content through a different lineage (#3973, #3998, #4024) and its wording is the later evolution, so no fix is lost. - main-only edits that are still live are preserved; files main changed only on paths that develop has since reorganised (agentConfig/ -> capability/, agents/ -> agents/[agentId]/) resolve to their new locations. - v2.6.1_001_remove_human_interaction.sql stays deleted: #4024 already folded its body into v2.7.0_merged_migrations.sql. - VERSION stays v2.7.0, the release being prepared. The merged tree is identical to develop's tip; this commit exists only to rejoin the two histories so the release PR merges without conflicts. Verified: python syntax across merged modules, and the test_model_management_service / test_config_sync_service suites.
* merge v2.6.1 hotfix release from hotfix/v2.6.1 (#3971)
* 🐛 Fix(evaluation): run trials in runtime service (#3954)
* fix(evaluation): run trials in runtime service
Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.
Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5
* test(evaluation): stub config thread manager
Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.
Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5
* test(evaluation): stub runtime jwt helper
* test(evaluation): cover trial proxy error paths
* Fix/override delete (#3958)
* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)
* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key
* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation
* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)
* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it
* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)
* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"
This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.
* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.
---------
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)
* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)
* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests
Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.
Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.
* fix(hitl): guarantee observer chunks precede human_interaction in DB event order
When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.
Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.
Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.
* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths
Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.
Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.
* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety
Fix 4 real issues flagged by github-code-review:
1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
reset on every drain, meaning a model that kept producing chunks could
stall the worker forever. deadline is now computed once at entry and
the sleep call clips to hard_deadline - now.
2. _emit_in_flight Event bridges the async emit path and the worker's
idle poll. Without this, buffer-empty = 'persisted' was confused with
buffer-empty = 'taken for emit but still in run_blocking queue'. The
worker now checks both 'buffer empty for settle_ms' AND 'no emit in
flight' before deciding the async side is truly idle.
3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
Previously the async loop drained the buffer, decided it was not yet
due, then put everything back. That transiently-empty window (16 us
normally, arbitrarily long under GIL/GC/preemption) was enough for
the worker's 20 ms poll to mis-fire. We now peek (read count, no
drain) and only take_chunks when we actually intend to persist.
4. emit_chunks wrapped in try/except that puts drained chunks back into
the shared buffer before re-raising, and finish() wraps its flush call
in try/except: pass. Guarantees (a) no chunk loss on DB failure and
(b) the terminal human_run row is always written even if the flush
step fails.
Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
test_application_execute_attempt.py drives the full consumer loop
through _flush_if_due (peek → take → begin/end_emit) and the final
flush, verifying that every patch line added in application.py is hit.
* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected
When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.
* perf(hitl): replace polling with SSE subscription and move snapshot off the write path
Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
subscription to the backend's /{run_id}/events SSE stream. Discovery is
now one-shot: conversationId change and the agent stream pause
(isRunning true→false), the exact moment a HITL run is most likely to
exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
concurrent callers and returns cached state; (2) 3s minimum interval
so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
to prevent EventSource from reconnecting forever against COMPLETED runs.
Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
before processing each request, and expire_waiting() remains the
periodic scheduler sweep.
Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.
* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false
After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).
The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.
Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.
Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.
* test(hitl): update mock from snapshot to light_snapshot after read-path refactor
test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.
* test(hitl): raise diff coverage above the 90% merge gate
Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.
* style(hitl): unify comment style across HITL changes
Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.
* style(hitl): unify comment style across HITL changes
Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.
* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate
SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.
* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm
openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.
* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"
This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.
* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)
* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)
* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.
* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format
---------
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* [codex] fix(agent): silently retry transient model errors (#3965)
* fix(agent): retry transient model failures silently
* fix(model): support OpenAI httpx2 timeout client
* test(model): add deterministic OpenAI-compatible mock
* fix(agent): keep stream runtime within line budget
* Fix: AIDP knowledge base bug fix (#3967)
* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection
* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change
---------
Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>
* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)
* fix(agent): enforce explicit CodeAgent termination
* fix(test): restore CodeAgent CI compatibility
* merge(main): merge v2.6.1 hotfixes from hotfix/v2.6.1 (#3983)
* 🐛 Fix(evaluation): run trials in runtime service (#3954)
* fix(evaluation): run trials in runtime service
Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.
Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5
* test(evaluation): stub config thread manager
Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.
Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5
* test(evaluation): stub runtime jwt helper
* test(evaluation): cover trial proxy error paths
* Fix/override delete (#3958)
* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)
* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key
* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation
* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)
* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it
* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)
* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"
This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.
* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.
---------
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)
* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)
* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests
Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.
Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.
* fix(hitl): guarantee observer chunks precede human_interaction in DB event order
When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.
Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.
Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.
* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths
Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.
Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.
* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety
Fix 4 real issues flagged by github-code-review:
1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
reset on every drain, meaning a model that kept producing chunks could
stall the worker forever. deadline is now computed once at entry and
the sleep call clips to hard_deadline - now.
2. _emit_in_flight Event bridges the async emit path and the worker's
idle poll. Without this, buffer-empty = 'persisted' was confused with
buffer-empty = 'taken for emit but still in run_blocking queue'. The
worker now checks both 'buffer empty for settle_ms' AND 'no emit in
flight' before deciding the async side is truly idle.
3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
Previously the async loop drained the buffer, decided it was not yet
due, then put everything back. That transiently-empty window (16 us
normally, arbitrarily long under GIL/GC/preemption) was enough for
the worker's 20 ms poll to mis-fire. We now peek (read count, no
drain) and only take_chunks when we actually intend to persist.
4. emit_chunks wrapped in try/except that puts drained chunks back into
the shared buffer before re-raising, and finish() wraps its flush call
in try/except: pass. Guarantees (a) no chunk loss on DB failure and
(b) the terminal human_run row is always written even if the flush
step fails.
Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
test_application_execute_attempt.py drives the full consumer loop
through _flush_if_due (peek → take → begin/end_emit) and the final
flush, verifying that every patch line added in application.py is hit.
* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected
When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.
* perf(hitl): replace polling with SSE subscription and move snapshot off the write path
Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
subscription to the backend's /{run_id}/events SSE stream. Discovery is
now one-shot: conversationId change and the agent stream pause
(isRunning true→false), the exact moment a HITL run is most likely to
exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
concurrent callers and returns cached state; (2) 3s minimum interval
so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
to prevent EventSource from reconnecting forever against COMPLETED runs.
Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
before processing each request, and expire_waiting() remains the
periodic scheduler sweep.
Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.
* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false
After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).
The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.
Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.
Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.
* test(hitl): update mock from snapshot to light_snapshot after read-path refactor
test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.
* test(hitl): raise diff coverage above the 90% merge gate
Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.
* style(hitl): unify comment style across HITL changes
Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.
* style(hitl): unify comment style across HITL changes
Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.
* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate
SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.
* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm
openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.
* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"
This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.
* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)
* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)
* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.
* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format
---------
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* [codex] fix(agent): silently retry transient model errors (#3965)
* fix(agent): retry transient model failures silently
* fix(model): support OpenAI httpx2 timeout client
* test(model): add deterministic OpenAI-compatible mock
* fix(agent): keep stream runtime within line budget
* Fix: AIDP knowledge base bug fix (#3967)
* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection
* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change
* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)
* fix(agent): enforce explicit CodeAgent termination
* fix(test): restore CodeAgent CI compatibility
* cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981)
* Fix StopAsyncIteration leak in execute_attempt finally block
Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors.
Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate.
* Fix HITL form not appearing until page refresh
Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload.
Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh.
* Fix SSE chunk/HITL event ordering race under real server load
Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order.
Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order.
* fix(hitl): recover pending form when SSE delivery stalls silently
Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones.
Changes:
- Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls
- Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it
- Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop
* fix(hitl): stop parked human-input waits from consuming scheduler slots
Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed.
Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments.
Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots.
Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43.
* test(scheduler): fix flaky waiting-concurrency assertion on fast event loops
Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently.
Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.
* Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980)
The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which
also carried two unrelated upstream changes that this release branch never
had:
- #3909 Knowledge base interface optimization (AIDP UI refactor)
- #3930 support AIDP knowledge file deletion and download
Both are removed here so the release line keeps only the bug fix.
- Restore the AIDP frontend components to their pre-refactor layout and
drop the helper modules only the refactor used:
AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils.
- Drop the #3930 document remove/download endpoints from services/api.ts
and the AIDP translations that only those screens referenced.
- Keep the fix itself unchanged: knowledge-base scoped Channels and
KnowledgeFiles/History paths, all-status document listing, keyword
search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the
upload-triggered polling.
Verified:
- pytest test/ext_components/aidp -q -> 583 passed
- frontend `npm run type-check` (tsc --noEmit) -> no errors
---------
Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>
Co-authored-by: chase <byzhangxin11@126.com>
* merge(main): merge v2.6.1 hotfixes from hotfix/v2.6.1 (#3993)
* 🐛 Fix(evaluation): run trials in runtime service (#3954)
* fix(evaluation): run trials in runtime service
Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.
Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5
* test(evaluation): stub config thread manager
Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.
Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5
* test(evaluation): stub runtime jwt helper
* test(evaluation): cover trial proxy error paths
* Fix/override delete (#3958)
* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)
* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key
* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation
* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)
* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it
* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)
* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"
This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.
* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.
---------
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)
* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)
* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests
Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.
Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.
* fix(hitl): guarantee observer chunks precede human_interaction in DB event order
When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.
Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.
Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.
* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths
Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.
Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.
* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety
Fix 4 real issues flagged by github-code-review:
1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
reset on every drain, meaning a model that kept producing chunks could
stall the worker forever. deadline is now computed once at entry and
the sleep call clips to hard_deadline - now.
2. _emit_in_flight Event bridges the async emit path and the worker's
idle poll. Without this, buffer-empty = 'persisted' was confused with
buffer-empty = 'taken for emit but still in run_blocking queue'. The
worker now checks both 'buffer empty for settle_ms' AND 'no emit in
flight' before deciding the async side is truly idle.
3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
Previously the async loop drained the buffer, decided it was not yet
due, then put everything back. That transiently-empty window (16 us
normally, arbitrarily long under GIL/GC/preemption) was enough for
the worker's 20 ms poll to mis-fire. We now peek (read count, no
drain) and only take_chunks when we actually intend to persist.
4. emit_chunks wrapped in try/except that puts drained chunks back into
the shared buffer before re-raising, and finish() wraps its flush call
in try/except: pass. Guarantees (a) no chunk loss on DB failure and
(b) the terminal human_run row is always written even if the flush
step fails.
Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
test_application_execute_attempt.py drives the full consumer loop
through _flush_if_due (peek → take → begin/end_emit) and the final
flush, verifying that every patch line added in application.py is hit.
* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected
When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.
* perf(hitl): replace polling with SSE subscription and move snapshot off the write path
Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
subscription to the backend's /{run_id}/events SSE stream. Discovery is
now one-shot: conversationId change and the agent stream pause
(isRunning true→false), the exact moment a HITL run is most likely to
exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
concurrent callers and returns cached state; (2) 3s minimum interval
so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
to prevent EventSource from reconnecting forever against COMPLETED runs.
Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
before processing each request, and expire_waiting() remains the
periodic scheduler sweep.
Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.
* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false
After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).
The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.
Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.
Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.
* test(hitl): update mock from snapshot to light_snapshot after read-path refactor
test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.
* test(hitl): raise diff coverage above the 90% merge gate
Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.
* style(hitl): unify comment style across HITL changes
Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.
* style(hitl): unify comment style across HITL changes
Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.
* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate
SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.
* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm
openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.
* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"
This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.
* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)
* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)
* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.
* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format
---------
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* [codex] fix(agent): silently retry transient model errors (#3965)
* fix(agent): retry transient model failures silently
* fix(model): support OpenAI httpx2 timeout client
* test(model): add deterministic OpenAI-compatible mock
* fix(agent): keep stream runtime within line budget
* Fix: AIDP knowledge base bug fix (#3967)
* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection
* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change
* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)
* fix(agent): enforce explicit CodeAgent termination
* fix(test): restore CodeAgent CI compatibility
* cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981)
* Fix StopAsyncIteration leak in execute_attempt finally block
Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors.
Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate.
* Fix HITL form not appearing until page refresh
Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload.
Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh.
* Fix SSE chunk/HITL event ordering race under real server load
Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order.
Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order.
* fix(hitl): recover pending form when SSE delivery stalls silently
Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones.
Changes:
- Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls
- Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it
- Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop
* fix(hitl): stop parked human-input waits from consuming scheduler slots
Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed.
Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments.
Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots.
Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43.
* test(scheduler): fix flaky waiting-concurrency assertion on fast event loops
Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently.
Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.
* Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980)
The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which
also carried two unrelated upstream changes that this release branch never
had:
- #3909 Knowledge base interface optimization (AIDP UI refactor)
- #3930 support AIDP knowledge file deletion and download
Both are removed here so the release line keeps only the bug fix.
- Restore the AIDP frontend components to their pre-refactor layout and
drop the helper modules only the refactor used:
AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils.
- Drop the #3930 document remove/download endpoints from services/api.ts
and the AIDP translations that only those screens referenced.
- Keep the fix itself unchanged: knowledge-base scoped Channels and
KnowledgeFiles/History paths, all-status document listing, keyword
search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the
upload-triggered polling.
Verified:
- pytest test/ext_components/aidp -q -> 583 passed
- frontend `npm run type-check` (tsc --noEmit) -> no errors
* fix(agent): accept reasoning-prefixed code actions (#3990)
* refactor: remove human interaction features and related configurations (#3988)
* refactor: remove human interaction features and related configurations
* refactor: remove human interaction features and related configurations
---------
Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>
Co-authored-by: Dallas98 <40557804+Dallas98@users.noreply.gi…
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Merge latest develop into main for the v2.7.0 release. No conflicts
(develop synchronized with main via #4047).
Changes
v2.7.0 release (58 commits)
Includes v2.6.1 main hotfixes (enforce explicit CodeAgent termination
and silent recovery #3969, plus the v2.6.1 hotfix merges #3971/#3983/#3993)
Agent Workbench launch flow, model config redesign, tenant-scoped model
management, evaluation and paged agent list fixes, and more
SQL migration consolidation into v2.7.0_merged_migrations.sql
Notes
v2.7.0 tag: (updated after merge)
Docker image build triggered