Skip to content

merge(main): merge v2.7.0 release from develop - #4044

Merged
jeffwu-1999 merged 61 commits into
mainfrom
develop
Sep 30, 2026
Merged

jeffwu-1999 merged 61 commits into
mainfrom
develop

Conversation

@jeffwu-1999

@jeffwu-1999 jeffwu-1999 commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary
Merge latest develop into main for the v2.7.0 release. No conflicts
(develop synchronized with main via #4047).

Changes
v2.7.0 release (58 commits)
Includes v2.6.1 main hotfixes (enforce explicit CodeAgent termination
and silent recovery #3969, plus the v2.6.1 hotfix merges #3971/#3983/#3993)
Agent Workbench launch flow, model config redesign, tenant-scoped model
management, evaluation and paged agent list fixes, and more
SQL migration consolidation into v2.7.0_merged_migrations.sql

Notes
v2.7.0 tag: (updated after merge)
Docker image build triggered

cj2026-bit and others added 30 commits September 17, 2026 15:41
* feat: support AIDP knowledge file management

Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: Codex

* fix: unblock AIDP PR quality checks

Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: Codex

* fix: reduce Sonar duplication and stabilize web install

Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: Codex

* test: improve AIDP file operation coverage

* fix: support unicode filenames in AIDP mock download

* fix: align AIDP document deletion with local KB

* fix: stream AIDP document downloads

* fix: preserve AIDP permission test compatibility

* fix: simplify AIDP document download streaming

* fix(aidp): align document removal request with API

* chore: remove unrelated assistant ui dependency change

* test(aidp): align mock document id type
* Feature: 知识库界面优化

* Fix: 修复知识库门禁问题

---------

Co-authored-by: hzw <hzw@qq.com>
* perf(frontend): enable Turbopack for local development

Enable Turbopack in the custom development server and stabilize its configuration dependencies.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* perf(frontend): upgrade Next.js for faster Turbopack

Upgrade Next.js and React to 16.3.5 and 19.2.3, migrate lint and proxy configuration, and resolve React 19 type compatibility.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* fix(frontend): silence Next HMR upgrade logging

Allow Next.js to handle its own HMR WebSocket upgrades without emitting a misleading proxy log.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* fix(frontend): migrate Ant Design 6 component APIs

Replace deprecated modal, drawer, and alert props with their Ant Design 6 equivalents. Mount hidden memory forms before use and make memory configuration card spacing explicit.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* build(web): align image runtime with frontend dependencies

Use Node 22 for the web image build and runtime stages so freshly resolved dependencies meet their engine requirements.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5
…troller, and break adapter on terminal human_run (#3948)

* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests

Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.

Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.

* fix(hitl): guarantee observer chunks precede human_interaction in DB event order

When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.

Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.

Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.

* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths

Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.

Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.

* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety

Fix 4 real issues flagged by github-code-review:

1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
   reset on every drain, meaning a model that kept producing chunks could
   stall the worker forever. deadline is now computed once at entry and
   the sleep call clips to hard_deadline - now.

2. _emit_in_flight Event bridges the async emit path and the worker's
   idle poll. Without this, buffer-empty = 'persisted' was confused with
   buffer-empty = 'taken for emit but still in run_blocking queue'. The
   worker now checks both 'buffer empty for settle_ms' AND 'no emit in
   flight' before deciding the async side is truly idle.

3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
   Previously the async loop drained the buffer, decided it was not yet
   due, then put everything back. That transiently-empty window (16 us
   normally, arbitrarily long under GIL/GC/preemption) was enough for
   the worker's 20 ms poll to mis-fire. We now peek (read count, no
   drain) and only take_chunks when we actually intend to persist.

4. emit_chunks wrapped in try/except that puts drained chunks back into
   the shared buffer before re-raising, and finish() wraps its flush call
   in try/except: pass. Guarantees (a) no chunk loss on DB failure and
   (b) the terminal human_run row is always written even if the flush
   step fails.

Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
  hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
  try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
  test_application_execute_attempt.py drives the full consumer loop
  through _flush_if_due (peek → take → begin/end_emit) and the final
  flush, verifying that every patch line added in application.py is hit.

* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected

When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.

* perf(hitl): replace polling with SSE subscription and move snapshot off the write path

Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
  subscription to the backend's /{run_id}/events SSE stream. Discovery is
  now one-shot: conversationId change and the agent stream pause
  (isRunning true→false), the exact moment a HITL run is most likely to
  exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
  adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
  our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
  concurrent callers and returns cached state; (2) 3s minimum interval
  so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
  not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
  to prevent EventSource from reconnecting forever against COMPLETED runs.

Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
  WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
  must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
  writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
  calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
  before processing each request, and expire_waiting() remains the
  periodic scheduler sweep.

Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.

* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false

After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).

The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.

Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.

Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.

* test(hitl): update mock from snapshot to light_snapshot after read-path refactor

test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.

* test(hitl): raise diff coverage above the 90% merge gate

Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate

SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.
* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)

* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key

* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation

* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)

* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it

* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)

* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"

This reverts commit 2376aa5.

* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* fix(evaluation): support multi-agent task filtering

Default the evaluation task list to all tenant agents, add searchable multi-select filtering with one JSON agent_ids parameter, and preserve the active filter after task creation.

Co-authored-by: Codex <noreply@tool-provider.com>

Generated-by: gpt-5-codex

* refactor(evaluation): simplify agent id validation

Use Pydantic strict JSON validation for the agent_ids query parameter while preserving the all-agents default and stable deduplication.

Co-authored-by: Codex <noreply@tool-provider.com>

Generated-by: gpt-5-codex
* fix(newchat): reject attachments larger than 10 MB

Validate attachments before upload in the newchat page and show a localized size error so oversized files never enter the upload queue. Cover the exact 10 MB boundary.

Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5

* fix(i18n): clarify local A2A development URLs

区分本地直启、本地容器和 k8s 启动时使用的 A2A 基础 URL。

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol

* fix(i18n): sync English local A2A URL hint

Clarify the local, local-container, and k8s A2A base URLs in English.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol

* fix(newchat): limit attachments to 50 files

Track pending newchat attachments per runtime and reject the 51st file with a localized error message. Releasing slots on removal or successful upload keeps subsequent selections available.

Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5

---------

Co-authored-by: Codex <noreply@openai.com>
* fix(agent-version): return not found for absent version

Reject direct lookups for absent version metadata so the existing API mapping returns HTTP 404 rather than 200 null.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* refactor(agent): remove legacy draft creation endpoint

Remove the unused draft allocation API and its frontend client.\nMigrate creation coverage to the existing update endpoint.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* test(agents): expose stable variable name locator

Add an explicit test contract for the agent variable-name field so browser flows do not depend on localized placeholder text.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-version): scope version lookup to tenant

Filter version metadata by tenant ID to prevent cross-tenant lookups and cover the predicate with a regression test.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex
* [codex] fix(agent): silently retry transient model errors (#3965)

* fix(agent): retry transient model failures silently

* fix(model): support OpenAI httpx2 timeout client

* test(model): add deterministic OpenAI-compatible mock

* fix(agent): keep stream runtime within line budget

* fix(agent): enforce explicit CodeAgent termination

* fix(test): restore CodeAgent CI compatibility

* test(agent): keep develop isolation checks compatible

* fix(agent): hide protocol repair generation stream

* fix(ci): satisfy CodeAgent quality gate

* test(model): cover retry classification branches
…ncurrency (#3977)

* Fix StopAsyncIteration leak in execute_attempt finally block

Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors.

Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate.

* Fix HITL form not appearing until page refresh

Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload.

Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh.

* Fix SSE chunk/HITL event ordering race under real server load

Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order.

Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order.

* fix(hitl): recover pending form when SSE delivery stalls silently

Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones.

Changes:
- Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls
- Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it
- Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop

* fix(hitl): stop parked human-input waits from consuming scheduler slots

Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed.

Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments.

Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots.

Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43.

* test(scheduler): fix flaky waiting-concurrency assertion on fast event loops

Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently.

Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.
* feat:新增需求设计开发阶段测试用例及脚本生成skill

* feat:新增a2a测试mock服务

* feat:一键部署mock服务

* test:ad test
* fix: standardize tenant resource limit errors

* fix: roll back rejected tenant user registrations

* fix: stabilize tenant limit test imports

* test: import HTTPStatus in northbound tests

* test: cover tenant resource limit error paths

* config: make tenant resource limits configurable
* refactor(resource): unify repository resource cards

Use the shared resource grid and card shell for Agent, MCP, and Skill repository listings while preserving their existing search, filtering, pagination, and actions.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* refactor(resource): share resource create cards

Route Agent, MCP, and Skill create-card entry points through the shared CreateResourceCard and use the shared grid for Agent and Skill mine lists.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* refactor(resource): share mine resource card shells

Apply the shared ResourceCard shell to Agent, MCP, and Skill mine-resource cards while retaining their existing content and actions.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* refactor(resource): align public repository card actions

Use the shared card's stacked footer layout for Agent, MCP, and Skill repository listings so metadata and actions remain consistent across resource types.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* refactor(agent-space): split tab content components

Move repository, my-agent, and review-center tab content behind same-level components and expand the card grids to four desktop columns.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* refactor(agent-space): move my agent view

Place the complete my-agent view beside the route page and preserve its resource-card grid behavior.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* fix(agent-space): resolve moved my agent imports

Update relative component imports after moving the my-agent view and keep layout regression coverage aligned.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* refactor(agent-space): remove legacy tab views

Delete duplicate repository and review tab implementations after moving them into dedicated components.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* refactor(agent-space): localize detail dialogs

Move repository copy and detail dialogs into the repository tab and version detail into the my-agent tab.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* refactor(agent-space): localize tab state and queries

Move list, filter, pagination, detail, and mutation state into each tab view. Keep tab counts on the page and preserve filters across tab switches while disabling inactive queries.

Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5

* refactor(agent-space): use shared cards and pagination

Render repository listings with ResourceCard and delegate server pagination to ResourceCardGrid. Rename the view to space.tsx and remove the unused agent card adapter and pagination controls.

Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5

* refactor(agent-space): streamline repository card actions

Open repository details from the card, move copying into the action menu, show versions as title badges, and place authors in the card footer.

Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5

* feat(agent-space): adapt repository cards to the viewport

Use Ant Design breakpoints and measured card-region height to select the server page size. Constrain ResourceCardGrid rows to the available viewport space.

Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5

* fix(agent-space): clamp card descriptions to available height

Reduce repository-card description lines as card rows shrink and make the shared card description overflow-safe.

Co-authored-by: Codex <noreply@openai.com>
Generated-by: gpt-5

* fix(agent-space): fit repository cards within viewport

Remove duplicate tab spacing and reserve the responsive page bottom padding when sizing the card grid.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): match repository card minimum height

Use the shared card minimum height when calculating responsive rows so cards and pagination remain within the viewport.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* refactor(agent-space): reposition repository card actions

Move copying to the card footer and show download counts beside repository names.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): defer four-column repository layout

Keep repository cards in three columns through xl and use four columns only at the xxl breakpoint.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): keep lg repository cards in two columns

Limit the repository grid to two columns through lg, then use three columns at xl and four at xxl.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(resource-card): align small-screen grid breakpoint

Start the two-column resource grid at 576px so CSS matches Ant Design responsive state and pagination capacity.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): emphasize repository copy action

Use a bordered default button for copying repository agents from the card footer.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): strengthen repository copy button

Increase the copy action size and use the primary button treatment for a clear visual affordance.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): simplify my agent toolbar

Remove redundant import and create-agent toolbar actions, align tag filtering with search, and reuse the shared resource pagination component.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol

* fix(agent-space): reorganize my agent card actions

Move evaluation into the card overflow menu and keep editing in the lower-right action area. Reposition version, lifecycle status, repository status, and creation date for clearer scanning.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol

* fix(agent-space): consolidate repository status badge

Show the listed version status beside the agent name and remove the redundant Hub badge.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol

* fix(agent-space): use blue badge for listed agents

Match the listed repository status badge to the former Hub blue styling.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol

* fix(agent-space): clarify current version label

Change the agent card version label to Current Version and add a primary color indicator for the active version.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol

* fix(agent-space): align lifecycle badge with more menu

Place the published or draft badge beside the card menu so both share the top-right header alignment.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): place lifecycle badge below more menu

Stack the published or draft badge below the top-right menu to match the card header layout.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): unify repository version display

Show the repository card version with the current-version label and typography used by my agent cards.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): unify my agent action button style

Use the repository card text-button treatment for editable and read-only my agent card actions.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(agent-space): fill my agent grid to viewport

Measure the available viewport height and distribute my agent cards across responsive grid rows.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* refactor(mcp-space): split tab content into components

Move MCP repository, mine, and review content into focused components and keep shared actions in a controller. Match agent-space page spacing and tab styling while preserving existing interactions.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* refactor(mcp-space): keep three tab components

Consolidate the MCP tab content into space, my-mcp, and review-center files while keeping the page focused on tab selection and preserving shared interactions.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* refactor(skill-space): split three tab views

Move repository, mine, and review state and actions into their corresponding components. Keep the route focused on tab selection and align its spacing and tabs with the other spaces.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* fix(mcp-space): refine repository card actions

Move downloads beside the more menu, show the publisher in the metadata row, and make installation the sole footer action. Open the detail modal when the card itself is clicked.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5-codex

* fix(mcp-space): align repository card footer

Match the agent repository card layout by placing the publisher and install action in the same inline footer row.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5-codex

* fix(agent-space): align repository download action

Place the repository download count in the top-right action area so it remains adjacent to the conditional more menu.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* fix(agent-space): align repository search height

Use the default Ant Design input height so the repository search field matches the adjacent tag filter button.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-space): match repository copy button styling

Use the same transparent text-button treatment as My Agent card actions for repository copy controls, and add a regression assertion.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5-codex

* fix(mcp-space): align search filter control heights

Match the MCP search input and status select height to Ant Design's default tag filter button height.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol
Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(mcp-space): match repository action button style

Use the same borderless text-button treatment as the Agent repository card while keeping installed repository entries disabled.

Co-authored-by: Cursor <noreply@cursor.com>

Generated-by: gpt-5.6-sol

* refactor(mcp-space): align responsive repository cards

Align Agent and MCP repository layouts, and simplify My MCP card actions for direct editing and compact controls.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-space): align repository status badges

Place the repository listing status beside the lifecycle state on My Agent cards.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(resource): top-align card header actions

Align multi-line header action groups from their top edge so adjacent controls stay level.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-space): center repository header actions

Keep the download indicator vertically aligned with the adjacent more menu in repository cards.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* fix(mcp-space): align mine card header actions

Keep the more menu beside the health-check action while the review badge stays below it.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* fix(mcp-space): group mine card header actions

Place health check and more actions on the first header row, with the review badge beneath them.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* fix(mcp-space): right-align mine review badge

Align the review status badge below the more action at the card header's right edge.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* refactor(mcp-space): unify card grid pagination

Use ResourceCardGrid for repository and mine MCP cards so pagination follows the shared responsive layout.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* fix(resource): always show card pagination

Keep the shared pagination control visible for single-page card lists.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* fix(resource): reserve space for card pagination

Include the always-visible pagination area in responsive card grid height calculations.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* docs(skill-repository): document card design

Document the approved Skill repository card layout and implementation plan.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* refactor(skill-repository): align cards with resource grid

Use responsive ResourceCardGrid pagination and streamline Skill repository and My Skills card actions.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(skill-repository): refine card metadata layout

Square Skill search controls, move filtering beside search, reserve tag space, and show repository authors.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(skill-repository): enforce card metadata layout

Apply square controls without relying on CSS ordering and display the repository submitter when author is absent.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(skill-repository): match agent search controls

Use the Agent repository search radius and default filter button styling in both Skill tabs.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(skill-repository): use default component radius

Remove the Skill page radius override so Ant Design controls match the Agent repository.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(skill-repository): render listing tags directly

Render repository tags as direct ResourceCard content while preserving empty tag space when no tags exist.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(skill-repository): show managed Skill tags

Populate My Skill cards from structured tag assignments with a legacy tag fallback.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(skill-repository): use default description lines

Remove the two-line override from Mine Skill cards so descriptions use the shared ResourceCard three-line default.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* refactor(resource-cards): align default action styles

Use default small text-button styling in Agent and MCP cards, keep the Agent page full height, and rename the Chinese MCP page title.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* test(resource-cards): remove temporary layout tests

Remove the Agent and Skill layout test files created during the refactor-card work.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* test(resource-cards): remove remaining temporary tests

Remove the remaining Agent and MCP layout test files requested for the refactor-card branch.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* docs(skill-repository): remove temporary plans

Remove the Skill repository design and implementation plans created for this branch.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

* test(skill-repository): isolate managed tag lookup

Mock tag assignment lookups in Skill repository service unit tests so they do not connect to PostgreSQL.\n\nCo-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5

---------

Co-authored-by: Codex <noreply@openai.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
* 🐛 Fix(evaluation): run trials in runtime service (#3954)

* fix(evaluation): run trials in runtime service

Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub config thread manager

Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub runtime jwt helper

* test(evaluation): cover trial proxy error paths

* Fix/override delete (#3958)

* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)

* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key

* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation

* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)

* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it

* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)

* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"

This reverts commit 2376aa5.

* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)

* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)

* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests

Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.

Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.

* fix(hitl): guarantee observer chunks precede human_interaction in DB event order

When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.

Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.

Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.

* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths

Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.

Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.

* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety

Fix 4 real issues flagged by github-code-review:

1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
   reset on every drain, meaning a model that kept producing chunks could
   stall the worker forever. deadline is now computed once at entry and
   the sleep call clips to hard_deadline - now.

2. _emit_in_flight Event bridges the async emit path and the worker's
   idle poll. Without this, buffer-empty = 'persisted' was confused with
   buffer-empty = 'taken for emit but still in run_blocking queue'. The
   worker now checks both 'buffer empty for settle_ms' AND 'no emit in
   flight' before deciding the async side is truly idle.

3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
   Previously the async loop drained the buffer, decided it was not yet
   due, then put everything back. That transiently-empty window (16 us
   normally, arbitrarily long under GIL/GC/preemption) was enough for
   the worker's 20 ms poll to mis-fire. We now peek (read count, no
   drain) and only take_chunks when we actually intend to persist.

4. emit_chunks wrapped in try/except that puts drained chunks back into
   the shared buffer before re-raising, and finish() wraps its flush call
   in try/except: pass. Guarantees (a) no chunk loss on DB failure and
   (b) the terminal human_run row is always written even if the flush
   step fails.

Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
  hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
  try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
  test_application_execute_attempt.py drives the full consumer loop
  through _flush_if_due (peek → take → begin/end_emit) and the final
  flush, verifying that every patch line added in application.py is hit.

* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected

When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.

* perf(hitl): replace polling with SSE subscription and move snapshot off the write path

Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
  subscription to the backend's /{run_id}/events SSE stream. Discovery is
  now one-shot: conversationId change and the agent stream pause
  (isRunning true→false), the exact moment a HITL run is most likely to
  exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
  adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
  our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
  concurrent callers and returns cached state; (2) 3s minimum interval
  so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
  not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
  to prevent EventSource from reconnecting forever against COMPLETED runs.

Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
  WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
  must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
  writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
  calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
  before processing each request, and expire_waiting() remains the
  periodic scheduler sweep.

Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.

* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false

After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).

The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.

Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.

Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.

* test(hitl): update mock from snapshot to light_snapshot after read-path refactor

test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.

* test(hitl): raise diff coverage above the 90% merge gate

Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate

SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.

* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm

openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b93 and later accidentally reverted by commit d086da2.

* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"

This reverts commit 8e06ce3.

* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.

* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* [codex] fix(agent): silently retry transient model errors (#3965)

* fix(agent): retry transient model failures silently

* fix(model): support OpenAI httpx2 timeout client

* test(model): add deterministic OpenAI-compatible mock

* fix(agent): keep stream runtime within line budget

* Fix: AIDP knowledge base bug fix (#3967)

* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection

* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change

* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)

* fix(agent): enforce explicit CodeAgent termination

* fix(test): restore CodeAgent CI compatibility

* cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981)

* Fix StopAsyncIteration leak in execute_attempt finally block

Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors.

Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate.

* Fix HITL form not appearing until page refresh

Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload.

Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh.

* Fix SSE chunk/HITL event ordering race under real server load

Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order.

Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order.

* fix(hitl): recover pending form when SSE delivery stalls silently

Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones.

Changes:
- Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls
- Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it
- Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop

* fix(hitl): stop parked human-input waits from consuming scheduler slots

Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed.

Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments.

Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots.

Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43.

* test(scheduler): fix flaky waiting-concurrency assertion on fast event loops

Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently.

Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.

* Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980)

The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which
also carried two unrelated upstream changes that this release branch never
had:

  - #3909 Knowledge base interface optimization (AIDP UI refactor)
  - #3930 support AIDP knowledge file deletion and download

Both are removed here so the release line keeps only the bug fix.

  - Restore the AIDP frontend components to their pre-refactor layout and
    drop the helper modules only the refactor used:
    AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils.
  - Drop the #3930 document remove/download endpoints from services/api.ts
    and the AIDP translations that only those screens referenced.
  - Keep the fix itself unchanged: knowledge-base scoped Channels and
    KnowledgeFiles/History paths, all-status document listing, keyword
    search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the
    upload-triggered polling.

Verified:
  - pytest test/ext_components/aidp -q  ->  583 passed
  - frontend `npm run type-check` (tsc --noEmit)  ->  no errors

* fix(agent): accept reasoning-prefixed code actions (#3990)

* refactor: remove human interaction features and related configurations (#3988)

* refactor: remove human interaction features and related configurations

* refactor: remove human interaction features and related configurations

---------

Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>
Co-authored-by: Dallas98 <40557804+Dallas98@users.noreply.github.com>
Co-authored-by: chase <byzhangxin11@126.com>
…st suite (#4002)

- quota: reject negative GB/MB inputs and warning>=critical threshold pairs
  (previously persisted as 200 with semantically inverted config)
- api-key: mask access_key in /api-keys and /user/tokens list responses;
  only create/refresh may return the full secret once
- tag: translate the DB trigger's 'Tag assignment limit exceeded' (without
  the 'Resource ' prefix) into a structured 409, and flush replacement
  deletes before inserts so a full tag replacement no longer trips the
  assignment-capacity trigger
- model: propagate ValueError from create/batch_create so duplicate display
  names map to 409 and malformed batch entries to 422 instead of 500;
  include model_type in /model/llm_list items; tolerate quick-config
  entries without model_repo in get_model_name_from_config
- test: align model service tests with the ValueError passthrough contract

Co-authored-by: chase <byzhangxin11@126.com>
…ry (#4004)

The guard rejected private/reserved IP literals and DNS names resolving
into non-routable ranges before fetching {base_url}/models. Internal and
IP-based deployments (including private-network OpenAI-compatible
endpoints) were blocked from batch model discovery. Remove the
_reject_non_public_ip / _validate_provider_base_url /
_assert_resolved_ips_public chain and the related tests; keep the
ssl_verify TLS toggle, which is independent of the guard.

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
…e authorization (#3798)

* ✨ Feature: Pass user context to MCP tools and A2A agents for tool-side authorization

* test: update NexentAgent exact-call assertions for user_context kwarg

* test: cover user_context injection branches for codecov patch coverage

* docs: remove internal tool user context design

* fix: support legacy agent run info context

* fix: hide injected MCP context from tool signatures

* fix: hide injected MCP context from action traces

* Revert "fix: hide injected MCP context from action traces"

This reverts commit de25e57.

* fix: hide injected context from model tool metadata
* bugfix:支持思考挡位配置

* bugfix: support model reasoning effort configuration

* test: improve reasoning configuration coverage

* test: complete reasoning coverage

* fix: resolve CI quality gate failures

* test: complete coverage for gate branches

* bugfix: support configurable model reasoning effort

* test: complete reasoning coverage

* test: cover reasoning override branch

* bugfix: support model reasoning effort settings

* bugfix: complete reasoning capability matching

* test: align reasoning capability assertions

* test: close reasoning coverage gaps

---------

Co-authored-by: hzw <hzw@qq.com>
…3996)

* fix(file): reject invalid upload destinations at the API boundary

An unknown destination value previously reached the service layer, where
it raised a bare Exception and surfaced as HTTP 500 with a generic
"File upload error." message. Constrain the form field to the documented
"local"/"minio" values so FastAPI rejects unknown destinations with 422
before any business logic runs.

Co-authored-by: ZCode <noreply@zcode.ai>
Generated-by: new-provider/glm-5.3

* fix(data-process): call module-level get_task_details in the details route

The route called service.get_task_details, a method that has never
existed on DataProcessService (regression introduced by the coding-style
refactor in 4a70a99), so every request to GET /tasks/{task_id}/details
failed with AttributeError and returned 500. Call the data_process.utils
implementation that the celery support introduced instead.

Co-authored-by: ZCode <noreply@zcode.ai>
Generated-by: new-provider/glm-5.3

* fix(file): stop the upload conflict scan from matching its own batch

The filename conflict scan in upload_files_impl calls list_files after the
current batch's lifecycle rows have been created, and list_files has merged
durable lifecycle rows into the file list since the lifecycle ledger landed.
The scan therefore counted the upload's own row as an existing document and
renamed every first upload of a name to <name>_1.

Exclude the current batch's file_ids when collecting existing names. ES
entries carry no file_id and are still treated as pre-existing documents, so
genuine conflicts keep resolving to _1 and within-batch duplicates keep
getting sequential suffixes.

Covered by a unit test whose list_files fixture exposes only the batch's own
row: the uploaded name is kept unchanged and no filename rewrite is recorded.
The shared upload mock harness is extracted so the new test does not
duplicate the existing conflict-resolution fixtures.

* fix(evaluation): allow tenant administrators to delete evaluation runs

DELETE /agent-evaluations/{id} only allowed the run's creator (160208 /
HTTP 403). When the creator account was deleted, the run could not be
removed until the 30-day cleanup, and runs with NULL created_by could not
be deleted by anyone.

After the creator check, also allow callers whose tenant role is in
consts.const.CAN_EDIT_ALL_USER_ROLES (SU/ADMIN/SPEED/ASSET_OWNER). The role
is resolved inside the service via database.user_tenant_db (unknown callers
fall back to USER, fail-closed), keeping the endpoint signature unchanged.
Error copy updated in error_message.py and the zh/en locales; code 160208
stays the same.

Covered by three new unit tests (admin may delete others' runs, DEV cannot,
unknown role fails closed) on top of the existing creator-only test.

* test(evaluation): stub database.user_tenant_db in the pure-logic harness

CI only failed on test_evaluation_pure_logic: the harness reloads
services.agent_evaluation_service, whose new
'from database.user_tenant_db import get_user_tenant_by_user_id' pulled the
real module outside the harness's stub set; the real module then executed
'from database.db_models import TenantGroupInfo' against this file's
two-class db_models stub and raised ImportError at setup for all 27 tests.

Register database.user_tenant_db in the harness (same pattern as the other
database submodules) so the reload stays fully stubbed.

---------

Co-authored-by: ZCode <noreply@zcode.ai>
* bugfix: remove automatic model renaming

* fix: restore reasoning compatibility for test gate

* test: cover reasoning wire adapters

* fix: validate provider hosts before reasoning mapping

---------

Co-authored-by: hzw <hzw@qq.com>
Non-array skill_tags values persisted in the skill_t JSON column caused
the "t.tags || []).forEach is not a function" error on the agents page
when the skill selection dialog computed its tag filter set.

- Backend _to_dict now normalizes skill_tags into a validated string list
- Frontend fetchSkills reuses normalizeTags for the same defense in depth
- Add malformed-tags coverage to test_skill_db.py
…3978)

* fix(resource-manage): paginate knowledge base list in quota settings modal

The quota settings modal rendered the knowledge base breakdown table
without pagination, which froze the UI on deployments with many KBs.
Use a 10-row small pagination with a total count hint (zh/en locales).

* fix(resource-manage): harden platform quota panel on slow and large deployments

Three related issues in the SU platform quota overview:

- The panel never listened to QUOTA_USAGE_CHANGED_EVENT, so quota
  changes made elsewhere did not refresh it until reopen.
- With data still loading, the header rendered misleading
  \"0 B / Unlimited\" placeholders; on a 786-tenant deployment the
  overview aggregation is slow enough that users saw quota values
  apparently reset to zero. Show a Card loading state until real
  data arrives.
- toQuotaInput could floor a nonzero finite quota below 1 MB down
  to 0; clamp the MB view to at least 1 MB.

Also add destroyOnHidden to the system settings modal so it always
remounts with fresh data.

* fix(resource-manage): stop bare 0 leak in quota fair share line

The fair share reference rounded hard_limit_gb / kb_count with
Math.round, so a 2 GB limit across many KBs produced 0. Because the
render condition used the value directly, React printed the bare
number 0 between the usage tag and the per-KB breakdown heading.

Keep the raw ratio and format it to 2 decimals when it is not an
integer (e.g. 0.17 GB/KB), and guard the JSX with != null so a null
value never renders.
…elopers (#4008)

- Add RBAC checks to all /model/* endpoints via permissions.depends.require
- Restrict cross-tenant /manage/* endpoints to the SU role
- Gate model page buttons by model:update / model:delete permissions
- Remove the /models menu entry from the DEV role via a new migration
- Keep model:read so agent editors can pick admin-configured models

Co-authored-by: chase <byzhangxin11@126.com>
…options (#4013)

* bugfix: preserve AIDP permissions when submitting collapsed advanced options

* test: add AIDP regression evidence screenshot

* chore: remove local regression screenshot artifact

---------

Co-authored-by: hzw <hzw@qq.com>
charmingchi1 and others added 19 commits September 28, 2026 14:05
* feat(logging): 新增模型调用日志分类与路由能力

1.  新增model_call日志分类,将SDK模型层日志与系统日志分离
2.  配置运行时服务日志同时包含runtime和model_call分类
3.  为core_agent、openai_llm等模型相关日志添加专用日志器
4.  新增.gitignore规则忽略.serena目录
5.  移除run_agent.py中多余的DEBUG日志级别设置
6.  完善.env.example日志配置注释
7.  新增完整的模型日志路由测试用例

* feat(logging): route context_evidence logger into model_call file

- Add "context_evidence" to MODEL_CALL_LOGGERS so the run-level
  "Agent loop context evidence" record lands in the model_call
  category file (console + file_model_call, propagate=False)
- Extend whitelist coverage and routing regression tests

Co-authored-by: ZCode <noreply@zcode.ai>
Generated-by: glm-5.3-flash

* refactor(test): deduplicate model_call routing regression tests

- Merge test_dictconfig_routes_model_records_to_dedicated_file and
  test_configure_logging_routes_model_records_to_dedicated_file into
  one parametrized test (ids: dictconfig / configure_logging)
- Shared log+assert body extracted to _log_and_assert_routing helper;
  per-path setup moved to module-level _apply_* functions

Co-authored-by: ZCode <noreply@zcode.ai>
Generated-by: glm-5.3-flash

* feat(logging): land model input/output body records in model_call file

- Pin the MODEL_CALL_LOGGERS whitelist to DEBUG in both the dictConfig
  "loggers" section and configure_logging bind/unbind, and widen the
  file_model_call handler to DEBUG, so MODEL INPUT PARAMETERS logged at
  DEBUG on model_call.core_agent is no longer dropped by root INFO
- Add the symmetric MODEL OUTPUT body record in core_agent
- Merge the core_agent module logger into the model_call.core_agent
  namespace so every core-agent record routes to the model_call file
  instead of runtime; verified end-to-end (model records land only in
  model_call, runtime records unaffected, 41 logging tests pass)

Co-authored-by: ZCode <noreply@zcode.ai>
Generated-by: glm-5.3-flash

* test: 增加核心代理日志记录相关的测试用例并添加调试日志

在 core_agent 的运行逻辑中添加调试日志以记录新任务,同时为日志记录功能新增多项测试,包括验证模型输出是否写入 model_call 日志、日志记录器名称是否保持在指定命名空间、日志参数是否正确记录到文件和面板等,确保日志功能符合预期且敏感数据被正确脱敏。

* fix(core_agent): 修复模型输出日志记录时未读取最新值的问题

调整日志记录的位置,确保在记录模型输出前已完成 model_output 的赋值,避免记录到空值或旧值

* refactor(logging): 将模型调用日志级别调整为跟随根日志级别

调整模型相关日志的级别设置,不再单独固定为DEBUG,而是跟随根日志级别。同时修改SDK中的模型输入、输出和任务日志从DEBUG级别改为INFO,确保在默认根日志级别为INFO时这些关键模型日志能正常输出。更新测试用例以匹配新的日志级别配置,同时修改日志配置文件中的相关注释和设置,使日志行为更符合预期,避免关键模型日志在默认配置下被丢弃。

---------

Co-authored-by: ZCode <noreply@zcode.ai>
* feat: redesign model configuration page (v0 design) - inline default-model slot grid replaces the DefaultModelDialog (priority badges, per-slot connectivity dots and hints, configured counter, one-click verify), model library section keeps the existing table, filters, add/edit dialogs and all data flows

* feat: redesign model library as custom list (v0 design) - replaces the antd Table with lightweight rows (name + default badge, type badge, provider, connectivity dot with spinner, icon actions with tooltips), custom pagination, and a simplified toolbar (search + type + provider). Moves filtering/paging state into ModelLibraryList; drops the context/max-output columns and status filter per the design

* feat: batch edit / delete by connection group (v0 design) - new ModelManagerDialog groups the library by connection (source + API key + base URL); batch edit rotates key/URL for a whole group via partial updates, batch delete supports per-model checkboxes with group select-all. Extracts shared default-slot cleanup so deleted models always vacate their slots

* feat: one-click verify checks default-slot models only (v0 design) - the button lives in the default-config section, so it now delegates to the selection-based verifier instead of probing the whole library; non-default models stay checkable via the per-row action

* feat: rebuild the model config redesign on the project's shadcn/ui primitives - the first pass used antd components whose visual language diverges from the v0 design. Adds radix-based Select/Card/Label ui components and rewrites ModelSlotSelect, ModelLibraryList and ModelManagerDialog on them (quiet borders, status dots inside selects, custom pagination, connection-group cards), matching the v0 aesthetic

* fix: drop the per-slot status line under the select - the v0 design conveys connectivity via the dot inside the select only, and the raw status value (available) leaked untranslated text

* style: widen the gap between the status dot and model name in slot selects - the radix trigger renders selected content with its own tight spacing, wrapping dot+name in one flex span keeps dropdown and trigger consistent

* style: align section titles with the page header - add px-2 to the redesign content to match CARD_HEADER.PADDING (8px), so 默认配置 lines up with 模型设置

* feat: full-height layout for the models page - the default-config card keeps its natural height, the library section fills the remaining height with the list as an internal scroll area (toolbar and pagination stay fixed); rows keep adaptive height and 8 items per page as the default

* fix: legend shows all three status dots with labels (available/unavailable/not-detected) instead of a single green dot labeled as connected

* style: legend keeps a single green dot, reworded to 'green means connected'

* style: widen the model library row columns (model 256px / type 160px / source 128px / actions 128px) so the flex status column no longer leaves a wide gap before the actions

* style: cap the models page content column at 1440px - the redesign was authored for a ~1150px column; stretched to the shared 1920px container the library rows showed wide empty gaps. Scoped to this page only

* style: prettier

* style: widen models page content column to 1600px

* style: widen models page content column to 1760px

* feat: v0-style add-model dialog (single / batch tabs) and reorder buttons - single tab: provider / type / name / display name / base url / key; batch tab: provider + key + fetch list with checkboxes (all selected by default, per-row type badges). No per-row capacity editing or pre-submit connectivity gate (v0 logic) - models land not_detected and are checked/edited from the library. Add-model button moves to the first position; sync-ModelEngine opens the batch tab. The old ModelAddDialogV2 stays for edit mode.

* fix: batch add no longer pre-selects all fetched models - user picks what to import

* feat: search filter for the batch-add fetched model list - filters display only, selections survive searching; toggle-all applies to visible rows

* feat: per-row advanced settings in batch add - gear button on each fetched model opens a dialog with capacity fields (context/input/output/reserve tokens) and inference params (temperature, top_p, enable_thinking, custom params via ModelAdvancedSettings). Overrides apply on submit; a dot indicator marks rows with overrides

* fix: per-row advanced settings now loads real inference specs so temperature/top_p/thinking/custom-params actually render, matching the old dialog's field set (capacity + inference, no display name)

* feat: batch add auto-fills capacity from suggestCapacity on fetch, per-row + batch connectivity check (optional, no submit gate) - suggestions are local catalog lookups applied automatically; user-modified overrides win over suggestions; blue dot only marks user edits

* fix: advanced settings pre-fills catalog suggestions - key={settingsRowId} forces remount so useState initializes from the correct row's suggestions instead of sticking with the first mount's empty state

* fix: remove duplicated capacity fields in per-row settings - ModelAdvancedSettings in override mode already renders capacity + inference, so the manual capacity grid was redundant; suggestions now flow directly into the component value (snake_case keys)

* fix: hide tokenizer_family from the per-row advanced settings

* feat: advanced settings (capacity + inference params) in the single-add tab - gear button opens the same RowSettingsDialog used by batch add; overrides apply on submit

* feat: batch add provider dropdown includes a custom option - custom requires Base URL (validated on fetch), modelFactory set to OpenAI-API-Compatible

* fix: capacity coverage warning now references the redesigned UI (edit button instead of removed manage button)

* fix: simplify capacity warning text to avoid JSX parser edge case with special characters

* feat: require connectivity check before batch add + remove capacity coverage banner - submit blocked until all selected models pass the probe (same gate as the old dialog); banner referenced removed UI and the gate makes it redundant

* fix: remove dangling JSX fragments from banner deletion

* fix: show human-readable type labels (not raw ids like llm/vlm) in batch delete rows

* fix: embedding models group with their parent connection - strip /embeddings suffix when computing the group key so BAAI/bge-m3 joins the SiliconFlow group instead of forming its own

* fix: show '已配置' instead of misleading '未设置' when the list API doesn't return the key value

* fix: remove redundant API key display from connection group cards - every model in the list necessarily has a key configured, so showing a label for it is noise; only the URL line remains

* fix: also strip trailing slashes in connection grouping - v1/ and v1 were different keys after /embeddings removal

* style: rename 选整条 to 全选

* feat: batch delete models are collapsed by default with expand/collapse toggle (same as batch edit)

* fix: batch add carries the verified connectivity status into the created record - the gate requires all selected models to pass the probe, so connectStatus: available is sent instead of defaulting to not_detected

* feat: per-model display name in batch add settings dialog - optional field at the top, empty means auto-generate (model name + random suffix)

* feat: v0-style model edit dialog - single-page form with basic info (display name, readonly model name, type/provider badges, API key, Base URL) and expandable advanced config (capacity + inference params). Replaces the old tab-based ModelAddDialogV2 for edit mode; connectivity check intentionally dropped (use library row check / one-click verify)

* fix: hide tokenizer_family from the edit dialog advanced config

* style: edit dialog advanced config now uses v0-style fields with hints (如 131072 placeholders, range descriptions) instead of the shared ModelAdvancedSettings component

* feat: add deep thinking toggle and custom params section to the edit dialog advanced config

* style: remove descriptive hint text from edit dialog advanced config fields

* fix(model): probe image-generation (vlm2) models via /images/generations

Image-generation models are not served on the chat-completions endpoint,
so the shared VLM chat probe reported them as "model does not exist".
Split vlm2 out of the chat-probe group: the free provider-catalog check
runs first, then a minimal real generation against
{base_url}/images/generations (2xx = connected, 60s default timeout).

* fix(model): carry capacity/inference params through add and edit flows

The add dialogs merged buildInferenceParamsPayload output (snake_case)
into create params, but addCustomModel's buildCapacityRequestBody only
reads camelCase keys - catalog capacity suggestions were silently dropped
on create (rows landed with NULL capacity). Mirror the explicit mapping
ModelAddDialogV2 does: capacity fields to camelCase, max_output mirrored
into max_tokens, capacitySource "operator" when capacity is sent.

The edit dialog also only pre-filled temperature/top_p/custom params -
its init passed an empty spec list to advancedSettingsValueFromRecord, so
nothing loaded. Build the value directly from the model record (capacity,
inference params, enable_thinking, custom params) so the advanced config
reflects the stored configuration, and re-send capacitySource/maxTokens
on save for parity with the old edit dialog.

frontend-only companion to the vlm2 probe fix: unify the advanced-config
section into a shared ModelAdvancedConfig component (v0 field grid) used
by both the edit dialog and the add-dialog advanced settings; add
per-row type editing in batch add and type editing + in-dialog
connectivity probe in the edit dialog (probe carries inference params so
invalid custom params fail at probe time); align the default badge with
the display-name line in the library list.

* style: prettier modelService after develop merge

* feat(model): reasoning effort / budget controls in the v0 advanced config

Port the reasoning-effort feature (#3953 on develop) onto the v0-style
ModelAdvancedConfig so the models-page dialogs expose the same controls
as the legacy ModelAdvancedSettings: below the deep-thinking switch,
render the effort-level select or the budget-tokens input when the
catalog-declared capability provides one (budget takes precedence over
effort, mirroring the payload builder). Toggling thinking off clears
both keys so they are dropped from the wire payload.

Capability wiring:
- edit dialog: seeded from the model record, refreshed (debounced) via
  suggestCapacity when the URL or type changes
- batch add: captured per row from the existing suggestCapacity lookup
  on fetch and passed into the per-row settings dialog
- single add: looked up (debounced) while the settings dialog is open

* test(model): cover the vlm2 probe network-error branch

Codecov patch coverage flagged the except path of
_image_generation_connectivity_check as uncovered; add a transport-error
test so the v0 PR diff clears the 90% patch target.

* refactor(model): dedupe type maps and reasoning resolution for Sonar gate

The SonarCloud quality gate failed on new-code duplication (5.3% > 3%)
and cognitive complexity (18 > 15 in ModelAdvancedConfig). Consolidate the
duplicated blocks:

- new modelTypeUi.ts holds the single copy of TYPE_LABEL_KEY_MAP,
  TYPE_BADGE_CLASS, STATUS_DOT_CLASS, typeLabel and useTypeOptions
  (previously copied across four dialogs) plus a shared ProviderSelect
  for the add-dialog tabs
- resolveReasoningControls / clampReasoningBudget extracted from
  ModelAdvancedSettings and reused by ModelAdvancedConfig, replacing the
  ported inline block (fixes the cognitive-complexity failure and the
  nested-ternary warnings)
- normalizeUrlForGrouping drops the backtracking regex Sonar flagged
- clean unused imports (Alert, Badge, DialogFooter, ChevronUpIcon,
  useMemo) and add a keyboard stopPropagation to the per-row type editor
  wrapper span

* feat(model): probe gate and auto-detection in the single-add form

The batch tab gates submit on a passing connectivity probe and auto-fills
capacity / reasoning capability from suggestCapacity; the single-add form
did neither - models landed with a hardcoded 4096 max_tokens placeholder
and unverified credentials.

- add an in-form 检测连通性 button with inline result; submit is blocked
  until the probe passes (same gate as batch) and the verified status is
  carried into the created record
- run the catalog lookup (debounced) whenever name + URL are filled:
  capacity prefill and reasoning capability flow into the gear dialog as
  suggestions (user overrides still win) and a compact "已自动识别" hint
  shows what was detected (context window / max output / thinking mode)
- any name / type / URL / key / settings change resets the probe
- extract capacitySuggestionToSettings shared by both tabs

* fix(model): disable browser autofill on the single-add form inputs

* fix(model): reset manual advanced settings when the single-add model name changes

* fix(model): seed enable_thinking for LLMs in the v0 advanced config

The payload builder drops reasoning_effort / reasoning_budget_tokens when
enable_thinking is not strictly true, but the v0 switch renders undefined
as ON - a user who only changed the effort level (never toggling the
switch) saved nothing. Materialize the default (true) for LLMs the same
way ModelAdvancedSettings does.

* refactor(model): express the default-model slots as a compact seed table

buildModelSlots' ten inline slot literals read as 21-46 line duplicated
blocks to Sonar (37.6% duplication on the file, the last chunk keeping
the PR above the 3% gate). Move the per-slot metadata into a
MODEL_SLOT_SEEDS table whose rows differ on every line and expand it in
buildModelSlots; behaviour is unchanged.

* chore: restore tsconfig.json after a local next build rewrote it

* refactor(model): shrink remaining Sonar duplication after the develop merge

Sonar's copy-paste detector normalizes string literals, so the slot seed
objects still matched each other in bulk (25.6% on ModelSlotSelect) and
the #4009 strip/seed effects I mirrored into ModelAdvancedConfig
duplicated ModelAdvancedSettings.

- ModelSlotSelect: split the long i18n defaults into SLOT_LABELS /
  SLOT_HINTS records with a compact tuple table for the structural fields
- extract useReasoningFormEffects (capability-gated visibility + legacy
  strip + enable_thinking seed) shared by both advanced-config components

* fix(model): gate the probe key fallback on the stored endpoint

The temporary-healthcheck fallback substituted a stored api_key while the
caller still controlled base_url, so any authenticated tenant member could
exfiltrate another model's stored key by probing it against their own
server (the key rides the Authorization header to the caller-chosen URL).

Only borrow the stored key when the probe targets the stored endpoint
(exact base_url match after stripping trailing slashes). Edit-dialog
probes with an untouched URL keep working; a changed URL with an empty
key now honestly fails instead of silently sending the old key to the
new address.

* fix(model): stop auto-suffixing display names in the v0 add dialogs

develop #4009 removed the automatic model renaming from the legacy V2
dialog but our v0 ModelAddDialog still appended a random 5-char suffix to
every batch-imported display name (kimi-k3j72bi). Use the clean model
name as the default display name, matching develop; a collision with an
existing row surfaces as a per-row failure the user can resolve with the
gear's display-name field instead of silently creating suffixed
duplicates.

* feat(model): make the model name editable in the v0 edit dialog

The edit dialog kept the model name read-only; allow editing it like
the type. The connectivity probe uses the edited name, the reasoning
capability lookup re-runs (debounced) when it changes, and saving a
changed name sends model_name plus a not_detected status reset so the
renamed record must be re-verified.

* fix(model): address review findings in the v0 model-config dialogs

Selective fixes for the Copilot review findings (outdated or intentional
items skipped):

- batch add: keep the dialog open when every create fails instead of
  closing with only a warning; changing the shared API key or base URL
  now clears all per-row probe results (they were measured with the
  previous credentials)
- batch delete: require an explicit confirmation before the destructive
  batch (states the count and the default-slot vacation)
- batch-add rows: render the selectable row as a div instead of a button
  — the row hosts a Select and icon Buttons, and nesting interactive
  elements inside <button> is invalid HTML; keyboard toggle preserved
- slot selects: associate the visible label with the select trigger
  (htmlFor/id) for screen readers
- remove the dead capacity-coverage fetch chain (state, request-id ref
  and type import) left over from the removed warning banner

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
* fix: preserve agent trace hierarchy and improve monitoring metadata

Trace the full agent lifecycle, sandbox callbacks, child agents, and conversation title generation. Use user emails and agent display names, filter noisy spans, and add observability regression coverage.

Include the current Docker build adjustments.

* Trace sandbox warmup with agent monitoring metadata

* Isolate lazy human interaction service in stop endpoint test

---------

Co-authored-by: hhhhsc <name>
* feat: enforce knowledge base resource limits

* test: cover knowledge resource limit branches

* fix: align knowledge resource limits with latest develop

* config: expose knowledge resource limits via environment

* test: improve knowledge limit patch coverage
* feat: add agent sharing backend foundation

* feat: expose northbound API base URL

* feat: guide users to published agent usage

* feat: add agent usage guide modal

* feat: support scoped runtime rate limits

* feat: add authenticated agent share session APIs

* feat: isolate agent share run identities

* feat: protect agent share run lifecycle

* feat: add authenticated agent share page

* feat: protect agent share responses from caching

* feat: secure all agent share responses

* feat: rate limit agent share runs

* test: cover isolated agent share runtime sessions

* test: cover agent share revocation during runs

* fix: sanitize agent share management errors

* test: cover agent share session history isolation

* test: keep shared sessions out of conversation lists

* feat: link northbound guide to API key settings

* feat: localize agent share page feedback

* feat: retry unavailable A2A guide settings

* feat: restore published agent usage guide targets

* feat: complete agent share guide lifecycle

* test: lock agent publish guide outcomes

* feat: guard and restore agent share pages

* feat: improve agent guide accessibility

* chore: align agent share validation style

* fix: accept uuid share ids in token resolution

* feat: proxy northbound api and relocate runtime config

* test: add agent usage guide component coverage

* chore: ignore local openspec skill copies

* fix: route agent share APIs to runtime

* fix: localize shared agent input

* fix: redact agent share tokens from access logs

* fix: apply share token redaction to uvicorn access logs

* fix: configure runtime access log redaction

* fix: preserve agent share login return

* fix: support share runs without history

* fix: correct northbound usage guide calls

* feat: guide published agents through card menu

* feat: share agents through chat deep links

* fix: open agent deep links in current chat

* fix: open agent deep links directly

* fix: restore agent repository access and usage guide deep-link parsing

* fix: wait for agent usage guide activation

* fix: refresh editable agents after publish

* fix: stabilize agent guide checks

* refactor: remove standalone agent share runtime

* fix: satisfy runtime and architecture unit tests

* fix: remove unrelated deployment changes

* test: cover agent share runtime paths

* refactor: focus agent usage guide on frontend

* fix: use published agent name in northbound guide

* fix: stabilize northbound frontend proxy

* remove .agents
…rtain scenarios (#4022)

* bugfix: fix infinite loop when agent configuration update fails in certain scenarios

* bugfix: fix infinite loop when agent configuration update fails in certain scenarios

---------

Co-authored-by: hzw <hzw@qq.com>
* fix(agent): preserve model attempt streaming across semantic retries

(cherry picked from commit e38fd50)

* feat(agent): make strict output repair opt-in per agent

* fix(agent): continue bare output and reinforce action protocol

* test(agent): keep context reminder fixture lint clean

* Restore legacy code action routing and accept final strict attempt

* Gate empty model response retries on strict code action mode

* Show a quiet final hint for legacy empty model responses

* test(agent): adapt protocol regression tests to develop

* fix(agent): align protocol tests and preserve guardrail user boundary
…, bump version (#4024)

* merge v2.6.1 hotfix release from hotfix/v2.6.1 (#3971)

* 🐛 Fix(evaluation): run trials in runtime service (#3954)

* fix(evaluation): run trials in runtime service

Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub config thread manager

Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub runtime jwt helper

* test(evaluation): cover trial proxy error paths

* Fix/override delete (#3958)

* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)

* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key

* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation

* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)

* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it

* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)

* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"

This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.

* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)

* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)

* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests

Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.

Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.

* fix(hitl): guarantee observer chunks precede human_interaction in DB event order

When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.

Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.

Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.

* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths

Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.

Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.

* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety

Fix 4 real issues flagged by github-code-review:

1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
   reset on every drain, meaning a model that kept producing chunks could
   stall the worker forever. deadline is now computed once at entry and
   the sleep call clips to hard_deadline - now.

2. _emit_in_flight Event bridges the async emit path and the worker's
   idle poll. Without this, buffer-empty = 'persisted' was confused with
   buffer-empty = 'taken for emit but still in run_blocking queue'. The
   worker now checks both 'buffer empty for settle_ms' AND 'no emit in
   flight' before deciding the async side is truly idle.

3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
   Previously the async loop drained the buffer, decided it was not yet
   due, then put everything back. That transiently-empty window (16 us
   normally, arbitrarily long under GIL/GC/preemption) was enough for
   the worker's 20 ms poll to mis-fire. We now peek (read count, no
   drain) and only take_chunks when we actually intend to persist.

4. emit_chunks wrapped in try/except that puts drained chunks back into
   the shared buffer before re-raising, and finish() wraps its flush call
   in try/except: pass. Guarantees (a) no chunk loss on DB failure and
   (b) the terminal human_run row is always written even if the flush
   step fails.

Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
  hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
  try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
  test_application_execute_attempt.py drives the full consumer loop
  through _flush_if_due (peek → take → begin/end_emit) and the final
  flush, verifying that every patch line added in application.py is hit.

* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected

When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.

* perf(hitl): replace polling with SSE subscription and move snapshot off the write path

Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
  subscription to the backend's /{run_id}/events SSE stream. Discovery is
  now one-shot: conversationId change and the agent stream pause
  (isRunning true→false), the exact moment a HITL run is most likely to
  exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
  adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
  our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
  concurrent callers and returns cached state; (2) 3s minimum interval
  so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
  not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
  to prevent EventSource from reconnecting forever against COMPLETED runs.

Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
  WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
  must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
  writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
  calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
  before processing each request, and expire_waiting() remains the
  periodic scheduler sweep.

Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.

* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false

After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).

The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.

Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.

Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.

* test(hitl): update mock from snapshot to light_snapshot after read-path refactor

test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.

* test(hitl): raise diff coverage above the 90% merge gate

Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate

SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.

* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm

openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.

* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"

This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.

* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.

* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* [codex] fix(agent): silently retry transient model errors (#3965)

* fix(agent): retry transient model failures silently

* fix(model): support OpenAI httpx2 timeout client

* test(model): add deterministic OpenAI-compatible mock

* fix(agent): keep stream runtime within line budget

* Fix: AIDP knowledge base bug fix (#3967)

* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection

* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change

---------

Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>

* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)

* fix(agent): enforce explicit CodeAgent termination

* fix(test): restore CodeAgent CI compatibility

* merge(main): merge v2.6.1 hotfixes from hotfix/v2.6.1 (#3983)

* 🐛 Fix(evaluation): run trials in runtime service (#3954)

* fix(evaluation): run trials in runtime service

Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub config thread manager

Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub runtime jwt helper

* test(evaluation): cover trial proxy error paths

* Fix/override delete (#3958)

* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)

* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key

* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation

* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)

* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it

* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)

* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"

This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.

* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)

* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)

* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests

Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.

Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.

* fix(hitl): guarantee observer chunks precede human_interaction in DB event order

When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.

Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.

Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.

* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths

Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.

Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.

* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety

Fix 4 real issues flagged by github-code-review:

1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
   reset on every drain, meaning a model that kept producing chunks could
   stall the worker forever. deadline is now computed once at entry and
   the sleep call clips to hard_deadline - now.

2. _emit_in_flight Event bridges the async emit path and the worker's
   idle poll. Without this, buffer-empty = 'persisted' was confused with
   buffer-empty = 'taken for emit but still in run_blocking queue'. The
   worker now checks both 'buffer empty for settle_ms' AND 'no emit in
   flight' before deciding the async side is truly idle.

3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
   Previously the async loop drained the buffer, decided it was not yet
   due, then put everything back. That transiently-empty window (16 us
   normally, arbitrarily long under GIL/GC/preemption) was enough for
   the worker's 20 ms poll to mis-fire. We now peek (read count, no
   drain) and only take_chunks when we actually intend to persist.

4. emit_chunks wrapped in try/except that puts drained chunks back into
   the shared buffer before re-raising, and finish() wraps its flush call
   in try/except: pass. Guarantees (a) no chunk loss on DB failure and
   (b) the terminal human_run row is always written even if the flush
   step fails.

Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
  hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
  try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
  test_application_execute_attempt.py drives the full consumer loop
  through _flush_if_due (peek → take → begin/end_emit) and the final
  flush, verifying that every patch line added in application.py is hit.

* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected

When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.

* perf(hitl): replace polling with SSE subscription and move snapshot off the write path

Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
  subscription to the backend's /{run_id}/events SSE stream. Discovery is
  now one-shot: conversationId change and the agent stream pause
  (isRunning true→false), the exact moment a HITL run is most likely to
  exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
  adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
  our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
  concurrent callers and returns cached state; (2) 3s minimum interval
  so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
  not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
  to prevent EventSource from reconnecting forever against COMPLETED runs.

Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
  WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
  must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
  writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
  calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
  before processing each request, and expire_waiting() remains the
  periodic scheduler sweep.

Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.

* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false

After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).

The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.

Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.

Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.

* test(hitl): update mock from snapshot to light_snapshot after read-path refactor

test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.

* test(hitl): raise diff coverage above the 90% merge gate

Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate

SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.

* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm

openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.

* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"

This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.

* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.

* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* [codex] fix(agent): silently retry transient model errors (#3965)

* fix(agent): retry transient model failures silently

* fix(model): support OpenAI httpx2 timeout client

* test(model): add deterministic OpenAI-compatible mock

* fix(agent): keep stream runtime within line budget

* Fix: AIDP knowledge base bug fix (#3967)

* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection

* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change

* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)

* fix(agent): enforce explicit CodeAgent termination

* fix(test): restore CodeAgent CI compatibility

* cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981)

* Fix StopAsyncIteration leak in execute_attempt finally block

Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors.

Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate.

* Fix HITL form not appearing until page refresh

Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload.

Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh.

* Fix SSE chunk/HITL event ordering race under real server load

Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order.

Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order.

* fix(hitl): recover pending form when SSE delivery stalls silently

Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones.

Changes:
- Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls
- Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it
- Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop

* fix(hitl): stop parked human-input waits from consuming scheduler slots

Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed.

Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments.

Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots.

Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43.

* test(scheduler): fix flaky waiting-concurrency assertion on fast event loops

Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently.

Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.

* Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980)

The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which
also carried two unrelated upstream changes that this release branch never
had:

  - #3909 Knowledge base interface optimization (AIDP UI refactor)
  - #3930 support AIDP knowledge file deletion and download

Both are removed here so the release line keeps only the bug fix.

  - Restore the AIDP frontend components to their pre-refactor layout and
    drop the helper modules only the refactor used:
    AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils.
  - Drop the #3930 document remove/download endpoints from services/api.ts
    and the AIDP translations that only those screens referenced.
  - Keep the fix itself unchanged: knowledge-base scoped Channels and
    KnowledgeFiles/History paths, all-status document listing, keyword
    search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the
    upload-triggered polling.

Verified:
  - pytest test/ext_components/aidp -q  ->  583 passed
  - frontend `npm run type-check` (tsc --noEmit)  ->  no errors

---------

Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>
Co-authored-by: chase <byzhangxin11@126.com>

* merge(main): merge v2.6.1 hotfixes from hotfix/v2.6.1 (#3993)

* 🐛 Fix(evaluation): run trials in runtime service (#3954)

* fix(evaluation): run trials in runtime service

Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub config thread manager

Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub runtime jwt helper

* test(evaluation): cover trial proxy error paths

* Fix/override delete (#3958)

* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)

* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key

* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation

* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)

* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it

* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)

* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"

This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.

* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)

* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)

* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests

Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.

Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.

* fix(hitl): guarantee observer chunks precede human_interaction in DB event order

When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.

Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.

Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.

* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths

Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.

Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.

* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety

Fix 4 real issues flagged by github-code-review:

1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
   reset on every drain, meaning a model that kept producing chunks could
   stall the worker forever. deadline is now computed once at entry and
   the sleep call clips to hard_deadline - now.

2. _emit_in_flight Event bridges the async emit path and the worker's
   idle poll. Without this, buffer-empty = 'persisted' was confused with
   buffer-empty = 'taken for emit but still in run_blocking queue'. The
   worker now checks both 'buffer empty for settle_ms' AND 'no emit in
   flight' before deciding the async side is truly idle.

3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
   Previously the async loop drained the buffer, decided it was not yet
   due, then put everything back. That transiently-empty window (16 us
   normally, arbitrarily long under GIL/GC/preemption) was enough for
   the worker's 20 ms poll to mis-fire. We now peek (read count, no
   drain) and only take_chunks when we actually intend to persist.

4. emit_chunks wrapped in try/except that puts drained chunks back into
   the shared buffer before re-raising, and finish() wraps its flush call
   in try/except: pass. Guarantees (a) no chunk loss on DB failure and
   (b) the terminal human_run row is always written even if the flush
   step fails.

Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
  hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
  try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
  test_application_execute_attempt.py drives the full consumer loop
  through _flush_if_due (peek → take → begin/end_emit) and the final
  flush, verifying that every patch line added in application.py is hit.

* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected

When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.

* perf(hitl): replace polling with SSE subscription and move snapshot off the write path

Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
  subscription to the backend's /{run_id}/events SSE stream. Discovery is
  now one-shot: conversationId change and the agent stream pause
  (isRunning true→false), the exact moment a HITL run is most likely to
  exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
  adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
  our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
  concurrent callers and returns cached state; (2) 3s minimum interval
  so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
  not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
  to prevent EventSource from reconnecting forever against COMPLETED runs.

Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
  WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
  must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
  writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
  calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
  before processing each request, and expire_waiting() remains the
  periodic scheduler sweep.

Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.

* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false

After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).

The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.

Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.

Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.

* test(hitl): update mock from snapshot to light_snapshot after read-path refactor

test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.

* test(hitl): raise diff coverage above the 90% merge gate

Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate

SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.

* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm

openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.

* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"

This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.

* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.

* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* [codex] fix(agent): silently retry transient model errors (#3965)

* fix(agent): retry transient model failures silently

* fix(model): support OpenAI httpx2 timeout client

* test(model): add deterministic OpenAI-compatible mock

* fix(agent): keep stream runtime within line budget

* Fix: AIDP knowledge base bug fix (#3967)

* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection

* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change

* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)

* fix(agent): enforce explicit CodeAgent termination

* fix(test): restore CodeAgent CI compatibility

* cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981)

* Fix StopAsyncIteration leak in execute_attempt finally block

Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors.

Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate.

* Fix HITL form not appearing until page refresh

Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload.

Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh.

* Fix SSE chunk/HITL event ordering race under real server load

Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order.

Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order.

* fix(hitl): recover pending form when SSE delivery stalls silently

Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones.

Changes:
- Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls
- Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it
- Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop

* fix(hitl): stop parked human-input waits from consuming scheduler slots

Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed.

Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments.

Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots.

Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43.

* test(scheduler): fix flaky waiting-concurrency assertion on fast event loops

Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently.

Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.

* Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980)

The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which
also carried two unrelated upstream changes that this release branch never
had:

  - #3909 Knowledge base interface optimization (AIDP UI refactor)
  - #3930 support AIDP knowledge file deletion and download

Both are removed here so the release line keeps only the bug fix.

  - Restore the AIDP frontend components to their pre-refactor layout and
    drop the helper modules only the refactor used:
    AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils.
  - Drop the #3930 document remove/download endpoints from services/api.ts
    and the AIDP translations that only those screens referenced.
  - Keep the fix itself unchanged: knowledge-base scoped Channels and
    KnowledgeFiles/History paths, all-status document listing, keyword
    search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the
    upload-triggered polling.

Verified:
  - pytest test/ext_components/aidp -q  ->  583 passed
  - frontend `npm run type-check` (tsc --noEmit)  ->  no errors

* fix(agent): accept reasoning-prefixed code actions (#3990)

* refactor: remove human interaction features and related configurations (#3988)

* refactor: remove human interaction features and related configurations

* refactor: remove human interaction features and related configurations

---------

Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.co…
…time (#4016)

* Fix: keep ingested AIDP documents visible when the channel history is used

The document list switched data sources instead of combining them: while the
resolved channel directory was empty the list was served from the knowledge-base
scoped listing, and as soon as one upload landed in that directory the all-status
history took over. The history only covers the resolved channel directory, so
every file ingested into another directory disappeared from the list right after
an upload — reported on 2.6.1 for files that were created before the upgrade.

Merge the two sources instead:

- The knowledge-base scoped listing guarantees membership, the history supplies
  the live statuses, and a file only the history knows about (still uploading,
  or failed before ingestion) is kept exactly as reported.
- Items are matched through every identity they expose, so a file that one
  payload describes with a uuid and the other with an ino number stays one row.
- Files taken from the listing are marked COMPLETED, which is all that listing
  ever returns.

The listing is read in pages of 100, capped at 20 pages, and a failing page
degrades to the files read so far (logged) instead of failing the request.

Verified: pytest test/ext_components/aidp -q -> 586 passed, including new
coverage for ingested files outside the channel directory, a history-only
upload, and a failing listing.

* Fix: read AIDP file history across pages so bursts of uploads stay visible

The history endpoint is paginated and sorts files that are still being processed
to the front. Reading only the first page therefore drops exactly the files the
document list exists to show: more simultaneous uploads than fit in one page push
the rest out of view.

- Page the history (body `page`, one-based): the walk stops at an empty page, at
  an explicit "no next page" signal, or on a page that adds nothing new. The last
  one keeps a build that ignores `page` from looping over the same files, and
  logs that everything beyond page one stays invisible.
- Cap the walk at 20 pages and log when the cap truncates, so one list request
  cannot turn into an unbounded number of upstream calls.
- Trust only unambiguous pagination signals: this endpoint family reports
  `total_count` as the size of the current page elsewhere, so it is not read as a
  grand total.
- The mock now mirrors the real contract: `page` in the body, ten entries per
  page (tunable through POST /_mock/history-page-size), files that are still
  being processed listed first, and `total_count`/`next_link` in the response. A
  status pinned through POST /_mock/doc-status also stays PROCESSING instead of
  advancing to COMPLETED, so that stage can be inspected locally.

Verified: pytest test/ext_components/aidp -q -> 588 passed, including a burst of
uploads spanning three history pages; end-to-end against the mock -> 10/10 checks
covering page order, page metadata and the last page.

* Fix: show AIDP files still processing first and keep listed file metadata

Order the AIDP document list so files that are still uploading or extracting sit above the finished ones, which keeps an upload's progress on page 1 instead of behind files that were already ingested. An unrecognised stage counts as open as well, so it neither stops the polling nor drops behind the ingested rows.

Merge the channel history with the document listing field by field: the history still supplies the live status, but a payload that omits metadata (name, size, creation time) no longer blanks out the fields the listing already carries.

* Fix: keep the creation time AIDP history reports as created_at

The history endpoint reports the creation time under the canonical created_at (ISO string), but normalization only read first_upload_time / create_time and overwrote created_at with a failed conversion, which left the column empty for every history row. Both spellings are now accepted, numeric or ISO, for created_at and updated_at.

* Fix: read the AIDP document creation time through every spelling reported

A blank field means 'not reported' in AIDP payloads, so an empty create_time used to shadow a populated canonical created_at and the history rows lost the column. Blank values are skipped now, the listing's upload stamp and the canonical created_at are both accepted, and update_time closes the chain for files AIDP has not finished registering yet.
* 调整mcp添加mcp和编辑页面,与mcp配置页面保持一致;
mcp容器化启动点击后就展示日志,配置页面和仓库页面都修改;
修复mcp服务更新后,无法刷新工具的问题;

* mcp仓库增加分页;
fix test case

* fix test case

* fix test case

* Update custom.json

* 修复本地镜像上传部署日志显示问题

* fix bug

* fix bug

* add test case

* add test case

* 删除mcp市场

* 优化前端页面

* Update custom.json

* Normalize custom locale file endings

* Fix MCP registry endpoint typing

* Remove obsolete MCP registry frontend code

* 修改本地镜像上传MCP错误提示

* 修复mcp提交审核后不及时显示问题

* 修复api转mcp服务切换租户添加会导致原有工具列表被清空的问题

* 修复API转MCP后工具列表未同步

* 恢复工具手动刷新逻辑

* 避免API转MCP工具因扫描失败被隐藏

* docs(agents): define bundled official agent design

Document industry-specific official agent packages, reserved platform templates, startup synchronization, repository visibility, and tenant copy permissions.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs(agents): add bundled agent implementation plan

Break the approved industry-specific official agent design into test-first backend, repository, frontend, and deployment tasks.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs(agents): add Chinese design and plan

Provide Chinese versions of the approved bundled official agent design and implementation plan.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(agents): add official bundle sync foundation

Support multi-profile official agent bundles with safe ZIP loading and idempotent synchronization into the reserved platform repository.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(agents): expose official listings in repository

Merge selected official listings into repository reads, keep platform templates read-only through tenant scoping, and trigger synchronization at startup or through the deployment script.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(agents): mark official listings in clients

Show official repository entries as read-only templates in the UI and inject selected bundle directories into Docker and Kubernetes deployments.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(deploy): select official agent profiles offline

Allow offline package builds to retain only explicitly selected official agent profiles and emit the runtime official-agent configuration.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agents): isolate bundle skill validation

Resolve official bundle skill entries with a local model so repository service test doubles and runtime imports cannot corrupt the Pydantic forward reference.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(deploy): sync official agents during installation

Run official agent synchronization after Docker and Kubernetes deployments become ready, while retaining startup synchronization as an idempotent fallback.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(deploy): let users select agent categories

Fetch the official agent catalog during installation, scan its category directories, and persist the user's multi-category selection for Docker and Kubernetes deployment.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(agents): install official repository dependencies

Route official repository imports through tenant-level MCP, Skill, and knowledge-base preparation while preserving ordinary imports. Add ZIP and directory bundle loading and retain the current Agent service with only the required skill-link helpers.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agents): resolve bundles by repository agent name

Use the official repository name as the bundle directory key and restore the normal listing content value. Keep bundle lookup aligned with the contract that the bundle folder and root Agent name are identical.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(deploy): separate official agent deployment

Move official Agent resource selection and synchronization out of the main Nexent installation flow. Add a standalone Docker/Kubernetes/local deployment entry, keep official repository visibility independent of startup profiles, and preserve tenant-level installation for user copies.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(deploy): mark official agent entry executable

Allow the standalone official Agent deployment entry to run directly in Git-based deployment environments.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(deploy): accept host paths for official agents

Copy local official Agent resources through the running config container and explicitly synchronize the mounted container directory. Remove the conflicting read-only submount so Windows host paths can be used after Nexent is deployed.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(deploy): preserve container paths in git bash

Disable MSYS path conversion for Docker and Kubernetes sync arguments so mounted Linux paths are not rewritten into Windows paths.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(repository): resolve official listings during import

Allow repository prechecks and imports to fall back to the reserved official tenant so shared official listings can be copied by business tenants.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(repository): report official import requirements

Build official repository precheck items from the mounted bundle so copy dialogs show model, knowledge base, MCP, Skill, and tool availability instead of an empty list.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agents): keep official bundle keys stable

Store the Bundle directory key in official repository listings and repair existing official records during synchronization so prechecks and imports reload the correct mounted Bundle.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agents): resolve bundles inside profiles

Allow official installation and precheck to locate directory, JSON, and ZIP bundles nested under selected profile directories.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agents): ignore skill archives during bundle scan

Keep Skill ZIP payloads from being misclassified as standalone official Agent bundles during profile synchronization.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agents): diagnose official bundle lookup

Log the resolved official bundle path and validation target so runtime lookup failures can be distinguished from missing mounted resources.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(deploy): pin official bundle path in container

Override host-local OFFICIAL_AGENTS_PATH values from backend/.env with the mounted container path so repository prechecks can resolve official bundles.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(repository): choose tenant models during official copy

Show tenant-configured language and embedding models in the repository copy dialog and pass the selected IDs through the official install path.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(repository): preserve official flag for model selection

Expose the publisher tenant in repository summaries and use it as a fallback so official copy dialogs always show tenant model selectors.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(repository): preserve publisher on summary listings

Restore the publisher tenant on lightweight repository records so official listings reach the copy dialog with the official flag.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agents): apply official skill resolutions

Honor the copy dialog choice to reuse an existing tenant Skill or create a renamed Skill when installing an official agent.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): ensure dependencies before reuse

Prepare official knowledge bases and other dependencies even when the root agent already exists, so partial installs can be repaired instead of being skipped. Add runtime logs for official repository routing and bundle dependency counts.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agents): honor skill conflict resolutions

Evaluate official Skill reuse or rename choices before raising duplicate errors, so dependency preparation can continue to knowledge base creation and Agent import.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): repair reused knowledge references

Update existing official Agent tool instances after resolving a tenant knowledge base, so earlier partial installs no longer keep the bundle-local kb key.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): isolate user agent copies

Treat official dependencies as tenant-scoped while keeping copied Agents user-scoped. Create a uniquely named visible copy when another user already owns the bundle's root Agent.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): share tenant knowledge bases

Make official knowledge bases visible to every group in the tenant while keeping them read-only, and show bundle knowledge base display names instead of logical placeholders during import precheck.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(official-agents): label copies by email

Use the current account email in copied official Agent and Skill display names while retaining private technical identifiers for tenant uniqueness.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): hide official publisher label

Hide the publisher text below the official badge for official Agent listings while preserving author display for regular listings.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): reuse existing official embeddings

Hide embedding model selection when all official knowledge bases already exist in the tenant and let the backend reuse their configured models.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): block official copy without embeddings

Show a clear tenant embedding-model requirement and prevent official knowledge-base installation from starting when no usable embedding model is configured.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): validate official kb embedding model

Treat an official knowledge base as reusable only when its tenant embedding model still exists and is available, while leaving ordinary repository imports unchanged.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(deploy): remove official profile env setting

Keep profile selection in the standalone official-agent deployment flow instead of the main environment template.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(deploy): remove legacy official profile variable

Keep official agent profile selection exclusively in the standalone deployment script.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(k8s): remove unused market backend config

Remove the obsolete external market backend address from Kubernetes deployment values and rendering.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(docs): remove bundled agent plans from code repo

Keep bundled default agent design documents in the centralized nexent-doc archive instead of the business repository.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-import): use migrated agent management service

Replace stale services.agent_service imports in the official agent installer with the migrated agent management entry points so repository imports can complete after the agent service refactor.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-import): use migrated knowledge base service

Replace the removed vectordatabase_service import in official agent knowledge-base installation with the migrated knowledge-base and model management services.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): refresh mine agents after copy

Switch to the mine tab and refresh its query after a repository copy succeeds so the created or reused Agent is immediately visible.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): label copied agents with user email

Append the current user's email to official Agent display names for both new installs and idempotent retries.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-import): import model status enum

Allow repository import precheck to evaluate tenant model availability without raising a NameError.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): create unique copies on name conflicts

Match ordinary repository imports by always creating official Agent copies and generating tenant-scoped unique names while preserving tenant dependency reuse.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(official-agents): add default marketplace tags

Populate category tags in official agent bundle metadata so synchronized repository listings and local marketplace submissions have default labels.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): preserve default tags on import

Apply official Bundle category tags to the newly imported tenant Agent after resolving tenant-specific tag values. Keep ordinary repository imports unchanged.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): remove forced default tags

Keep official Agent tags empty so users choose from the tenant tag definitions when applying for listing.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(tags): allow tenant users to read tag definitions

Keep tag library mutations restricted to management roles while allowing authenticated tenant users to load tag choices for repository listing forms.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): render tag limit in hint

Pass the tag limit to the localized hint so the UI renders the numeric value instead of the raw interpolation placeholder.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: align regression tests with migrated services

Update stale mocks, assertions, and service paths after the management-service migration. Cover official agent knowledge-base reuse, skill resolution, MCP headers, tag reads, and asynchronous tool refresh.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs: add official agent deployment guide

Document online, local, offline, Docker, and Kubernetes deployment of official Agent bundles, including profile selection, synchronization, verification, and troubleshooting.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs: document agent export bundle conversion

Add the conversion workflow from Nexent-exported Agent JSON or ZIP to a profile-scoped official bundle, including Skill and knowledge-base seed handling.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: mock tenant MCP endpoint in tool tests

Patch the tenant-scoped local MCP endpoint used by get_all_mcp_tools instead of the obsolete base URL mock.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs: remove official agent deployment guide

Withdraw the deployment guide from the feature branch as requested.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: cover official agent synchronization

Add unit coverage for source-agent materialization, repository upsert, and failure isolation during official bundle synchronization.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: cover official repository import flow

Exercise official listing fallback lookup, bundle precheck, installation outcomes, and download accounting branches.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: cover official agent bundle loading

Exercise profile validation, safe archive handling, skill and knowledge-base loading, and directory bundle synchronization paths.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: stub model parameter filter for official agent tests

Keep the isolated consts.model test module compatible with the naming service import chain used in CI.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: cover official agent service edge cases

Add coverage for bundle discovery, nested archives, resource loading failures, partial KB repair, and existing-agent installation paths.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: cover tenant MCP routing branches

Exercise tenant refresh skips, app construction fallbacks, router authorization and lazy initialization, and MCP startup cleanup paths.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: cover repository import precheck branches

Test embedding model availability and official knowledge-base name fallback during repository import prechecks.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test: cover official agent bundle defaults

Verify official bundle card fields derive from the root agent and fall back to safe defaults when metadata is absent.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(mcp): avoid duplicate tool cleanup on delete

Clean up MCP tools exactly once during deletion and use the shared outer-apis name for OpenAPI services.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(deploy): remove obsolete agent TUI branches

Align common.sh with develop after official agent deployment was separated from the main installation flow.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(official-agents): keep local snapshots out of remote

Stop tracking the two local official agent snapshots while preserving the files for local deployment and testing.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(agents): remove unused attachment fallback

Revert the unrelated MinIO bucket fallback that is not used by official agent installation.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(deploy): remove redundant official agent directory setup

Keep official agent directory creation in the standalone deployment flow and restore deploy.sh to develop behavior.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(deploy): remove obsolete official agent env example

Restore the environment example after official agent deployment moved to the standalone flow.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(deploy): remove official agent helm wiring

Restore Kubernetes deployment templates to develop and keep official agent installation outside the main Nexent deployment flow.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(offline): remove obsolete official agent packaging

Restore offline package generation to develop and remove the unused official-agent profile filtering path.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(agent): manage official agent visibility

Add tenant-scoped visibility overrides and super-admin management APIs for official agent templates. Provide resource-management UI with visibility switch and global template deletion while preserving tenant copies.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* refactor(agent): remove tenant visibility controls

Keep official-agent management global and remove the tenant visibility table, migration, switch, and filtering. Super admins can still delete official templates while tenant copies and dependencies remain untouched.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): restrict official agent management to super admins

Use the authenticated global role for official agent management APIs and hide the management control from non-super-admin users.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): confirm official agent deletion in modal

Replace the inline confirmation with a descriptive modal before deleting the official template and bundle files.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): resolve official knowledge base conflicts

Allow official agent imports to reuse or create a tenant knowledge base when names collide, and pass the selected resolution through the repository copy flow. Keep ordinary repository imports on the existing index-name path and add regression coverage.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): separate knowledge base resolution panel

Render official knowledge base reuse or create-new choices as a sibling panel to skill conflicts and model selection.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): align knowledge base conflict styling

Match the official knowledge base conflict panel to the skill conflict layout with white cards and separate titles and options.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): show accurate dependency status

Prefer repository requirement reason codes over generic activation labels so existing knowledge bases with unavailable embedding models are not shown as unactivated.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): show knowledge base copy name

Display the expected copy name beside the create-new knowledge base option, matching the skill conflict experience.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): remove copy name parentheses

Display the knowledge base copy name without surrounding parentheses in the repository import dialog.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): treat duplicate skills as resolvable

Keep duplicate skills in the dedicated reuse-or-rename panel instead of also classifying them as unavailable dependencies.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): show duplicate knowledge bases as conflicts

Keep same-name knowledge bases in the abnormal dependency section with a conflict label, matching duplicate skill behavior while preserving reuse-or-create choices.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): avoid duplicate knowledge base status

Exclude knowledge bases requiring conflict resolution from the available dependency section so they appear only in the conflict section.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): align official import arguments

Use the authenticated tenant and user context when importing official agents instead of passing unsupported arguments to import_agent_impl.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(deploy): support Unicode official agent profiles

Allow official agent profile directory names such as Chinese names while retaining path traversal protection. Add coverage for Unicode profile parsing.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(deploy): support nested official agent profiles

Add a profile-root option so Hub layouts such as AgentsHub/行业智能体/金融 can be selected without treating the nested directory as a Git repository URL.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(deploy): limit Hub checkout to selected profiles

Skip whole-repository Git LFS checkout and sparsely materialize only the selected profile paths. This prevents unrelated unavailable LFS objects from blocking official agent deployment.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(deploy): materialize sparse official agent checkout

Update the worktree explicitly after configuring sparse checkout for no-checkout Hub clones, so nested Unicode profiles are available to the synchronization step.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-sync): avoid logging expected missing agents

Use a non-raising agent lookup during official bundle synchronization so first-time provisioning does not emit database error logs. Preserve the existing raising lookup for callers that require strict lookup semantics and add coverage for both paths.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(deploy): preserve official profile directories

Copy selected profile contents into an explicit profile directory in the target mount so synchronization can resolve paths such as /mnt/nexent/official-agents/金融.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): hide official take-down action from admins

Keep official agent management in the super-admin resource page and prevent tenant administrators from seeing the ordinary repository take-down action on official listings.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): preserve container paths in sync script

Clean selected profiles before copying and disable MSYS path conversion so Windows Git Bash does not leave stale nested bundles behind.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): convert Git Bash copy sources

Keep container paths protected from MSYS conversion while converting temporary local sources to Docker-readable Windows paths.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): publish Nexent as repository author

Use Nexent as the marketplace author for official bundles without changing the source agent snapshot metadata.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): restore official listing badge

Render the official badge in the unified agent repository card while preserving the existing version badge and listing behavior.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): hide activation for name conflicts

Keep skill and knowledge base name conflicts in the resolution flow instead of offering an activation action.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): prioritize dependency name conflicts

Detect skill duplicates by their precheck reason and show knowledge base conflicts before availability errors, so conflicting resources do not expose activation actions.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-repository): show conflict reason without activation

Use the conflict-prioritized reason label in the non-action status branch so duplicate knowledge bases are not shown as unavailable.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test(agent-repository): update knowledge resolution coverage

Align repository and official-agent tests with knowledge-base resolution payloads and the embedding availability checks introduced by official imports.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* test(agent-repository): cover official management branches

Add coverage for repository fallback errors, official listing management, bundle deletion, and invalid official bundle paths.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* chore(official-agents): exclude document writing snapshot

Remove the document writing official agent snapshot from version control while preserving the local deployment asset.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs(deployment): add official agent deployment guide

Document independent official-agent deployment from Agent Hub, local directories, and offline archives, including profile selection, verification, and troubleshooting.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs(deployment): document official agent deletion

Refine the official-agent guide to match the Chinese documentation style and explain super-admin template deletion, preserved tenant copies, and post-deletion verification.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs(deployment): align official agent page style

Remove the page navigation block so the official-agent deployment guide follows the existing deployment documentation style.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs(deployment): remove page metadata

Keep the official-agent deployment guide consistent with deployment documents that use a plain Markdown heading.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs(config): document official agent variables

Add the official-agent resource path and optional profile fallback to the environment template while keeping profile selection in the independent deployment command.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* refactor(mcp): move tenant URL builder to utils

Keep MCP configuration constants in consts while moving tenant-scoped URL construction into a utility module. Update all production callers and the focused test to use the new location.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* refactor(agent): unify official identity under system IDs

Use one system tenant and user identity for official agent synchronization and repository visibility. Keep deprecated aliases for compatibility while updating frontend detection and tests.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): use system identity for official agents

Replace the separate official tenant and user identifiers with the shared system identity and update all related repository, sync, frontend, and test references.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* docs(deployment): reorder official agent sections

Place profile redeployment guidance before troubleshooting and keep the section numbering sequential.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* refactor(agent): sync official bundles through local API

Replace the backend script entrypoint with a loopback-only synchronization API and make Docker and Kubernetes deployment scripts call it from inside the config container. Keep synchronization business logic in the service layer.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent): avoid MSYS path conversion in sync request

Use the service default official-agent directory instead of passing the container path through docker exec curl, which MSYS rewrites on Windows Git Bash.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-sync): prevent user-controlled bundle paths

Remove the base_dir query parameter from the internal synchronization API so bundle loading always uses the configured container path.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-sync): resolve profiles from mounted directories

Avoid concatenating the selected profile directly into a filesystem path and cover Unicode profile directory resolution.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(agent-sync): scan configured root before filtering profiles

Avoid recursively traversing a directory derived from user-selected profile input by scanning the configured bundle root once and filtering descendants by their top-level profile.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* feat(deploy): simplify bundled official agent installation

Read official bundles from the repository deploy directory and let operators select profiles interactively, while updating the documentation site navigation and deployment guide.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(official-agents): preserve knowledge base display names

Keep official knowledge base declarations when synchronizing repository snapshots so logical folder names remain internal references and user-facing display names are retained. Normalize single-object knowledge_base metadata and cover the regression with bundle and sync tests.

Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* feat(deploy): bundle official agents in main image

Include official Agent seed resources in the main image without auto-synchronizing them at startup. Let the deployment script fall back to the image-bundled resources when the host asset directory is unavailable.

Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* fix(official-agents): refresh repository metadata on resync

Refresh official repository card fields and snapshots even when the bundle version is unchanged, so edits to agent.json become visible after repeated deployment.

Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* Revert "feat(deploy): bundle official agents in main image"

This reverts commit 588a3ce.

* docs(installation): integrate official agent deployment

Move official Agent deployment instructions into the Chinese and English installation guides and update overview links so the deployment documentation has a single entry point.

Co-authored-by: Codex <noreply@openai.com>\nGenerated-by: gpt-5-codex

* feat(official-agents): add bundled official agent assets

Track the official Agent bundles and their document, Skill, and knowledge-base resources in the Nexent repository for deployment-time installation.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix(repository): align app merge with develop

Keep the official-agent and repository-icon endpoints while matching develop import ordering and repository payload serialization behavior.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* revert(docs): remove obsolete official agent links

Remove the official agent deployment navigation and overview links because the referenced documentation page is no longer present.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* merge: adopt latest develop agent repository UI

Use the post-PR3899 develop versions of the Agent page, repository listing modal, and repository API tests so subsequent develop changes are not reverted.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* fix

* test(official-agents): cover knowledge base declaration branches

Cover single-object knowledge base declarations and reject invalid metadata types in official bundle parsing.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex

* refactor(deploy): move official agents under deploy root

Relocate official Agent bundles to deploy/official-agents and update the deployment script and installation documentation to use the new repository path.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5-codex
* update version style

* 删除右下角“联系我们”,优化超级管理员界面
…上限;(3)应用评测:单个评测集大小上限 (#4025)

* feat: enforce MCP resource and request timeout limits

* test: load MCP error helpers in planning test shim

* test: improve MCP timeout coverage

* fix: distinguish generic and MCP timeouts

* fix: adapt MCP timeout handling to latest managed runtime

* test: cover managed MCP request timeouts

* config: make MCP limits and timeout configurable

* refactor: group MCP environment settings

* refactor: address MCP resource limit review feedback

* feat: enforce Skill tenant and upload limits

* test: update Skill quota database mocks

* test: cover Skill limit propagation branches

* feat: limit evaluation set Excel upload size

* test: provide tenant limit exception mock for skill db

* ci: rerun PR checks

* fix: limit MCP timeout to connection establishment

* fix: resolve Sonar reliability findings
…nd default directly to the Agent Workbench interface. (#4036)

* ♻️ Refactor: When launching the Agent Workbench, skip the home page and default directly to the Agent Workbench interface.

* ♻️ Refactor: When launching the Agent Workbench, skip the home page and default directly to the Agent Workbench interface.
* fix(frontend): move skill listing action to more menu

* fix(frontend): clarify skill listing action label
* update version style

* 删除右下角“联系我们”,优化超级管理员界面

* fix(frontend): keep agent conversations visible after returning

* fix: show repository status in paged agent list

* fix(frontend): move skill listing action to more menu

* fix: restrict agent repository review data in paged list

* fix(frontend): return from chat when thread reload fails

---------

Co-authored-by: Summer-Si <mingmingsu22@gmail.com>
… in the Agent Workbench; fixed the issue where agents were not fully displayed on the agent configuration page. (#4040)
* update version style

* 删除右下角“联系我们”,优化超级管理员界面

* fix(frontend): keep agent conversations visible after returning

* fix: show repository status in paged agent list

* fix(frontend): move skill listing action to more menu

* fix: restrict agent repository review data in paged list

* fix(frontend): return from chat when thread reload fails

* fix(frontend): handle agent name overflow and card tag layout

---------

Co-authored-by: Summer-Si <mingmingsu22@gmail.com>
Co-authored-by: panyehong <2655992392@qq.com>
#4008 closed a horizontal-privilege hole on /model/manage/* by restricting
the endpoints to the SU role. The whitelist did not distinguish a
cross-tenant call from one naming the caller's own tenant, so ADMIN users
lost the whole Models tab on /resource-manage: manage/list returned 403 and
the page rendered an empty table with no error, because the create, update,
delete, healthcheck and provider endpoints share the same guard.

The page always sends the caller's own tenant_id (UserManageComp falls back
to user.tenantId for non-SU), so rejecting ADMIN blocked no cross-tenant
access -- it only broke tenant admins managing their own models.

Replace the role whitelist with a role + tenant scope check: SU may target
any tenant, ADMIN only the tenant its token belongs to, and every other role
is rejected as before. ADMIN and SU share the same model:* permission seeds,
so the role itself still has to be checked -- permissions alone cannot
separate them.

The existing ADMIN cross-tenant test kept passing because it used a foreign
tenant_id, which is why the regression went uncaught. Add own-tenant
coverage for list and update, plus a DEV case proving require(model:read) is
not sufficient on its own to reach the manage surface.

Co-authored-by: jeffwu <meetjeff.wu@gmail.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
cj2026-bit and others added 4 commits September 30, 2026 10:44
…on for unauthorized users (#4042)

* 🐛 Fix(evaluation): surface run delete errors and hide button for unauthorized users

* ♻️ Refactor(evaluation): extract shared error-toast helper to fix Sonar duplication
* fix(model): let backfill swap auto-picked defaults for larger-context models

The default-model backfill runs after EVERY model creation, and a slot it
fills is treated as final. Batch adds create models one by one, so the
first-created model permanently occupied the slot before better candidates
landed -- the "available first, then larger context window" ranking never
got to compare across the batch. Observed live: a 5-model batch import
left a 256K-context model as the default LLM while two 1M-context models
arrived right after it.

Distinguish user choices from backfill placeholders via the config row's
user_id: the UI save path (set_single_config) stamps the acting user on
rows it writes, backfill-inserted rows leave it empty. Backfill now:

- never touches a slot whose row carries a user_id (user's explicit choice)
- re-evaluates a previously auto-configured slot on every create and swaps
  in the best candidate (available first, then larger context); the first
  user save flips the row to user-owned and locks it
- repairs dangling rows and fills never-configured slots as before

get_single_config_info now also returns the row's user_id for this
classification.

* fix(model): fill empty default slots only from newly created models

The backfill used to pick the best candidate from ALL live models of a slot's
type. When a user deliberately cleared a default slot and then added one new
model, the backfill resurrected an older, larger-context model they had
passed over -- silently overriding the clear.

Pass the ids of the models created by the current call into the backfill:
- an empty slot (never configured, or cleared by the user) is now filled
  only from those newly created models; if the call added none of the
  slot's type, the slot stays empty
- dangling rows still repair from the full pool (the previous choice is
  gone, so the best remaining replacement is appropriate)
- auto-configured slots keep re-evaluating among all candidates (the
  larger-context swap from the previous commit)
- user-configured slots remain locked

create_model_record returns only a bool, so the new ids are recovered via
display-name lookup (_ids_for_created_models); multi_embedding creates
include their embedding twin automatically.

* fix(model): freeze settled auto-defaults, swap only within an import session

The auto-slot swap from the earlier commit never expired: an auto-configured
default could be replaced by a better model at any later create, so adding
models months after an import could still move the default. Users expect a
default that has been sitting in the slot to stay put -- only the batch
import that is still in progress should converge on the best model.

Gate the swap on occupant freshness:

- swap candidates are the current occupant plus the models created in the
  current call; older models the user passed over are never resurrected
  through the swap path (previously the swap re-ranked ALL models of the
  type, so a cleared-then-refilled slot could drift back to an old giant)
- the swap only runs while the occupant was created within
  _AUTO_SLOT_SWAP_WINDOW (5 minutes) of the newest model in the current
  call -- batch imports create their rows seconds apart, so mid-batch
  convergence still works; an occupant from an earlier session is frozen
- timestamps come from the DB on both sides, so no clock/timezone skew;
  missing create_time disables the gate (permissive, legacy behaviour)
- empty-slot fill (new-only), dangling repair and user-choice locking are
  unchanged

* fix(model): never move occupied default slots; batches finalize once

The auto-slot swap was removed: a default slot occupied by ANY live model --
user-picked or system-backfilled -- is now never touched by later creates.
Users expect "the slot already has a model" to mean exactly that; swapping
in a better model months after an import silently moved defaults and
compounded with the empty-slot rules into hard-to-predict behaviour. The
previous commit's freshness window tried to reconcile this with mid-batch
convergence via heuristics; explicit batch context replaces it.

Batch imports now carry their own flow control:

- ModelRequest gains an optional skip_default_backfill flag (popped by the
  app layer before the dict reaches the service/DB layer, same contract as
  the accept-signal fields). The batch dialog marks every per-row create
  with it, so no row claims empty slots as it lands.
- A new POST /model/backfill_defaults endpoint finalizes the batch: it
  resolves the created display names to ids and runs the backfill ONCE with
  the whole batch as candidates, so an empty slot gets the best model of
  the batch (available first, then larger context) in a single decision.
- The user-facing batch dialog calls the finalize after its loop; the
  manage-tenant path keeps its existing per-row behaviour (its request
  model ignores the flag and its frontend service does not forward it).

Final slot semantics: occupied -> never touched; empty -> best of the
current call's new models (or all models for legacy callers); dangling ->
repaired from the full pool; user choice -> locked.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
The v2.6.1 hotfixes landed on main but were never merged back into
develop, so both branches evolved the same files independently. The
resulting divergence makes the v2.7.0 release PR (develop -> main)
conflict on 45 files. This merge restores the missing back-flow.

Resolution:
- Conflicted files take the develop side. develop already carries the
  hotfix content through a different lineage (#3973, #3998, #4024) and
  its wording is the later evolution, so no fix is lost.
- main-only edits that are still live are preserved; files main changed
  only on paths that develop has since reorganised (agentConfig/ ->
  capability/, agents/ -> agents/[agentId]/) resolve to their new
  locations.
- v2.6.1_001_remove_human_interaction.sql stays deleted: #4024 already
  folded its body into v2.7.0_merged_migrations.sql.
- VERSION stays v2.7.0, the release being prepared.

The merged tree is identical to develop's tip; this commit exists only
to rejoin the two histories so the release PR merges without conflicts.

Verified: python syntax across merged modules, and the
test_model_management_service / test_config_sync_service suites.
* merge v2.6.1 hotfix release from hotfix/v2.6.1 (#3971)

* 🐛 Fix(evaluation): run trials in runtime service (#3954)

* fix(evaluation): run trials in runtime service

Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub config thread manager

Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub runtime jwt helper

* test(evaluation): cover trial proxy error paths

* Fix/override delete (#3958)

* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)

* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key

* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation

* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)

* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it

* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)

* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"

This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.

* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)

* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)

* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests

Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.

Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.

* fix(hitl): guarantee observer chunks precede human_interaction in DB event order

When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.

Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.

Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.

* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths

Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.

Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.

* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety

Fix 4 real issues flagged by github-code-review:

1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
   reset on every drain, meaning a model that kept producing chunks could
   stall the worker forever. deadline is now computed once at entry and
   the sleep call clips to hard_deadline - now.

2. _emit_in_flight Event bridges the async emit path and the worker's
   idle poll. Without this, buffer-empty = 'persisted' was confused with
   buffer-empty = 'taken for emit but still in run_blocking queue'. The
   worker now checks both 'buffer empty for settle_ms' AND 'no emit in
   flight' before deciding the async side is truly idle.

3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
   Previously the async loop drained the buffer, decided it was not yet
   due, then put everything back. That transiently-empty window (16 us
   normally, arbitrarily long under GIL/GC/preemption) was enough for
   the worker's 20 ms poll to mis-fire. We now peek (read count, no
   drain) and only take_chunks when we actually intend to persist.

4. emit_chunks wrapped in try/except that puts drained chunks back into
   the shared buffer before re-raising, and finish() wraps its flush call
   in try/except: pass. Guarantees (a) no chunk loss on DB failure and
   (b) the terminal human_run row is always written even if the flush
   step fails.

Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
  hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
  try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
  test_application_execute_attempt.py drives the full consumer loop
  through _flush_if_due (peek → take → begin/end_emit) and the final
  flush, verifying that every patch line added in application.py is hit.

* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected

When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.

* perf(hitl): replace polling with SSE subscription and move snapshot off the write path

Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
  subscription to the backend's /{run_id}/events SSE stream. Discovery is
  now one-shot: conversationId change and the agent stream pause
  (isRunning true→false), the exact moment a HITL run is most likely to
  exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
  adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
  our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
  concurrent callers and returns cached state; (2) 3s minimum interval
  so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
  not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
  to prevent EventSource from reconnecting forever against COMPLETED runs.

Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
  WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
  must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
  writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
  calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
  before processing each request, and expire_waiting() remains the
  periodic scheduler sweep.

Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.

* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false

After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).

The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.

Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.

Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.

* test(hitl): update mock from snapshot to light_snapshot after read-path refactor

test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.

* test(hitl): raise diff coverage above the 90% merge gate

Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate

SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.

* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm

openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.

* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"

This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.

* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.

* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* [codex] fix(agent): silently retry transient model errors (#3965)

* fix(agent): retry transient model failures silently

* fix(model): support OpenAI httpx2 timeout client

* test(model): add deterministic OpenAI-compatible mock

* fix(agent): keep stream runtime within line budget

* Fix: AIDP knowledge base bug fix (#3967)

* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection

* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change

---------

Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>

* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)

* fix(agent): enforce explicit CodeAgent termination

* fix(test): restore CodeAgent CI compatibility

* merge(main): merge v2.6.1 hotfixes from hotfix/v2.6.1 (#3983)

* 🐛 Fix(evaluation): run trials in runtime service (#3954)

* fix(evaluation): run trials in runtime service

Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub config thread manager

Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub runtime jwt helper

* test(evaluation): cover trial proxy error paths

* Fix/override delete (#3958)

* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)

* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key

* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation

* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)

* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it

* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)

* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"

This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.

* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)

* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)

* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests

Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.

Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.

* fix(hitl): guarantee observer chunks precede human_interaction in DB event order

When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.

Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.

Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.

* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths

Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.

Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.

* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety

Fix 4 real issues flagged by github-code-review:

1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
   reset on every drain, meaning a model that kept producing chunks could
   stall the worker forever. deadline is now computed once at entry and
   the sleep call clips to hard_deadline - now.

2. _emit_in_flight Event bridges the async emit path and the worker's
   idle poll. Without this, buffer-empty = 'persisted' was confused with
   buffer-empty = 'taken for emit but still in run_blocking queue'. The
   worker now checks both 'buffer empty for settle_ms' AND 'no emit in
   flight' before deciding the async side is truly idle.

3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
   Previously the async loop drained the buffer, decided it was not yet
   due, then put everything back. That transiently-empty window (16 us
   normally, arbitrarily long under GIL/GC/preemption) was enough for
   the worker's 20 ms poll to mis-fire. We now peek (read count, no
   drain) and only take_chunks when we actually intend to persist.

4. emit_chunks wrapped in try/except that puts drained chunks back into
   the shared buffer before re-raising, and finish() wraps its flush call
   in try/except: pass. Guarantees (a) no chunk loss on DB failure and
   (b) the terminal human_run row is always written even if the flush
   step fails.

Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
  hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
  try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
  test_application_execute_attempt.py drives the full consumer loop
  through _flush_if_due (peek → take → begin/end_emit) and the final
  flush, verifying that every patch line added in application.py is hit.

* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected

When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.

* perf(hitl): replace polling with SSE subscription and move snapshot off the write path

Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
  subscription to the backend's /{run_id}/events SSE stream. Discovery is
  now one-shot: conversationId change and the agent stream pause
  (isRunning true→false), the exact moment a HITL run is most likely to
  exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
  adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
  our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
  concurrent callers and returns cached state; (2) 3s minimum interval
  so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
  not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
  to prevent EventSource from reconnecting forever against COMPLETED runs.

Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
  WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
  must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
  writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
  calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
  before processing each request, and expire_waiting() remains the
  periodic scheduler sweep.

Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.

* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false

After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).

The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.

Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.

Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.

* test(hitl): update mock from snapshot to light_snapshot after read-path refactor

test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.

* test(hitl): raise diff coverage above the 90% merge gate

Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate

SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.

* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm

openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.

* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"

This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.

* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.

* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* [codex] fix(agent): silently retry transient model errors (#3965)

* fix(agent): retry transient model failures silently

* fix(model): support OpenAI httpx2 timeout client

* test(model): add deterministic OpenAI-compatible mock

* fix(agent): keep stream runtime within line budget

* Fix: AIDP knowledge base bug fix (#3967)

* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection

* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change

* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)

* fix(agent): enforce explicit CodeAgent termination

* fix(test): restore CodeAgent CI compatibility

* cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981)

* Fix StopAsyncIteration leak in execute_attempt finally block

Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors.

Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate.

* Fix HITL form not appearing until page refresh

Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload.

Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh.

* Fix SSE chunk/HITL event ordering race under real server load

Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order.

Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order.

* fix(hitl): recover pending form when SSE delivery stalls silently

Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones.

Changes:
- Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls
- Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it
- Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop

* fix(hitl): stop parked human-input waits from consuming scheduler slots

Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed.

Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments.

Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots.

Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43.

* test(scheduler): fix flaky waiting-concurrency assertion on fast event loops

Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently.

Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.

* Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980)

The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which
also carried two unrelated upstream changes that this release branch never
had:

  - #3909 Knowledge base interface optimization (AIDP UI refactor)
  - #3930 support AIDP knowledge file deletion and download

Both are removed here so the release line keeps only the bug fix.

  - Restore the AIDP frontend components to their pre-refactor layout and
    drop the helper modules only the refactor used:
    AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils.
  - Drop the #3930 document remove/download endpoints from services/api.ts
    and the AIDP translations that only those screens referenced.
  - Keep the fix itself unchanged: knowledge-base scoped Channels and
    KnowledgeFiles/History paths, all-status document listing, keyword
    search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the
    upload-triggered polling.

Verified:
  - pytest test/ext_components/aidp -q  ->  583 passed
  - frontend `npm run type-check` (tsc --noEmit)  ->  no errors

---------

Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>
Co-authored-by: chase <byzhangxin11@126.com>

* merge(main): merge v2.6.1 hotfixes from hotfix/v2.6.1 (#3993)

* 🐛 Fix(evaluation): run trials in runtime service (#3954)

* fix(evaluation): run trials in runtime service

Route trial evaluations through the authenticated Config-to-Runtime proxy and use Config's manager only for creation-stage preparation. Keep Agent execution and evaluator scoring in Runtime.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub config thread manager

Keep pure-logic service import tests aligned with the Config and Runtime thread-manager split.

Co-authored-by: Codex <noreply@openai.com>

Generated-by: gpt-5

* test(evaluation): stub runtime jwt helper

* test(evaluation): cover trial proxy error paths

* Fix/override delete (#3958)

* Fix: override dialog only shows override values, not model defaults (deleted params no longer reappear)

* Fix: custom param deletion persists (null markers), per-agent capacity overrides take effect, and edit-dialog connectivity probe uses stored api_key

* Fix: rename ModelRequest.model_id to probe_model_id - model_dump() is spread into INSERT column lists, so a model_id field injected an explicit NULL primary key and broke model creation

* Fix: move probe_model_id to a dedicated ModelProbeRequest subclass - ModelRequest.model_dump() is spread into INSERT column lists, so any non-column field breaks model creation (Unconsumed column names)

* Fix: editing/adding a model no longer steals the occupied default-model slot - persistCustomLocalConfig now only writes the slot when it is empty (onboarding) or the submitted model already occupies it

* Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)

* Revert "Fix: remove persistCustomLocalConfig - the frontend-cached-config guard could still steal an occupied default slot when the cache was stale/empty. Default-slot writes now come only from the server (create-time backfill for empty/dangling slots)"

This reverts commit 2376aa50ccb0f170e5412a14c5aee33798ca1f06.

* Fix: VLM connectivity probe never found the local test image - the gateway adapter's relative dirname chain resolved two levels short of the package root, so every probe silently fell back to a public URL that is unreachable in offline deployments. Anchor both probe copies on nexent.__file__ so the path survives module moves.

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* cherry-pick: HITL bugfixes from PR #3948 into hotfix/v2.6.1 (#3959)

* Bubfix: guarantee event order, harden chunk buffer, SSE-subscribe controller, and break adapter on terminal human_run (#3948)

* fix(hitl): preserve correct event ordering between observer chunks and human_interaction requests

Root cause: the worker thread writes human_interaction events synchronously via
SQLAlchemy in ask_user, while model_output_thinking/parse observer messages
flow through the async consumer and are flushed only on a batched threshold
(32 chunks or 250ms). When the worker suspends before that flush fires,
human_interaction gains a lower event_seq number than the already-buffered
observer chunks, causing the SSE replay stream to show them in the wrong order.

Fix: replace the plain async-for consumer loop with a manual asyncio.wait
iterator using a 50ms timeout. Once the worker finishes producing model
output (i.e. right before ask_user), the loop times out and flushes any
buffered observer chunks to the DB first, guaranteeing they precede the
subsequent human_interaction row. Empty queue idle periods are essentially
zero-cost; overall DB write frequency stays on par with the original.

* fix(hitl): guarantee observer chunks precede human_interaction in DB event order

When the agent invokes ask_user, two independent write paths caused the
human_interaction row to be persisted BEFORE model_output_thinking / parse
chunks, breaking the SSE replay ordering: the worker thread writes HITL
events synchronously via SQLAlchemy, while observer messages flow through
the async consumer which only flushes on a batched threshold.

Fix: introduce a thread-safe shared chunk buffer on RuntimeInteractionPort
(port.add_chunk / port.take_chunks). The async consumer pushes every
processed chunk there; the worker thread calls flush_chunks_until_idle()
before dispatching any HITL event — it polls the shared buffer and waits
for the async loop to drain the observer queue (20ms idle window, 500ms
max wait), then persists every chunk in its own transaction. This
guarantees chunk event_seq < human_interaction event_seq regardless of
async scheduling latency.

Also fix ImportError: openai 2.50 removed the httpx2 module. OpenAIModel
now falls back from httpx2 to httpx at import time.

* test(hitl): cover shared chunk buffer, flush_chunks_until_idle, and httpx2-fallback paths

Add unit tests for the RuntimeInteractionPort thread-safe chunk buffer
and the flush_chunks_until_idle poll loop that guarantees observer
chunks precede human_interaction events in DB order. All five HITL
entry points (dispatch / boundary / receipt / finish / _wait_until_ready)
are verified to invoke the idle flush before opening their transaction.

Also add two tests for the openai_llm httpx2 → httpx ImportError fallback
introduced to support openai >= 2.50 where the httpx2 shim was removed:
one covers the fallback path, one confirms httpx2 still wins when present.

* fix(hitl): address 4 review comments — hard deadline, emit_in_flight, peek_chunks, try/except safety

Fix 4 real issues flagged by github-code-review:

1. Non-resettable hard_deadline in flush_chunks_until_idle — previously
   reset on every drain, meaning a model that kept producing chunks could
   stall the worker forever. deadline is now computed once at entry and
   the sleep call clips to hard_deadline - now.

2. _emit_in_flight Event bridges the async emit path and the worker's
   idle poll. Without this, buffer-empty = 'persisted' was confused with
   buffer-empty = 'taken for emit but still in run_blocking queue'. The
   worker now checks both 'buffer empty for settle_ms' AND 'no emit in
   flight' before deciding the async side is truly idle.

3. peek_chunks() replaces the take-put-back pattern in _flush_if_due.
   Previously the async loop drained the buffer, decided it was not yet
   due, then put everything back. That transiently-empty window (16 us
   normally, arbitrarily long under GIL/GC/preemption) was enough for
   the worker's 20 ms poll to mis-fire. We now peek (read count, no
   drain) and only take_chunks when we actually intend to persist.

4. emit_chunks wrapped in try/except that puts drained chunks back into
   the shared buffer before re-raising, and finish() wraps its flush call
   in try/except: pass. Guarantees (a) no chunk loss on DB failure and
   (b) the terminal human_run row is always written even if the flush
   step fails.

Tests added:
- 12 pure-mock unit tests in test_runtime_port_chunk_buffer.py cover
  hard_deadline, _emit_in_flight, peek_chunks, begin_emit/end_emit,
  try/except path, and every HITL entry-point's flush-before-transaction.
- 1 async execute_attempt integration test in new
  test_application_execute_attempt.py drives the full consumer loop
  through _flush_if_due (peek → take → begin/end_emit) and the final
  flush, verifying that every patch line added in application.py is hit.

* perf(hitl): stop polling while SSE stream is active, fallback to 5s when disconnected

When isRunning=true the EventSource already pushes human_interaction and
human_execution events in real time — the 1.5s polling loop duplicated that
work, hitting the DB and re-rendering the frontend for every tick. Disable
polling entirely while the SSE stream is alive, and drop to 5s intervals
only when the stream is closed (e.g. page load before the first run, or
after a run finishes) so we can still discover WAITING_HUMAN requests that
were created while the client was disconnected. Add isRunning to the useEffect
dependency array so the polling cadence resets immediately when the SSE
connection state changes.

* perf(hitl): replace polling with SSE subscription and move snapshot off the write path

Frontend — /conversation polling → /{run_id}/events SSE:
- Replace the 5s conversation snapshot polling with a native EventSource
  subscription to the backend's /{run_id}/events SSE stream. Discovery is
  now one-shot: conversationId change and the agent stream pause
  (isRunning true→false), the exact moment a HITL run is most likely to
  exist. The SSE stream then keeps run state live with native auto-reconnect.
- Add dual guards inside refresh() to absorb the thundering herd from
  adapter.onHumanInteractionEvent (fires once per HITL SSE chunk) plus
  our own SSE effect: (1) in-flight dedupe — one snapshot absorbs all
  concurrent callers and returns cached state; (2) 3s minimum interval
  so bursts after the in-flight resolves do not immediately re-hit DB.
- Use a runRef mirror so refresh() stays stable and downstream effects do
  not re-run on every snapshot.
- Detect terminal status inside SSE onmessage and proactively es.close()
  to prevent EventSource from reconnecting forever against COMPLETED runs.

Backend — snapshot off the write path:
- Add repository.read_only() context manager: plain SELECT without
  WITH FOR UPDATE, no transaction, no flush, no _expire scan. Pure reads
  must not contend with worker writes on the same row lock.
- Add service.light_snapshot() using read_only. Retain snapshot() as a
  writer-path API for any future lock-held callers.
- Route conversation_snapshot, run snapshot endpoint, and both snapshot
  calls inside stream_run() through light_snapshot.
- Move expiration to the writer path: decide() still calls _expire inline
  before processing each request, and expire_waiting() remains the
  periodic scheduler sweep.

Impact: conversation snapshot calls drop from 12+/min (polling) or
10+/s (burst from adapter + SSE) to at most one every 3s. Each call is
now two plain SELECTs instead of a lock-held transaction with a possible
write from _expire. Read and write paths are fully decoupled.

* fix(hitl): detect terminal human_run in adapter and break stream so isRunning flips false

After a HITL run reaches FAILED/COMPLETED, Assistant-UI's isRunning
stayed true — the stop button remained visible and new messages went
into the queue buffer instead of being sent normally. The root cause
is that isRunning is driven entirely by the ChatModelRun generator
lifetime, which only returns when the backend SSE HTTP connection
closes (reader.read() -> done=true).

The backend stream_run loop can hang on heartbeat even after the run
is terminal when the SSE was opened during WAITING_HUMAN with
attempt_active=true: the break condition requires both cursor >=
event_seq AND (terminal status OR WAITING_HUMAN with attempt_active=false
and empty rows). If continueHitl fires mid-flight with a stale
after_event, the cursor never catches up, so the SSE stays alive
forever and the generator never returns.

Stop depending on the backend closing first. Inside the adapter's SSE
chunk loop, detect a terminal human_run event (status in COMPLETED,
FAILED, STOPPED, EXPIRED), set a hitlTerminal flag, break the inner
for-loop, and let the outer while-loop exit via the same flag on the
next iteration. Assistant-UI sees the generator return and flips
isRunning false immediately.

Only affects HITL streams — the normal non-HITL agent path never
emits human_run events so this branch is never taken.

* test(hitl): update mock from snapshot to light_snapshot after read-path refactor

test_human_interaction_app.py still mocked service.snapshot after commit
288ae4e69 moved conversation_snapshot and the run snapshot endpoint to
service.light_snapshot (read-only path, no lock, no _expire). The fixture
return_value and the two assert_called_once_with/assert_not_called
assertions all referenced the old method name, causing CI to fail because
MagicMock.snapshot was never called.

* test(hitl): raise diff coverage above the 90% merge gate

Codecov reported 70.43% patch coverage (target 90%) because new error and race paths in the HITL changes had no tests. Add mocked unit tests for: leftover chunk flush in execute_attempt's finally block before the failed finish, CancelledError scope and stop-event fallbacks, RunTerminated finish race, recovery-required outcome, and chunk iterator aclose failure tolerance; runtime_port in-flight emit busy detection, chunk restoration when emit_chunks raises, and terminal status persistence on flush failure; light_snapshot/read_only service behavior with signed tenant and user scoping; and the httpx fallback when openai._base_client.httpx2 is absent. Measured locally with CI-equivalent per-file pytest isolation: patch coverage 202/202 = 100%.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* style(hitl): unify comment style across HITL changes

Merge explanatory inline comments into docstrings, keep single-line comments for inline notes, convert TypeScript block notes to JSDoc, and drop banner/separator lines. Comment-level changes only, no behavior change.

* refactor(hitl-test): dedupe fake port setup to satisfy SonarCloud duplication gate

SonarCloud failed the quality gate with new_duplicated_lines_density=5.2% (threshold 3%), caused solely by test_application_execute_attempt.py: the inline _Port stub in the flush test and the one in _run_execute_attempt duplicated ~69 lines (2 CPD blocks, 14.4% file density). Extract a shared _build_port_class/_make_port_factory plus a _patched_application context manager and _execute_attempt_args so both call sites reuse a single definition; drop dead code (last_flush, install/monkeypatches, unused imports) and fix the latent bare-contextmanager NameError by using contextlib.contextmanager. No behavioral change; all 8 tests pass.

* fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm

openai >= 2.50 removed the httpx2 shim from openai._base_client. The bare
import httpx2 causes ImportError in CI and on systems with recent openai
versions. This restores the try/except fallback introduced in PR #3948
commit 9521b934 and later accidentally reverted by commit d086da259.

* Revert "fix(sdk): restore httpx2 → httpx ImportError fallback in openai_llm"

This reverts commit 8e06ce348dfc8c34baf42c6bfa8211883d0a2c61.

* 🐛 Bugfix: Fixed an issue where the sandbox container user lacked the permissions to create folders and files. (#3963)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngi… (#3962)

* Fix: dispatch ModelEngine provider listing to the dedicated ModelEngineProvider - get_provider_models routed every provider through the OpenAI-compatible adapter, so ModelEngine batch import failed (wrong endpoint path /open/router/v1/models, self-signed cert, custom type taxonomy, missing per-model base_url). The dedicated class existed but was never wired in.

* chore: ModelEngine catalog base_url placeholder - preset public URL is wrong for private deployments, placeholder communicates the required /open/router/v1 path format

---------

Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>

* [codex] fix(agent): silently retry transient model errors (#3965)

* fix(agent): retry transient model failures silently

* fix(model): support OpenAI httpx2 timeout client

* test(model): add deterministic OpenAI-compatible mock

* fix(agent): keep stream runtime within line budget

* Fix: AIDP knowledge base bug fix (#3967)

* Fix: editing a ModelEngine model no longer flips ssl_verify to True - the update path only checked api_key emptiness while the create path also exempts open/router URLs (ModelEngine self-signed certs). The edit dialog prefills the real key and always submits it, so any edit silently broke connectivity. Exemption now checks the payload URL with a fallback to the stored record; batch-edit groups get the same protection

* refactor: extract MODEL_ENGINE_URL_MARKER constant (SonarCloud S1192) and use a placeholder domain in test fixtures - no behavior change

* [fix] enforce explicit CodeAgent termination and silent recovery (#3969)

* fix(agent): enforce explicit CodeAgent termination

* fix(test): restore CodeAgent CI compatibility

* cherry-pick: HITL reliability fixes from PR #3977 into hotfix/v2.6.1 (#3981)

* Fix StopAsyncIteration leak in execute_attempt finally block

Root cause: when the agent chunk stream exhausted normally, the finally block awaited the already-consumed anext_task, re-raising StopAsyncIteration which was not suppressed by the existing CancelledError handling. The leftover chunk flush was skipped, successful runs were marked as failed, and the claiming scheduler job logged errors.

Fix: reset anext_task to None before breaking out of the consumption loop so the finally block skips the await and always reaches the leftover flush and terminal finish() write. Tightened the regression test to assert that a normally exhausted stream does not leak StopAsyncIteration and that finish() is called. Also deduplicated the two _Port stub classes via a shared factory to satisfy the SonarCloud new_duplicated_lines_density gate.

* Fix HITL form not appearing until page refresh

Root cause: the frontend discovery chain rate-limited every refresh() with a 3s min interval and in-flight coalescing, silently dropping the critical human_run/human_interaction events that follow an ask_user suspension. The run event stream goes quiet afterwards, so nothing re-triggered the snapshot and the form only appeared after a manual page reload.

Fix: refresh() now takes a force flag that bypasses the throttle; force callers arriving while a snapshot is in flight are re-run via a trailing refreshRef invocation instead of being dropped. SSE human_run/human_interaction/human_decision/human_execution messages and the chat-adapter onHumanInteractionEvent callback now force refresh.

* Fix SSE chunk/HITL event ordering race under real server load

Root cause: chunk persistence and HITL event writes ran in two threads (async consumer via run_blocking on the control-io lane, worker thread synchronously) with seq assigned at DB row-lock acquisition time. Two race windows reordered messages on loaded servers but never locally: (1) flush_chunks_until_idle's 500ms hard deadline fired while the async drain was still in flight, so the HITL row committed before chunks produced earlier (form appearing before model output); (2) worker emit_chunks and the async _flush_if_due drained concurrently without mutual exclusion, so seq order followed lock acquisition instead of production order.

Fix: replace the begin_emit/end_emit Event with a shared threading.Lock and move take_chunks+emit_chunks into one atomic critical section (drain_and_emit) used by both the async consumer and the worker flush. The hard deadline may now only fire once the lock is free, guaranteeing in-flight drains commit before the caller writes its HITL transaction. Added regressions: flush waiting for an in-flight drain past its deadline, and concurrent drains preserving chunk production order.

* fix(hitl): recover pending form when SSE delivery stalls silently

Root cause: form discovery relied solely on a single EventSource plus refresh() with no fallback. A half-open connection (e.g. hung dev proxy) never raises an error event or reconnects, so human_interaction events are lost until a manual page refresh. A hung snapshot fetch could also keep refreshInFlight stuck forever, silently dropping every later refresh, including forced ones.

Changes:
- Poll the read-only snapshot every 5s while a run is active; the tick shares the refresh throttle and in-flight guard, so it adds no load while SSE delivery is healthy and discovers a pending form within 5s when the stream stalls
- Add a 15s AbortSignal timeout to human-interaction client requests so a hung fetch releases the in-flight guard instead of bricking it
- Wrap the human_run chunk JSON.parse in the chat adapter with try/catch: the stream loop has a finally but no catch, so a malformed payload would silently kill the whole chat stream read loop

* fix(hitl): stop parked human-input waits from consuming scheduler slots

Root cause: while a run waits for a human decision, its executor task parks inside _wait_until_ready and the lease renewal loop keeps the lease alive, so the run occupies one HITL_MAX_CONCURRENCY slot for up to HITL_WAIT_SECONDS (default 24h). With HITL enabled every non-debug chat is dispatched through this scheduler, so two unattended forms filled the default concurrency of 2 and froze all conversations: new agent/run streams only emitted heartbeats because READY runs were never claimed.

Fix: add LeaseScheduler.mark_waiting so executors can flag themselves as parked on external input. Slot capacity is now max_concurrency minus executing jobs only (running minus waiting), and the waiting flag is cleared in the job's finally block. RuntimeInteractionPort relays enter/exit of _wait_until_ready through a wait_reporter callback, covering resume, termination and lease loss paths. The reporter degrades safely: a stale SDK copy without mark_waiting falls back to slot-consuming waits, and call_soon_threadsafe is wrapped in a lambda because it does not forward keyword arguments.

Config: raise the env example defaults from 24h to 1h waits and concurrency 2 to 100, since waiting runs no longer consume execution slots.

Tests: new test_waiting_jobs_do_not_consume_concurrency; scheduler suite 14/14, HITL service 12 passed, runtime and app suites 43/43.

* test(scheduler): fix flaky waiting-concurrency assertion on fast event loops

Root cause: job 2's executor completed instantly after appending to started, so its done-callback could discard it from _running before the active_count == 2 assertion ran. On Linux CI the event loop schedules that callback first, making the test fail intermittently.

Fix: both executors now park on the shared gate via separate running events, so the assertion observes a stable running set instead of a transient window.

* Fix: drop unrelated AIDP interface refactor from the AIDP knowledge base fix (#3980)

The AIDP knowledge base fix reached hotfix/v2.6.1 through PR #3967, which
also carried two unrelated upstream changes that this release branch never
had:

  - #3909 Knowledge base interface optimization (AIDP UI refactor)
  - #3930 support AIDP knowledge file deletion and download

Both are removed here so the release line keeps only the bug fix.

  - Restore the AIDP frontend components to their pre-refactor layout and
    drop the helper modules only the refactor used:
    AidpKnowledgeBaseModalParts, useAidpGroupOptions, aidpUploadUtils.
  - Drop the #3930 document remove/download endpoints from services/api.ts
    and the AIDP translations that only those screens referenced.
  - Keep the fix itself unchanged: knowledge-base scoped Channels and
    KnowledgeFiles/History paths, all-status document listing, keyword
    search, status labels (UPLOADING / PROCESSING / EXTRACTING) and the
    upload-triggered polling.

Verified:
  - pytest test/ext_components/aidp -q  ->  583 passed
  - frontend `npm run type-check` (tsc --noEmit)  ->  no errors

* fix(agent): accept reasoning-prefixed code actions (#3990)

* refactor: remove human interaction features and related configurations (#3988)

* refactor: remove human interaction features and related configurations

* refactor: remove human interaction features and related configurations

---------

Co-authored-by: cj2026-bit <647646783@qq.com>
Co-authored-by: lijiayang619 <1170349871@qq.com>
Co-authored-by: ljy <ljy@DESKTOP-65OBISN.(none)>
Co-authored-by: bernard1234 <840646206@qq.com>
Co-authored-by: panyehong <91180085+YehongPan@users.noreply.github.com>
Co-authored-by: Jason Wang <56037774+JasonW404@users.noreply.github.com>
Co-authored-by: gs-aion <gs597153711@qq.com>
Co-authored-by: Dallas98 <40557804+Dallas98@users.noreply.gi…
@jeffwu-1999 jeffwu-1999 changed the title v2.7.0 merge(main): merge v2.7.0 release from develop Sep 30, 2026
@jeffwu-1999
jeffwu-1999 merged commit 519b9b1 into main Sep 30, 2026
6 of 8 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.