Skip to content

archived threads with a missing rollout file: content unrecoverable and unarchive fails; WSL-side writers store /mnt/... paths in the state DB and config.toml #49335

Description

@Cpixss400

What version of Codex CLI is running?

codex-cli 0.159.0 (standalone npm CLI), managed app-server daemon 0.159.0, ChatGPT Desktop (MSIX) 26.924.2738.0.

Which model were you using?

N/A — local codex doctor / state-DB diagnostics. (model = gpt-6-luna, provider openai.)

What platform is your computer?

Windows 11 x64, NTFS. CODEX_HOME = F:\CodexHome, also reachable as C:\Users\<user>\.codex (junction). runCodexInWindowsSubsystemForLinux = true, and a WSL-side agent uses the same folder as /mnt/f/CodexHome.

What issue are you seeing?

Two defects in the rollout ↔ state-DB inventory area that I could not find tracked. They are not the "valid paginated continuation" false positive — that class is already tracked in #41608 and I am not duplicating it here (our instance of it: 13 duplicate rollout thread ids, all from rollout-<ts>-<threadId>_<physicalRolloutId>.jsonl files whose session_meta.id equals the leading id, with history_base pointing at the retained segment; we did not create rows for them).

Defect A — archived threads whose rollout file is gone: content unrecoverable, and unarchive fails silently

  • state.rollout_db_parity reports stale rows = 393; 371 of them are archived threads whose threads.rollout_path points into …\archived_sessions\ while only 7 files remain in that directory.
  • Of those 371 rows, 207 have zero rows in thread_history_1.sqlite → thread_items, and 206 have no copy anywhere under CODEX_HOME (searched sessions/, archived_sessions/, _archive/, worktrees/, plus the drive recycle bin). codex cloud list → No tasks found.
  • These are real sessions, not samples: e.g. a HERMES — PRIMARY BUILDER FOR VECTOR SELF-UPGRADE PHASE 2–5 thread with tokens_used = 40,777,330 created 2026-08-17. Sum of tokens_used for the 206 rows ≈ 1.82e9.
  • codex unarchive <id> fails: Error: failed to unarchive session — no reason is surfaced, even with RUST_LOG=codex_core=debug,codex_app_server=debug (run against both a thread with no file and one with projected items but no file). Afterwards codex resume <id> / codex exec resume <id> <prompt> still refuses with thread/resume failed: session <id> is archived. Run codex unarchive <id> … (code -32600).
  • codex migrate-rollouts --json reports 912/912 already_paginated and never mentions these rows (it enumerates files on disk, so rows whose file is missing are invisible to it). codex doctor offers no repair for them. codex debug app-server has no read/inspect verb we could find.

Questions: is the raw rollout file authoritative for an archived session (with thread_items only a projection)? Is an official code path supposed to delete archived_sessions\*.jsonl (we found none, and there is no log trace of the deletion)? Should unarchive surface the real reason when the file is missing, and is there a supported recovery (or at least a way to hide/clean stale rows without hand-editing SQLite)?

Defect B — a WSL-side writer stores /mnt/... and /home/... paths in the state DB and in config.toml

  • 22 non-archived threads rows store rollout_path as /mnt/f/CodexHome/sessions/.... The files exist; the Windows side cannot use those paths, so the same file is simultaneously reported as "row points at an unusable file" and "file has no row". Desktop logs also show the app failing on them: state db discrepancy during list_threads_db: stale_db_path_retained (246 occurrences) and find_thread_path_by_id_str_in_subdir: falling_back (6).

  • The same class silently broke features when ~/.codex/config.toml was rewritten from the WSL side (each value below is a real line we found and corrected):

    Key Value found Effect on Windows
    mcp_servers.node_repl.command /mnt/c/Users/<user>/AppData/Local/OpenAI/Codex/runtimes/cua_node/<hash>/bin/node_repl.exe MCP server never started: MCP client for node_repl failed to start … No such file or directory (os error 2) — 44 occurrences on 09-28, 21 on 09-29, 0 after correction
    marketplaces.openai-bundled.source /mnt/f/CodexHome/.tmp/bundled-marketplaces/openai-bundled marketplace source unresolvable
    marketplaces.claude-cowork.source /home/<user>/.codex/plugins/cache/claude-cowork marketplace source unresolvable
    projectlessWorkspaceRoot /mnt/f/Documents/Codex/2026-08-16 projectless sessions point at a path Windows cannot open
    [projects."/mnt/f/kotr"], [projects."/mnt/f/KOTR"] both, in addition to the correct [projects."f:\\kotr"] duplicate entries for one folder, in two spellings

    After rewriting all six lines to Windows spelling and restarting the daemon, node_repl initializes (start_server_task … Service initialized as client) and today's failure count is 0. Request: normalize/validate path values against the platform that will consume them, or keep per-platform config.

What steps can reproduce the bug?

  1. Run codex doctor --json and inspect checks["state.rollout_db_parity"].details on a Windows install whose CODEX_HOME is also mounted into WSL (/mnt/<drive>/…), with archived threads present.
  2. For Defect A: pick any stale rows id whose rollout_path is under archived_sessions\ and whose file does not exist → codex unarchive <id> (fails with no reason) → codex exec resume <id> "Reply with exactly: OK" (refused due to archived state).
  3. For Defect B: grep config.toml for values starting /mnt/ or /home/; on Windows each such value fails with No such file or directory (os error 2); the threads rows exhibit the path mismatch in the same way.

Diagnostics used (all read-only; the live DBs were never modified)

codex doctor --json, codex migrate-rollouts --json (dry run), codex cloud list, codex unarchive, plus state_5.sqlite / thread_history_1.sqlite / logs_2.sqlite inspected from copies (sqlite3 backup API for consistent snapshots). We deliberately did not edit rollout_path values, reset migration cursors, delete stale rows, rename the paginated-continuation files, or remove any -wal, because the counters are not the bug.

Impact

  • Defect B is fully fixed locally by correcting the paths (no data loss).
  • Defect A is not recoverable locally: 207 archived conversations (≈1.82e9 tokens, mostly 2026-07/08) have neither a rollout file nor projected items; no cloud task exists; unarchive fails. Related symptom noise while the app runs: fs/getMetadata errors (5,292 on 09-27, 576 on 09-28, 630 on 09-29) as the desktop app stats missing rollout paths.

codex doctor excerpt (filing time)

{"rollout DB rows": "1271", "rollout DB active files": "918", "rollout DB active rows": "898",
 "rollout DB archived files": "7", "rollout DB archived rows": "373", "rollout DB stale rows": "393",
 "rollout DB missing active rows": "42", "rollout DB missing archived rows": "5",
 "rollout DB duplicate rollout thread ids": "13", "rollout DB archive mismatches": "0",
 "rollout DB duplicate DB paths": "0", "rollout DB malformed file names": "0",
 "rollout DB scan errors": "0", "rollout DB sources": "subagent:other=421, subagent:thread_spawn=379, vscode=247, exec=193, cli=31"}

The full codex doctor --json (21 KB, overallStatus: warning) contains no credentials (auth.credentials: stored API key = false, stored ChatGPT tokens = true, stored auth mode = chatgpt) and can be pasted on request.

Related: #41608 (valid paginated continuations counted as duplicates — same metric, different mechanism).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    CLIIssues related to the Codex CLIbugSomething isn't workingconfigIssues involving config.toml, config keys, config merging, or config updatessessionIssues involving session (thread) management, resuming, forking, naming, archivingwindows-osIssues related to Codex on Windows systems

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions