Skip to content

FIX: mark truncated LiteLLM responses and handle empty truncation - #2991

Open
Chen Yufeiyang (feiiiiii5) wants to merge 1 commit into
microsoft:mainfrom
feiiiiii5:codex/litellm-truncation
Open

Chen Yufeiyang (feiiiiii5) wants to merge 1 commit into
microsoft:mainfrom
feiiiiii5:codex/litellm-truncation

Conversation

@feiiiiii5

Copy link
Copy Markdown
Contributor

Description

Fixes #2878. LiteLLM responses stopped by max_tokens were saved as complete answers; an empty truncated response raised EmptyResponseException and aborted the turn.

This follows the existing OpenAI target behavior using shared helpers: warn on finish_reason="length", mark the returned piece truncated, and return an empty piece with response_error="empty" when no content remains. Partial content, usage, finish reason, and response cost are retained. Non-truncated empty responses still raise.

Tests and Documentation

  • pytest tests/unit/prompt_target/target/test_litellm_chat_target.py -q -rs: 84 passed, no skips. All 112 installed packages match uv.lock, including LiteLLM 1.102.1. This includes six local loopback wire tests through the real adapter; no external provider was used.
  • Both new truncation regressions fail with the source restored to current main (2596f34): the partial piece is unmarked and the empty case raises. Removing the warning separately makes its assertion fail.
  • Ruff check/format and scoped ty passed on both changed files. Full optional-dependency tests and live providers were not run.
  • No documentation change: this applies the established truncation contract to another target.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

LiteLLMChatTarget does not flag or survive output-token truncation

1 participant