Skip to content

feat(audio): support explicit PCM16 input without container probing - #3751

Merged
LauraGPT merged 2 commits into
mainfrom
codex/explicit-pcm-input-20261001
Oct 1, 2026
Merged

LauraGPT merged 2 commits into
mainfrom
codex/explicit-pcm-input-20261001

Conversation

@LauraGPT

@LauraGPT LauraGPT commented Oct 1, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Related to #3739; this does not claim to reproduce or diagnose the reported three-minute decoder stall.

  • Add opt-in input_format="pcm_s16le" for known mono signed little-endian PCM16 bytes. It normalizes samples directly, bypassing container detection/decoding, and rejects partial two-byte samples. The default auto byte-format behavior is unchanged.
  • Forward the option through data preparation, explicit byte batches, multimodal inputs, ordinary inference, VAD preprocessing and export input preparation. Include matching English/Chinese API examples.
  • Fix a pre-existing sample-rate propagation problem exposed by this path: VAD receives the source rate; ASR receives the rate of the already-resampled segments. Keep the original input rate stable across a batch so a prior ASR call cannot alter how subsequent input is loaded.
  • Run the new regression suite and existing audio-byte tests in both existing NumPy CI lanes; add their source/test trigger paths.
result = model.generate(input=pcm_bytes, input_format="pcm_s16le", fs=16000)

The format flag does not make other encodings valid PCM16. Do not use it for encoded WAV/MP3 bytes, floating-point PCM, other bit depths, big-endian audio or interleaved stereo. Streaming callers retain their cache/chunk/final-chunk handling.

Validation

  • Initial format regressions on unchanged main: 17 failed / 2 existing-behavior controls passed; implementation then passed the same tests.
  • Additional 8/48 kHz VAD tests covering arrays/PCM, single/batch, and constructor/per-call rates: 16 failed before the rate fix, then passed. They exercise actual preprocessing/resampling with recording model doubles, not downloaded acoustic models.
  • Final focused input, byte decoder, API documentation/signature and submodel tests: 93 passed, 0 skipped.
  • Full existing NumPy workflow test selection plus the added suites: 93 passed, 0 skipped per environment, on NumPy 1.26.4 and 2.4.0, each with PyTorch 2.10.0+cpu.
  • Independent final review found no remaining P1/P2 issues; git diff --check passes.

Coverage includes probe/decoder bypass, signed sample boundaries, malformed input, default behavior, batch/multimodal routing, nonempty VAD segments, streaming options, per-call reset, export preparation, and bilingual executable examples. Export preparation uses a capture boundary, not an ONNX export acceptance test. No model weights, GPU, model-quality evaluation, original reporter-environment reproduction, package release or website deployment is claimed.

Integration Status

Merged as bb303f85ea415e64d2188de6847ddcf2adc2593f; merge parents and tree were verified against the tested candidate 2d6cb0ed216254b2f7d046fb78027a3933e3ddc2.

  • All four exact-head hosted workflows passed: NumPy compatibility, KWS, MOSS, and product site.
  • The CI follow-up explicitly installs/checks FFmpeg. Downloaded JUnit artifacts from NumPy run 36888888302 confirm 93 tests, 0 failures/errors/skips in each lane. The older head had a missing-FFmpeg skip and is not used as zero-skip evidence.
  • Source-only change; not a new PyPI release. Issue Bug: model.generate(input=bytes) hangs for minutes on raw PCM chunks that trigger audio container false-positive #3739 remains open because its original decoder stall has not been reproduced.
  • All five exact-merge workflows passed: NumPy compatibility (36893811611), KWS (36893811800), MOSS (36893811789), product site including browser checks (36893811808), and API documentation (36893811759). Downloaded merge JUnit artifacts again confirm 93 tests and 0 failures/errors/skips per NumPy lane.

Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
Signed-off-by: LauraGPT <18321252+LauraGPT@users.noreply.github.com>
@LauraGPT
LauraGPT merged commit bb303f8 into main Oct 1, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant