Skip to content

fix(frontend): stop cmvn readers from crashing on blank lines or returning empty statistics - #3760

Open
Lesereingrape wants to merge 1 commit into
modelscope:mainfrom
Lesereingrape:contrib/cmvn-blank-line
Open

Lesereingrape wants to merge 1 commit into
modelscope:mainfrom
Lesereingrape:contrib/cmvn-blank-line

Conversation

@Lesereingrape

Copy link
Copy Markdown

Summary

Both cmvn readers in the installed package walk the file line by line and index the result without checking it: line_item[0] on a blank line raises IndexError: list index out of range, and the unbounded one-line lookahead lines[i + 1] raises the same error when a <AddShift> / <Rescale> tag is the last line. Because the parse is a pure line-matching loop, a file whose tags do not start a line on its own (which Kaldi does emit — the whole <Nnet> on one line) matched nothing and load_cmvn returned empty statistics with no error at all; the failure then only surfaced on the first inference as RuntimeError: The size of tensor a (N) must match the size of tensor b (0).

Fixes #3759

Three-line change per reader (funasr/frontends/wav_frontend.py:15, funasr/frontends/default.py:390):

lines = [line for line in f.readlines() if line.split()]          # skip blank/whitespace-only lines
line_item = lines[i + 1].split() if i + 1 < len(lines) else []     # bounded lookahead
if line_item and line_item[0] == "<LearnRateCoef>":
...
if not means_list or not vars_list:
    raise ValueError(f"No <AddShift>/<Rescale> statistics found in cmvn file: {cmvn_file}")

Skipping blank lines first (rather than guarding inside the loop) also makes <AddShift> + blank line + <LearnRateCoef> parse, which a per-line guard would silently drop. The ValueError follows the convention the sibling runtime loader already uses (runtime/python/onnxruntime/funasr_onnx/utils/frontend.py:148 names the offending cmvn file), and blank-line skipping matches how this repository reads other plain-text input (funasr/bin/realtime_ws.py:1363, funasr/utils/compute_det_ctc.py:59).

Type of change

  • Bug fix
  • Documentation
  • Example or demo
  • Runtime or deployment
  • Benchmark or evaluation
  • Model/training change

Validation

Measured on main @ 66d7a4c2, Python 3.13.7 / numpy 2.5.3 / torch 2.14.1+cpu, Windows, with PYTHONPATH pointed at the worktree (verified by printing funasr.__file__):

  • RED with only the two source files restored from origin/main and the new tests present: python -m pytest -q tests/test_numpy_compatibility.py -k cmvn → 2 failed, 3 passed.

  • GREEN with the fix: same command → 5 passed; whole file python -m pytest -q tests/test_numpy_compatibility.py → 9 passed, including the two pre-existing cmvn tests and test_frontend_applies_cmvn_and_preserves_padding.

  • Regression guard as a test of its own: test_cmvn_reader_loads_the_am_mvn_shipped_by_the_runtime asserts the am.mvn this repository ships under runtime/triton_gpu/model_repo_sense_voice_small/feature_extractor/ still loads as (2, 560) finite values — it passes on main too, so a file that parses correctly today provably does not change.

  • python -m compileall -q funasr/frontends tests/test_numpy_compatibility.py → exit 0.

  • Formatting: the repository is not black-clean at main (e.g. funasr/frontends/default.py docstring block at 392-395 is already reformatted by black --line-length=100), so black --check fails on both the pristine and the patched file. Checked instead that black --line-length=100 leaves every added line untouched — the only hunks it proposes in these files are pre-existing ones outside my diff.

  • python -m compileall funasr examples tests

  • Docs or links checked

  • Runtime/deployment command tested

User impact

Paraformer / SenseVoice users who point cmvn_file at a hand-edited or concatenated .mvn file — or at a Kaldi artifact written on a single line — currently get either a bare IndexError during frontend construction, or a frontend that builds fine and fails at the first model.generate() call. Both cases now either parse or raise an error that names the file.

Notes for reviewers

  • The same loop is duplicated in five more places under runtime/ (runtime/python/libtorch/, runtime/python/onnxruntime/, runtime/triton_gpu/, export_lfr_cmvn_pe_onnx.py). Those are standalone export/deployment helpers rather than the installed package, so this PR deliberately leaves them alone; the issue lists them.
  • Two of the three possible failure modes were silent or crash-only, so the ValueError is a behaviour change in one narrow case: a cmvn file that yielded empty statistics used to return them. Such a frontend could never produce features (it hit the RuntimeError above on the first call), so this only converts a downstream failure into an actionable one.
  • The practical limitation to note: a fork PR in this repository does not trigger workflow runs, so all evidence above is local, and the reproduction needs no model weights or downloads.
  • Prepared with an AI coding agent; a human maintainer at the PR author's side has not re-reviewed the diff beyond the checks listed above.

Both package cmvn readers index line_item[0] and lines[i + 1] without
checking, so a blank or whitespace-only line anywhere in the file
(trailing newline at EOF, a separator between sections) aborts frontend
construction with IndexError, and a file whose tags the walker does not
recognize returns empty statistics that only fail much later at inference.
Skip blank lines, bound the lookahead, and raise a ValueError naming the
file when neither <AddShift> nor <Rescale> parsed.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: cmvn readers raise IndexError on blank lines and return empty statistics without error

1 participant