Describe the bug
set_timesteps appends sigma_last behind the schedule. With final_sigmas_type="sigma_min" that is sigma(t=0), and the converted schedules (use_karras_sigmas, use_exponential_sigmas, use_beta_sigmas, use_lu_lambdas) already end at exactly that value, because _convert_to_karras and the other conversions ramp down to sigma_min = in_sigmas[-1], the smallest sigma of the training schedule. So the last two sigmas are identical and the final step has h = lambda_t - lambda_s0 = 0:
karras, 3 steps: timesteps [999, 593, 0] sigmas [157.4073, 5.9113, 0.0100, 0.0100]
linspace, 3 steps: timesteps [999, 666, 333] sigmas [157.4073, 9.4889, 1.4624, 0.0100]
The first-order and the midpoint second-order updates return the sample unchanged on that step, which is the legacy behavior sigma_min was kept for in #6477. The updates that divide by h return NaN for the whole batch instead:
DPMSolverMultistepScheduler, solver_order=3: multistep_dpm_solver_third_order_update computes r0, r1 = h_0 / h, h_1 / h and r0 / (r0 + r1), which is inf / inf, and (exp(-h) - 1) / h, which is 0 / 0. Same for dpmsolver++ and sde-dpmsolver++. The docstring recommends solver_order=3 for unconditional sampling.
DPMSolverMultistepScheduler, solver_order=2, solver_type="heun": (exp(-h) - 1) / h + 1 is 0 / 0.
UniPCMultistepScheduler, solver_order=3, lower_order_final=False: h_phi_k = h_phi_1 / hh - 1 is 0 / 0 and torch.linalg.solve propagates it into rhos_p.
lower_order_final only protects DPMSolverMultistepScheduler below 15 steps, euler_at_final is off by default, and final_sigmas_type="zero" is exempt because step already forces the first-order update for it. So with the default config plus solver_order=3, a converted schedule and final_sigmas_type="sigma_min", 14 steps work and 15 or more steps come back all-NaN.
The same duplicated sigma shows up in DPMSolverSinglestepScheduler, DEISMultistepScheduler and SASolverScheduler (which always append sigma_min) and in EulerDiscreteScheduler (also with the default linspace spacing, whose last timestep is 0), but their final updates do not divide by the step width, so there it only costs one model evaluation on a no-op step. #14887 is about the other end of the same line, final_sigmas_type="zero", where h is inf.
Reproduction
import torch
from diffusers import DPMSolverMultistepScheduler, UniPCMultistepScheduler
sample = torch.rand(1, 3, 8, 8, generator=torch.Generator().manual_seed(0))
def run(cls, num_inference_steps=20, **config):
scheduler = cls(final_sigmas_type="sigma_min", **config)
scheduler.set_timesteps(num_inference_steps)
x = sample * scheduler.init_noise_sigma
for t in scheduler.timesteps:
x = scheduler.step(0.1 * x, t, x).prev_sample
nan = torch.isnan(x).float().mean().item()
print(f"{cls.__name__:28s} {str(config):68s} last sigmas {scheduler.sigmas[-2:].tolist()} nan {nan:.0%}")
for sigmas in ["use_karras_sigmas", "use_exponential_sigmas", "use_beta_sigmas", "use_lu_lambdas"]:
run(DPMSolverMultistepScheduler, solver_order=3, **{sigmas: True})
run(DPMSolverMultistepScheduler, solver_order=2, solver_type="heun", **{sigmas: True})
if sigmas != "use_lu_lambdas":
run(UniPCMultistepScheduler, solver_order=3, lower_order_final=False, **{sigmas: True})
run(DPMSolverMultistepScheduler, solver_order=3) # no conversion: the final step is a real step
run(DPMSolverMultistepScheduler, solver_order=3, use_karras_sigmas=True, num_inference_steps=14) # < 15 steps: first order
for name, config in [("karras", dict(use_karras_sigmas=True)), ("none", {})]:
s = DPMSolverMultistepScheduler(final_sigmas_type="sigma_min", **config)
s.set_timesteps(3)
print(name, "timesteps", s.timesteps.tolist(), "sigmas", [round(v, 4) for v in s.sigmas.tolist()])
Logs
# diffusers 0.40.0 and main at e0abab8
DPMSolverMultistepScheduler {'solver_order': 3, 'use_karras_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
DPMSolverMultistepScheduler {'solver_order': 2, 'solver_type': 'heun', 'use_karras_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
UniPCMultistepScheduler {'solver_order': 3, 'lower_order_final': False, 'use_karras_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
DPMSolverMultistepScheduler {'solver_order': 3, 'use_exponential_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
DPMSolverMultistepScheduler {'solver_order': 2, 'solver_type': 'heun', 'use_exponential_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
UniPCMultistepScheduler {'solver_order': 3, 'lower_order_final': False, 'use_exponential_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
DPMSolverMultistepScheduler {'solver_order': 3, 'use_beta_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
DPMSolverMultistepScheduler {'solver_order': 2, 'solver_type': 'heun', 'use_beta_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
UniPCMultistepScheduler {'solver_order': 3, 'lower_order_final': False, 'use_beta_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
DPMSolverMultistepScheduler {'solver_order': 3, 'use_lu_lambdas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
DPMSolverMultistepScheduler {'solver_order': 2, 'solver_type': 'heun', 'use_lu_lambdas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 100%
DPMSolverMultistepScheduler {'solver_order': 3} last sigmas [0.1760101169347763, 0.010001329705119133] nan 0%
DPMSolverMultistepScheduler {'solver_order': 3, 'use_karras_sigmas': True} last sigmas [0.010001329705119133, 0.010001329705119133] nan 0%
karras timesteps [999, 593, 0] sigmas [157.4073, 5.9113, 0.01, 0.01]
none timesteps [999, 666, 333] sigmas [157.4073, 9.4889, 1.4624, 0.01]
System Info
- diffusers 0.40.0 (PyPI) and main at e0abab8
- torch 2.14.0+cpu, numpy 2.2.6
- Python 3.13, Linux
AI disclosure: I used an AI coding agent to help find this, to write the reproduction, the patch, the tests and this report. I have read and checked the report, the patch and the tests myself and I will answer questions personally.
Proposed fix
Two ways to go, and I would like your view before opening a PR:
- Take the zero-width final step as first order.
DPMSolverMultistepScheduler.step gets or self.sigmas[self.step_index + 1] == self.sigmas[self.step_index] in the lower_order_final condition, and UniPCMultistepScheduler.step sets this_order = 1 on such a final step. The first-order update returns the sample unchanged there, the same as the lower orders already do, so nothing that is finite today changes: I ran both schedulers over 572 configurations (orders 1 to 3, both solver types, both algorithm types, both final_sigmas_type values, all conversions, 10 and 20 steps, euler_at_final on and off) and every finite output is bit-identical, the 36 NaN configurations become finite. That is the patch below, on Nicholas022400701:fix/zero-width-final-step, with a regression test per scheduler that fails on main and passes with the patch. pytest tests/schedulers/test_scheduler_dpm_multi.py tests/schedulers/test_scheduler_unipc.py tests/schedulers/test_scheduler_dpm_multi_inverse.py: 144 passed, 1 skipped. ruff check, ruff format --check and python utils/check_copies.py are clean.
- Make the final step a real step: build the converted schedule with
num_inference_steps + 1 points and drop the last one when final_sigmas_type == "sigma_min", so the appended sigma_min lies one step below. That spends the last model evaluation usefully, but it changes the output of every sigma_min plus converted-schedule configuration, which is exactly what that option exists to preserve, so I did not go that way.
Patch
git diff against main (4 files, +70 -1), branch https://github.com/Nicholas022400701/diffusers/tree/fix/zero-width-final-step
diff --git a/src/diffusers/schedulers/scheduling_dpmsolver_multistep.py b/src/diffusers/schedulers/scheduling_dpmsolver_multistep.py
index edcfa45..12877a5 100644
--- a/src/diffusers/schedulers/scheduling_dpmsolver_multistep.py
+++ b/src/diffusers/schedulers/scheduling_dpmsolver_multistep.py
@@ -1235,11 +1235,14 @@ class DPMSolverMultistepScheduler(SchedulerMixin, ConfigMixin):
if self.step_index is None:
self._init_step_index(timestep)
- # Improve numerical stability for small number of steps
+ # Improve numerical stability for small number of steps. The converted sigma schedules (Karras, exponential,
+ # beta, Lu lambdas) already end at sigma_min, so with `final_sigmas_type="sigma_min"` the final step has zero
+ # width and the higher-order updates would divide by h = 0; take it as a first-order step as well.
lower_order_final = (self.step_index == len(self.timesteps) - 1) and (
self.config.euler_at_final
or (self.config.lower_order_final and len(self.timesteps) < 15)
or self.config.final_sigmas_type == "zero"
+ or bool(self.sigmas[self.step_index + 1] == self.sigmas[self.step_index])
)
lower_order_second = (
(self.step_index == len(self.timesteps) - 2) and self.config.lower_order_final and len(self.timesteps) < 15
diff --git a/src/diffusers/schedulers/scheduling_unipc_multistep.py b/src/diffusers/schedulers/scheduling_unipc_multistep.py
index 5c2cbcc..67aa655 100644
--- a/src/diffusers/schedulers/scheduling_unipc_multistep.py
+++ b/src/diffusers/schedulers/scheduling_unipc_multistep.py
@@ -1210,6 +1210,14 @@ class UniPCMultistepScheduler(SchedulerMixin, ConfigMixin):
else:
this_order = self.config.solver_order
+ # The converted sigma schedules (Karras, exponential, beta) already end at sigma_min, so with
+ # `final_sigmas_type="sigma_min"` the final step has zero width and the higher-order predictor would divide by
+ # h = 0; take it as a first-order step.
+ if self.step_index == len(self.timesteps) - 1 and bool(
+ self.sigmas[self.step_index + 1] == self.sigmas[self.step_index]
+ ):
+ this_order = 1
+
self.this_order = min(this_order, self.lower_order_nums + 1) # warmup for multistep
assert self.this_order > 0
diff --git a/tests/schedulers/test_scheduler_dpm_multi.py b/tests/schedulers/test_scheduler_dpm_multi.py
index 28c3547..bd36a5b 100644
--- a/tests/schedulers/test_scheduler_dpm_multi.py
+++ b/tests/schedulers/test_scheduler_dpm_multi.py
@@ -366,3 +366,30 @@ class DPMSolverMultistepSchedulerTest(SchedulerCommonTest):
def test_exponential_sigmas(self):
self.check_over_configs(use_exponential_sigmas=True)
+
+ def test_zero_width_final_step(self):
+ # the converted sigma schedules already end at sigma_min, so `final_sigmas_type="sigma_min"` repeats it and
+ # the final step has h = 0; the higher-order updates must not divide by it
+ scheduler_class = self.scheduler_classes[0]
+ for sigmas_kwarg in ["use_karras_sigmas", "use_exponential_sigmas", "use_beta_sigmas", "use_lu_lambdas"]:
+ for solver_order, solver_type in [(3, "midpoint"), (2, "heun")]:
+ scheduler_config = self.get_scheduler_config(
+ solver_order=solver_order,
+ solver_type=solver_type,
+ final_sigmas_type="sigma_min",
+ **{sigmas_kwarg: True},
+ )
+ scheduler = scheduler_class(**scheduler_config)
+ scheduler.set_timesteps(20)
+ assert scheduler.sigmas[-1] == scheduler.sigmas[-2]
+
+ model = self.dummy_model()
+ sample = self.dummy_sample_deter
+ for t in scheduler.timesteps[:-1]:
+ sample = scheduler.step(model(sample, t), t, sample).prev_sample
+ t = scheduler.timesteps[-1]
+ prev_sample = scheduler.step(model(sample, t), t, sample).prev_sample
+
+ msg = f"{sigmas_kwarg}, solver_order={solver_order}, solver_type={solver_type}"
+ assert torch.isfinite(prev_sample).all(), msg
+ assert torch.allclose(prev_sample, sample), msg
diff --git a/tests/schedulers/test_scheduler_unipc.py b/tests/schedulers/test_scheduler_unipc.py
index ac7e1d3..a641926 100644
--- a/tests/schedulers/test_scheduler_unipc.py
+++ b/tests/schedulers/test_scheduler_unipc.py
@@ -284,6 +284,37 @@ class UniPCMultistepSchedulerTest(SchedulerCommonTest):
assert abs(result_sum.item() - 315.5757) < 1e-2, f" expected result sum 315.5757, but get {result_sum}"
assert abs(result_mean.item() - 0.4109) < 1e-3, f" expected result mean 0.4109, but get {result_mean}"
+ def test_zero_width_final_step(self):
+ # the converted sigma schedules already end at sigma_min, so `final_sigmas_type="sigma_min"` repeats it and
+ # the final step has h = 0; the third-order predictor must not divide by it. The corrector of the final step
+ # is disabled so that the zero-width predictor step can be checked to leave the sample unchanged.
+ scheduler_class = self.scheduler_classes[0]
+ num_inference_steps = 20
+ for sigmas_kwarg in ["use_karras_sigmas", "use_exponential_sigmas", "use_beta_sigmas"]:
+ for solver_type in ["bh1", "bh2"]:
+ scheduler_config = self.get_scheduler_config(
+ solver_order=3,
+ solver_type=solver_type,
+ lower_order_final=False,
+ final_sigmas_type="sigma_min",
+ disable_corrector=[num_inference_steps - 2],
+ **{sigmas_kwarg: True},
+ )
+ scheduler = scheduler_class(**scheduler_config)
+ scheduler.set_timesteps(num_inference_steps)
+ assert scheduler.sigmas[-1] == scheduler.sigmas[-2]
+
+ model = self.dummy_model()
+ sample = self.dummy_sample_deter
+ for t in scheduler.timesteps[:-1]:
+ sample = scheduler.step(model(sample, t), t, sample).prev_sample
+ t = scheduler.timesteps[-1]
+ prev_sample = scheduler.step(model(sample, t), t, sample).prev_sample
+
+ msg = f"{sigmas_kwarg}, solver_type={solver_type}"
+ assert torch.isfinite(prev_sample).all(), msg
+ assert torch.allclose(prev_sample, sample), msg
+
class UniPCMultistepScheduler1DTest(UniPCMultistepSchedulerTest):
@property
Describe the bug
set_timestepsappendssigma_lastbehind the schedule. Withfinal_sigmas_type="sigma_min"that issigma(t=0), and the converted schedules (use_karras_sigmas,use_exponential_sigmas,use_beta_sigmas,use_lu_lambdas) already end at exactly that value, because_convert_to_karrasand the other conversions ramp down tosigma_min = in_sigmas[-1], the smallest sigma of the training schedule. So the last two sigmas are identical and the final step hash = lambda_t - lambda_s0 = 0:The first-order and the midpoint second-order updates return the sample unchanged on that step, which is the legacy behavior
sigma_minwas kept for in #6477. The updates that divide byhreturn NaN for the whole batch instead:DPMSolverMultistepScheduler,solver_order=3:multistep_dpm_solver_third_order_updatecomputesr0, r1 = h_0 / h, h_1 / handr0 / (r0 + r1), which isinf / inf, and(exp(-h) - 1) / h, which is0 / 0. Same fordpmsolver++andsde-dpmsolver++. The docstring recommendssolver_order=3for unconditional sampling.DPMSolverMultistepScheduler,solver_order=2, solver_type="heun":(exp(-h) - 1) / h + 1is0 / 0.UniPCMultistepScheduler,solver_order=3, lower_order_final=False:h_phi_k = h_phi_1 / hh - 1is0 / 0andtorch.linalg.solvepropagates it intorhos_p.lower_order_finalonly protectsDPMSolverMultistepSchedulerbelow 15 steps,euler_at_finalis off by default, andfinal_sigmas_type="zero"is exempt becausestepalready forces the first-order update for it. So with the default config plussolver_order=3, a converted schedule andfinal_sigmas_type="sigma_min", 14 steps work and 15 or more steps come back all-NaN.The same duplicated sigma shows up in
DPMSolverSinglestepScheduler,DEISMultistepSchedulerandSASolverScheduler(which always appendsigma_min) and inEulerDiscreteScheduler(also with the defaultlinspacespacing, whose last timestep is 0), but their final updates do not divide by the step width, so there it only costs one model evaluation on a no-op step. #14887 is about the other end of the same line,final_sigmas_type="zero", wherehisinf.Reproduction
Logs
System Info
AI disclosure: I used an AI coding agent to help find this, to write the reproduction, the patch, the tests and this report. I have read and checked the report, the patch and the tests myself and I will answer questions personally.
Proposed fix
Two ways to go, and I would like your view before opening a PR:
DPMSolverMultistepScheduler.stepgetsor self.sigmas[self.step_index + 1] == self.sigmas[self.step_index]in thelower_order_finalcondition, andUniPCMultistepScheduler.stepsetsthis_order = 1on such a final step. The first-order update returns the sample unchanged there, the same as the lower orders already do, so nothing that is finite today changes: I ran both schedulers over 572 configurations (orders 1 to 3, both solver types, both algorithm types, bothfinal_sigmas_typevalues, all conversions, 10 and 20 steps,euler_at_finalon and off) and every finite output is bit-identical, the 36 NaN configurations become finite. That is the patch below, onNicholas022400701:fix/zero-width-final-step, with a regression test per scheduler that fails on main and passes with the patch.pytest tests/schedulers/test_scheduler_dpm_multi.py tests/schedulers/test_scheduler_unipc.py tests/schedulers/test_scheduler_dpm_multi_inverse.py: 144 passed, 1 skipped.ruff check,ruff format --checkandpython utils/check_copies.pyare clean.num_inference_steps + 1points and drop the last one whenfinal_sigmas_type == "sigma_min", so the appendedsigma_minlies one step below. That spends the last model evaluation usefully, but it changes the output of everysigma_minplus converted-schedule configuration, which is exactly what that option exists to preserve, so I did not go that way.Patch
git diff against main (4 files, +70 -1), branch https://github.com/Nicholas022400701/diffusers/tree/fix/zero-width-final-step