Repository navigation
Expose backend options in the Python runtime - #23576
Conversation
🔗 Helpful Links🧪 See artifacts and rendered test results at hud.pytorch.org/pr/pytorch/executorch/23576
Note: Links to docs will display an error until the docs builds have been completed. ✅ No FailuresAs of commit 49f1807 with merge base d87fb11 ( This comment was automatically generated by Dr. CI and updates every 15 minutes. |
|
why do we need two paths for backend option setting, instead of only |
|
Because backends read options from two different places in C++, and neither replaces the other.
Some keys are meant for only one of the two. The Torch-TensorRT delegate, for example, accepts So this mirrors the C++ API, where both exist: |
ce2a946 to
c093a37
Compare
Gasoonjia
left a comment
There was a problem hiding this comment.
LGTM, thanks! A couple of small suggestions:
-
Type errors in
to_backend_options: any value that isn't bool/int goes throughpy::cast<std::string>, so a float orNoneraises an opaque pybindcast_errorthat doesn't name the offending key, andbytesis silently accepted as a string. The same applies to out-of-range ints and to a non-dict per-backend value (e.g.{"XnnpackBackend": 1}). Could we check the types explicitly and raiseTypeError/ValueErrorthat names the key? -
Tests: consider adding cases for an unsupported value type (e.g. float) and for
get_optionon an unknown backend. Intest_load_options_reach_the_backend, the mode=1 case doesn't by itself show that the option was applied; the mode=3 failure is what proves the plumbing. Splitting it into two tests or adding a comment would make that clearer.
Backends take options in two ways: process-wide through
executorch::runtime::set_option, and per load through the
LoadBackendOptionsMap that Module::load passes to every method. The C++,
Android and Apple APIs expose both. The Python runtime exposed neither, so
a Python program could not turn on an option such as a shared scratch
buffer, or pass a load-time option to a delegate.
Add the same two paths to the Python runtime:
runtime = Runtime.get()
runtime.backend_registry.set_option("XnnpackBackend", {"weight_cache_enabled": True})
program = runtime.load_program(
"model.pte",
backend_options={"XnnpackBackend": {"workspace_sharing_mode": 1}},
)
BackendRegistry.set_option and get_option wrap the C++ free functions.
Runtime.load_program takes backend_options, keyed by backend name, and
every method loaded from that program receives them, as with
Module::load. Option values are bools, ints or strings, with the same
length limits as the C++ API. Any other value type raises TypeError and an
out of range value raises ValueError, naming the option.
c093a37 to
49f1807
Compare
Problem
Backends take options in two ways. Some options apply to the whole process, set with
executorch::runtime::set_option. Others apply to one program, passed when it loads with aLoadBackendOptionsMap, asModule::loaddoes. The C++, Android and Apple APIs expose both. The Python runtime exposes neither, so Python code cannot turn on a backend option, such as a shared scratch buffer, or pass a load-time option to a delegate.Change
Add the same two paths to the Python runtime:
BackendRegistry.set_optionandget_optionwrap the C++ free functions of the same names.Runtime.load_programtakesbackend_options, keyed by backend name. Every method loaded from that program receives them, as withModule::load.TypeError, and an out of range value raisesValueError, naming the option.Test plan
New tests, using XNNPACK so they run in the CPU jobs:
set_optionthenget_optionreturns the new value, and the original value is restored afterwards.set_optionandget_option) raise errors.workspace_sharing_moderuns and matches eager.