Repository navigation
Conversation
Two things MSL does differently from every other backend. It has no work-item builtins: the grid dimensions arrive as kernel attributes. Without a branch of its own Metal fell through to the host definitions of get_group_id() and friends, which name iBlock and nBlocks, variables that only exist in the host-side loop. And it requires every kernel parameter to carry an attribute, so the sector index cannot be passed by value. GPUCA_KRNLGPU_DEF therefore gets two hooks, GPUCA_KRNL_SECTOR_ARG and GPUCA_KRNL_GRID_ARGS, which the Metal source fills in with a buffer and the four grid attributes. Both default to what the signature had, so CUDA, HIP and OpenCL generate exactly the same entry point as before. Kernel list diagnostics: 408 to 96, and the translation unit 1108 to 998. The 96 left are the 48 kernels that take arguments, which still need an answer for how Metal passes them.
Member
Author
|
@davidrohr any objections? in particular on how the GPUCA_KRNL_SECTOR_ARG and GPUCA_KRNL_GRID_ARGS were implemented? |
davidrohr
reviewed
Oct 8, 2026
| #endif | ||
| #define GPUCA_KRNLGPU_DEF(x_class, x_attributes, x_arguments, ...) \ | ||
| GPUg() void GPUCA_ATTRRES(GPUCA_M_STRIP(x_attributes)) GPUCA_M_CAT(krnl_, GPUCA_M_KRNL_NAME(x_class))(GPUCA_CONSMEM_PTR int32_t _iSector_internal GPUCA_M_STRIP(x_arguments)) | ||
| GPUg() void GPUCA_ATTRRES(GPUCA_M_STRIP(x_attributes)) GPUCA_M_CAT(krnl_, GPUCA_M_KRNL_NAME(x_class))(GPUCA_CONSMEM_PTR GPUCA_KRNL_SECTOR_ARG GPUCA_M_STRIP(x_arguments) GPUCA_KRNL_GRID_ARGS) |
Collaborator
There was a problem hiding this comment.
Hm, but this means that you pass in local and global id and size as argument to the kernel function.
However, in OpenCL / CUDA / HIP, these varaibles are available everywhere, without being passed in.
I.e., they are also available in subfunctions. And I don't want to pass them in explicitly to each place where they are used. Is this somehow possible with metal?
| // grid dimensions the same way, so the backend gets to shape both ends of the | ||
| // parameter list. | ||
| #ifndef GPUCA_KRNL_SECTOR_ARG | ||
| #define GPUCA_KRNL_SECTOR_ARG int32_t _iSector_internal |
Collaborator
There was a problem hiding this comment.
I don't understand why you need a special treatment for the sector variable in metal?
The sector variable is a normal variable, which is passed in like any other parameter to function calls.
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Two things MSL does differently from every other backend.
It has no work-item builtins: the grid dimensions arrive as kernel attributes.
Without a branch of its own Metal fell through to the host definitions of
get_group_id() and friends, which name iBlock and nBlocks, variables that only
exist in the host-side loop.
And it requires every kernel parameter to carry an attribute, so the sector
index cannot be passed by value.
GPUCA_KRNLGPU_DEF therefore gets two hooks, GPUCA_KRNL_SECTOR_ARG and
GPUCA_KRNL_GRID_ARGS, which the Metal source fills in with a buffer and the four
grid attributes. Both default to what the signature had, so CUDA, HIP and OpenCL
generate exactly the same entry point as before.
Kernel list diagnostics: 408 to 96, and the translation unit 1108 to 998. The 96
left are the 48 kernels that take arguments, which still need an answer for how
Metal passes them.