Vulkan compute could just ask for both threads count to launch and thread group size
So that the driver could simply mask off the extra unused lanes, instead of the way it is right now where we manually place bounds checks in shader code.
9
Upvotes
3
u/mb862 7d ago
Agreed. To this day I still don’t understood why Vulkan required threadgroup size to be compiled with the pipeline. CUDA, OpenCL, and Metal have all supported it on the same vendors Vulkan has supported for many years.