Bring PyTorch operators to Apple Silicon with native kernels.
GPU kernel development
Implement native Metal kernels and MPS dispatch for PyTorch operators on Apple Silicon.
When to use it
Use when adding MPS support, implementing Metal shaders, or porting CUDA kernels to Apple Silicon.
Give it a PyTorch project and an operator to add or migrate; it updates dispatch, writes the Metal kernel, and implements the host-side stub.
What you provide
No additional actions listed in the analysis.
The workflow implements and runs Metal kernels on Apple Silicon.
Ninja is used to compile the torch_cpu target after implementation.
Python is used to run the specific operator test or the full MPS test suite.
NumPy enables the optional strategy of comparing individual GPU-kernel output with a CPU reference.