Lightlines
By pytorch

metal-kernel

pytorch

Bring PyTorch operators to Apple Silicon with native kernels.

GPU kernel development

What it does

Implement native Metal kernels and MPS dispatch for PyTorch operators on Apple Silicon.

When to use it

Use when adding MPS support, implementing Metal shaders, or porting CUDA kernels to Apple Silicon.

How to use it

Give it a PyTorch project and an operator to add or migrate; it updates dispatch, writes the Metal kernel, and implements the host-side stub.

What you provide

  • operator to add or migrate
  • An existing project

Access · 0

No additional actions listed in the analysis.

What you need · 4

The workflow implements and runs Metal kernels on Apple Silicon.

Ninja is used to compile the torch_cpu target after implementation.

Python is used to run the specific operator test or the full MPS test suite.

NumPy enables the optional strategy of comparing individual GPU-kernel output with a CPU reference.


About this skill

Visibility
Public
Repository
pytorch/pytorch
Created
Oct 8, 2026
Updated
Oct 8, 2026
Files
1