Embodied AI Glossary中文

Operator / Kernel

算子Advanced

The basic computational building block of a neural network — like matmul or softmax — and its hardware implementation.

An operator is the smallest unit of computation in a deep learning framework — for example, matrix multiplication, convolution, normalization, or softmax. A model's forward pass is a computation graph made of operators wired together. A “kernel” is the actual implementation code for a given operator on a specific piece of hardware, such as a CUDA kernel that runs on a GPU. How fast a model runs at inference depends heavily on how well these kernels are written, and on whether several small operators can be “fused” into a single kernel to cut down on memory reads and writes. This fusing, and picking the fastest kernel, is largely what TensorRT and torch.compile do; when engineers talk about “operator support” while porting a model to a Chinese domestic chip, they mean whether that hardware has a matching implementation for each operator the model uses.

ExampleFlashAttention fuses several steps inside attention — matrix multiply, softmax, then multiplying by the value matrix — into a single CUDA kernel, substantially cutting memory access and speeding up inference on long sequences.

Also called
kernel, CUDA kernel
Related
CUDA · FlashAttention · NVIDIA TensorRT · torch.compile · Compute Architecture for Neural Networks (Huawei Ascend, CANN) · CUDA Graphs
Sources
PyTorch 文档:Custom C++ and CUDA Operators (Chinese)

See it in the full glossary →