L1 Loss
L1 损失CommonA loss function that measures error as the absolute value of the difference between a prediction and the ground truth.
L1 loss takes the absolute value of the difference between each prediction and its ground truth and sums it up; averaged per sample, it's called mean absolute error (MAE). Compared with L2 loss (mean squared error), which squares the error, L1 penalizes large errors linearly rather than blowing them up quadratically, so it's less thrown off by outliers — but its gradient has constant magnitude, so it doesn't automatically shrink as the prediction nears the target. From a probabilistic view, minimizing L1 loss corresponds to maximum likelihood estimation under the assumption that errors follow a Laplace distribution, and the optimal solution trends toward the median rather than the mean. In embodied AI, L1 is common for continuous action regression: ACT uses L1 for reconstructing action sequences, with the authors reporting it models action sequences more precisely than the more common L2; OpenVLA-OFT also switched to L1 regression to output continuous actions directly.
ExampleOpenVLA-OFT replaced OpenVLA's approach of generating discrete action tokens one at a time with parallel decoding, action chunking, and L1 regression for continuous action output, raising action-generation throughput 26-fold and lifting the average success rate across four LIBERO task suites from 76.5% to 97.1%.
- Also called
- Mean Absolute Error, MAE, Absolute Error Loss
- Related
- Mean Squared Error · Loss Function · Continuous Action Regression · Action Chunking with Transformers · OpenVLA-OFT · Action Chunking
- Sources
- Linear regression: Loss(Google Machine Learning Crash Course)
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)
Fine-Tuning Vision-Language-Action Models: Optimizing Speed and Success (OpenVLA-OFT, arXiv 2502.19645)