normalized Dynamic Time Warping
归一化动态时间规整nDTWAdvancedA 0-to-1 navigation metric for how closely an agent's path matches the shape and order of a reference path.
nDTW was proposed by Ilharco, Baldridge, and colleagues in 2019 to evaluate agents that navigate by following language instructions. Success Rate only checks the endpoint, and SPL only checks whether the path avoided needless detours — neither cares whether the agent actually followed the route the instruction described. nDTW borrows Dynamic Time Warping (DTW) from time-series analysis, which aligns two sequences of different lengths in order and sums the distance between aligned points, to compute the minimum cumulative distance between the agent's path and the reference path. That distance is then passed through an exponential function to normalize it to a 0–1 score, where higher is better: it penalizes deviation smoothly and stays sensitive to the order in which places are visited. The authors' human evaluation found it tracks human rankings better than other metrics. The same paper also proposes SDTW, which scores 0 on failed episodes and nDTW on successful ones. It is commonly reported alongside NE, SR, and SPL on R2R, R4R, and VLN-CE.
ExampleAn instruction says to enter the kitchen before going to the bedroom, but the agent walks straight to the bedroom — it reaches the right endpoint, so Success Rate still counts it as a success, but because the route skipped the kitchen, nDTW would score it noticeably lower.
- Also called
- nDTW, normalized DTW
- Related
- Navigation Error / Oracle Success Rate / Trajectory Length · Success weighted by Path Length · Vision-and-Language Navigation · Room-to-Room · Success Rate
- Sources
- General Evaluation for Instruction Conditioned Navigation using Dynamic Time Warping (arXiv 1907.05446)
Beyond the Nav-Graph: Vision-and-Language Navigation in Continuous Environments (VLN-CE)