Critical analysis of datasets for sign language translation
IntroductionIn recent years, significant progress has been made in Machine Translation (MT), including multilingual and low-resource settings. However, Sign Language Translation (SLT) remains underdeveloped, largely due to the scarcity of high-quality datasets and the overreliance on a few small, widely used benchmarks. This study aims to critically assess the datasets most commonly used in SLT research to determine whether their characteristics may lead to overfitting and misleading evaluation results.MethodsWe then conduct a detailed empirical study comparing training and test set similarity for PHOENIX14T, CSL-Daily, and LSE-Health. Using both gloss-based (TwoStream-SLT) and gloss-free (GFSLT-VLP) models, we evaluate the extent to which models memorize training data and how this affects BLEU scores.ResultsOur analysis reveals that PHOENIX14T exhibits substantial overlap between training and test sets, leading to inflated BLEU scores and can even mask signs of overfitting. CSL-Daily shows less overlap and more robust generalization. We also show that a small subset of “training-like” sentences disproportionately contributes to BLEU scores.DiscussionWe recommend that future SLT research move away from overused benchmarks and adopt larger, more diverse datasets such as How2Sign, CSL-News, and FLEURS-ASL. We also advocate for a shift toward gloss-free approaches and more careful interpretation of evaluation metrics, especially in low-resource settings.

Putting sign language AI into users’ hands
Introducing sign-language-to-text (SL2T), our breakthrough model powering new sign language features for Deaf and hard of hearing users.
Collection
Sources for the Aug. 13 Sensemaker brief on Google DeepMind's SL2T sign-language-to-text release and its low-stakes deployment boundary.