A Video-Based Reverse Dictionary for Sign Language Using Gesture Similarity

Batyrbek Orazumbekov, Daniyal Bayanov, Aruzhan Kaltay, Anara Sandygulova


Abstract
Sign language recognition systems are usually modeled as classification systems that map gesture videos to pre-defined glosses. But these systems do not allow similarity searches, where a user can search for similar gestures without knowing the corresponding gloss. This paper presents a pose-based video-to-video search framework for isolated signs, which acts as a reverse gesture dictionary. The system employs keypoints on the skeletal structure instead of RGB images. Two architectures are proposed for modeling temporal information: an encoder with self-attention in a Transformer architecture and a Spatial-Temporal Graph Convolutional Network (ST-GCN). The embedding space is optimized using metric learning objectives, including supervised contrastive learning and ArcFace angular margin loss. The performance of the retrieval system is evaluated on the WLASL dataset using ranking metrics like Recall@K and mean Average Precision (mAP). Experiments reveal that the temporal modeling using the Transformer architecture is an improvement over the graph-based modeling approach in the low-shot learning scenario. The attention-based temporal pooling approach further enhances the ranking quality, with the best-performing model achieving an mAP of 0.237 on the WLASL validation set. Cross-dataset evaluation on a 226-label AUTSL dataset reveals non-trivial generalization performance on the unseen dataset, despite training only on the WLASL dataset.
Anthology ID:
2026.signlang-1.41
Volume:
Proceedings of the LREC 2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion
Month:
May
Year:
2026
Address:
Palma, Mallorca (Spain)
Editors:
Eleni Efthimiou, Stavroula-Evita Fotinea, Thomas Hanke, Julie A. Hochgesang, Johanna Mesch, Marc Schulder
Venues:
SignLang | WS
SIG:
Publisher:
ELRA Language Resources Association (ELRA)
Note:
Pages:
398–407
Language:
External URL:
https://lrec.elra.info/lrec2026-ws-signlang-41
DOI:
10.63317/54ysrywg3ktr
Bibkey:
Cite (ACL):
Batyrbek Orazumbekov, Daniyal Bayanov, Aruzhan Kaltay, and Anara Sandygulova. 2026. A Video-Based Reverse Dictionary for Sign Language Using Gesture Similarity. In Proceedings of the LREC 2026 12th Workshop on the Representation and Processing of Sign Languages: Language in Motion, pages 398–407, Palma, Mallorca (Spain). ELRA Language Resources Association (ELRA).
Cite (Informal):
A Video-Based Reverse Dictionary for Sign Language Using Gesture Similarity (Orazumbekov et al., SignLang 2026)
Copy Citation: