To address the issues of scarce annotated samples and insufficient feature representation capability in sign language recognition tasks, this study proposes a pre-training method for sign language visual features based on self-supervised learning. By constructing a spatial-temporal joint augmentation mechanism and a contrastive learning framework, spatiotemporal feature representation is optimized under unsupervised conditions. A dual-branch encoding architecture integrating C3D and Transformer is designed to balance local motion dynamics and global semantic dependencies. An occlusion reconstruction auxiliary task is introduced to enhance the model’s ability to recover key motion regions and missing information. After pre-training, a frozen transfer strategy is adopted to migrate features to downstream sign language recognition tasks, combined with BiLSTM-based temporal modeling and CTC-based classification decoding to build the overall system. Experimental results on the RWTH-PHOENIX-Weather 2014T dataset demonstrate a recognition accuracy of 85.7% and a character error rate reduced to 10.3%, significantly outperforming traditional transfer learning approaches and validating the effectiveness and generalization advantages of the proposed self-supervised pre-training mechanism in low-annotation resource scenarios.
Key words
self-supervised learning /
sign language recognition /
feature pre-training /
temporal modeling /
visual encoding
{{custom_sec.title}}
{{custom_sec.title}}
{{custom_sec.content}}
References
[1] 张国晨,温瑞,井正宏,等.基于多模态传感器与深度学习的双向智能手语识别手套研究[J].物联网技术,2025,15(19):22-25+29.
[2] 吕宇堃. 基于深度学习的手语识别综述[J].信息与电脑,2025,37(17):54-56.
[3] 马旭冉,陈圣贤,李熙文,等.基于改进YOLOv5算法的智能手语翻译方法研究[J].电脑知识与技术,2025,21(24):36-39.
[4] 董欣. AI手语主播在新闻报道中的应用探究[J].采写编,2025(8):10-12.
[5] 佟婷,王书芹,潘佳雯,等.融合VGG16和关键点信息的手语识别研究[J].软件工程,2025,28(8):22-25+47.
[6] 刘开林,侯永宏,郭子慧.用于连续手语识别的语义引导时空表征学习[J].华中科技大学学报(自然科学版),2025,54(2):161-167.
[7] 郑淇阳,简彩仁.基于多模态的手语视频识别[J].物联网技术,2025,15(23):18-24.
[8] 张磊,王振宇,连帅帅,等.基于深度学习的手语翻译:过去、现状与未来[J].计算机应用研究,2025,42(8):2241-2254.
[9] 郭乐铭. 连续手语识别的视觉模型研究[D].天津:天津理工大学,2024.
[10] 顾楠. 基于注意力自监督对比学习的线索语手形特征提取研究[D].天津:天津大学,2021.