针对手语识别任务中标注样本稀缺与特征表达能力不足的问题,提出一种基于自监督学习的手语视觉特征预训练方法。通过构建空间—时间联合增强机制与对比学习框架,实现了在无监督条件下的时空特征表征优化;设计融合C3D与Transformer的双分支编码结构,以兼顾局部动作动态与全局语义依赖关系;引入遮挡重建辅助任务,提高模型对关键动作区域与缺失信息的恢复能力。在预训练完成后,采用冻结式迁移策略,将特征迁移至下游手语识别任务,并结合BiLSTM时序建模与CTC分类解码实现整体系统构建。实验结果表明,该方法在RWTH-PHOENIX-Weather 2014T数据集上识别准确率达到85.7%,字符错误率降至10.3%,较传统迁移学习方案显著提升,验证了所提自监督预训练机制在低标注资源场景下的有效性与泛化优势。
Abstract
To address the issues of scarce annotated samples and insufficient feature representation capability in sign language recognition tasks, this study proposes a pre-training method for sign language visual features based on self-supervised learning. By constructing a spatial-temporal joint augmentation mechanism and a contrastive learning framework, spatiotemporal feature representation is optimized under unsupervised conditions. A dual-branch encoding architecture integrating C3D and Transformer is designed to balance local motion dynamics and global semantic dependencies. An occlusion reconstruction auxiliary task is introduced to enhance the model’s ability to recover key motion regions and missing information. After pre-training, a frozen transfer strategy is adopted to migrate features to downstream sign language recognition tasks, combined with BiLSTM-based temporal modeling and CTC-based classification decoding to build the overall system. Experimental results on the RWTH-PHOENIX-Weather 2014T dataset demonstrate a recognition accuracy of 85.7% and a character error rate reduced to 10.3%, significantly outperforming traditional transfer learning approaches and validating the effectiveness and generalization advantages of the proposed self-supervised pre-training mechanism in low-annotation resource scenarios.
关键词
自监督学习 /
手语识别 /
特征预训练 /
时序建模 /
视觉编码
Key words
self-supervised learning /
sign language recognition /
feature pre-training /
temporal modeling /
visual encoding
{{custom_sec.title}}
{{custom_sec.title}}
{{custom_sec.content}}
参考文献
[1] 张国晨,温瑞,井正宏,等.基于多模态传感器与深度学习的双向智能手语识别手套研究[J].物联网技术,2025,15(19):22-25+29.
[2] 吕宇堃. 基于深度学习的手语识别综述[J].信息与电脑,2025,37(17):54-56.
[3] 马旭冉,陈圣贤,李熙文,等.基于改进YOLOv5算法的智能手语翻译方法研究[J].电脑知识与技术,2025,21(24):36-39.
[4] 董欣. AI手语主播在新闻报道中的应用探究[J].采写编,2025(8):10-12.
[5] 佟婷,王书芹,潘佳雯,等.融合VGG16和关键点信息的手语识别研究[J].软件工程,2025,28(8):22-25+47.
[6] 刘开林,侯永宏,郭子慧.用于连续手语识别的语义引导时空表征学习[J].华中科技大学学报(自然科学版),2025,54(2):161-167.
[7] 郑淇阳,简彩仁.基于多模态的手语视频识别[J].物联网技术,2025,15(23):18-24.
[8] 张磊,王振宇,连帅帅,等.基于深度学习的手语翻译:过去、现状与未来[J].计算机应用研究,2025,42(8):2241-2254.
[9] 郭乐铭. 连续手语识别的视觉模型研究[D].天津:天津理工大学,2024.
[10] 顾楠. 基于注意力自监督对比学习的线索语手形特征提取研究[D].天津:天津大学,2021.