新闻教学多任务场景下大语言模型的拟态评估实证

郑小英, 杨云, 李江, 杨志奇

电脑与电信 ›› 2025 ›› Issue (11) : 48-54.

电脑与电信 ›› 2025 ›› Issue (11) : 48-54.
大模型与知识图谱

新闻教学多任务场景下大语言模型的拟态评估实证

  • 郑小英1,2, 杨云1, 李江1, 杨志奇3
作者信息 +

Empirical Evaluation of Mimetic Assessment for Large Language Models in News Teaching Multitasking Scenarios

  • ZHENG Xiao-ying1,2, YANG Yun1, LI Jiang1, YANG Zhi-qi3
Author information +
文章历史 +

摘要

大语言模型为评价智能化提供了可能,但从技术可行性到实际应用仍存在差距,目前对其在综合场景评价的稳定性尚缺乏系统验证。以新闻专业教学为场景,构建了从理论应用、结构分析到脚本创作三个认知层次递进的评价任务,选取了ChatGLM-4-9B、Qwen2.5-14B和Baichuan2-13B三个参数规模相近的主流中文大语言模型,通过让每个模型对同一作业进行 100 次独立重复评分,系统验证了模型拟态评估的稳定性表现。研究发现:不同模型稳定性差异显著;评分稳定性与任务复杂度呈显著负相关;评语的语义稳定性与评分波动一致,不同模型形成独特的评价话语体系;语义空间分析揭示模型在复杂任务中呈现差异化的适应策略。探索了大语言模型在教育评价中的适用边界,为新闻教育智能化应用提供了实证依据和实践指导。

Abstract

Large Language Models have made intelligent assessment possible,but a gap remains between technical feasibility and practical application.Currently,there is a lack of systematic verification of their stability in comprehensive scenario-based evaluation.Using journalism education as the context,this study constructs evaluation tasks progressing through three cognitive levels-from theoretical application and structural analysis to script creation-and selects three mainstream Chinese Large Language Models with comparable parameter scales:ChatGLM-4-9B,Qwen2.5-14B,and Baichuan2-13B.By having each model independently score the same assignment 100 times,we systematically verify the stability performance of model-based mimetic assessment.The findings reveal that:stability varies significantly across different models;scoring stability shows a significant negative correlation with task complexity;the semantic stability of evaluation comments aligns with scoring fluctuations,with different models forming distinctive evaluative discourse systems;and semantic space analysis reveals that models exhibit differentiated adaptation strategies in complex tasks. This study explores the applicable boundaries of Large Language Models in educational assessment,providing empirical evidence and practical guidance for the intelligent application of journalism education.

关键词

大语言模型 / 教育评价 / 评分稳定性 / 新闻教学 / 拟态评估

Key words

Large Language Models / educational assessment / scoring stability / journalism education / mimetic evaluation

引用本文

导出引用
郑小英, 杨云, 李江, 杨志奇. 新闻教学多任务场景下大语言模型的拟态评估实证[J]. 电脑与电信. 2025(11): 48-54
ZHENG Xiao-ying, YANG Yun, LI Jiang, YANG Zhi-qi. Empirical Evaluation of Mimetic Assessment for Large Language Models in News Teaching Multitasking Scenarios[J]. Computer & Telecommunication. 2025(11): 48-54
中图分类号: TP18    G649.1   

参考文献

[1] 李运福. 人工智能赋能高等教育评价改革的国际借鉴[J].电化教育研究,2025,46(2):32-40.
[2] 邓筱小. 课程思政视域下新闻传播类专业课程教学模式与评价体系改革研究[J].新闻研究导刊,2023,14(5):74-77.
[3] 赵平,胡咏梅."双减"背景下中小学教师减负:问题、成因与对策[J].首都师范大学学报(社会科学版),2023(5):151-161.
[4] Mizumoto A,Eguchi M.Exploring the Potential of Using an AI Language Model for Automated Essay Scoring[J].Research Methods in Applied Linguistics, 2023,2(2):100050.
[5] Yavuz F,Çelik Ö,Yavaş Çelik G.Utilizing Large Language Models for EFL Essay Grading:An Examination of Reliability and Validity in Rubric-Based Assessments[J]. British Journal of Educational Technology,2025,56(1):150-166.
[6] Seo H,Hwang T,Jung J,et al.Large Language Models as Evaluators in Education:Verification of Feedback Consistency and Accuracy[J].Applied Sciences,2025,15(2):671.
[7] 马思腾,时广军,王琦.人工智能技术反噬教育公平现象及其矫正[J].中国远程教育,2025(7):98-114.
[8] Lu C,Cutumisu M.Integrating Deep Learning into an Automated Feedback Generation System for Automated Essay Scoring[J].International Educational Data Mining Society, 2021.
[9] Ridley R,He L,Dai X Y,et al.Automated Cross- Prompt Scoring of Essay Traits[J].Proceedings of the AAAI Conference on Artificial Intelligence,2021,35(15):13745-13753.
[10] Cai Y,Liang K,Lee S,et al.Rank-then- Score:Enhancing Large Language Models for Automated Essay Scoring[EB/OL].2025:arXiv:2504.05736. https://arxiv.org/abs/2504.05736[LinkOut]
[11] Hussein M A,Hesham A,Nassef M.A Trait-Based Deep Learning Automated Essay Scoring System with Adaptive Feedback[J].International Journal of Advanced Computer Science and Applications,2020,11(5):287-293.
[12] Taghipour K,Ng H T.A Neural Approach to Automated Essay Scoring[C]//Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing.Association for Computational Linguistics,2016: 1882-1891.
[13] Tang X,Chen H,Lin D,et al.Harnessing LLMS for Multi-Dimensional Writing Assessment:Reliability and Alignment with Human Judgments[J].Heliyon,2024, 10(14):e34262.
[14] Kim S,Jo M.Is GPT-4 Alone Sufficient for Automated Essay Scoring?:A Comparative Judgment Approach Based on Rater Cognition[C]//Proceedings of the Eleventh ACM Conference on Learning@Scale.ACM, 2024:315-319.
[15] Nguyen H T,Goebel R,Toni F,et al.Black-Box Analysis:GPTS across Time in Legal Textual Entailment Task[EB/OL].2023:arXiv:2309.05501.https://arxiv.org/abs/2309.05501.
[16] Kim S,Kim S.Can Language Models Evaluate Human Written Text?Case Study on Korean Student Writing for Education[J].CoRR, 2024.
[17] Ji Z,Lee N,Frieske R,et al.Survey of Hallucination in Natural Language Generation[J].ACM Computing Surveys,2023,55(12):1-38.
[18] Chen L,Zaharia M,Zou J.How Is ChatGPT's Behavior Changing over Time?[J].Harvard Data Science Review,2024,6(2).
[19] Quah B,Zheng L,Sng T J H,et al.Reliability of ChatGPT in Automated Essay Scoring for Dental Undergraduate Examinations[J].BMC Medical Education, 2024,24(1):962.
[20] Jauhiainen J S,Guerra A G.Evaluating Students'Open-Ended Written Responses with LLMS:Using the RAG Framework for GPT-3.5,GPT-4,Claude-3,and Mistral-Large[J].Advances in Artificial Intelligence and Machine Learning,2024,4(4):3097-3113.
[21] Nunes D,Primi R,Pires R,et al.Evaluating GPT- 3.5 and GPT-4 Models on Brazilian University Admission Exams[EB/OL].2023:arXiv:2303.17003.https://arxiv.org/abs/2303.17003
[22] Yang K,Raković M,Gašević D,et al.Does ThePrompt-Based Large Language Model Recognize Students'Demographics andIntroduce Bias inEssay Scoring?[C]//Artificial Intelligence in Education.Cham: Springer,2025:75-89.
[23] Santos R,Silva J R,Gomes L,et al.Advancing Generative AI for Portuguese with Open Decoder Gerv á sio PT[C]//Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages@LREC-COLING 2024.Association for Computational Linguistics,2024:16-26.
[24] Choi J H,Hickman K E,Monahan A,et al.ChatGPT Goes to Law School[J].SSRN Electronic Journal,2022,71(3):387.[LinkOut]
[25] Zhou Y,Liu X,Ning C,et al.MultifacetEval: Multifaceted Evaluation to Probe LLMS in Mastering Medical Knowledge[EB/OL].2024:arXiv:2406.02919. https://arxiv.org/abs/2406.02919.
[26] Su J,Yan Y,Gao Z,et al.CAFES:A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring[EB/OL].2025:arXiv:2505.13965. https://arxiv.org/abs/2505.13965.

基金

云南省教育厅知识工程与智能教育科技创新团队; 师范生A1素养分层培育与教学应用能力提升研究(LSJG202502)

Accesses

Citation

Detail

段落导航
相关文章

/

〈 〉