Empirical Evaluation of Mimetic Assessment for Large Language Models in News Teaching Multitasking Scenarios

ZHENG Xiao-ying, YANG Yun, LI Jiang, YANG Zhi-qi

Computer & Telecommunication ›› 2025 ›› Issue (11) : 48-54.

Computer & Telecommunication ›› 2025 ›› Issue (11) : 48-54.

Empirical Evaluation of Mimetic Assessment for Large Language Models in News Teaching Multitasking Scenarios

  • ZHENG Xiao-ying1,2, YANG Yun1, LI Jiang1, YANG Zhi-qi3
Author information +
History +

Abstract

Large Language Models have made intelligent assessment possible,but a gap remains between technical feasibility and practical application.Currently,there is a lack of systematic verification of their stability in comprehensive scenario-based evaluation.Using journalism education as the context,this study constructs evaluation tasks progressing through three cognitive levels-from theoretical application and structural analysis to script creation-and selects three mainstream Chinese Large Language Models with comparable parameter scales:ChatGLM-4-9B,Qwen2.5-14B,and Baichuan2-13B.By having each model independently score the same assignment 100 times,we systematically verify the stability performance of model-based mimetic assessment.The findings reveal that:stability varies significantly across different models;scoring stability shows a significant negative correlation with task complexity;the semantic stability of evaluation comments aligns with scoring fluctuations,with different models forming distinctive evaluative discourse systems;and semantic space analysis reveals that models exhibit differentiated adaptation strategies in complex tasks. This study explores the applicable boundaries of Large Language Models in educational assessment,providing empirical evidence and practical guidance for the intelligent application of journalism education.

Key words

Large Language Models / educational assessment / scoring stability / journalism education / mimetic evaluation

Cite this article

Download Citations
ZHENG Xiao-ying, YANG Yun, LI Jiang, YANG Zhi-qi. Empirical Evaluation of Mimetic Assessment for Large Language Models in News Teaching Multitasking Scenarios[J]. Computer & Telecommunication. 2025(11): 48-54

References

[1] 李运福. 人工智能赋能高等教育评价改革的国际借鉴[J].电化教育研究,2025,46(2):32-40.
[2] 邓筱小. 课程思政视域下新闻传播类专业课程教学模式与评价体系改革研究[J].新闻研究导刊,2023,14(5):74-77.
[3] 赵平,胡咏梅."双减"背景下中小学教师减负:问题、成因与对策[J].首都师范大学学报(社会科学版),2023(5):151-161.
[4] Mizumoto A,Eguchi M.Exploring the Potential of Using an AI Language Model for Automated Essay Scoring[J].Research Methods in Applied Linguistics, 2023,2(2):100050.
[5] Yavuz F,Çelik Ö,Yavaş Çelik G.Utilizing Large Language Models for EFL Essay Grading:An Examination of Reliability and Validity in Rubric-Based Assessments[J]. British Journal of Educational Technology,2025,56(1):150-166.
[6] Seo H,Hwang T,Jung J,et al.Large Language Models as Evaluators in Education:Verification of Feedback Consistency and Accuracy[J].Applied Sciences,2025,15(2):671.
[7] 马思腾,时广军,王琦.人工智能技术反噬教育公平现象及其矫正[J].中国远程教育,2025(7):98-114.
[8] Lu C,Cutumisu M.Integrating Deep Learning into an Automated Feedback Generation System for Automated Essay Scoring[J].International Educational Data Mining Society, 2021.
[9] Ridley R,He L,Dai X Y,et al.Automated Cross- Prompt Scoring of Essay Traits[J].Proceedings of the AAAI Conference on Artificial Intelligence,2021,35(15):13745-13753.
[10] Cai Y,Liang K,Lee S,et al.Rank-then- Score:Enhancing Large Language Models for Automated Essay Scoring[EB/OL].2025:arXiv:2504.05736. https://arxiv.org/abs/2504.05736[LinkOut]
[11] Hussein M A,Hesham A,Nassef M.A Trait-Based Deep Learning Automated Essay Scoring System with Adaptive Feedback[J].International Journal of Advanced Computer Science and Applications,2020,11(5):287-293.
[12] Taghipour K,Ng H T.A Neural Approach to Automated Essay Scoring[C]//Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing.Association for Computational Linguistics,2016: 1882-1891.
[13] Tang X,Chen H,Lin D,et al.Harnessing LLMS for Multi-Dimensional Writing Assessment:Reliability and Alignment with Human Judgments[J].Heliyon,2024, 10(14):e34262.
[14] Kim S,Jo M.Is GPT-4 Alone Sufficient for Automated Essay Scoring?:A Comparative Judgment Approach Based on Rater Cognition[C]//Proceedings of the Eleventh ACM Conference on Learning@Scale.ACM, 2024:315-319.
[15] Nguyen H T,Goebel R,Toni F,et al.Black-Box Analysis:GPTS across Time in Legal Textual Entailment Task[EB/OL].2023:arXiv:2309.05501.https://arxiv.org/abs/2309.05501.
[16] Kim S,Kim S.Can Language Models Evaluate Human Written Text?Case Study on Korean Student Writing for Education[J].CoRR, 2024.
[17] Ji Z,Lee N,Frieske R,et al.Survey of Hallucination in Natural Language Generation[J].ACM Computing Surveys,2023,55(12):1-38.
[18] Chen L,Zaharia M,Zou J.How Is ChatGPT's Behavior Changing over Time?[J].Harvard Data Science Review,2024,6(2).
[19] Quah B,Zheng L,Sng T J H,et al.Reliability of ChatGPT in Automated Essay Scoring for Dental Undergraduate Examinations[J].BMC Medical Education, 2024,24(1):962.
[20] Jauhiainen J S,Guerra A G.Evaluating Students'Open-Ended Written Responses with LLMS:Using the RAG Framework for GPT-3.5,GPT-4,Claude-3,and Mistral-Large[J].Advances in Artificial Intelligence and Machine Learning,2024,4(4):3097-3113.
[21] Nunes D,Primi R,Pires R,et al.Evaluating GPT- 3.5 and GPT-4 Models on Brazilian University Admission Exams[EB/OL].2023:arXiv:2303.17003.https://arxiv.org/abs/2303.17003
[22] Yang K,Raković M,Gašević D,et al.Does ThePrompt-Based Large Language Model Recognize Students'Demographics andIntroduce Bias inEssay Scoring?[C]//Artificial Intelligence in Education.Cham: Springer,2025:75-89.
[23] Santos R,Silva J R,Gomes L,et al.Advancing Generative AI for Portuguese with Open Decoder Gerv á sio PT[C]//Proceedings of the 3rd Annual Meeting of the Special Interest Group on Under-resourced Languages@LREC-COLING 2024.Association for Computational Linguistics,2024:16-26.
[24] Choi J H,Hickman K E,Monahan A,et al.ChatGPT Goes to Law School[J].SSRN Electronic Journal,2022,71(3):387.[LinkOut]
[25] Zhou Y,Liu X,Ning C,et al.MultifacetEval: Multifaceted Evaluation to Probe LLMS in Mastering Medical Knowledge[EB/OL].2024:arXiv:2406.02919. https://arxiv.org/abs/2406.02919.
[26] Su J,Yan Y,Gao Z,et al.CAFES:A Collaborative Multi-Agent Framework for Multi-Granular Multimodal Essay Scoring[EB/OL].2025:arXiv:2505.13965. https://arxiv.org/abs/2505.13965.

Accesses

Citation

Detail

Sections
Recommended

/

〈 〉