doi: 10.18178/ijiet.2026.16.9.2697
EduDesignEval: A Structured Evaluation Framework for Educational Program Designer Using Large Language Models
- 1Department of Computer Engineering, Astana IT University, Astana, Kazakhstan
- 2Department of Computational and Data Science, Astana IT University, Astana, Kazakhstan
- 3Department of Information Technologies, Kyiv National University of Construction and Architecture, Kyiv, Ukraine
- 4Department of Social Disciplines, Astana IT University, Astana, Kazakhstan
- 5Department of Physics and Computer Science, Shakarim State University, Semey, Kazakhstan
- Manuscript receivedJanuary 16, 2026
- revisedMarch 4, 2026
- acceptedMarch 23, 2026
- publishedSeptember 15, 2026
Abstract
The rapid advancement of Large Language Models (LLMs) has created unprecedented opportunities for transforming educational program design. However, the lack of standardized evaluation frameworks raises concerns regarding the content quality and pedagogical effectiveness. This study proposes EduDesignEval, a comprehensive evaluation framework designed to assess LLM-generated educational content across multiple dimensions of quality. The framework defines three metrics: groundedness for measuring factual accuracy and hallucination mitigation, coherence for evaluating logical flow and thematic integrity, and fluency for assessing linguistic naturalness and readability. These metrics are synthesized into a unified score that provides an adaptable, weighted assessment based on specific educational requirements. We first validated the EduDesignEval framework through metric definitions and demonstrated its utility by employing ten publicly available LLM models across two critical tasks: text generation for learning goals and course descriptions, and multi-label classification for professional standards and competency frameworks, using a dataset derived from over 8,200 educational programs across 150 Kazakhstani institutions. Our fine-tuned model has demonstrated promising zero-shot classification performance with high groundedness while maintaining balanced coherence and fluency, outperforming larger models utilizing one-shot prompting. Preliminary results show that big models excel in groundedness, while smaller models show limitations in classification due to redundancy and low coherence. EduDesignEval bridges gaps in literature by providing multi-level quality control, fostering human-Artificial Intelligence (AI) collaboration, and promoting equitable, adaptive education. Ultimately, this framework ensures reliable LLM integration, enhancing program design efficiency and pedagogical integrity for global educational systems.
Keywords
- large language models
- educational program design
- evaluation framework
- curriculum development
- ai in education
- text generation
- multi-label classification
How to Cite
Ayanbek Serikov, Andrii Biloshchytskyi, Aidos Mukhatayev, and Balgyn Zheldybayeva, "EduDesignEval: A Structured Evaluation Framework for Educational Program Designer Using Large Language Models," International Journal of Information and Education Technology, vol. 16, no. 9, pp. 2387-2397, 2026. https://doi.org/10.18178/ijiet.2026.16.9.2697
Copyright & License
Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).