IJIET 2026 Vol.16(9): 2378-2386
doi: 10.18178/ijiet.2026.16.9.2696
doi: 10.18178/ijiet.2026.16.9.2696
Integrating Transformer-based and Embedding Models into Rasa NLU for Vietnamese University Support System
Le Ba Cuong *, Le Anh Tien, and Pham Van Huong
Information Technology Department, Academy of Cryptography Techniques, Hanoi, Vietnam
Email: cuonglb304@gmail.com (L.B.C.); tienla@actvn.edu.vn (L.A.T.); huongpv@gmail.com (P.V.H.)
*Corresponding author
Email: cuonglb304@gmail.com (L.B.C.); tienla@actvn.edu.vn (L.A.T.); huongpv@gmail.com (P.V.H.)
*Corresponding author
Manuscript received December 9, 2025; revised January 16, 2026; accepted April 21, 2026; published September 10, 2026
Abstract—This paper presents a Vietnamese university support chatbot developed using the Rasa Natural Language Understanding (NLU) framework, integrating Transformer-based and embedding models, including PhoBERT, FastText, Multilingual BERT (mBERT), and additional baseline methods such as Support Vector Machine (SVM) and Naive Bayes. The system is trained on a domain-specific dataset consisting of 99 intents and 1773 annotated examples covering academic and administrative queries. To ensure reliable evaluation, all models are assessed using 5-fold cross-validation. Experimental results show that PhoBERT achieves the best performance with an average accuracy of approximately 90.5% and an F1-Score of 90.1%, significantly outperforming both traditional machine learning methods and multilingual Transformer models. Among baseline approaches, SVM demonstrates strong performance, highlighting the effectiveness of classical models under limited data conditions. Further analysis using confusion patterns reveals that most misclassifications occur between semantically similar intents, emphasizing the challenges of fine-grained intent classification in Vietnamese. The results confirm that language-specific pretraining plays a crucial role in improving performance in low-resource settings. This study provides an empirical evaluation of multiple modeling approaches under consistent experimental conditions and demonstrates the potential of Transformer-based models for Vietnamese university support systems, while highlighting limitations related to dataset size and intent overlap.
Keywords—chatbot, education, transformer, artificial intelligence, natural language processing, Bidirectional Encoder Representations from Transformers (BERT), rasa
Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).
Keywords—chatbot, education, transformer, artificial intelligence, natural language processing, Bidirectional Encoder Representations from Transformers (BERT), rasa
Cite: Le Ba Cuong, Le Anh Tien, and Pham Van Huong, "Integrating Transformer-based and Embedding Models into Rasa NLU for Vietnamese University Support System," International Journal of Information and Education Technology, vol. 16, no. 9, pp. 2378-2386, 2026.
Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).