Nizam Radwan

Author directory

2026

This study compared machine learning, transformer fine-tuning, prompt-based LLM, and novel ensemble approaches for automated essay scoring in a Canadian large-scale provincial assessment. Fine-tuned transformers achieved the highest reliability, followed by machine learning models. Meanwhile, prompt-based LLMs provided greater explainability, and ensemble architectures highlighted opportunities for balancing reliability and explainability.