2026-07-01
2026-03-27
2025-12-30
Manuscript received April 16, 2026; accepted June 22, 2026; published July 27, 2026.
Abstract—Automated Essay Scoring (AES) using Large Language Models (LLMs) presents a significant challenge due to its unstable behavior, which oscillates between excessive leniency on subjective criteria, such as argumentative depth, and undue severity on formal errors. This misalignment with nuanced human cognition undermines the pedagogical purpose of automated assessment. To address this issue, this paper proposes a hybrid agentic framework to calibrate algorithmic judgment, introducing the Hybrid Agentic Assessment System (SAHA) as its proof-of-concept implementation. The architecture of SAHA goes beyond the monolithic application of LLMs by employing multiple AI agents, each with a distinct evaluative persona. These agents engage in a deliberative debate to analyze the text from different perspectives, modeling the cognitive conflict inherent in human evaluation. This multi-agent deliberation aims to produce a more balanced and robust assessment. A key innovation of this framework is the targeted integration of human expertise, identifying specific points of disagreement among agents. This approach seeks to ensure that technology serves as an intelligent support in the educational process. Keywords—Automated Essay Scoring (AES), Large Language Models (LLMs), agentic Artificial Intelligence (AI), human-AI collaboration, algorithmic judgment Cite: Daisy C. A. Silva and Sérgio S. C. Silva, "Calibrating Algorithmic Judgment: A Hybrid Agentic Framework for Subjective Competency Assessment with Human-AI Collaboration," International Journal of Learning and Teaching, Vol. 12, No. 3, pp. 221-225, 2026. Copyright © 2026 by the authors. This is an open access article distributed under the Creative Commons Attribution License which permits unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited (CC BY 4.0).