Introducing Automated Text Scoring
Since 2025, NZQA has used Automated Text Scoring (ATS) to help mark Literacy Writing (US 32405) assessments. ATS is used alongside human marking, quality checks, and review processes.
ATS uses artificial intelligence (AI) including machine learning and natural language processing to support assessment marking.
Why we introduced AI marking
We introduced ATS as part of a wider project to improve result turnaround times.
Schools told us that faster results for Literacy, Te Reo Matatini, Numeracy and Pāngarau assessments help them better support students before their next assessment opportunity.
In 2025, ATS and other process improvements helped NZQA return results to schools and students between 3.5 and 4.5 weeks sooner than in 2024.
Testing, research and human markers
Before we introduced ATS, we carried out extensive pilot testing.
Substantial human marking is used to check and verify results, especially when a student’s score is near the Achieved and Not Achieved boundary.
We also commissioned independent reviews of our ATS marking model. These reviews helped us to understand how confident NZQA, schools, students and the public can be in the results produced.
On this page
Independent reviews of NZQA's ATS marking model
In late 2025, NZQA commissioned 2 independent analyses of its ATS model.
NZQA created the model in 2024, trialled it with student assessments completed in September 2025, and used it to assess the Literacy Writing standard in May 2026.
NZCER
The New Zealand Council for Educational Research (NZCER) reviewed how closely the model’s scores match human markers. The review looked at whether scores were accurate, consistent and unbiased compared to human marking.
ClearPoint
ClearPoint reviewed the technical design of the ATS tool. It assessed whether the tool was suitable for its intended purpose.
NZCER review
NZCER evaluated the model's:
- performance
- bias and fairness
- readiness for marking Literacy Writing assessments.
It also looked at whether the model could be used to mark other assessment standards.
NZCER's findings
"When ATS was used as the production marking model, it demonstrated strong operational performance."
"Overall, the ATS model is considered fit for continued operational use in marking the Literacy Writing assessment. In AE1 2026, when ATS was used as the production marking model, it demonstrated strong operational performance. At the rubric element level, exact match between ATS and human scores was 79.6%, exact-or-adjacent agreement reached 99.8%, and the mean Quadratic Weighted Kappa (QWK) was 0.763. At the total score level, 84.5% of ATS scores were within one point of the corresponding human score, and the QWK was 0.872. Performance in AE2 2025 was similarly strong."
ClearPoint review
ClearPoint investigated the model's:
- architectural soundness
- process validation
- readiness for use.
ClearPoint also assessed the model's safety measures and safeguards, how well it performed and was tested, and the quality of its documentation and knowledge management practices.
ClearPoint's findings
"NZQA has moved from a proof-of-concept, piloted alongside human marking, to a compact, cloud-native production platform."
"Working with AWS Professional Services, NZQA has moved from a proof-of-concept, piloted alongside human marking, to a compact, cloud-native production platform that aligns with the AWS Well-Architected Framework and follows sound, industry-standard engineering practice. The solution completed its first live marking run in AE1 2026.
The recommendations from the 2025 review have largely been addressed.
ClearPoint is of the opinion that pragmatic decisions have been taken, with effort focused where it matters most and other areas appropriately left to mature as the solution's requirements evolve and become clearer.
The result is a lean and adaptable platform. ClearPoint has identified some areas for development, ideally addressed before the platform scales to additional models or assessment events; these are set out in the findings and recommendations in this report."
Download the August 2026 ClearPoint review report [PDF, 3.2 MB]
Download the March 2026 ClearPoint review report [PDF, 3.3 MB]
Review findings and recommendations
NZCER found that NZQA’s ATS model produces results that closely match those of human markers. The review also found that the model scores students equitably across gender, ethnicity, and socio-economic indicators.
ClearPoint found that NZQA has developed an accurate and consistent ATS model. This was the main focus of the model’s initial development and evaluation before it was implemented.
Both sets of reports make some recommendations. We are addressing these recommendations as we continue to improve the ATS model in preparation for each assessment event.