The Texas Education Agency announced September 4, 2026, that a manual rescoring of short typed answers on STAAR reading exams gave additional credit to about 27,200 students and raised the A-F accountability ratings for 11 districts and campuses. The review covered roughly 1.6 million exams, and about 1.7 percent of them saw score improvements, though not every increase changed a school's rating, TEA reported.
The rescoring targeted text-entry responses, which TEA officials described as brief typed answers such as numbers, words, or short phrases. The agency stressed that none of those answers had been graded by its automated scoring engine, which handles longer written responses. TEA said content experts looked for responses that may not have fit the original scoring rules, citing an example where a student who wrote ferst instead of first could now receive credit. No student's score was lowered, the agency said.
As a result of the rescoring, five districts and nine campuses saw their A-F ratings rise, according to the Texas Tribune. The affected districts were Chester, Crandall, Daingerfield-Lone Star, Floydada Collegiate, Forney, Fredericksburg, Galveston, Jim Hogg County, Odem-Edroy, Southwest, and Walnut Bend ISDs. In Galveston ISD, Central Middle School's rating climbed six points to 75, moving from a D into a better band. Southwest ISD's overall rating rose from 79 to 80, pushing it from a C to a B.
How other states manage automated scoring
Texas is not alone in turning to computers to help grade student writing. Several other states have used automated scoring engines for state assessments, often with a hybrid model that routes some responses to human reviewers.
Tennessee uses Pearson's machine learning technology for its TCAP English language arts exams. A 2022 Tennessee study of 56,000 essay scores found high agreement between the algorithm and human scorers, with no performance gaps across student subgroups, according to the Tennessee State Board of Education. That same Pearson technology has been used in statewide testing in Texas, Utah, Ohio, and Massachusetts since 2010.
West Virginia's experience with automated scoring for its WESTEST 2 writing assessment produced a 2011 comparability study. The study found that human-to-engine exact agreement rates (41 percent) were nearly identical to human-to-human rates (42 percent), and exact or adjacent agreement rates were 88 percent for the engine and 87 percent for humans. However, for 8 of 10 grade levels, the engine assigned slightly lower average scores than trained human scorers, though the difference was described as practically insignificant, amounting to 2 percent to 5 percent of available points, per the West Virginia Department of Education.
The American Institutes for Research developed an automated essay scoring engine used in Arizona's AzMERIT, Ohio's OST, Utah's SAGE, and other state assessments. A 2018 AzMERIT evaluation found machine-to-final-human exact agreement rates averaged 0.75 across prompts, with a quadratic weighted kappa of 0.68, comparable to human-human agreement rates. Low-confidence responses routed for human review showed lower exact agreement, averaging 0.59, confirming the value of confidence-based routing, the study reported.
The Smarter Balanced Assessment Consortium field-tested automated scoring across 665 ELA and literacy short-text constructed-response items. Only 39 percent of those items met all criteria for automated scoring, 26 percent were deemed unsuitable, and 28 percent needed further review for subgroup performance. For longer essay items, 67 percent met all criteria for the Organization and Purpose trait. The consortium found that mathematics items performed considerably better than ELA items, likely because math responses are less open-ended and use narrower lexicons.
What research says about automated scoring accuracy
TEA's own 2024 technical report on hybrid scoring for STAAR reading language arts exams concluded that the hybrid design provides accurate, reliable, and fair scoring and that all items met full performance criteria on the random sample. But the report also found that the engine performed poorly on low-confidence responses, with exact agreement 15 percent lower than on random samples. It flagged areas for future refinement including condition-code routing and reprogramming timing, according to the Texas Education Agency and Cambium Assessment, Inc.
A peer-reviewed study of the AzMERIT automated scoring engine, published by the American Institutes for Research, found that machine-to-final-human exact agreement averaged 0.75 across prompts, with a range of 0.65 to 0.83. That matched or exceeded human-human agreement rates of 0.71 exact and 0.65 quadratic weighted kappa. The study also confirmed that low-confidence responses routed for human review had lower agreement, supporting the hybrid model's design.
The Smarter Balanced field test results, also cited in peer-reviewed contexts, showed that automated scoring suitability varied widely by item type. For ELA short-text items, only 39 percent met all automated scoring criteria, while 74 percent of mathematics items did. The consortium attributed the difference to ELA responses being more open-ended with broader vocabularies, which makes reliable scoring harder for algorithms.
Another peer-reviewed study examined the effect of automated essay scoring on test equating errors in mixed-format tests using a BLSTM model. It found that equating errors from automated scoring were not significantly different from human-rater equating errors, with a small effect size (Cliff's Delta = -0.18). The study was limited to about 1,200 test-takers, so generalizability is constrained.
Unverified and tracking
In January 2025, the Dallas Morning News reported that Dallas ISD had raised concerns after roughly 43 percent of more than 4,600 STAAR responses the district submitted for rescoring showed score improvements. The report suggested those results raised questions about the accuracy of automated scoring. According to the Dallas Morning News, TEA responded that the improved essays represented about 3 percent of Dallas ISD's 2024 written submissions and that fewer than 1 percent of all written responses statewide had been submitted for rescoring. TEA has not separately confirmed these figures through its own published data.
