SchoolDecisionby Formative Spaces, Inc.
School DecisionThe Newsroom
SATURDAY, SEPTEMBER 12, 2026
Beyond the headline
SCHOOLDECISION.COM/NEWSROOM

TEA manual rescoring of STAAR text-entry items improves scores for 27,200 students, raises A-F ratings for 11 districts and campuses

Texas regraded brief typed answers on reading exams, finding 1.7% of reviewed responses deserved more credit. The error rate affected accountability ratings for 11 schools and districts.

The Texas Education Agency announced September 4, 2026, that a manual rescoring of short typed answers on STAAR reading exams gave additional credit to about 27,200 students and raised the A-F accountability ratings for 11 districts and campuses. The review covered roughly 1.6 million exams, and about 1.7 percent of them saw score improvements, though not every increase changed a school's rating, TEA reported.

27,200Number of students whose scores improved after TEA manually rescored text-entry responses on STAAR reading exams. [1]

The rescoring targeted text-entry responses, which TEA officials described as brief typed answers such as numbers, words, or short phrases. The agency stressed that none of those answers had been graded by its automated scoring engine, which handles longer written responses. TEA said content experts looked for responses that may not have fit the original scoring rules, citing an example where a student who wrote ferst instead of first could now receive credit. No student's score was lowered, the agency said.

As a result of the rescoring, five districts and nine campuses saw their A-F ratings rise, according to the Texas Tribune. The affected districts were Chester, Crandall, Daingerfield-Lone Star, Floydada Collegiate, Forney, Fredericksburg, Galveston, Jim Hogg County, Odem-Edroy, Southwest, and Walnut Bend ISDs. In Galveston ISD, Central Middle School's rating climbed six points to 75, moving from a D into a better band. Southwest ISD's overall rating rose from 79 to 80, pushing it from a C to a B.

How other states manage automated scoring

Texas is not alone in turning to computers to help grade student writing. Several other states have used automated scoring engines for state assessments, often with a hybrid model that routes some responses to human reviewers.

Tennessee uses Pearson's machine learning technology for its TCAP English language arts exams. A 2022 Tennessee study of 56,000 essay scores found high agreement between the algorithm and human scorers, with no performance gaps across student subgroups, according to the Tennessee State Board of Education. That same Pearson technology has been used in statewide testing in Texas, Utah, Ohio, and Massachusetts since 2010.

West Virginia's experience with automated scoring for its WESTEST 2 writing assessment produced a 2011 comparability study. The study found that human-to-engine exact agreement rates (41 percent) were nearly identical to human-to-human rates (42 percent), and exact or adjacent agreement rates were 88 percent for the engine and 87 percent for humans. However, for 8 of 10 grade levels, the engine assigned slightly lower average scores than trained human scorers, though the difference was described as practically insignificant, amounting to 2 percent to 5 percent of available points, per the West Virginia Department of Education.

The American Institutes for Research developed an automated essay scoring engine used in Arizona's AzMERIT, Ohio's OST, Utah's SAGE, and other state assessments. A 2018 AzMERIT evaluation found machine-to-final-human exact agreement rates averaged 0.75 across prompts, with a quadratic weighted kappa of 0.68, comparable to human-human agreement rates. Low-confidence responses routed for human review showed lower exact agreement, averaging 0.59, confirming the value of confidence-based routing, the study reported.

The Smarter Balanced Assessment Consortium field-tested automated scoring across 665 ELA and literacy short-text constructed-response items. Only 39 percent of those items met all criteria for automated scoring, 26 percent were deemed unsuitable, and 28 percent needed further review for subgroup performance. For longer essay items, 67 percent met all criteria for the Organization and Purpose trait. The consortium found that mathematics items performed considerably better than ELA items, likely because math responses are less open-ended and use narrower lexicons.

What research says about automated scoring accuracy

TEA's own 2024 technical report on hybrid scoring for STAAR reading language arts exams concluded that the hybrid design provides accurate, reliable, and fair scoring and that all items met full performance criteria on the random sample. But the report also found that the engine performed poorly on low-confidence responses, with exact agreement 15 percent lower than on random samples. It flagged areas for future refinement including condition-code routing and reprogramming timing, according to the Texas Education Agency and Cambium Assessment, Inc.

A peer-reviewed study of the AzMERIT automated scoring engine, published by the American Institutes for Research, found that machine-to-final-human exact agreement averaged 0.75 across prompts, with a range of 0.65 to 0.83. That matched or exceeded human-human agreement rates of 0.71 exact and 0.65 quadratic weighted kappa. The study also confirmed that low-confidence responses routed for human review had lower agreement, supporting the hybrid model's design.

The Smarter Balanced field test results, also cited in peer-reviewed contexts, showed that automated scoring suitability varied widely by item type. For ELA short-text items, only 39 percent met all automated scoring criteria, while 74 percent of mathematics items did. The consortium attributed the difference to ELA responses being more open-ended with broader vocabularies, which makes reliable scoring harder for algorithms.

Another peer-reviewed study examined the effect of automated essay scoring on test equating errors in mixed-format tests using a BLSTM model. It found that equating errors from automated scoring were not significantly different from human-rater equating errors, with a small effect size (Cliff's Delta = -0.18). The study was limited to about 1,200 test-takers, so generalizability is constrained.

Unverified and tracking

In January 2025, the Dallas Morning News reported that Dallas ISD had raised concerns after roughly 43 percent of more than 4,600 STAAR responses the district submitted for rescoring showed score improvements. The report suggested those results raised questions about the accuracy of automated scoring. According to the Dallas Morning News, TEA responded that the improved essays represented about 3 percent of Dallas ISD's 2024 written submissions and that fewer than 1 percent of all written responses statewide had been submitted for rescoring. TEA has not separately confirmed these figures through its own published data.

Analysis

By the School Decision Newsroom, written after the reporting above was filed.

The rescored items are exact-match scored, not graded by the automated engine at all

Text-entry items are scored by a computer checking whether the typed string exactly equals the correct answer. "Ferst" does not match "first," so the student gets zero credit. The automated scoring engine, which uses natural language processing, handles a different category: short and extended constructed responses, where students write paragraphs or explanations. TEA's own scoring matrix confirms this split. The rescoring fixed a rigidity problem in exact-match rules. It did not touch, or validate, the engine that grades longer written answers.

The reassuring 1.7% and the alarming 35-43% measure different item types

TEA points to the 1.7% statewide improvement rate as evidence the system catches errors. That figure covers text-entry items only. The engine-graded constructed responses that Dallas ISD flagged showed 35 to 43 percent of submitted responses improving across two consecutive years. TEA correctly notes that districts cherry-pick which responses to rescore, inflating that rate. Both points are fair. But no one has published a statewide random-sample rescoring rate for engine-graded items specifically. Parents trying to assess whether the automated engine is reliable have no apples-to-apples number for the item type that matters most.

Final A-F ratings land after the September 11 appeals deadline

Preliminary 2026 ratings were released August 14 on TXschools.gov. Districts and charter schools have until 5 p.m. on September 11, 2026 to file appeals. TEA will automatically update the 11 affected districts and campuses without requiring them to appeal. But any other district believing a rescoring should have changed its rating must file by that deadline. Final ratings, reflecting all appeals outcomes, will be posted on TXschools.gov and TEAL afterward. Parents comparing schools should treat current posted ratings as preliminary until that final release.

Sources

  1. The Texas Tribune. 11 Texas districts, campuses to see A-F ratings increase View
  2. Texas Education Agency. STAAR Hybrid Scoring Key Questions View
  3. Texas Education Agency. Scoring Process for STAAR Constructed Responses View
  4. Texas Education Agency / Cambium Assessment, Inc.. Texas STAAR RLA Spring 2024 Administration: Automated Scoring Methods and Results View
  5. Tennessee State Board of Education. TCAP Automated Scoring Workshop View
  6. West Virginia Department of Education. Findings from the 2011 West Virginia Online Writing Scoring Comparability Study View
  7. American Institutes for Research (via DOI). Evaluating the Performance of Automated Essay Scoring in AzMERIT Assessments View
  8. Smarter Balanced Assessment Consortium. Smarter Balanced Automated Scoring Research Studies (Field Test) View
  9. Educational Research and Reviews (ERIC). Automated Essay Scoring Effect on Test Equating Errors in Mixed-format Test View
  10. The Dallas Morning News. Dallas ISD raises questions about automated STAAR test scoring View
  11. Tomball ISD / STAAR Redesign Scoring Guide. New Question Types Scoring and Reporting Guide - RLA View
  12. The Dallas Morning News. Dallas ISD asked Texas to rescore thousands of STAAR tests. About one-third went up. View
  13. Texas Education Agency. 2026 A-F Accountability Ratings Appeals Process and Timeline View
TEA manual rescoring of STAAR text-entry items improves scores for 27,200 students, raises A-F ratings for 11 districts and campuses | School Decision