Illinois Draws the Line Inside the Task: AI May Support a Teacher Evaluation but It May Not Score One
Public Act 104-0565 splits a single institutional process into machine-eligible and human-only components and legislates the boundary between them. It passed 55-0 and 113-0. The same week, a peer-reviewed experiment found that the design of an AI report moves teacher judgment when the student work is unchanged.
- Governance signal. Illinois enacted Public Act 104-0565 on July 10, barring an evaluator from using an artificial intelligence tool to assign a numerical score or qualitative rating for any component of a teacher's evaluation, or for any evaluation task requiring professional judgment. AI may still support the evaluator in administrative tasks. It takes effect January 1, 2027.
- The line is inside the task. This is neither a ban nor a role definition. The statute splits a single process into machine-eligible and human-only components and assigns each to a different kind of actor. It passed the Senate 55-0 and the House 113-0.
- Second signal. Hawaii Governor Josh Green signed Senate Bill 3001, the Artificial Intelligence Disclosure and Safety Act, on July 13. It requires conversational AI operators to disclose non-human status to minors each session or hourly in continuous conversation, and to maintain protocols for suicidal ideation and self-harm prompts.
- Key research finding, peer-reviewed. A controlled experiment published in Frontiers in Psychology found that the design of an AI text-detection report shifted teachers' evaluative judgments of student writing when the writing itself was unchanged. Stronger numerical warnings paired with visually salient risk cues moved judgment the most.
- Evidence infrastructure, preprint. A July preprint releases a labeled dataset of 1,639 K-12 explanations built to train auditors that detect pedagogical risk across five dimensions, including factual precision and ideological bias. The audit instrument is arriving before any audit requirement.
- Evidence gap. Illinois protected teachers from algorithmic judgment. No state has extended the same protection to students, and no causal evidence exists on whether AI-assisted evaluation of either group is accurate, fair, or improves anything.
Framing
For three years, the governing question in public education has been what artificial intelligence may do to students and for teachers. On July 10, Illinois asked a different question, and answered it. Public Act 104-0565 amends the Evaluation of Certified Employees Article of the School Code to prohibit an evaluator from using an artificial intelligence tool to assign a numerical score or qualitative rating for any component of a teacher's evaluation, or for any evaluation task that requires professional judgment. The same statute expressly permits an artificial intelligence tool to support the evaluator in administrative tasks. It passed the Senate 55-0 and the House 113-0, and it takes effect on January 1, 2027.
Read that carefully, because the construction is more sophisticated than the headlines suggest. This is not a ban. It is not a definition of who counts as an employee. Illinois took a single institutional process, cut it into two parts, and assigned each part to a different kind of actor. Scheduling the observation, assembling the evidence file, formatting the write-up, tracking the timeline: machine-eligible. Deciding what the evidence means about a professional's practice: human evaluators only. Earlier state action drew lines around categories and roles. This one draws a line inside a task, at the exact point where information becomes judgment. That is the first defensible seam any state has written into K-12 personnel law, and it is the seam every district will eventually have to draw for itself in a dozen other processes.
The research surfacing this same week pushes against that seam from the opposite side, and it should unsettle anyone who assumes the human half is safe simply because a human occupies it. A peer-reviewed experiment published in Frontiers in Psychology tested what happens when teachers evaluate student writing while holding an AI text-detection report. The text under review never changed. The report design did. Stronger numerical warnings, paired with visually salient risk cues, shifted teachers' evaluative judgments and intervention-oriented responses, with the perceived likelihood of AI authorship emerging as a plausible pathway. The human was in the loop the entire time. The human was also steered.
Put the two together and the week produces an uncomfortable structural finding. Illinois has protected the professional judgment of adults from being replaced by an algorithm, in a statute that passed without a single dissenting vote in either chamber. No state has extended that protection to students, who are graded, flagged, referred, and disciplined on the strength of algorithmic outputs every day. The best available evidence shows that the human reviewing those outputs is measurably influenced by their presentation. Districts have been treating human-in-the-loop as the answer. This week's evidence says human-in-the-loop is the beginning of the question, and that the design of what you put in front of the human is itself a governance decision no board has voted on.
Top Research and Policy Signals
1. Illinois Bars AI From Assigning Teacher Evaluation Ratings, and Draws the Line Inside the Task
Source type. State legislation, enacted. Illinois Public Act 104-0565 (Senate Bill 2909), 104th General Assembly. Governor approved July 10, 2026. Effective January 1, 2027.
Illinois General Assembly. (2026). Senate Bill 2909, Public Act 104-0565: School Code, teacher evaluation, artificial intelligence (104th General Assembly).
Senate Bill 2909 amends the Evaluation of Certified Employees Article of the Illinois School Code. It prohibits an evaluator from using an artificial intelligence tool to assign a numerical score or qualitative rating for any component of a teacher's evaluation. It extends that prohibition to any evaluation task that requires professional judgment. It then carves out what remains permitted: an artificial intelligence tool may be used to support the evaluator in administrative tasks. Capitol News Illinois additionally reports that teachers may not use AI to satisfy their own portion of the evaluation, and that the law leaves AI use in other areas of administrative and instructional work untouched.
The vote record matters as much as the text. The Senate passed it 55-0 on April 16. The House passed it 113-0 on May 27. Chief sponsorship was bipartisan, with Senator Christopher Belt as lead sponsor and Senator Chapin Rose as a chief co-sponsor. Representative Mary Beth Canty, the chief House sponsor, framed the rationale narrowly and precisely in a public statement: she supports AI for basic organization and for streamlining simple aspects of modern work, but the technology is not capable of effectively carrying out judgment-based tasks of this complexity. That is a statement about capability boundaries, not about technology risk in general, and it is why the statute reads as a seam rather than a wall.
Leadership implication. This is the most transferable governance instrument produced by any state this year, and it does not require Illinois residency to use. Direct your cabinet to inventory every institutional process in which an AI output currently influences a determination about a person, including teacher evaluation, student discipline referrals, special education eligibility, attendance intervention, and hiring screens, and then, in writing, split each one into machine-eligible administrative components and human-only judgment components. Boards that adopt this seam before a state imposes it will control where the line falls; boards that wait will inherit someone else's line.
2. A Peer-Reviewed Experiment Finds AI Report Design Moves Teacher Judgment When the Student Work Is Identical
Source type. Peer-reviewed journal article. Frontiers in Psychology, Volume 17 (2026). Controlled single-stimulus experiment.
Automation bias in teachers' evaluation of student writing: Effects of algorithmic warnings and visual risk cues in AI detection reports. (2026). Frontiers in Psychology, 17. [Author names not confirmed from full-text access; article title, journal, volume, year, and DOI verified.]
Most research on AI text detection has examined whether the detectors are accurate, how often they produce false positives, whether they are fair across student groups, and what policy should say about them. This study asked a different and more operationally useful question: does the design of the report itself change what the teacher concludes, holding the student's writing constant?
It does. In a controlled single-stimulus experiment, the strength of the numerical warning and the visual salience of the risk cue shifted teachers' self-reported evaluative judgments and their intervention-oriented responses. An exploratory mediation analysis indicated that the perceived likelihood of AI authorship was a plausible psychological pathway. The authors situate the work in higher education, which is a real limitation on direct K-12 generalization and should be stated plainly rather than glossed over. The mechanism, however, is a general property of how people read algorithmic risk output under uncertainty, and integrity tools, early-warning dashboards, and reading-risk screeners hand K-12 teachers exactly this class of artifact.
Leadership implication. Human review is not a control unless you govern what the human is shown. Require vendors to disclose, before contract signature, how risk scores are displayed, whether thresholds and color cues are configurable by the district, and whether staff can see the underlying evidence before seeing the score. Then write into local procedure that no disciplinary, integrity, or placement determination rests on a tool output alone, and train evaluators specifically on automation bias rather than assuming professional experience inoculates them against it.
3. Hawaii Enacts Disclosure and Crisis-Protocol Requirements for Conversational AI Used by Minors
Source type. State legislation, enacted. Hawaii Senate Bill 3001, Artificial Intelligence Disclosure and Safety Act. Signed by Governor Josh Green on July 13, 2026.
Hawaii State Legislature. (2026). Senate Bill 3001, C.D. 1: Artificial Intelligence Disclosure and Safety Act (Thirty-Third Legislature, 2026). [URL resolves but returned no extractable text; provisions below are drawn from the Transparency Coalition legislative update of July 17, 2026, which is verified.]
Senate Bill 3001 requires operators of conversational artificial intelligence services to disclose that a user is interacting with a machine rather than a person. For minors, the disclosure must be repeated each session or at least hourly during continuous conversation. The Act establishes additional protections for minor account holders against manipulative engagement techniques and sexually explicit content, requires operator tools that let parents and guardians manage screen time and account settings, and requires operators to maintain protocols responding to user prompts involving suicidal ideation or self-harm. Annual reporting to the state Behavioral Health Administration begins January 1, 2028.
The Act does not regulate schools or mention them. It regulates the companion and conversational products that students already use outside instructional time and, in many districts, inside it. Hawaii joins a fast-moving national pattern: the Transparency Coalition counted 78 chatbot bills across 27 states six weeks into the 2026 session, and Rhode Island enacted a therapy chatbot ban and a chatbot self-harm safety measure in June.
Leadership implication. Student-facing conversational AI is being regulated as a consumer product while districts are deploying it as an instructional one, and the district assumes the duty of care in the gap. Confirm which tools in your environment meet a conversational AI definition, whether they carry crisis-response protocols, and who on your staff is notified when a self-harm signal is generated. If a vendor cannot answer that last question with a named human role and a response time, that is a procurement finding, not a technical detail.
4. A New K-12 Audit Dataset Arrives Before Any Audit Requirement Exists
Source type. Preprint, not yet published in final form. arXiv, July 2026. Accepted at the IEEE International Carnahan Conference on Security Technology (ICCST 2026), scheduled October 14, 2026.
Irigoyen, J., Daza, R., Jurado, F., Fierrez, J., Tolosana, R., Ortigosa, A., Blas, E., & Morales, A. (2026). AIriskEval-edu: New dataset for risk assessment in AI-mediated K-12 educational explanations [Preprint]. arXiv:2607.01934.
This work releases a dataset built to train and evaluate language-model auditors that assess pedagogical risk in K-12 instructional content. It comprises 1,639 explanations drawn from 170 curated ScienceQA questions spanning science, language arts, and social sciences. For each question, the dataset pairs one explanation written by a human teacher with eleven explanations generated by model-simulated teacher profiles, each associated with a distinct pedagogical risk. The rubric covers five dimensions: factual precision, depth and completeness, focus and relevance, student-level appropriateness, and ideological bias. A further 785 explanations carry structured explainability annotations that localize and describe the risk, produced semi-automatically with expert teacher validation.
The authors report validation experiments comparing frontier proprietary models with a locally deployable Llama 3.1 8B model, testing whether supervised fine-tuning enables a local model to approach frontier performance while keeping educational audit data within the institution. That privacy-preserving framing is the part districts should notice. The paper has cleared conference peer review and is scheduled for presentation in October, so it should be treated as accepted work in preprint form rather than as a published finding.
Leadership implication. Every state model policy and district AI plan issued this year asks whether instructional content is appropriate, but none specifies how that question is answered at scale. This is the beginning of an answer, and it is a research artifact rather than a product. Ask your instructional and technology leads to review the five risk dimensions and map them against your current review process for AI-generated materials, because the gap between the two lists reflects your actual exposure.
5. Teachers Tie AI Instructional Quality to Professional Development, Not to the Tools
Source type. Peer-reviewed journal article. Smart Learning Environments (Springer), 2026. Survey of 532 school teachers.
Mah, D. K., Gross, N., & Egloffstein, M. (2026). Artificial intelligence in K-12 instruction: The role of teacher professional development. Smart Learning Environments. [Author list and DOI verified from publisher page; full findings not extracted from full text.]
This survey of 532 school teachers examines perspectives on instructional quality with AI, the new teaching and learning opportunities teachers perceive, and the necessity of professional development. It sits within an emerging line of work that extends established frameworks, such as technological pedagogical content knowledge, to account for AI in instruction.
Reported here with appropriate restraint: the study is peer-reviewed, K-12, and recent, and its central organizing claim is that professional development is the mediating variable between AI availability and instructional quality. That finding is consistent with and reinforces the transmission result featured in Edition 24, in which institutional AI readiness reached student AI literacy via aggregated teacher capability rather than via infrastructure or attitude. Two independent studies now point to the same lever.
Leadership implication. The budget implication is unglamorous and unavoidable. If professional development is the mediating variable, then a procurement line that funds licenses without funding role-differentiated capability building is not an investment; it is a deferred liability. When you build next year's budget, tie every AI license request to a named professional development commitment and a named owner, and decline the ones that arrive without both.
Emerging Strategic Themes
Theme 1. The seam between support and judgment. Illinois did not choose between permitting and prohibiting. It split one process into machine-eligible and human-only components and legislated the boundary between them. This is the most portable governance pattern to emerge in 2026, and it applies far beyond teacher evaluation. Districts should expect the same seam to be demanded next in student discipline, special education eligibility, and hiring.
Theme 2. Asymmetric protection. Adults in school buildings now have a statutory shield against algorithmic determination in at least one state. Students have none anywhere. Students are the population subject to the highest volume of automated flagging, scoring, and referral, and they are the population with the least procedural recourse. That asymmetry is a defensible policy choice only if a board has actually made it, and no board has.
Theme 3. Interface design as an ungoverned governance surface. The Frontiers experiment locates real decision-making power in something no policy document addresses: how a score is displayed. Threshold placement, color, and warning language change what professionals conclude about identical work. Districts govern which tools they buy and who may use them. They do not govern the screen, and the screen is where judgment gets moved.
Theme 4. Audit capability outrunning audit mandate. A labeled K-12 pedagogical risk dataset with expert-validated annotations now exists in the research literature and is designed to run on locally hosted models for privacy. No state requires such an audit, and no district procurement standard references one. The capability is arriving ahead of the requirement, which is the reverse of the usual sequence and a rare opening for districts that move early.
What Was Not Found
Adoption is outpacing evidence. This section names the gap precisely, as of this week's window.
- No causal evidence exists that removing AI from teacher evaluation improves evaluation accuracy, fairness, teacher retention, or instructional quality. Illinois acted on a capability argument rather than an outcome study, and the sponsor said so directly. That is a legitimate basis for legislation, but districts adopting the same seam voluntarily should understand they are adopting a reasoned position rather than an evidence-backed one.
- No peer-reviewed study in this window measures automation bias in K-12 teachers specifically. The Frontiers experiment is peer-reviewed, well-designed, and situated in higher education. The mechanism is plausible in K-12, and the artifacts are nearly identical. Still, the direct K-12 replication does not exist, and no study measures whether bias effects differ when the flagged student is an English learner, has a disability, or attends a high-poverty school. Those are precisely the students most often flagged.
- No study establishes whether the disclosure and crisis-protocol requirements now enacted in Hawaii and Rhode Island reduce harm to minors. These are new statutory instruments with no record of efficacy. Districts should not treat vendor compliance with them as evidence of student safety.
- The AIriskEval-edu dataset provides an audit instrument but no field evidence. There is no published study of what a pedagogical risk audit finds when run against the AI-generated materials actually circulating in a real district, and therefore no baseline against which any district can judge its own exposure.
- Elementary grades, literacy instruction, and non-STEM subject areas remain thinly represented across everything reviewed this week. The evaluation and integrity research concentrates on writing and secondary or postsecondary contexts. A district making elementary literacy decisions on the strength of this evidence base is extrapolating, and should name that in its board documentation.
- Adoption continues to outpace evidence, and the gap this week is specific rather than general. Illinois wrote a boundary that takes effect in five months. The peer-reviewed evidence on whether humans hold that boundary under algorithmic pressure suggests they do not, at least partially. The audit tooling that would let a district check its own compliance exists as a July preprint and nowhere in the procurement market. Districts are being asked to govern a seam that research has just shown is porous, using instruments that are not yet purchasable.
Novo Executive Summary
Illinois has produced the most useful governance artifact of the year, and its value is structural rather than jurisdictional. By prohibiting artificial intelligence from assigning any component rating in a teacher evaluation while expressly permitting it to carry administrative load, Public Act 104-0565 legislates a seam between machine support and human determination that every district will eventually need in a dozen other processes. The peer-reviewed evidence arriving the same week complicates the assumption underlying that seam, showing that presenting an algorithmic score can shift professional judgment even when the underlying work is unchanged. Read together, they say something districts have been slow to hear: naming a human decision-maker is the start of governance architecture, not its completion, and the interface that human is handed is itself an ungoverned lever. Meanwhile, the protection runs one direction only, shielding adults from algorithmic determination while students remain subject to it without equivalent procedural standing. Dr. Reginald Griffin works with district leadership teams through Novo Innovative Pathways to build exactly this architecture: mapping where machine output intersects with human determination, defining role-based AI literacy for the people making those decisions, and writing procurement and review standards that hold when a statutory deadline arrives.
Watch This Week
- Illinois districts have until January 1, 2027 to bring evaluation procedures and evaluator training into compliance with Public Act 104-0565. Watch whether the Illinois State Board of Education issues implementation guidance, and whether it defines evaluation tasks requiring professional judgment.
- Whether any state introduces the Illinois seam on the student side, extending professional-judgment protection to discipline, eligibility, or placement determinations.
- California's Legislature returns from summer recess on August 3 with roughly 30 AI bills live, including AB 1159 on student privacy protections for school-purposed digital operators and AB 2392 on generative AI procurement standards in public higher education.
- New York Governor Kathy Hochul has until December 31 to act on the kids chatbot safety bill and the AI training data transparency act passed in June.
- New Jersey S 4469 and A 5184, which would require boards of education to adopt AI policy and direct the state Department of Education to establish a model policy, both sitting in the Education Committee.
- New York City's final AI guidance, promised for later this summer, following the July procurement pause reported in Edition 24. Whether it names a vetting standard is the test.
- Presentation of the AIriskEval-edu work at ICCST 2026 on October 14, and whether any state or district procurement standard begins referencing pedagogical risk auditing.
Sources
Governance and Policy
Note. Legislative history, vote counts, sponsors, and effective date verified via LegiScan: legiscan.com
Szalinski, B. (2026, July 13). Pritzker signs new laws on birth control, AI regulations, play-based learning. Capitol News Illinois. capitolnewsillinois.com [Verified.]
Hawaii State Legislature. (2026). Senate Bill 3001, C.D. 1: Artificial Intelligence Disclosure and Safety Act (Thirty-Third Legislature, 2026). capitol.hawaii.gov [URL resolves; no extractable text returned. Provisions sourced from the verified Transparency Coalition update below.]
Barcott, B. (2026, July 17). AI legislative update: July 17, 2026. Transparency Coalition. transparencycoalition.ai [Verified.]
Research, Peer-Reviewed
Note. Article, journal, volume, year, and DOI verified. Author names not confirmed from full-text access and are therefore not listed.
Mah, D. K., Gross, N., & Egloffstein, M. (2026). Artificial intelligence in K-12 instruction: The role of teacher professional development. Smart Learning Environments. doi.org/10.1186/s40561-026-00442-4
Note. Publisher page verified. Full text not extracted; findings reported at the level supported by the abstract and publisher summary.
Research, Preprint (Not Peer-Reviewed in Final Published Form)
Note. URL verified; title confirmed. Accepted at IEEE ICCST 2026, scheduled October 14, 2026. Conference peer review cleared; not yet published in final form.
AI in Public Education Brief is published weekly by Novo Innovative Pathways. For district advisory engagements, contact Dr. Reginald Griffin through Novo Innovative Pathways.
If your district cannot say today, process by process, where machine output stops and human determination begins, then it has not drawn the seam Illinois just legislated. The Novo 10-Domain Readiness Brief is where that boundary, and the named human who owns it, get written down.
Schedule a Readiness Conversation