When Artificial Intelligence Masks Readiness: A Practitioner Action Research Study in an Elementary Mathematics Methods Course

Denise Polojac-Chenoweth (Hillsborough County Public Schools)

Generative artificial intelligence (AI) has rapidly shifted higher education landscapes, providing preservice teachers with immediate access to instructional resources, lesson planning templates, and automated feedback. While these tools offer efficiency and support iterative planning, their unstructured use presents a significant challenge for mathematics teacher education. Methods courses are explicitly designed to make preservice teachers’ mathematical reasoning, pedagogical content knowledge, and instructional decision-making visible. When AI generates sophisticated instructional artifacts, instructors may have greater difficulty determining candidates’ independent understanding and readiness for classroom practice.

This tension raises an urgent question for mathematics teacher educators: How can preparation programs maintain assessment validity when AI blurs the distinction between scaffolded support and cognitive substitution? If submitted work no longer reflects independent thinking, we cannot accurately identify misconceptions, provide targeted feedback, or evaluate teaching competence.

Effective mathematics teaching requires educators to interpret student thinking and make real-time instructional decisions (Ball et al., 2008). Likewise, the Association of Mathematics Teacher Educators (2017) emphasizes preparing candidates to engage students in meaningful mathematical learning. Despite growing scholarship on the opportunities and challenges of generative AI in higher education (Kasneci et al., 2023), relatively little attention has been given to how AI affects the assessment of preservice teacher readiness within mathematics methods courses. This gap is important because methods courses provide a critical opportunity to evaluate candidates' instructional decision-making before student teaching. This practitioner action research study examines how unstructured AI use affected assessment validity in an Elementary Mathematics Methods III course, and details how an intentional assessment redesign restored instructional visibility while supporting responsible AI integration.

Context and Problem of Practice

This study was situated within Elementary Mathematics Methods III, the capstone of a three-course elementary mathematics methods sequence. While Methods I and II focus on foundational number concepts, algebraic thinking, and introductory lesson design, Methods III emphasizes standards-aligned instruction, analysis of student thinking, STEM integration, and reflective practice. Taken immediately before candidates begin their final full-time student teaching internship, this course serves as the program's critical gatekeeper to assess readiness for professional practice.

Midway through the semester, several clear indicators suggested an over-reliance on generative AI tools. Student submissions began featuring highly polished, homogeneous pedagogical language that differed noticeably from our active classroom discussions. Multiple written assignments contained identical phrasing and identical lesson plan formatting. Most concerning, candidates struggled to explain or defend their written work when questioned. For instance, several students prominently used phrases like "productive struggle," "student-centered discourse," and "mathematical agency" in their text, yet they could not articulate how those concepts informed their lesson design or supported student understanding.

Informal conversations with the 25 enrolled preservice teachers revealed highly inconsistent expectations across their preparation sequence. Students described prior courses where AI tools were either permitted unconditionally or largely unmonitored, creating programmatic confusion regarding appropriate use. These observations raised the central problem of practice: How can mathematics teacher educators accurately assess and support learning when submitted work masks a candidate's independent pedagogical reasoning?

Practitioner Action Research Design

To address this problem, I initiated a practitioner action research study midway through the semester to redesign course assessments. Rather than banning AI—which is both impractical and counterproductive—the intervention distinguished between AI as cognitive substitution and AI as a tool for professional critique and refinement.

This inquiry was guided by three research questions:

  1. How did unstructured AI use affect the assessment of preservice teacher readiness?
  2. How did assessment redesign affect the visibility of preservice teachers’ mathematical and pedagogical reasoning?
  3. What tensions emerged as preservice teachers responded to revised expectations for AI use?

The study involved 25 elementary education majors. Data sources included my instructor field notes, oral lesson defense notes, student assessment artifacts, and informal written student feedback. I analyzed this qualitative data inductively, coding for recurring patterns related to assessment validity, instructional reasoning, and student emotional responses to changing expectations.

While external state certification examinations assess baseline mathematical content knowledge, methods courses uniquely reveal a candidate's responsive teaching practices. Therefore, I replaced assignments vulnerable to AI-generated responses with four high-visibility, multimodal tasks requiring real-time demonstrations of reasoning:

  • In-Class Performance Tasks: Problem-solving sessions completed without digital devices.
  • Handwritten Mathematical Work: Documenting step-by-step mathematical modeling and anticipating student misconceptions by hand.
  • Collaborative Analysis with Individual Accountability: Group analysis of student work samples paired with individual written rationales.
  • Oral Lesson Defenses connected to submitted instructional plans.

To model responsible integration, I designed one structured AI assignment. Candidates used AI to generate a standard STEM lesson plan, utilized course frameworks and state standards to critique its limitations, and submitted a revised version tracking their own professional modifications. This positioned AI as an assistant for initial content generation but left the cognitive load of professional judgment entirely on the candidate.

Findings and Discussion

Finding 1: AI Can Mask Readiness

Unstructured AI use obscured evidence of candidates' mathematical and pedagogical understanding, limiting my ability to assess readiness accurately. Once AI scaffolds were removed via the redesigned assessments, clear differences in demonstrated readiness emerged. Across the oral lesson defenses, several candidates who had submitted polished written lesson plans struggled to explain their instructional decisions or justify how planned modifications supported conceptual understanding. For example, when asked to defend an AI-generated task on fraction division, three candidates could not identify the underlying mathematical misconceptions the task was designed to address, suggesting that unstructured AI use had masked important gaps in pedagogical understanding.

Finding 2: Structured AI Use Can Strengthen Learning

When AI was intentionally constrained as a tool for critique rather than production, candidates engaged deeper with instructional design. In the structured STEM assignment, evaluating AI-generated outputs forced candidates to actively apply course frameworks and state standards. Student feedback suggested that identifying weaknesses in AI-generated lessons required more rigorous mathematical reasoning than generating lessons from scratch because candidates had to justify instructional decisions using course frameworks and state standards.

Finding 3: Program Coherence Matters

The assessment redesign prompted strong reactions from students. Several expressed frustrations that AI was accepted elsewhere but penalized or heavily scrutinized here, making the revised expectations feel unfair. This pattern highlighted a critical need for program-wide coherence regarding AI policies. Fragmented expectations across a teacher preparation program create learner confusion and undermine the collective validity of our programmatic assessments.

Implications and Recommendations

These findings suggest that teacher preparation programs bear an ethical responsibility to ensure that institutional certification reflects demonstrated classroom readiness. Based on this study, I offer four recommendations for mathematics teacher educators:

  1. Develop Program-Wide AI Expectations: Establish consistent, clear expectations across all sequence methods courses to eliminate student confusion.
  2. Design Assessments that Make Thinking Visible: Incorporate performance tasks, oral explanations, and handwritten work alongside digital submissions to ensure instructional visibility.
  3. Teach Candidates to Evaluate AI Critically: Intentionally integrate assignments that require candidates to critique, debug, and revise AI outputs using established mathematics education frameworks.
  4. Maintain Assessment Integrity: Ground high-stakes certification and readiness decisions in independent, verified demonstrations of pedagogical reasoning rather than unverified digital artifacts.

Conclusion

Artificial intelligence is reshaping teacher education, but uncritical use threatens assessment integrity in mathematics methods coursework. This practitioner action research study demonstrates that intentional assessment design can restore instructional visibility while supporting the responsible integration of AI.

Rather than excluding AI from mathematics teacher education, these findings support its intentional, developmentally appropriate use after candidates have demonstrated independent mathematical and pedagogical reasoning. As mathematics teacher educators, our responsibility extends beyond supporting success in university coursework; it includes ensuring that future teachers can demonstrate the mathematical understanding and instructional judgment required for effective classroom practice.

References

Association of Mathematics Teacher Educators. (2017). Standards for preparing teachers of mathematics. AMTE.

Ball, D. L., Thames, M. H., & Phelps, G. (2008). Content knowledge for teaching: What makes it special? Journal of Teacher Education, 59(5), 389–407.

Kasneci, E., Sessler, K., Küchemann, S., Bannert, M., Dementieva, D., Fischer, F., Gasser, U., Groh, G., Günnemann, S., Hüllermeier, E., Krusche, S., Kutyniok, G., Michaeli, T., Nerdel, C., Pfeiffer, F., Poquet, O., Sailer, M., Schmidt, A., Seidel, T., … Kasneci, G. (2023). ChatGPT for good? On opportunities and challenges of large language models for education. Learning and Individual Differences, 103, Article 102274.