CAIRE’s Academic Highlights
Bridging the gap between artificial intelligence and real-world classrooms. Our research provides evidence-based insights and solutions that empower educators to enhance instructional quality.
Track 1: AI-Assisted Instruction & Teacher Support
SciEval: A Benchmark for Automatic Evaluation of K–12 Science Instructional Materials
Manual evaluation of AI-generated science materials is difficult to scale. We introduce SciEval, a benchmark dataset of 273 materials evaluated across 13 criteria. Our results show that domain-aligned fine-tuning of LLMs yields significant gains in automated pedagogical evaluation.
DrawSim-PD: Simulating Student Science Drawings to Support NGSS-Aligned Teacher Diagnostic Reasoning
To address privacy restrictions on sharing student work, we present DrawSim-PD, a generative framework that simulates NGSS-aligned student science drawings with controllable imperfections. We release a corpus of 10,000 artifacts to overcome data scarcity in visual assessment research.
The Transformative Collaboration of Human Intelligence and Artificial Intelligence in Designing Knowledge-in-Use Science Assessment for Learning
This study investigates the collaboration between human experts and GPT-4 to design NGSS-aligned, 3D science assessments. Using a design-based research approach, we demonstrate that principled human scaffolding—through structured prompts and iterative expert evaluation—enables AI to co-produce high-quality, equitable tasks. This work offers a transferable refinement framework, positioning generative AI as a collaborative design partner rather than a mere automated tool.
Track 2: Learning Sciences & Science Learning
Manufacturing authenticity as part of written PBL curriculum: Contrived versus spontaneous events
Project-based learning (PBL) centers on authenticity, yet prepackaged curricula struggle to predict genuine classroom events. Using Portraiture, this study examines a third-grade bilingual class to contrast pre-planned lessons with a spontaneous departure caused by a spring snowstorm. Findings reveal that while contrived events support learning, spontaneous events uniquely enable students to actively craft authentic disciplinary tools. We highlight the critical role of teacher expertise in seizing these moments and advocate for trusting teachers to adapt curricula for authentic engagement.
Adapting scientific modeling practice for promoting elementary students’ productive disciplinary engagement
This study explores how elementary school teachers adapted online modeling activities to promote students’ Productive Disciplinary Engagement (PDE). Using a collective case study, we investigated three fourth-grade teachers during the 2020–2021 school year through professional learnings, observations, interviews, and student artifacts. Four adaptations increased students’ PDE: leveraging technology tools, maximizing student-centered choice and epistemic agency, incorporating family and community resources, and using students’ diverse knowledge and expertise. These findings identify effective strategies for promoting student engagement and scientific modeling in online learning environments.
Transforming standards into classrooms for knowledge-in-use: An effective and coherent project-based learning system
Global science education reform calls for developing students’ knowledge-in-use by integrating core ideas and scientific practices to make sense of phenomena or solve problems. This paper presents an iterative design process for developing a standards-aligned, coherent learning system using a project-based learning approach. Developed through a five-year NSF-funded project, the system includes four consecutive high school chemistry curriculum and instruction materials, assessments, and professional learning. The theory-driven, empirically validated system can inform teachers and researchers in transforming science standards into curriculum materials that support students’ knowledge-in-use development.
Track 3: Responsible and Critical AI
Refusing Educational Technology: Artificial Intelligence, Inequity, and the Problem of Critical Optimism
Artificial intelligence (AI)—like many technologies before—comes with hope and worry about its implications for education. In this essay, using generative AI as a case, we appraise the ideological field of positions that stakeholders take in relation to educational technology overall. Drawing on science and technology studies, our analysis shows that education’s default position is optimism, even among those genuinely “critical” of technology’s capacity to amplify inequity. Refusal is not widely regarded as a legitimate response to educational technology. We argue, in contrast, that given the potential for scalable harm from unproven technologies like generative AI, the most rational position with new technology is robust caution with a readiness to refuse. We conclude by detailing three forms of refusal for educational stakeholders to consider: curricular, bureaucratic, and administrative.
Culturally and linguistically “Blind” or Biased? Challenges for AI Assessment of Models with Multiple Language Students
Investigating AI’s role in educational assessments, this study compares AI- provided and teacher scores of hand-drawn scientific models by Multilingual Language Learners (MLLs) in elementary classrooms. Using Convolutional Neural Networks (CNN) for scoring, we aligned AI assessments with those of experienced teachers. The results show moderate agreement (Kappa = 0.326), with AI favoring mid-range scores, while teachers provided a broader score spectrum. This suggests AI’s consistency may miss the interpretive nuances teachers offer. The study emphasizes careful AI integration to support the diverse assessments of MLLs, though it notes the limitations of a small sample size and the opaque AI scoring rationale. Our findings advocate for combining AI’s analytical strengths with teacher expertise to enhance equitable, effective educational assessments.
Can we and should we use artificial intelligence for formative assessment in science?
In this commentary, we respond to Zhai et al. (2022), who present automated assessment as a solution to limited assessment time in middle school science classrooms. Drawing on our expertise in science assessment, machine learning, artificial intelligence, and culturally relevant and linguistically responsive pedagogy, we highlight significant limitations of using AI for formative assessment, particularly for students from nondominant cultural and linguistic backgrounds. We discuss whether AI can effectively assess students’ emergent sensemaking, whether it should be used for formative assessment, and how it can be used more effectively.