As artificial intelligence (AI) becomes a bigger part of education, the challenge is no longer whether students should use it, but how teachers can design learning that preserves rigor while building AI literacy. This three-part series explores the research, classroom practice, and practical strategies behind using AI to deepen thinking rather than replace it.
Developing educational AI pedagogy is beginning to point the way, but only practical classroom experience can tell us whether we’re on track. In the first article of this series, we argued that balancing AI literacy with academic rigor is not a compromise but a design challenge. Rather than asking whether students should use AI, we asked a different question: Could we design learning experiences where AI required students to think more, not less?
The only way to answer that question was to try it. Across physics, mathematics, and sociales, we designed AI-supported learning experiences that shared one principle: the AI would not do the thinking for students. Instead, it would ask better questions, challenge assumptions, and continually return students to their teachers. These classroom examples show what happened when that principle met practice.
Student and Class Examples
The following student vignettes are drawn from actual Flint.ai sessions, an educational AI tool that can be customized and adapted to diverse frameworks. The three student examples are from the AP Physics 1 final project at CNG. Also included are additional use cases of Flint.ai along with ChatGPT in math and sociales classes.
In the physics project, students selected a real-world phenomenon, developed a driving question, created an assessment rubric, and designed a learning product that was both personalized to their interests and readiness levels and aligned to NGSS Science and Engineering Practices and AP Physics 1 standards. The AI was programmed as a Socratic coach, trained on real student conference transcripts and conversations to mimic the conferencing style of their teacher. It would not give answers; it asked questions. Password-gated checkpoints, issued by the teacher, were required at each stage before a student could advance. Students had to leave the AI and return to their teacher, so the design created more occasions for that relationship, not fewer.
Student 1: Grade 11 | 1.9 overall | Spacecraft reentry and splashdown
| Phenomenon Spacecraft capsule splashdown |
Entry point
Student 1 opened their session carrying a low grade and personal hardships, alongside a clear reason to care: they wanted to reach AP Physics C and study engineering. "I think once I get it, everything else will fall into place." They also held a textbook misconception that space has no gravity. The AI did not correct them. It asked what keeps the Moon in orbit.
The Socratic arc
What followed was a correction they performed on themself. The AI only asked questions about the forces on a fuel-spent rocket, about whether a gravity assist could work if gravity vanished, until they arrived at the idea on their own, "They're kind of infinitely 'falling'... they orbit so fast that it allows them to not actually fall into the earth." The AI's reply, "Now that is AP Physics thinking." They went on to trace the energy transformations of reentry, revise their earlier claim that the heat came from a chemical reaction once the AI pressed them on it, and narrow their phenomenon over six iterations to the physics of a capsule striking the ocean.
What the session reveals
This is the tool at its most consequential. A student, in a high-stakes final, was most likely to fail, but instead left having built a rigorous understanding through their own reasoning. Not despite the AI's refusal to answer, but because of it. A checkpoint partway through sent them to find their teacher. That conversation happened.
Student 2: Grade 11 | 3.7 overall | Olympic Trap clay shooting recoil
| Phenomenon Shotgun recoil in Olympic Trap shooting |
Entry point
Student 2 entered with a 3.7 and a phenomenon rooted in their identity: the recoil they feel every tournament as a competitive Olympic Trap shooter. "It's part of my sport and identity." For a strong student, the risk is that the AI simply validates. It did not.
The Socratic arc
They identified Newton's Third Law and conservation of momentum early, and the AI immediately stress-tested them, asking whether the system was truly isolated in a real shot. Their answer was the kind of model-aware reasoning the Advanced Placement (AP) guidelines call Advanced. They named gravity, air resistance, and their own body's energy absorption, explained why conservation holds only approximately at the instant of firing, and traced where the recoil energy goes: a little into backward motion, most into internal energy as tissue deforms, some stored and released as heat. "My body acts like a shock absorber," they wrote. The most sophisticated moment came in their rubric, where they distinguished between evidence that is comprehensive and evidence that is integrated, tying each piece to the central claim.
What the session reveals
That is advanced thinking about scientific argumentation, and they reached it through Socratic pressure, not instruction. The tool did not let the student coast on his competence; it raised the ceiling.
Student 3: Grade 11 | 2.9 overall | Guitar strings and simple harmonic motion
| Phenomenon Guitar string vibration, standing waves, resonance |
Entry point
Student 3 brought their full grade profile to their first message and flagged that their low Argument from Evidence score was an outlier, "I HAVE TO RAISE MY GRADE for #3, I only have 1 grade, that's also why it's so low." The AI recognized the difference between a low score and a low ceiling and reoriented the session around making this project their chance to show that practice. Asked what had been hardest, they answered: everything. The AI did not accept the deflection; it pressed for specifics until they named something workable.
The Socratic arc
They chose guitar strings, not the obvious choice for demonstrating SHM, but the right one, because they already understood the instrument intuitively. From there, they did the conceptual work largely on their own, tracing the energy chain from finger to string to sound and heat, and identifying string tension as the restoring force that makes the motion harmonic. They connected standing waves, nodes, and antinodes to a phenomenon they had lived with for years without naming it as physics. Their rubric named exact concepts at each level, and the AI noted it was "a rubric a teacher could actually use."
What the session reveals
The student came in thinking about their grade. They left having done genuine wave analysis grounded in something they cared about. The AI honored their motivation without letting the session stay there, redirecting a grade-focused start into authentic engagement.
Outcomes
Three students, three grades, three phenomena, three different content areas within one course. The same Socratic partner met each of them differently, surfacing misconceptions for one, stress-testing strong reasoning for another, and redirecting strategic self-awareness for the third, while holding the same content standard throughout. None of the three was handed an answer. Each was handed a better question, and each, at the password checkpoint, was handed back to their teacher.
In sociales, students worked in groups to design a conversational AI bot that supports women facing violence in Colombia. Using prior research on Colombian institutions, women’s rights, conflict resolution, and protection routes, they programmed the bot to provide empathetic guidance and explain available resources. The project required students to balance accurate institutional information with sensitive, human-centered communication.
Sociales Class: Grade 10 | Gender roles, violence prevention, and institutional trust in Colombia
Entry point
Students entered with strong academic performance and a clear social concern: how to make institutional information feel emotionally accessible to vulnerable users. The assignment was not framed as "using AI to get answers," but as designing a conversational bot that could responsibly guide women facing violence in Colombia to formal routes of assistance. The AI was positioned not as a tutor but as a toolmaker, something the student had to shape, constrain, and evaluate.
From the beginning, students realized that factual correctness alone would not be enough. They had to decide how the bot should speak, what tone it should use, when it should provide institutional guidance, and where its limits needed to be explicit. The project immediately moved them into questions of empathy, ethics, and public trust.
The arc
The intellectual demand emerged through design decisions rather than content retrieval. A search engine could provide definitions of psychological or economic violence, but it could not force the student to decide how a frightened user might interpret the bot's response. Students had to anticipate real human situations, distinguish between legal information and emotional support, and determine when the bot should recommend institutions such as the Defensoría del Pueblo, the Personería, or Comisarías de Familia.
The Spanish-language dimension mattered deeply. Students worked carefully on tone, realizing that wording that sounded technically correct could still feel cold, judgmental, or unsafe in context. The AI interaction pushed them toward perspective-taking: not simply identifying violence categories, but considering how women in Colombia might experience fear, shame, economic dependence, or distrust of institutions. The assignment demanded that they translate civic knowledge into culturally situated, emotionally responsible communication.
A representative moment
“I learned that AI does not generate responses automatically on its own, but rather builds its answers based on how I formulate the question. If I ask something very broad, the response is more general; if I provide clear instructions, the result becomes more precise. I also realized that the tone, wording, and level of detail I use guide the AI: asking for something formal, brief, or emotional completely changes what it produces. In other words, the quality of the response depends greatly on the clarity and intention behind the instructions I give.”
What the session reveals
This assignment shows AI increasing intellectual demand by shifting the difficult work toward judgment, mediation, and ethical design. Students were not rewarded for generating fast answers; they were required to evaluate how language, institutional knowledge, and emotional tone interact in real social contexts.
AI did not replace disciplinary thinking; it created a structure in which disciplinary thinking became necessary. The bot could only function responsibly if the student understood Colombian institutions, gender violence, conflict resolution, and human rights protections well enough to design meaningful interactions around them.
The cultural and linguistic specificity was essential. This was not a generic chatbot about "women's issues." It was grounded in Colombian legal structures, Spanish-language communication norms, and culturally specific tensions around gender, authority, and institutional trust. That specificity transformed the AI task from content production into authentic social reasoning.
In the mathematics projects, students investigated real-world systems under pressure through mathematical modeling, statistical reasoning, and AI-supported critique. Using real, simulated, or AI-assisted realistic data, students explored topics connected to climate change, social systems, media, and public belief while building models, interpreting graphs, revising conclusions, and critiquing representations. The AI tools were designed as thinking partners rather than answer generators, using Socratic prompting, reflection, critique cycles, and teacher-guided checkpoints to support visible reasoning, revision, and collaborative sense-making. In a separate session, the efficacy of using an AI tool, Flint, as a substitute teacher during a teacher's absence was tested.
Accelerated Geometry Class: Grade 9 | Statistical Inquiry (One-Variable Statistics)
The student began with a highly ambitious and emotionally engaging conspiracy claim about industries manipulating collective memory through the Mandela Effect. They were confident in the idea and brought strong curiosity, creativity, and cultural examples into the investigation, but initially approached the task rhetorically rather than statistically, relying on anecdote and causal assumption rather than measurable variables or defensible evidence. The student had chosen the topic independently after exploring different possible claims and was motivated by genuine personal interest. At the start of the interaction, the AI acted as a statistical reasoning partner: helping the student identify numerical and categorical variables, clarify what could realistically be measured, and rethink the difference between proving a claim and investigating a pattern. Rather than shutting the idea down, the AI validated the student's interest while redirecting attention toward realistic data, overlap, variation, interpretation, and limitations.
The arc
The AI moved the student forward through repeated cycles of support, critique, revision, and reinterpretation. It helped them brainstorm datasets, interpret graphs, understand measures of center and spread, and analyze box plots, histograms, overlap, variability, outliers, and standard deviation, while consistently refusing to provide certainty where the mathematics could not justify it. When the student overstated conclusions or blurred correlation with causation, the AI redirected back toward statistical defensibility and limitations.
At one point, the student jokingly remarked that "the AI is in on the conspiracy" because multiple AI systems resisted unsupported causal conclusions in similar ways. While humorous, the comment revealed an important epistemic tension. The student was beginning to recognize that AI systems were not neutral validators of any claim, but tools operating within evidentiary and interpretive constraints. Rather than dismissing the comment, the interaction became another opportunity to distinguish between speculation, interpretation, correlation, and defensible statistical argument.
Across the session, the student gradually shifted from trying to "prove" a conspiracy toward understanding how graphs and representations shape audience interpretation. The most significant development was epistemic rather than procedural: the student began recognizing statistics not as neutral truth, but as interpretive tools that can strengthen, complicate, or distort a narrative depending on how they are used.
A representative moment
“I kind of thought the AI was part of the conspiracy at first because it kept pushing back on everything I said. But then I realized that was kind of the point. The more we worked on the graphs, the more I noticed how easy it was to shape what people believe just by changing what you show, hide, or emphasize. At some point the conspiracy stopped being just the topic — it became the way we were presenting the data.”
What the session reveals
This interaction reveals AI functioning as a sustained epistemic critique partner. The mathematics remained rigorous throughout: distributions, measures of center and spread, overlap, variability, multiple representations, and statistical communication, but the deeper learning emerged through revision, critique, and increasingly sophisticated reasoning about evidence and representation. The student also began recognizing that bias can emerge not only from conclusions, but from datasets, omitted categories, selective graphing, framing choices, and representation itself.
Mathematics Student: Grade 10 | End Behavior of Polynomial functions
Entry point
This lesson came out of a pretty normal teacher problem: I was absent and didn't want my students to just have a study hall day or a busywork packet. Instead, I used a Flint AI lesson on multiplicity of zeros as part of my sub plans. My goal was to keep students actively thinking about math even while I wasn't physically in the room.
Most students came in comfortable with factoring polynomials, especially pulling out GCFs and recognizing perfect square trinomials. Where they struggled was connecting multiplicity to graph behavior. A lot of students could factor correctly, but then misread multiplicities from exponents or confused the signs of zeros.
The arc
Throughout the lesson, Flint consistently pushed students to reason through mistakes instead of immediately giving answers. When students got multiplicities or graph behavior wrong, the AI usually responded with a question like, "What's the exponent on that factor?" or "Is that odd or even?" That kept the thinking on the student instead of the AI doing the work for them.
One student showed this progression really clearly. They could factor correctly, but sometimes mixed up multiplicity and graph behavior. At one point, they said both zeros would "bounce" because they overlooked the exponent on x³. Instead of correcting them directly, Flint guided them back to the factor, and they eventually self-corrected, "Oh, so one crosses and one is tangent."
A representative moment
One of the best moments came during the challenge problem, where students had to work backwards from x-intercepts to build a polynomial equation. The student got stuck when Flint introduced the point (1, 216) to solve for the leading coefficient and finally admitted, “I don’t understand.”
Instead of just solving it for them, Flint slowed things down and changed its approach. It used a “recipe” analogy to explain how each x-intercept creates a factor, modeled the first few steps, and then gradually handed the work back to them. Eventually, they built the factored form themself and solved for the coefficient through guided questioning and self-correction. What stood out to me was that the AI didn’t remove the struggle. Instead, it supported them through it. The breakthrough came because they stayed engaged long enough to make sense of the structure themself, not because the answer was simply given to them.
What the session reveals
This session completely changed the way I think about AI during teacher absences. Normally, when a teacher is out, expectations naturally drop: students get review work, independent practice, or just a quieter day. Instead, this lesson kept students actively engaged in mathematical thinking throughout.
What impressed me most was how much visibility Flint gave me into student thinking. During the lesson, it sent notifications when students repeatedly struggled with concepts such as multiplicity and correctly identifying zeros. Afterward, I could see how long students spent on the activity, their engagement levels, and where they were strong conceptually versus procedurally.
Instead of coming back to a completed worksheet, I came back to actual formative assessment data. Flint was identifying misconceptions, tracking persistence, and helping me see patterns across the class. Even though I was physically absent, I honestly felt more connected to my students' thinking than I normally would after a traditional sub day.
What we learned
The worry behind this whole project was that integrating AI would mean lowering the bar. What we found was closer to the opposite. Across six sessions in three subjects, a high and rigorous bar was held throughout, and what changed was not its height but where it was applied. The AI met a struggling student, a confident one, a strategically anxious one, a class designing for vulnerable users, and a student chasing a conspiracy theory, all at different starting points, and did different kinds of work with each of them. But it held every one of them to the same standard: a defensible claim, evidence tied to an argument, and reasoning the student could explain in their own words. The differentiation happened in the questioning, not in the expectations. That is a distinction worth sitting with, because it is the whole game. Good teaching has always meant meeting students where they are without moving the finish line, and a well-designed AI task can do exactly that, at a scale and patience no single teacher can sustain across thirty students at once.
In the mathematics examples, AI did not make the work less human; it made the relational nature of mathematics more visible. Students were not simply calculating or producing answers; they were negotiating meaning: deciding what counted as evidence, what a representation revealed or concealed, and how their interpretations changed when questioned by peers, teachers, or AI. What surprised us most was that AI interactions often pushed students back toward one another rather than into isolation. After moments of insight, students regularly sought out classmates or teachers to share discoveries, compare interpretations, or test ideas aloud. The rigor came not from the tool alone, but from the classroom culture and task design surrounding it: mathematics as a co-constructed practice of interpretation, critique, revision, and shared sense-making.
The physics project examples provide a practical answer to the cognitive-debt problem. The MIT researchers found that students who let AI do the thinking disengaged, while those who did their own thinking first and brought AI in afterward maintained cognitive rigor. Our checkpoint designs enforced that sequence by construction. A student could not advance until they had reasoned through the current step in their own words, and even then, they had to leave the screen and get a password from their teacher before going on. The tool was built so that the only way through it was to think. The password was never really a security measure; it was a decision about who holds authority in the room. Every checkpoint pulled the student out of the conversation with the machine and back into a conversation with a person, and more than once, those check-ins surfaced exactly the kind of teaching moment no software could have produced.
This is where the frameworks stopped being abstract for us. In SAIL terms, these sessions lived at the Ideate and Evaluate levels: the AI helped students generate and pressure-test their own thinking, but it never authored their work, and the assignment forbade AI-generated content in the final product. That is the same line UNESCO draws when it describes students as critical evaluators and co-creators rather than passive consumers. The "teacher-AI-student" dynamic that the framework names was not a metaphor in our classrooms; it was literally the structure of the checkpoints. The principle on paper was that AI should enhance human judgment rather than replace it. The design is what lets a student feel the difference.
So the deeper realization, the one that reframed the work for us, is that standards and AI literacy were never actually in tension. A student who interrogates an AI's questions, defends a claim under pressure, corrects their own misconception out loud, and builds a model they can explain is demonstrating content mastery and AI literacy in the same act. We did not have to choose between them. We had to design a task in which they were the same thing.
Looking across these classrooms, what stands out is not the technology but the consistency of the pattern. Regardless of subject, student readiness, or starting point, the AI never lowered the standard. Instead, it differentiated through questioning while keeping the intellectual destination the same. The result was not less rigorous learning, but more visible thinking. If the same design holds across physics, statistics, and a Spanish-language civics project, the question for any teacher is not whether it fits in their subject, but where in the subject matter it could go.
Those experiences also changed the way we think about AI as educators. They challenged an assumption we started with: that taking cognitive load off students helps them, when in fact it can let them skip the metacognition that often matters most. And they clarified which design decisions made the difference. The technology itself was never the determining factor. The instructional choices surrounding it were.
In the final article, we reflect on what these experiences taught us and share practical principles for designing AI-supported lessons that strengthen, rather than weaken, student learning.
Jacob LaPlante is an Adavance Placement science teacher, science department lead, and AI team lead at Colegio Nueva Granada in Bogotá, Colombia. He has over 10 years of experience teaching science and math in the United States, Vietnam, Bahrain, and Colombia, with a current focus on integrating AI into international curricula. For eight years, he has been working with technology development policies, classroom integration, and professional development.
LinkedIn: https://www.linkedin.com/in/jacob-laplante-2a451682/
Jason Boll is a math educator, varsity basketball coach, and student leadership mentor at Colegio Nueva Granada. He has eight years of teaching experience in Taiwan and Colombia, working with students in both academic and advisory settings. His work focuses on creating engaging, student-centered learning experiences that emphasize collaboration, critical thinking, and meaningful relationships.
LinkedIn: https://www.linkedin.com/in/jason-boll-17b133239
Andrew Cajina is an international mathematics educator currently teaching at Colegio Nueva Granada in Bogotá, Colombia, with teaching experience across Canada, South Korea, Nicaragua, and Colombia. His work focuses on inquiry-driven mathematics learning, mathematical modeling, and critical evaluation of complex real-world systems. He is currently finishing a master’s focused on ethical AI integration in education that supports deeper learning.
LinkedIn: https://www.linkedin.com/in/andrew-cajina-5874791b3/
Manuela Peñalosa is a social studies teacher and member of the AI core team at Colegio Nueva Granada. She has more than 10 years of experience teaching social studies across Chile, Costa Rica, Panama, and Colombia. In addition, she has over six years of experience designing and facilitating professional learning experiences focused on the pedagogical use of digital tools. She currently provides virtual professional development in artificial intelligence for educators throughout Latin America.
LinkedIn: https://www.linkedin.com/in/manuela-peñalosa-632672b8