By Sal Khan, Founder and CEO of Khan Academy

Last week, our efficacy and research team flagged a new NBER working paper to me: “One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment,” by Philip Oreopoulos and Nina Low. Up front: Khan Academy was not involved in designing or running this study. It’s independent research, which is exactly why it’s worth sharing. This is my read on what’s interesting in it. Adjusting for the implied effect of a full year of active participation, a secondary analysis on below-grade-level students shows the gain was 0.14 SD (standard deviations). That’s the number I want to anchor on. The paper notes that 0.14 SD is in the range of experimental evidence of math practice software, which is in the range of 0.05 to 0.20 SD. This result is lower than the very strong 0.45 effect size we saw in the India study done by one of the same researchers, but the context here was very different. First, the student population in the recent study are below grade level in math. Second, the control group included usage of other ed tech tools. So the 0.14 SD could be interpreted as the benefit of Khan Academy and Khanmigo above and beyond the alternative ed tech tools, and helpful to students who begin the year below grade level.
The setup
Researchers at the University of Toronto and Charles River Associates ran this as a true randomized controlled trial (RCT) across 18 middle schools over two school years (school year ’24-25 and school year ’25-26) in Hamilton County, Tennessee. An RCT is generally considered the gold standard of evidence. Instead of comparing schools or students who happened to opt in, an RCT randomly assigns who gets the program and who doesn’t, which is the best approach we have for isolating cause and effect. Students in the study were below grade level and participating in a daily math intervention program. More than a quarter of the study sample had a special education designation. One knock on ed tech is that it rarely works for this student group. That’s another reason that I was interested in what this study would find. Can we help this student group make gains?
Why the control group matters so much here
Here’s the detail that matters most for interpreting that 0.14SD number: the control group was what the researchers call “business as usual.” This means they continued to use whatever was in place before the study. It isn’t Khan Academy compared to “no technology” or a traditional classroom. For most of these schools, business as usual meant other online tools like Waggle, IXL, Zearn, and DeltaMath. Of the 18 schools, 14 used at least one of these tools in their control condition.
So this study isn’t asking “does Khan Academy beat doing nothing.” It’s measuring Khan Academy against what the district was already running in their intervention period, which in most schools included other edtech products. In my opinion, that’s a much higher bar, and it’s the one that actually matters for a district deciding where to spend its budget. Put another way, this isn’t the total benefit of Khan Academy versus nothing. It is trying to measure the incremental benefit of Khan Academy over the existing interventions.
The headline number, and the one I think is the real one
Using the standard, most conservative comparison (Intent to Treat: everyone assigned to control vs. everyone assigned to Khan Academy, regardless of how much they actually used it), the combined two-year effect was about 0.06 SD, with Year 2 alone at 0.08 SD once implementation matured. Year 1 showed no reliable effect, which the paper attributes to a slow start: rostering delays and a mid-year testing-system switch.
But here’s the number I think is the more important read of the program’s actual impact: 0.14 SD.
The study has a “once in, always in” rule. Once a student was in the study, they were always in the study. But roughly two in five students actually graduated out of the intervention program during the study. They were still included in the calculation even though they weren’t there. When the researchers isolated only the students who stayed in the intervention program the entire second year, the treatment group using Khan Academy gained 0.14 SD compared to the control group. I would argue that it could have been even higher if the ⅖ of students who were seeing the most gain and graduated out instead stayed in the program and kept up that acceleration.
What the study leads with
I should be clear that a finding the authors put front and center isn’t this one. The title of the study is “One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment.” It was a study that intended to examine the effects of Khanmigo. However, students used Khanmigo infrequently, which matches what we’ve been saying publicly for a while. The paper attributes low Khanmigo use to behavioral barriers around help-seeking, which is compounded by a non-proactive interface, rather than to any limitation in what Khanmigo itself could do. We have observed this too, and it’s exactly why we rebuilt the classroom experience so that Khanmigo is part of practice rather than something students have to go and open. All our school district partners will use the redesigned experience this school year.
Why I’m sharing this
This is one of many studies on Khan Academy, and like several of them, we were not involved in the design or execution. A rigorous RCT showing a 0.14 SD effect for students who are below grade level and who got sustained access, measured against what the district was already running rather than against nothing, in a population that’s historically hard to move, is a genuinely strong result. I think the 0.06 and 0.08 headline numbers actually understate it.
Full paper here if you work in ed policy, edtech, or just want to see what rigorous evidence on ed tech looks like right now: https://www.nber.org/papers/w35620

