Goblins
Try Free Forever

Why Adaptive Learning Doesn't Work

Sawyer Altman

Written by Sawyer Altman, CEO

Kaitlyn LiPuma
"With adaptive tools, it takes kids on this demoralizing side of, "Oh you don't know this, but do you know this, oh you still don't know that, do you know this" and then all of a sudden they're at a second grade level and now they don't want to participate because the program's calling them stupid."
Kaitlyn LiPuma, math department chair at DREAM Charter

We are in the midst of a national math crisis, and districts are spending billions of dollars on 15-year-old technology that does not work. Not a single company out of the long array billing themselves as "adaptive learning" platforms has presented compelling evidence that they boost math outcomes. Some are lacking studies entirely, and the ones that show even modest effects sizes generally pay for their studies themselves and hire the researchers:

  • i-Ready — Its main math studies are nonrandomized (an easy sign of low-quality) and were conducted by Curriculum Associates or research groups working with it, and in a 2025 study, two years of recommended use produced no statistically significant benefit.
  • IXL — Its strongest study, funded by them, found small improvement on a single screener, which did not generalize to state tests, and they extrapolate large claims from results that are not statistically significant.
  • Renaissance Freckle — Renaissance paid for an observational study that found tiny effects in elementary school and no effect in middle school, even measuring with their own assessment, FastBridge aMath.
  • DreamBox — DreamBox’s evidence is narrow and sensitive to how the data are analyzed. In its best-known early experiment , a full-sample analysis initially found no significant effect. A different adjusted analysis barely produced one. A later, much larger trial found a benefit for younger students but essentially no effect on the state test for grades 3–5.
  • Zearn — Zearn’s largest and best-designed experiment did not produce statistically significant results on either of its two preregistered tests using the Texas state assessment. The study’s authors explicitly say it did not provide confirmatory evidence that Zearn improves grade-level achievement.

The list goes on. In fact, there is mounting evidence that they are counterproductive. Not only do they take up precious classtime, but as we'll see, they hurt students' math mindsets, making it harder to engage them in class.

We believe classtime and screentime are both more precious resources than ever, and as budgets are under strain, vendors and admin need to be able to justify the quality of the pedagogy they deliver. The first step to doing so is to understand why the current paradigm is failing.

Adaptive Tools' Core Ideas

Edtechies have been chasing the dream of personalized learning for decades.

Women using PLATO IV terminals
The PLATO IV terminal launched in 1972;

Their rallying idea was that the original sin of the public classroom is age-grading: rather than grouping kids by level, we group by age (something we stole from the Prussians), and therefore kids sit in the same class all at different levels, forcing teachers be the glue that keeps everyone on-task and growing. Early advocates claimed that personal computing meant that learning could finally be mastery-based — kids can be shown content, and assessed in real-time, providing "feedback" that teaches the system about the students' level and thereby keeping them in their zone of proximal development. Early attempts, like the PLATO IV, a computer terminal launched in 1972, attempted to do just this, but its rules-based engine for selecting topics was not quite smart enough yet.

Benjamin Bloom's presented his 2-sigma problem in the 80's, further animating this quest. Public figures like Sal Khan quotes Bloom constantly by name, having given a 2023 TED talk titled, "The Two Sigma Solution" (where he launched Khanmigo), and trendy programs like Alpha School also regularly allude to it. (However dubious the study may have been).

Then, about 15 years ago, it seemed like the missing piece had arrived. Machine learning and item response theory could supply the feedback loop that PLATO's rules engine lacked. All you needed was right and wrong answers; the algorithm would handle the rest.

The premise was an engineer’s dream: let’s just model all of the prerequisites in one big tree, score each leaf based on whether the student got it right or wrong, and keep running this in a loop, always serving the lowest-scored topic until the student has mastered the whole subject.

It's a beautiful design. It's also built on two assumptions about how math learning works that are simply wrong.

#1 — Adaptive tools do not know anything about students' process.

Adaptive tools ask students to pick from multiple-choice options or match answers in a textbox, yielding a right or wrong answer, which tells you almost nothing about a math error. Say a student simplifies -3(x - 3) and writes -3x + 6. The tool marks it wrong and assumes, "oh, this student doesn't get distributing over a negative". It will then force them to do more problems in that domain, and over thousands of subsequent problems, via a vague sense that this student is "struggling", it'll progressively pull them back all the way to 3rd-grade multiplication.

This is a fundamental problem of how adaptive tools assess. The nature of the student's misconception is not as simple as "distributing negatives", and yet that's what gets dinged. Think of this like trying to identify a face in a low-resolution image; there's not enough information to draw conclusions.

A better intervention would be to circle the math error for the student, affirm their grasp of the concept, but pump the breaks on the arithmetic. "Nice! But what did you do with those 3's?" You can't deliver that intervention if you can't see the work itself.

Missing process also limits teachers and staff from relying on adaptive tools as regular formative assessment. They can sense broad strokes over time from a lot of data, but they cannot report in real-time who in a class is struggling specifically with which concept, meaning teachers must get this information from additional assessment (more class time, and more teacher time spent grading).

#2 — Adaptive tools just pick problems and tell students they're wrong; they do not coach students.

It's a common misconception that Bloom's 2-sigma intervention was all about selecting the ultimate path; actually, the real magic was that a tutor provided corrective instruction in real-time.

Adaptive learning only automates, however crudely, the assessment-and-routing half. But the corrective teaching half is far more than right and wrong, and it's where the main effect size comes from.

The student who gets stuck at the bottom of Kaitlyn's "demoralizing slide" is stuck because the ladder they're given to climb back up has no rungs. "You're wrong, but good luck finding out why and how to fix it".

This is what leads to the common observation by teachers and admins is that with time, students resort to clicking aimlessly on their UIs, with no strategy, in loops of trial-error (6? No. 7? Still no. Ok 8?). Struggling students realize that, whether they bother to spend the energy thinking or not, they will only get out "wrong" responses, not enough information to actually learn from. Many will end up clicking just to avoid getting caught for being disengaged.

As a result, adaptive tools impinge on motivation, especially among the most vulnerable.

As evidenced by the efficacy data, adaptive tools are at best an unproductive suck of classroom time, where the classroom when run by a great teacher, is a clear site for student growth. But it's not enough that they are unproductive; they actually cause harm.

Adaptive tools provide students with the impression that they are constantly being punished instead of supported. In some programs, like IXL, this punishment borders on Sisyphean: IXL's SmartScore drops every time you get a problem wrong, meaning you owe the program more solves to climb back up. Compound this with the frustration of the common failure mode, when a student enters 0.75 when the program wanted 3/4, so the string doesn't match. IXL can't tell, so the score drops, and now the student is further behind, still confused, and saddled with the pangs of injustice.

This only decays a student's self-image as a math-learner, making them less likely to engage during regular classtime. And because they're often sold as tools for filling gaps and catching up struggling students, the students most frequently routed to adaptive tools are disproportionately the ones already at the bottom of the distribution — the kids most likely to have decided that they are "not a math person". This convinces them fully.

Goblins is raising the bar.

A tool that consumes the scarcest resources in the building — class time and screen time — should have a much high evidence bar, and none of these tools come close to beating normal classtime.

So what does have a higher evidence bar? What made Bloom's and subsequent studies so successful is just-in-time scaffolding based on a student's precise error. And consistently since, students given grade-level work with real-time feedback based on their authentic process outperform naive gap-filling.

This is exactly what Goblins does. If you're ready to inspire love of math learning and finally boost outcomes, .