Skip to content

AI tutor, course, or ChatGPT: what the evidence says

If you want to learn something today, three options present themselves: a structured course, a chat assistant, or one of the new AI tutors. The research for choosing between them is older than all three. It starts in 1984 with the most famous result in education research, and it has had a sharp update in the last two years.

What four decades of tutoring research says

One-to-one tutoring with mastery checking is the most effective form of instruction ever measured. Software that imitates its moves gets most of the way there. The first randomized trials of AI tutors match or beat live classrooms, in less time. And a chat assistant, by default, makes none of the moves that produce those results.

1984: the bar gets set

Benjamin Bloom reported that students tutored one-to-one, with mastery checks before moving on, performed about two standard deviations above a conventional classroom. The average tutored student outscored 98 percent of the class. He called it the 2 sigma problem, and the problem was cost: nobody can afford a tutor per student. Four decades of educational software descend from that framing. Later reviews put the average human-tutoring effect lower, around 0.8 standard deviations, which is still among the largest effects known for any intervention in education.

Software tutors got close

VanLehn’s 2011 review compared human tutors against intelligent tutoring software, mostly in math, physics, and programming: humans at d = 0.79, software at d = 0.76. A 2016 meta-analysis of 50 evaluations put the median software-tutor effect at 0.66 standard deviations over conventional instruction. The efficiency version of the result exists too: a Carnegie Mellon statistics course with practice and per-idea checking built in taught a semester’s material in half the time, with equal or better outcomes.

The catch: each of these systems took years to build for one subject, so they exist for a handful of school domains and almost nothing else. Removing that constraint is what the current generation of tools is attempting.

2023 onward: the first AI-tutor trials

A randomized trial at Harvard (194 students, published in Scientific Reports) compared an AI tutor against the same material taught in an active, well-regarded physics class. The AI-tutored students learned more than twice as much, in less time, and reported being more engaged. Two details matter before generalizing. The tutor was deliberately constrained: it revealed one step at a time, refused to hand over full solutions, and pushed students to attempt an answer before seeing one. And it ran on the instructor’s own vetted activities rather than on open-ended chat. Early trials elsewhere point the same direction, with the usual caution that this literature is young.

Why a chat assistant is not a tutor by default

A plain chat assistant answers what you ask, which is the opposite of most tutoring moves. It does not check what you already know. It does not sequence ideas so foundations come first. It does not make you produce answers, and it does not return to an idea after you get it right, which is where retention comes from. It keeps no record of what you have proven. Ask it to explain, and it explains: fluent input that feels like learning and mostly is not. The Harvard tutor produced its results precisely by constraining a chat model until it behaved like a tutor. The pedagogy was the product; the model was the material.

What to demand from anything calling itself a tutor

A checklist, in the spirit of the comparison framework:

  • It checks what you know before teaching you.
  • It sequences: foundations before the things that depend on them.
  • It makes you produce answers, not just pick them.
  • It returns to ideas after you succeed, and asks more the second time.
  • It can show you, per idea, what you have proven and what is still open.
  • It is honest about its evidence. Anyone claiming learner outcomes should be able to say who measured them, and how.

A structured course gives you sequence without adaptation: everyone gets the same path at the same pace. An assistant gives you adaptation without sequence: infinite patience, no plan. Tutoring, human or software, is the combination, and the combination is what four decades of results keep rewarding.

Where Ulern stands

We build Ulern, a learning system, and the checklist above is also our own bar made public. Ulern opens with a calibration check, sequences foundations first, makes you produce at the level you aim for, returns after success, and tracks every learning target as proven or still open. What we do not yet have is the outside evidence: no external trial, internal numbers only. By the standard of the last bullet, that sentence belongs in this post.

← All posts