Skip to content

How to compare learning methods fairly

Comparisons between ways of learning usually fail at the unit. A flashcard deck counts cards, a language app counts words and streaks, a lecture course counts hours, a tutor counts sessions. None of these are the same size, so the numbers slide past each other. The units can be made comparable, though, by asking three questions of any method. This post is the ruler, with the published numbers filled in.

Three questions make any method comparable

First: how many minutes does the method spend per idea it actually verifies, rather than merely presents? Second: at what level is that verification, recognizing an answer, producing it, or applying it to a new case? Third: does it verify every idea, or sample a few on an exam? Time per idea means little until you know what was proven.

Covered is not verified

A lecture covers five to ten concepts in fifty minutes, which sounds efficient until you apply question one: how many of those concepts does it verify? At presentation time, zero. Verification is deferred to an exam that samples some of them weeks later, mostly at the memorization level. Video scales the same property to millions of people: among those who enroll in massive open online courses, about 3 percent finish. Coverage without verification is exposure, and exposure fades.

The numbers, on one scale

Method One unit Time per unit What is verified
Spaced-repetition decks (Anki, SuperMemo) a flashcard ~1.2–1.8 min per card per year, recurring recall of that card; every card
Duolingo a word or pattern ~34 hours ≈ one university semester of beginner Spanish recognition and translation, by placement test
Lecture course a concept ~5–10 min of exposure, plus exam preparation a sample of concepts, 93% at memorization level
MOOC video a concept ~10–15 min of exposure nothing, for the ~97% who stop
Human tutor a concept ~5–10 min, informal checking whatever the tutor probes; unrecorded
Classic software tutors (math, programming) a skill step a few minutes per skill application of that skill; every skill. About as effective as human tutors
Ulern a learning target: an idea plus the level you aim for ~4–9 min per target, one-time (internal figure) production at the target’s level; every target

Three of these numbers deserve their footnotes. The spaced-repetition figure is SuperMemo’s own long-run estimate: 200 to 300 items per year for each minute of daily review. The Duolingo figure comes from an independent 2012 study the company commissioned: 26 to 49 hours of app use matched one college semester of Spanish by placement score, with large variance between learners. The Ulern figure is our own internal measurement, described below, and is the only row not backed by an outside study.

Reading the table honestly

The units are different sizes. A flashcard is smaller than a concept, and one application-level idea decomposes into several recall-level cards, so compare orders of magnitude, not decimals.

The recall row’s cost recurs. A deck’s minutes per card repeat for every year you keep the material alive. That is also its strength: durable retention comes from spaced re-checking, and decks automate exactly that. A one-time mastery session, in any system, still needs later re-checks to last.

Course hours buy more than verified ideas: breadth, projects, peers, a credential. The table prices none of that, and is not an argument against any of it.

Level is the multiplier that matters most. Recognizing a vocabulary word and correctly deciding whether a statistical method applies to a new problem can both be “one unit verified”, and they are not the same achievement. A method that verifies only recognition is not faster than one that verifies application. It is measuring a different thing.

Where Ulern stands

We build Ulern, a learning system, so read this section knowing that.

Our row comes from internal test runs: about three activities per target, and sessions of about seven targets in 30 to 60 minutes, which is where 4 to 9 minutes per target comes from. Every target’s final check asks you to produce at the target’s own level, and targets are checked one by one rather than sampled. We have not yet run an external trial, which is why our row carries the label it does. The framework is the point of this post. Apply the three questions to us with the same skepticism you would apply to anyone else’s table.

← All posts