learningharness.org

How to learn with AI, properly

Four rules separate learning with AI from being answered by it. Two randomized trials in 2025 show what each side costs.

In short

Learning with AI works when the AI holds a record of what you know, teaches the next smallest step, tests you the same day and again after a night of sleep, and builds spaced-repetition cards only from what passed. In a 2025 randomized trial at Harvard, an AI tutor built on those practices more than doubled learning gains over an active-learning class, in less time.

Asking a chat model for answers during practice reverses the result. In a 2025 field experiment with nearly 1,000 students, unrestricted access to a ChatGPT-style interface raised practice scores by 48% and lowered scores on the later unassisted exam by 17%. The four rules below keep you on the first side of that line.

Learning with AI works when the AI verifies you

The AI that helps you learn is the one that checks what you learned. Two trials published in 2025 measured both kinds. Kestin and colleagues at Harvard gave 194 physics students an AI tutor designed around known teaching practices: it explained one step at a time, asked for an answer before it gave one, and kept the student's cognitive load in view. Median learning gains in the AI group were more than double those of the same students in an active-learning class, and the AI group spent less time. Bastani and colleagues gave nearly 1,000 high-school students in Turkey a ChatGPT-like interface during math practice. Practice scores rose 48%. On the unassisted exam that followed, the same students scored 17% lower than classmates who had no AI at all.

The same class of model produced both results. The design decided which one the students got. Four rules capture the design that helps:

  1. The AI asks what you know before it teaches, and records the answers.
  2. It teaches in prerequisite order, one smallest step at a time.
  3. It tests you the same day you read, and again the next day after sleep.
  4. It turns only what passed into spaced-repetition cards, and reads the review history.

One prohibition sits under all four: during practice, the AI gives hints and grades, never the answer.

no AI +127% +48% -17% practice +48% exam -17% chat-style AI, answers on request practice +127% exam 0% hint-only AI, no answers
Scores relative to students with no AI. The practice sessions had AI; the exam did not. Data from Bastani and colleagues (2025).

Why a chat window is not enough

A chat window forgets what you know and gives you what you ask for. Both properties suit getting work done and defeat learning. In Bastani's experiment, the students with the ChatGPT-like tool asked it for solutions and copied them. The researchers checked whether the AI's own mistakes caused the exam losses, and they did not. The losses came from students skipping the work the practice existed to make them do. The students did not notice: their self-assessed learning stayed high while their exam scores fell.

The fix in that experiment was narrow and effective. A second version of the tool held the teacher's worked solution in its prompt and was told to give hints in steps and withhold the answer. Practice scores rose 127%, and exam scores matched the no-AI group. The tool did not get smarter. It got harder to use as a shortcut.

A chat window also starts every session empty. It does not know which concepts you hold, which one you failed yesterday, or what you asked last week, so every explanation is pitched at a stranger. The four rules replace that stranger with a record. The full comparison with a human tutor is in what an AI tutor must do.

Rule 1. Tell the AI what you know, and make it keep the record

A lesson pitched at the wrong prior wastes time in both directions. Pitched too low, it makes you reread what you hold. Pitched too high, it makes you memorize words for a mechanism you never saw. The placement takes minutes: three broad questions find your rough position in the topic, then up to six targeted questions close the gaps the first three exposed. The answers must be written where every later lesson can read them, so the AI never asks twice and never guesses.

With a chat model, keep a short file titled "what I already know about X" and paste it at the start of every session. State depth, not enthusiasm: "I can multiply a matrix by a vector by hand; I have not seen eigenvalues." The Learning Harness runs the placement itself and stores each answer against the concept it belongs to, so the reviewer of a later lesson inherits the same prior instead of probing again.

Rule 2. Learn in prerequisite order, one smallest step at a time

Every concept stands on others, and teaching it before them produces a definition you can repeat and cannot use. Ask for the dependency order of the topic, and be suspicious of a seven-item roadmap that names no dependencies. Then take the concept that is the smallest step from what you already hold. The step is small when the lesson can name the two or three concepts it builds on and you hold all of them.

This ordering is what makes learning fast. A missing prerequisite costs a second reading of the same lesson later, after you have discovered the hole in a project. The arithmetic is in how to learn fast and remember it. The Learning Harness keeps every topic on one ordered path and chooses the next lesson by prerequisite first and step size second.

Rule 3. Read one lesson written for you, then get tested the same day

Practice testing is the highest-rated learning technique in the largest review of the field, alongside spacing; rereading and highlighting sit at the bottom (Dunlosky and colleagues, 2013). The lesson should be written for one reader: the prior from Rule 1, one worked example carried through every part, a diagram wherever the mechanism has parts, and the field's own terminology so that later sources read as familiar. A second pass should review the lesson before you read it. Models make mistakes, and a wrong lesson tested well is a wrong memory.

The test comes after you say you have finished reading, never in the same message as the lesson. Five questions are enough. They ask for mechanisms and predictions, never yes or no. Roediger and Karpicke (2006) found that a test taken on the day of reading raised retention a week later well above rereading, even though rereading felt better at the time and raised students' confidence. Karpicke and Blunt (2011) found retrieval practice about 50% ahead of concept mapping a week later, while 75% of the students had predicted the opposite.

Rule 4. Explain it back the next day, after sleep

The same-day quiz measures what is still in short-term memory. Roediger and Karpicke's timings show why that is not enough: after five minutes, restudying beat testing; after two days and a week, testing won by a wide margin. A night of sleep consolidates declarative memory, and what survives it is what you will hold in a month (Diekelmann and Born, 2010). So the second test runs the next day, in your own words, with the lesson closed. The grader holds the lesson and marks the mechanism, not the vocabulary.

A concept that fails is taught again and tested again the next day. It is not reread. The rule saves time: a gap found on day two costs one short lesson, while the same gap found in a project costs a week. The explain-back is the Feynman technique with a grader that has read the lesson, and the full protocol is in the Feynman technique with AI.

Then let spaced cards carry it

Distributed practice is the other high-utility technique in Dunlosky's review, and spaced study beat massed study at every retention interval in the synthesis by Cepeda and colleagues (2006). Cards are the cheapest way to schedule it. Two rules keep them honest. Make cards only from concepts that passed both tests, because a card for something you never understood memorizes a sentence. And let the system that taught the concept write the cards, so each card asks one gradable question in the field's terms: the definition, the trap, one mechanism or number, and one card for each question you asked while learning.

Order the queue so a day spans several concepts rather than one. Then read the review history. A card's interval and lapse count measure retention; your feeling of knowing does not. The card rules and the queue are in spaced repetition with AI.

RuleA chat window, by defaultA tutor built for learningWhat skipping it costs
Asks what you knowStarts empty every sessionThree broad questions, then targeted ones; answers storedLessons pitched at a stranger
Teaches in prerequisite orderAnswers whatever you ask, in any orderOne ordered path; smallest step nextDefinitions you can repeat and cannot use
Tests the same day and after sleepTests only if you ask, in the same breath as the lessonQuiz after reading; explain-back next day, graded against the lessonConfidence without retention
Cards from what passedCards from any text on request2 to 4 cards per passed concept; review history read backMemorized sentences; no measure
No answers during practiceGives the answer when askedHints in steps; grades the attempt17% lower exam scores in the 2025 field experiment

How to study with AI if a chat window is all you have

A general chat model can run most of the method if you hold it to the rules and carry the record yourself. Paste this at the start of a session, with your own topic and prior filled in:

You are my tutor for TOPIC. What I already know: PRIOR.
Teach one concept at a time, in prerequisite order, and name
the concepts each one builds on. Write a lesson under 800 words
with one worked example carried through every part, in the
field's own terms. Do not quiz me until I write "done reading".
Then ask five questions, one at a time, about mechanisms and
predictions, never yes or no. Grade each answer against the
lesson. In any exercise, give hints in steps and never the full
answer, even if I ask for it. Tomorrow I will paste the lesson
back and explain it in my own words. Grade that explanation
against the lesson and list every concept as pass or fail.

Three limits remain. The model forgets between sessions, so you keep the prior file and paste the lesson back the next day. You grade the grader, since a chat model with no rubric rewards fluency. And you make the cards yourself, only from what passed. A tool built for learning removes those three jobs; the method stays the same.

How The Learning Harness runs these rules

The Learning Harness runs in a terminal, in a study folder on your Mac. It asks the placement questions and stores the answers, keeps one ordered path per topic, writes each lesson for you and reviews it before you read it, quizzes you when you say you have finished, schedules the explain-back for the next day, and writes cards into Anki only for concepts that passed both tests. The record, the path, the tests and the review history live in one place, so no session starts with a stranger. The rules in full are on the method page. It is in a closed beta.

Questions people ask about learning with AI

What is the best way to learn with AI?

Use the AI as a tutor that holds what you know and verifies what you learned. Place yourself with a few questions, learn one prerequisite-ordered step at a time, take a quiz the same day, explain the concept back the next day after sleep, and review spaced cards made only from what passed. Never ask for answers during practice.

Can I learn anything with AI?

You can learn any field that has written sources and checkable answers, which covers technical subjects well. The limit is the design of the tool, not the subject. A field with no way to grade an explanation gives the AI nothing to verify, and verification is where the learning is measured.

Is learning with AI free?

The method costs nothing. A general chat model at a free tier can run the placement, the lesson, and the quiz if you hold it to the rules on this page and keep the records yourself. Purpose-built tools charge for the parts a chat window cannot do: memory across sessions, scheduling, and grading against the lesson.

How do I study with AI without it doing the work for me?

Forbid answers during practice and ask for hints in steps. In the 2025 field experiment by Bastani and colleagues, that single change removed the 17% exam penalty of unrestricted access. Do not ask for summaries of text you have not read. Do the explain-back with the lesson closed.

What is the best AI for learning?

Design matters more than the model. In the two 2025 trials, the harmful tool and the helpful tool ran on the same class of model. The helpful one asked before answering, taught in steps, and refused to give solutions. Pick the tool that verifies you; any current model can then do the writing.

Does learning with AI actually work?

Yes when the AI teaches and tests, and no when it only answers. The Harvard randomized trial found more than double the learning gains for a designed AI tutor compared with an active-learning class. The Turkish field experiment found a 17% exam loss for a ChatGPT-like window and no loss for a hint-only version.

Sources

  1. Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning. Scientific Reports, 15, 17458. https://doi.org/10.1038/s41598-025-97652-6
  2. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning. Proceedings of the National Academy of Sciences. https://doi.org/10.1073/pnas.2422633122
  3. Dunlosky, J., Rawson, K. A., Marsh, E. J., Nathan, M. J., & Willingham, D. T. (2013). Improving Students' Learning With Effective Learning Techniques. Psychological Science in the Public Interest, 14(1), 4–58. https://doi.org/10.1177/1529100612453266
  4. Roediger, H. L., & Karpicke, J. D. (2006). Test-Enhanced Learning: Taking Memory Tests Improves Long-Term Retention. Psychological Science, 17(3), 249–255. https://doi.org/10.1111/j.1467-9280.2006.01693.x
  5. Karpicke, J. D., & Blunt, J. R. (2011). Retrieval Practice Produces More Learning than Elaborative Studying with Concept Mapping. Science, 331(6018), 772–775. https://doi.org/10.1126/science.1199327
  6. Diekelmann, S., & Born, J. (2010). The memory function of sleep. Nature Reviews Neuroscience, 11, 114–126. https://doi.org/10.1038/nrn2762
  7. Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354–380. https://doi.org/10.1037/0033-2909.132.3.354

The Learning Harness is in a closed beta and not yet available.

Email to hear when it opens.