Researchers have tested study methods for over a century, and a few results hold up across hundreds of experiments and real classrooms. This page lists them, rates the evidence, and ends with what each means for Kizuki.
Some numbers below are effect sizes, written as g or d. An effect size of 0.5 means the average learner who used the method scored about half a "spread" above the average learner who didn't, which moves a middle-of-the-class student to roughly the 69th percentile. Around 0.2 is small, 0.5 is medium, and 0.8 is large.
Strong means several meta-analyses (studies that pool many experiments) agree and classroom studies confirm it. Moderate means consistent lab results but fewer classroom tests, or limits on when it works. Weak means few studies or open disputes.
The best single starting point is Dunlosky et al. (2013), which rated ten common study methods. Improving Students' Learning With Effective Learning Techniques, Psychological Science in the Public Interest 14(1).
Trying to recall something without looking, whether by answering questions, explaining from memory, or self-quizzing, beats reading it again. In Roediger and Karpicke's study, students who took a recall test remembered more a week later than students who restudied, though the restudy group felt more confident.
The same amount of study, split across days, beats one long session. Cepeda et al. measured how long the gap should be: for a test a few weeks away, the best gap is about 20% of the time until the test; for a test a year away, about 5 to 10%. Too short a gap wastes the review. Too long costs less.
This method joins the first two: in each session, keep recalling an item until you get it right, then come back to it in later sessions. Rawson and Dunlosky showed large gains in real college courses, with students keeping material weeks later.
Students who prepare to teach learn more than students who study to take a test, and those who then teach learn more still. Kobayashi pooled 28 studies: g = 0.35 for preparing to teach, g = 0.56 for preparing and then teaching. Interactive teaching, where the "student" asks questions, beat lecturing to a camera. Roscoe and Chi found that tutors often slip into "knowledge-telling" (repeating what the book says). The gains come from "knowledge-building," which questions from the student trigger.
While reading or working a problem, you explain to yourself why each step follows or how a new fact connects to what you know. Chi et al. found that students prompted to self-explain understood a biology text far better.
For each stated fact, you ask "why would this be true?" and answer it.
Answering questions before studying, even with wrong guesses, helps you learn the answers once you see them, as long as you do see them.
Mixing different kinds of problems or concepts in one session, instead of finishing one kind before the next, helps you learn to tell them apart. Learners consistently rate blocked study as better, even when their test scores say the reverse.
Bjork's framework explains why the methods above feel worse while working better. Rereading feels fluent, and fluency feels like learning. Recall, spacing, and mixing feel hard, and that effort builds lasting memory, as long as the learner can overcome it.
Learners judge their own learning poorly, and they overrate it most right after studying. Dunlosky and Rawson found that overconfident students stopped studying too early and scored lower. Two findings help:
Showing the correct answer after a test makes testing work better and stops errors from sticking. Feedback about the task ("this contradicts page 4") helps; praise does little.
Studying several varied examples of an abstract idea helps you recognize it later. One example is weak; people latch onto its surface details.
For beginners, studying a solved problem step by step beats solving problems unaided, because working memory is small and searching for a solution uses it up. As learners gain skill, the benefit fades and can reverse (the "expertise reversal" effect), so support should fade too.
Words with a relevant diagram beat words alone. Decorative images hurt.
Mastery learning requires students to pass a check on one unit before starting the next, with extra help for those who don't. Kulik et al. pooled 108 studies and found an average gain of about half a spread (d ≈ 0.5), larger for weaker students. Bloom's famous "two sigma" claim for tutoring has not held up at that size.
Sleep after learning helps keep new memories, and sleeping between two study sessions helps relearning. Mazza et al. found that people who slept between sessions needed fewer tries to relearn and remembered more six months later.
The table lists what Kizuki does today and ideas it could add. Ideas are not built until a spec says so.
Kizuki's rule: the model never writes facts in its own words. It picks sentence labels, question kinds, and the user's own words; code writes the questions from templates. Ideas that would break this rule are marked (breaks the rule).
| Method | Kizuki today | Feature idea |
|---|---|---|
| Practice testing | Does it. Teach-back is recall, and the material stays hidden until you finish. | Built in 1.0.0. |
| Spacing | Does it. | Let the user enter an exam date and set gaps to about 10–20% of the time left (Cepeda). Even gaps are fine; no need for a clever curve. |
| Successive relearning | Does it. After misses, Kizuki offers up to 3 tries in the same sitting. | Built in 1.0.0. Only the first try sets the schedule, so practice never pushes a review later. |
| Learning by teaching | Does it, in the interactive form, the strongest kind. | The contradiction and gap questions push "knowledge-building." Track when a teach-back mostly repeats material sentences word for word (code can measure overlap) and ask the user to say it differently. |
| Self-explanation | Partly. | Add a "why" question kind with a template: "The material says: "[S3]". Why does that follow?" The model picks S3; the user supplies the reason. |
| Asking "why?" | Could add with the same template. | Best for users with some background; offer it after a concept's first clean review, not on first contact. |
| Pretesting | Could add. | Before the user reads a section, show only the confirmed concept names and ask "What do you think [concept] means?" (template), then show quotes. Record guesses but never schedule from them. |
| Interleaving | Could add. | Mix due concepts from different sections in one review, within the prerequisite order. Add a "compare" kind: "The material says "[S3]" about A and "[S7]" about B. How do they differ?" Code can find look-alike pairs with the existing meaning search. |
| Calibration | Could add. | Ask "How well do you know this? (1–4)" before each review and chart confidence against catches over time. Schedule high-confidence misses sooner, since corrections stick best there. All code. |
| Feedback | Does it. Every question quotes the material. | Show the quote next to the user's own words after each answer, so the correction is specific. |
| Recording catches | Does it. | Catches feed calibration and relearning. Repeated catches on one sentence could link to that passage by location. |
| Concrete examples | Could add, limited. | Quote example sentences from the material, chosen by the model by label. Ask the user for their own example and check it against quotes with existing question kinds. Having the model invent examples (breaks the rule). |
| Worked examples | Could add, limited. | Link to worked examples in the material by location. Fade support: material visible on first teach-back, hidden on later reviews. Generating solved problems (breaks the rule). |
| Words plus pictures | Could add. | Show figures from the source file (slide, PDF page) next to the quoted passage, by location. Model-drawn diagrams or captions (break the rule). |
| Prerequisites | Does it. Earlier concepts block later ones. | Mastery research supports this. Define "learned" as a clean teach-back after a delay, not right after first study, since same-day success overstates learning. |
| Sleep | Could add. | Schedule the first review for the next day at the earliest, never the same evening. All code. |
| Rereading, highlighting | Not offered. | Keep it that way. |
The ideas that fit best with the least risk are calibration ratings, exam-date gaps, and next-day first reviews: they are all plain code. "Why," "compare," and pretest questions need only new templates and question kinds. Examples, worked problems, and pictures stay limited to what the material already contains.