Kizuki
    Preparing search index...

    How people learn

    Researchers have tested study methods for over a century, and a few results hold up across hundreds of experiments and real classrooms. This page lists them, rates the evidence, and ends with what each means for Kizuki.

    Some numbers below are effect sizes, written as g or d. An effect size of 0.5 means the average learner who used the method scored about half a "spread" above the average learner who didn't, which moves a middle-of-the-class student to roughly the 69th percentile. Around 0.2 is small, 0.5 is medium, and 0.8 is large.

    Strong means several meta-analyses (studies that pool many experiments) agree and classroom studies confirm it. Moderate means consistent lab results but fewer classroom tests, or limits on when it works. Weak means few studies or open disputes.

    The best single starting point is Dunlosky et al. (2013), which rated ten common study methods. Improving Students' Learning With Effective Learning Techniques, Psychological Science in the Public Interest 14(1).

    Trying to recall something without looking, whether by answering questions, explaining from memory, or self-quizzing, beats reading it again. In Roediger and Karpicke's study, students who took a recall test remembered more a week later than students who restudied, though the restudy group felt more confident.

    • Evidence: strong. Rowland's meta-analysis found g = 0.50 against restudying; Adesope et al. pooled 272 effects and found practice tests beat restudying and every other comparison. The effect holds in classrooms (Yang et al. 2021).
    • Why it works: each recall strengthens the path back to the memory and shows you what you don't know.
    • Roediger & Karpicke (2006), Test-Enhanced Learning; Rowland (2014), meta-analysis; Adesope, Trevisan & Sundararajan (2017), Rethinking the Use of Tests.

    The same amount of study, split across days, beats one long session. Cepeda et al. measured how long the gap should be: for a test a few weeks away, the best gap is about 20% of the time until the test; for a test a year away, about 5 to 10%. Too short a gap wastes the review. Too long costs less.

    • Evidence: strong. Cepeda et al. (2006) pooled 317 experiments. Latimier et al. (2021) found spaced recall beat massed recall by g = 0.74, but found little evidence that growing gaps (1, 3, 7 days) beat even gaps.
    • Why it works: recalling something you have half forgotten takes effort, and that effort makes the memory last.
    • Cepeda et al. (2006), Distributed practice in verbal recall tasks; Cepeda et al. (2008), A temporal ridgeline of optimal retention; Latimier, Peyre & Ramus (2021), meta-analysis.

    This method joins the first two: in each session, keep recalling an item until you get it right, then come back to it in later sessions. Rawson and Dunlosky showed large gains in real college courses, with students keeping material weeks later.

    • Evidence: strong for facts and definitions; less tested for deep understanding.
    • Why it works: you get both the recall effect and the spacing effect, and nothing stays wrong.
    • Rawson, Dunlosky & Sciartelli (2013), The power of successive relearning.

    Students who prepare to teach learn more than students who study to take a test, and those who then teach learn more still. Kobayashi pooled 28 studies: g = 0.35 for preparing to teach, g = 0.56 for preparing and then teaching. Interactive teaching, where the "student" asks questions, beat lecturing to a camera. Roscoe and Chi found that tutors often slip into "knowledge-telling" (repeating what the book says). The gains come from "knowledge-building," which questions from the student trigger.

    While reading or working a problem, you explain to yourself why each step follows or how a new fact connects to what you know. Chi et al. found that students prompted to self-explain understood a biology text far better.

    • Evidence: moderate. Bisra et al. (2018) pooled 64 studies: g = 0.55. Dunlosky rated it moderate because it takes time and the prompts matter.
    • Why it works: it makes you fill in the steps the text skipped.
    • Chi et al. (1994), Eliciting self-explanations; Bisra et al. (2018), meta-analysis.

    For each stated fact, you ask "why would this be true?" and answer it.

    • Evidence: moderate. It works best when you already know something about the topic, and it has been tested mostly on lists of facts, not long texts.
    • Why it works: it ties the new fact to things you already know.
    • Dunlosky et al. (2013), above.

    Answering questions before studying, even with wrong guesses, helps you learn the answers once you see them, as long as you do see them.

    • Evidence: moderate. Pan and Carpenter reviewed over 60 papers and found the effect is real but varies with procedure. The gain is clearest for the exact questions asked; spillover to other material is mixed.
    • Why it works: a failed guess makes you notice and care about the answer when it arrives.
    • Kornell, Hays & Bjork (2009), Unsuccessful retrieval attempts enhance subsequent learning; Pan & Carpenter (2023), review.

    Mixing different kinds of problems or concepts in one session, instead of finishing one kind before the next, helps you learn to tell them apart. Learners consistently rate blocked study as better, even when their test scores say the reverse.

    • Evidence: moderate. Brunmair and Richter pooled 59 studies: g = 0.42 overall, 0.34 for math, and larger for visual categories. For word lists, blocking won (g = -0.39), and for prose texts the result was unclear. Interleaving helps most when the things mixed look alike but differ.
    • Why it works: side-by-side contrast teaches you which features matter.
    • Kornell & Bjork (2008), Is spacing the enemy of induction?; Brunmair & Richter (2019), meta-analysis.

    Bjork's framework explains why the methods above feel worse while working better. Rereading feels fluent, and fluency feels like learning. Recall, spacing, and mixing feel hard, and that effort builds lasting memory, as long as the learner can overcome it.

    Learners judge their own learning poorly, and they overrate it most right after studying. Dunlosky and Rawson found that overconfident students stopped studying too early and scored lower. Two findings help:

    • Judging your learning after a delay, by trying to recall, is far more accurate than judging right away (Nelson & Dunlosky 1991).
    • Errors made with high confidence are easier to fix once corrected than low-confidence errors (the hypercorrection effect; Butterfield & Metcalfe 2001).
    • Evidence: moderate to strong for the overconfidence and delay findings.
    • Dunlosky & Rawson (2012), Overconfidence produces underachievement; Nelson & Dunlosky (1991), delayed judgments of learning; Butterfield & Metcalfe (2001), hypercorrection.

    Showing the correct answer after a test makes testing work better and stops errors from sticking. Feedback about the task ("this contradicts page 4") helps; praise does little.

    Studying several varied examples of an abstract idea helps you recognize it later. One example is weak; people latch onto its surface details.

    • Evidence: moderate.
    • Why it works: varied examples show which features are the concept and which are incidental.
    • Rawson, Thomas & Jacoby (2015), The power of examples.

    For beginners, studying a solved problem step by step beats solving problems unaided, because working memory is small and searching for a solution uses it up. As learners gain skill, the benefit fades and can reverse (the "expertise reversal" effect), so support should fade too.

    Words with a relevant diagram beat words alone. Decorative images hurt.

    Mastery learning requires students to pass a check on one unit before starting the next, with extra help for those who don't. Kulik et al. pooled 108 studies and found an average gain of about half a spread (d ≈ 0.5), larger for weaker students. Bloom's famous "two sigma" claim for tutoring has not held up at that size.

    Sleep after learning helps keep new memories, and sleeping between two study sessions helps relearning. Mazza et al. found that people who slept between sessions needed fewer tries to relearn and remembered more six months later.

    • Rereading. Students use it the most, and it does the least. A second read soon after the first adds little (Callender & McDaniel 2009, The limited benefits of rereading). It feels productive because familiar text reads smoothly. Dunlosky: low.
    • Highlighting and underlining. No reliable benefit, and it can hurt by pulling attention to single facts over connections. Dunlosky: low.
    • Summarizing. Helps only students already trained to summarize well. Dunlosky: low.
    • Learning styles. The idea that you learn best when teaching matches your "style" (visual, auditory) has no support. Pashler et al. found almost no studies with the right design, and those that had it found no effect (Pashler et al. 2008, Learning styles: concepts and evidence). People do have preferences; matching them doesn't improve learning.
    • Cramming. It works for tomorrow's test and fails for next month. Spacing beats it at any delay longer than a day or two.
    • Judging learning by feel. Fluency and confidence right after study mislead; delayed recall is a better gauge.

    The table lists what Kizuki does today and ideas it could add. Ideas are not built until a spec says so.

    Kizuki's rule: the model never writes facts in its own words. It picks sentence labels, question kinds, and the user's own words; code writes the questions from templates. Ideas that would break this rule are marked (breaks the rule).

    Method Kizuki today Feature idea
    Practice testing Does it. Teach-back is recall, and the material stays hidden until you finish. Built in 1.0.0.
    Spacing Does it. Let the user enter an exam date and set gaps to about 10–20% of the time left (Cepeda). Even gaps are fine; no need for a clever curve.
    Successive relearning Does it. After misses, Kizuki offers up to 3 tries in the same sitting. Built in 1.0.0. Only the first try sets the schedule, so practice never pushes a review later.
    Learning by teaching Does it, in the interactive form, the strongest kind. The contradiction and gap questions push "knowledge-building." Track when a teach-back mostly repeats material sentences word for word (code can measure overlap) and ask the user to say it differently.
    Self-explanation Partly. Add a "why" question kind with a template: "The material says: "[S3]". Why does that follow?" The model picks S3; the user supplies the reason.
    Asking "why?" Could add with the same template. Best for users with some background; offer it after a concept's first clean review, not on first contact.
    Pretesting Could add. Before the user reads a section, show only the confirmed concept names and ask "What do you think [concept] means?" (template), then show quotes. Record guesses but never schedule from them.
    Interleaving Could add. Mix due concepts from different sections in one review, within the prerequisite order. Add a "compare" kind: "The material says "[S3]" about A and "[S7]" about B. How do they differ?" Code can find look-alike pairs with the existing meaning search.
    Calibration Could add. Ask "How well do you know this? (1–4)" before each review and chart confidence against catches over time. Schedule high-confidence misses sooner, since corrections stick best there. All code.
    Feedback Does it. Every question quotes the material. Show the quote next to the user's own words after each answer, so the correction is specific.
    Recording catches Does it. Catches feed calibration and relearning. Repeated catches on one sentence could link to that passage by location.
    Concrete examples Could add, limited. Quote example sentences from the material, chosen by the model by label. Ask the user for their own example and check it against quotes with existing question kinds. Having the model invent examples (breaks the rule).
    Worked examples Could add, limited. Link to worked examples in the material by location. Fade support: material visible on first teach-back, hidden on later reviews. Generating solved problems (breaks the rule).
    Words plus pictures Could add. Show figures from the source file (slide, PDF page) next to the quoted passage, by location. Model-drawn diagrams or captions (break the rule).
    Prerequisites Does it. Earlier concepts block later ones. Mastery research supports this. Define "learned" as a clean teach-back after a delay, not right after first study, since same-day success overstates learning.
    Sleep Could add. Schedule the first review for the next day at the earliest, never the same evening. All code.
    Rereading, highlighting Not offered. Keep it that way.

    The ideas that fit best with the least risk are calibration ratings, exam-date gaps, and next-day first reviews: they are all plain code. "Why," "compare," and pretest questions need only new templates and question kinds. Examples, worked problems, and pictures stay limited to what the material already contains.