← SAMs, Version Two
Teaching · DE Biology 101 · For teachers · Companion to SAMs, Version Two

The Evidence Behind SAMS2: What Supports Student-Authored Modules, What Does Not, and Why the Design Changed

A SAM rests on the two best-supported study strategies there are, and the review of the research moved the weight of the practice onto them.

I teach a dual-enrollment introductory biology course to high school juniors and seniors. For several years my students have built Student-Authored Modules: collections of their own questions, answered and explained, scored through a conversation with me about why they chose those questions. In September 2026 I reviewed about a hundred studies to test that practice against the research. What follows is the reasoning that came out of it, element by element, and the changes it led to.

The practice, in research terms

A student works from teacher-provided scaffolding: lecture slides, textbook clicker items, the course's Checkpoint question bank, and questions asked in class that nobody could answer. The student finds a gap between what they understand and what the course expects, writes an item aimed at that gap, answers it, and explains the reasoning. Items accumulate into a collection for each unit, revised across the year. The collection is scored as a whole, together with a short one-on-one account of why those questions were chosen. Where a student lacks the background to find a gap, I seed the item. AI may help with finding the gap, as long as the student can say how they used it.

In the field's terms, that is a scaffolded student-question-generation practice with a self-explanation component and an oral metacognitive assessment. Each of those parts has its own research literature.

Idea behind a SAMWhat the field calls itWhere the evidence stands
Answering your own questionsRetrieval practice; the testing effectStrongest in the set. Classroom g = 0.50 across 222 studies
Explaining the answerSelf-explanationStrong. Meta-analytic g = 0.55
Writing the questionStudent question generationStrong with stems, weak without; mixed against good alternatives
Asking “why” of factsElaborative interrogationModerate, and thin for complex content
Knowing what you don't knowMetacognitive monitoring; calibrationMonitoring can be improved; the link to learning rests on few studies
Talking it through with meOral assessmentSmall literature, including two biology studies, favorable

A SAM is an individual practice, not a group one, so it sits beside the well-known undergraduate pedagogies rather than competing with them. Active learning as a whole raises exam performance by about half a standard deviation and cuts failure rates from roughly 34 to 22 percent (Freeman et al., 2014). Peer Instruction gets much of its effect from discussion: in introductory genetics, students who talked after voting did better on a new question of the same kind (Smith et al., 2009). Both findings matter here, because SAMS2 borrows the talk.

What the evidence supports

Answering the questions, later and more than once — strong

Yang et al. (2021) pooled 222 classroom studies with 48,478 students and found a benefit of g = 0.499 for practice quizzing; high school was the strongest level (g = 0.655). Four moderators shape SAMS2 directly. Quizzing after instruction (g = 0.536) far outperformed quizzing before it (g = 0.186). The effect grew with repetition, from 0.444 for one test to 0.642 for three or more. A quiz format matching the exam helped. Feedback probably helped, though another meta-analysis found little difference (Adesope et al., 2017).

Two cautions travel with those numbers. Transfer to material that was never quizzed is real but smaller (Pan & Rickard, 2018), so quizzing part of a unit does little for the rest. And laboratory effects run far larger than classroom ones, which is why I quote the classroom figures. On the useful side, a high school study of weekly science quizzes tested a month later found interleaved quizzing beat blocked quizzing (Sana & Yan, 2022). A collection revised across a year is already an interleaving schedule.

The explanation underneath — strong, as self-explanation

Bisra et al. (2018), 64 studies and 5,917 participants, found self-explanation at g = 0.55. The effect roughly halves against an active comparison, and the kind of prompt matters a great deal: conceptual prompts reached g = 0.87 while metacognitive prompts reached 0.19. Education level did not moderate the effect, which lets me read the undergraduate literature across to my students. One design principle from that literature became a SAMS2 rule: prompt students to explain why a misconception is wrong (Rittle-Johnson, Loehr & Durkin, 2017).

The teacher-seeded SAM — strong, and the active ingredient

Rosenshine, Meister and Chapman (1996) reviewed 26 intervention studies on teaching students to generate questions. Generic question stems produced an effect of 1.12; no procedural prompt produced 0.14. King (1990) reached the same conclusion directly, with guided reciprocal questioning beating unguided, and King (1994) found stems that reach into prior knowledge worked better than stems confined to the lesson. I had treated seeded SAMs as a fallback for students who could not find their own question. The evidence says they are the mechanism. A student writing with no stem is in the 0.14 condition.

The conversation as the assessment — moderate, and the most interesting part

In biology, oral assessment produced higher scores than written assessment of the same questions, with no group of students disadvantaged (Huxham, Campbell & Westwood, 2012). An optional verbal final in introductory biology predicted later performance, though the students who opted in were not representative, which is an argument for making the conversation universal (Luckie et al., 2013).

Written reflection, by contrast, under-delivers. Upper-division biology students who recognized their strategies were ineffective kept using them, mainly to avoid the discomfort of changing (Dye & Stanton, 2017). Introductory genetics students did not become more accurate predictors of their own performance across a semester, and more frequent written reflection did not by itself go with better grades (Knight et al., 2022). Metacognitive exam-preparation assignments raised scores only for students below the median entrance score and barely moved their confidence (Angell et al., 2024). A live conversation removes the private exit that a worksheet leaves open.

Scoring the collection — consistent with the evidence, by analogy

No study scores a collection of student-written items as a unit. The closest evidence, 632 biochemistry students, found that gains tracked question quality rather than question count (Hilton et al., 2022). Scoring the collection and its rationale rather than tallying items fits that finding.

What a SAM replaces — strong

Rereading real textbook chapters rarely improved performance (Callender & McDaniel, 2009), and a major review rates rereading and highlighting low-utility (Dunlosky et al., 2013). It is nevertheless what students do: 84 percent of undergraduates at a selective university listed rereading as a strategy and 55 percent ranked it first, while 1 percent ranked self-testing first (Karpicke, Butler & Roediger, 2009). Students also misjudge which strategies work; in one set of studies 90 percent learned more from spaced practice while 72 percent believed massed practice had worked better (Bjork, Dunlosky & Kornell, 2013). The honest comparison for a SAM is therefore not another clever intervention. It is the rereading the student would otherwise do. Teachers should also expect students to report that SAMs feel less effective than rereading, because retrieval lowers the sense of fluency even as it raises retention (Roediger & Karpicke, 2006).

AI-assisted SAMs — untested

I found no peer-reviewed study of learning outcomes when undergraduates use a large language model to write their own study questions. In introductory physics, AI-generated practice problems were judged poorly by the AI itself on Bloom level and difficulty, and a few were off-topic entirely (Geisler & Kortemeyer, 2026). Since the gains in Hilton et al. tracked the quality of the student's own authoring, an AI that supplies the item removes the part that predicted learning. My rule follows from that: AI may help a student find the gap, but the item and the reasoning are the student's, and the conversation covers what the AI gave and what the student added.

Where the evidence pushed back

Writing the question is not what teaches

In three controlled experiments, generating questions matched answering ready-made questions, and took about twice as long (Weinstein, McDermott & Roediger, 2010). A randomized university study found generating and testing both beat restudying, with no advantage for generating (Ebersbach, Feierabend & Barzagar Nazari, 2020). The benefit of self-generated questions stayed at the cognitive level of the questions written; detail-level questions produced no conceptual gain (Bugg & McDaniel, 2012).

The strongest observational finding for authoring came from PeerWise, and its own developers tested it across three courses: authoring plus GPA explained 31.0 percent of exam variance and GPA alone explained 30.6 percent (Denny et al., 2011). Selection explains most of it. Stronger students choose to author. That is also the likeliest explanation of the classroom observation that started my SAMs: in a 300-level microbiology course at George Mason, the student who wrote and shared questions earned the top grade. It remains a good reason to have tried the practice. It is not evidence for the mechanism.

An imagined audience adds nothing on paper

My original prompt asked for two sentences aimed at someone who had the question wrong. Lachner, Jacob and Hoogerheide (2021) tested writing for a fictitious confused student against plain self-explanation. In a real course, within-subject and counterbalanced with 50 students, self-explanation won on transfer (d = 0.59). Their meta-analysis found explaining to others works when spoken (g = 0.34) and not when written (g = −0.07). The audience belongs in the conversation, which is where SAMS2 puts it. The written line becomes the supported version: why the right answer is right, and why the tempting wrong one fails.

“Why” questions have limits with complex content

Elaborative interrogation earned a moderate rating, mostly from studies of discrete facts tested within minutes (Dunlosky et al., 2013). The representative college biology study, 294 introductory students reading about digestion with regular “why” prompts, found 76 versus 69 percent against rereading (Smith, Holliday & Austin, 2010). The benefit also runs larger for students who already know more. That prior-knowledge gradient matches what I see when a student cannot write a SAM at all: the problem is a knowledge gap, not effort, and the fix is more input and a seeded question.

Question quality does not improve by practice alone

In a senior cell biology laboratory, the quality of students' questions barely improved across a semester of practice and did not track final grades; the authors concluded that improvement may require a specific instructional intervention (Keeling, Polacek & Ingram, 2009). Student-written questions can reach the standard of expert items, but only with item-writing training and heavy curation, about 120 of more than 1,000 candidates in one biochemistry course (Huang et al., 2021). SAMS2 answers both with stems, seeded items, a required item at Checkpoint level, and my curation of anything that reaches a game.

Level matters

In middle school science, application-level quizzes improved both definition and application exam items, while definition-level quizzes did nothing for application items (McDaniel et al., 2013). Practice transfers downward, not upward. Flashcards on functional groups are a useful floor. One application-level SAM per chapter does more than five recall items, so SAMS2 requires one at the level of that chapter's Checkpoint.

Why I stopped calling it a concept inventory

In biology education research a concept inventory is a validated instrument: distractors drawn from documented student reasoning, validity evidence from interviews and expert panels, psychometric characterization, and cohort-level use. A student's question collection is none of those, so I now call it a personal item bank. The real resemblance is worth keeping: when students build a question out of the Checkpoint options they almost chose, they are doing what inventory developers do when they surface what students actually believe (Garvin-Doxas & Klymkowsky, 2008).

Getting students to talk

If spoken explanation carries the effect, the classroom has to produce a great deal of it, and anyone who has run discussion sections knows that is the hard part. The biology education literature is specific about why.

For the conversation that is scored, the oral-assessment reviews are consistent. No study they found reported a reliability coefficient; fixed time, standard questions, a rubric and assessor practice helped; anxiety was the largest variable and fell with practice and clear expectations (Nallaya et al., 2024; Stephenson, Johnson-Glauch & Cruchley, 2025). A fixed structure in which the student chooses one item and the assessor chooses another led students to prepare for understanding rather than recall (Iannone & Simpson, 2012).

Games

A review of 93 studies found Kahoot! generally improved learning performance, attitudes and classroom dynamics, with time pressure among the reported problems (Wang & Tahir, 2020). The design problem for SAMS2 is the scoring: speed points reward the fastest recall and leave no room for the talk the evidence favors. The Wager Round keeps what students like about game play and changes the incentive: teams bet on their confidence, which is a spoken calibration judgment, and a named team member has to explain before the reveal. Kahoot stays in the room as an occasional treat.

What SAMS2 changes, and the evidence for each

ChangeEvidence
A seed SAM for every learning target and one stem list used all yearRosenshine et al., 1996; King, 1990, 1994
Class SAMs return as a game one to two weeks later; students retake their own items cold before each CheckpointYang et al., 2021; Sana & Yan, 2022
At least one SAM per chapter at Checkpoint levelBugg & McDaniel, 2012; McDaniel et al., 2013
The written line asks why the tempting wrong answer failsBisra et al., 2018; Rittle-Johnson et al., 2017
Every SAM is spoken to a partner and a team before the oralLachner et al., 2021; Smith et al., 2009
I curate which student items reach a gameHuang et al., 2021
No reflection worksheet; the conversation carries the metacognitionDye & Stanton, 2017; Knight et al., 2022; Angell et al., 2024
The oral has a fixed opening, a four-line rubric, and rehearsalsNallaya et al., 2024; Stephenson et al., 2025; Iannone & Simpson, 2012
Students predict their Checkpoint score before each CheckpointOsterhage et al., 2019
Question prompts, warned reporters, assigned wrong answers, and a rotating Questioner cardKnight et al., 2013, 2015; Cooper et al., 2018; Tanner, 2013

What nobody has tested

Several pieces of SAMS2 have no direct research behind them. Nobody has studied scoring why a student chose their questions, as opposed to the questions themselves; student-generated questions combined with an oral assessment in biology; a collection revised across a full year rather than within a semester; or AI assistance with a required spoken account of its use. And the setting itself is nearly absent from the literature. The federal evidence review of dual enrollment addresses access, credit accumulation and degree attainment, and its one estimate of college academic achievement was null (What Works Clearinghouse, 2017). I found no study of teaching in a college biology course taught in a high school, by a high school teacher, to high school students. That is why I describe SAMS2 to my students as an experiment, and why I tell them what I will watch and that we will change what does not work.

References