Features

Every feature starts with a finding, a quote, and a link you can check.

The point of this page is not to sound mysterious. It is to show the chain from research to product: what the study says, what Vrenberg does with it, and where the source lives.

Retrieval practice

Mastery modeling

Vrenberg models mastery at the individual rule level — the actual unit the MBE examines — rather than collapsing performance into subject-level averages that obscure precisely the weaknesses that matter.

The foundational insight behind Vrenberg's architecture comes from Roediger and Karpicke's landmark 2006 study in Psychological Science. Their experimental design was precise: participants studied prose passages and were assigned to either repeated study conditions or repeated testing conditions, then assessed at intervals of five minutes, two days, and one week. At five minutes, the study group performed marginally better. At one week, the testing group retained approximately fifty percent more material. The inversion was unambiguous — the act of retrieval restructured the memory trace in ways that passive re-exposure categorically could not. Vrenberg treats that finding not as a pedagogical suggestion but as a structural constraint on every interaction the platform permits.

Two years later, Karpicke and Roediger escalated the evidence in a Science publication that crystallized the distinction between repeated retrieval and repeated study. Students who practiced retrieval — recalling material from memory without notes — demonstrated dramatically superior long-term retention compared to students who restudied the same material an equivalent number of times. The paper's central finding was striking in its directness: repeated studying produced essentially no measurable benefit for long-term retention, whereas repeated retrieval markedly enhanced it. This is the empirical basis for Vrenberg's decision to make every engagement with the platform an act of recall, never an act of reading.

Robert Bjork's theoretical framework provides the mechanistic explanation. His 1994 theory of disuse introduces two independent memory parameters: storage strength, which accumulates monotonically and never decreases, and retrieval strength, which fluctuates with recency and context. The critical insight is that the most effective learning occurs when retrieval strength is low — when recall is effortful and uncertain — because precisely that difficulty produces the largest increments to both storage and retrieval strength. Bjork termed these conditions "desirable difficulties," and they explain why Vrenberg deliberately resurfaces rules when the candidate is beginning to forget them rather than while the material is still fresh.

Dunlosky, Rawson, Marsh, Nathan, and Willingham's 2013 monograph in Psychological Science in the Public Interest provided the most comprehensive comparative evaluation of learning techniques to date. After reviewing decades of experimental evidence across ten strategies, they assigned only two techniques to the highest utility category: practice testing and distributed practice. Highlighting, rereading, summarization, keyword mnemonics, and imagery use all received low or moderate ratings. The meta-analytic conclusion was definitive: the techniques that feel productive — rereading, highlighting, note-taking — produce the weakest retention. Vrenberg's entire methodology is built exclusively on the two techniques that survived Dunlosky's evaluative framework.

The granularity question is critical for bar examination preparation specifically. The MBE does not test seven subjects. It tests individual doctrinal rules — the distinction between an invitee and a licensee in premises liability, the precise elements of promissory estoppel, the specific conditions under which a dying declaration is admissible. A candidate who scores seventy percent in Evidence and sixty-two percent in Torts has learned nothing actionable from those numbers. Vrenberg tracks mastery across 1,000+ individual rules because that is the resolution at which the exam operates and therefore the resolution at which preparation must be calibrated.

Pyc and Rawson's 2009 research in the Journal of Memory and Language demonstrated that the retrieval effort hypothesis holds even when controlling for study time: items that require more effortful retrieval produce stronger subsequent retention, not because more time was spent but because the cognitive demand of retrieval itself strengthens the memory trace. Vrenberg applies this directly. When a rule is answered correctly with high confidence, its review interval extends. When retrieval fails or hesitates, the interval contracts. The system is not scheduling review arbitrarily — it is modulating difficulty to maintain the zone where retrieval effort is highest and learning is most efficient.

Elaborative interrogation

Curriculum tutor

The Vrenberg tutor constrains every response to the doctrinal universe — citing the governing rule, naming the authority, and refusing to generate outside the black-letter canon.

Benjamin Bloom's 1984 paper in Educational Researcher identified what he called the two-sigma problem: students who received individual tutoring performed two standard deviations above students in conventional classroom instruction. That effect size is extraordinary — it means the average tutored student outperformed ninety-eight percent of the conventionally taught group. The problem Bloom identified was not whether tutoring works but whether its effects could be achieved at scale. Vrenberg's curriculum tutor is an attempt to resolve that problem for a specific, well-bounded domain: the rules tested on the bar examination.

The mechanism through which tutoring achieves its effects is not passive information delivery. Chi, Bassok, Lewis, Reimann, and Glaser's 1989 study in Cognitive Science established the self-explanation effect: students who generated explanations for themselves while working through examples learned significantly more than students who did not. The self-explanation process forces the learner to identify gaps in understanding, connect new information to existing knowledge structures, and resolve inconsistencies — all metacognitive operations that passive listening does not reliably engage. The Vrenberg tutor is designed to provoke precisely this kind of elaborative processing by responding with questions, counterexamples, and doctrinal distinctions rather than summaries.

Hattie and Timperley's 2007 meta-analysis in the Review of Educational Research synthesized over five hundred studies on feedback and concluded that feedback is among the most powerful single influences on achievement — but only when it operates at the right level. They distinguished four levels: task, process, self-regulation, and self. Feedback at the task and process levels produced the largest effects; feedback at the self level ("you're smart") produced negligible or negative effects. Vrenberg's tutor operates exclusively at the task and process levels — naming the specific rule, identifying the specific analytical error, and specifying what the correct reasoning chain looks like.

The doctrinal grounding constraint is not a limitation but a design decision rooted in Dunlosky's finding that elaborative interrogation — the process of asking and answering "why" and "how" questions about factual material — receives a moderate utility rating precisely because its effectiveness depends on the accuracy and specificity of the explanations generated. An ungrounded tutor that generates plausible but doctrinally imprecise answers would actively harm learning by creating false retrieval targets. Vrenberg's tutor cites the Restatement section, the Model Penal Code provision, or the Federal Rule because the candidate's next retrieval attempt must recover the correct rule, not a confident approximation.

Butler and Roediger's 2008 research demonstrated that feedback after retrieval practice amplifies the testing effect itself. Their experiment showed that when students received corrective feedback after multiple-choice tests, the positive effects of testing were enhanced and the negative effects — the tendency for attractive distractors to create false memories — were eliminated. This finding has direct implications for the tutor's architecture: every wrong answer is an opportunity for corrective elaboration, and the tutor is optimized to deliver that correction immediately, specifically, and in the doctrinal language the candidate will need to reproduce on the examination.

Formative assessment

Essay calibration

Vrenberg returns rule-by-rule scoring, a model answer calibrated to minimum competence, and structural diagnostics — within seconds, not days.

Black and Wiliam's 1998 review article "Inside the Black Box" in Phi Delta Kappan synthesized over two hundred and fifty studies and concluded that formative assessment — assessment whose primary purpose is to inform the next instructional decision rather than to assign a summative grade — is among the most effective interventions in education. The effect sizes they reported ranged from 0.4 to 0.7 standard deviations, which places formative assessment among the largest-impact educational strategies ever documented. The critical condition is that the feedback must reach the learner while the performance is still cognitively accessible. A grade returned three weeks after submission is summative by default, regardless of the rubric that produced it.

Shute's 2008 comprehensive review in the Review of Educational Research refined the conditions under which formative feedback produces its effects. She identified four critical properties: feedback must be nonevaluative (focused on the work, not the person), supportive (constructive rather than punitive), timely (delivered close to the performance), and specific (identifying precisely what needs to change). Generic comments like "needs more analysis" satisfy none of these criteria. Vrenberg's essay grading system is architected around Shute's framework: each essay receives a rule-by-rule breakdown that names which doctrinal rules were present, which were absent, which were incorrectly stated, and how each rule should have been deployed in the specific fact pattern.

The temporal dimension of feedback is not a convenience factor — it is a cognitive one. Butler and Roediger's 2008 research demonstrated that feedback delivered immediately after a retrieval attempt produces different learning outcomes than delayed feedback, particularly when the initial response was incorrect. Immediate corrective feedback prevents the incorrect response from consolidating into long-term memory, whereas delayed feedback must compete with an already-encoded error. For essay writing specifically, this means that a candidate who receives feedback within minutes of completing a Contracts essay can rewrite the analysis while the fact pattern, their reasoning process, and their structural decisions are all still active in working memory. A candidate who waits weeks cannot.

Sadler's 1989 theory of formative assessment introduced the concept of the "reference level" — the standard against which the learner's work is compared. Without a concrete reference level, feedback becomes uninterpretable: the learner knows they fell short but not of what. Vrenberg addresses this directly by returning a model answer with every graded essay. The model answer is not aspirational — it is calibrated to the minimum-passing threshold, showing the candidate precisely what a competent response looks like for that specific prompt. The gap between the candidate's submission and the model answer constitutes the actionable feedback, and the rule-by-rule overlay makes the gap navigable.

Hattie and Timperley's feedback typology further supports the design. Their distinction between task-level and process-level feedback maps directly to essay scoring: task-level feedback tells the candidate which rules were missed (the "what"), while process-level feedback identifies structural deficiencies in the IRAC framework, misapplication of rules to facts, or failure to address counterarguments (the "how"). Vrenberg provides both simultaneously — the rule checklist addresses the task level, and the structural diagnostics address the process level — because Hattie's evidence shows that the combination produces effects neither level achieves independently.

Metacognitive scaffolding

Draft coach

The Vrenberg draft coach monitors each paragraph in real time — flagging missing sub-issues, analytical gaps, and structural omissions while the draft remains malleable.

Flower and Hayes's 1981 cognitive process model of writing, published in College Composition and Communication, reframed writing from a linear stage model (prewrite, write, revise) to a recursive, cognitively demanding process in which planning, translating, and reviewing operate simultaneously and compete for the same limited working memory resources. Their protocol analyses revealed that expert writers constantly monitor their own output during composition, detecting structural problems in real time rather than deferring all evaluation to a post-draft revision phase. The Vrenberg writing coach externalizes that monitoring function for candidates who have not yet developed the metacognitive infrastructure to perform it independently.

Chi's self-explanation research provides the theoretical mechanism. When learners are prompted to explain their reasoning during the learning process itself — not after, not before, but during — they generate deeper understanding and more durable knowledge representations. The writing coach applies this principle to legal writing by requiring the candidate to confront, in real time, whether the paragraph they just wrote actually accomplishes what it needed to accomplish: Did you state the rule? Did you apply it to the facts? Did you address the counterargument? These are self-explanation prompts embedded in the drafting workflow, not retrospective comments on a finished product.

Graham and Perin's 2007 meta-analysis Writing Next, commissioned by the Carnegie Corporation, evaluated the effect sizes of eleven instructional strategies for improving adolescent writing. Strategy instruction — explicitly teaching the cognitive and metacognitive processes involved in writing — produced the largest average effect size at 0.82 standard deviations. Process writing approaches and study of models also produced significant effects. Vrenberg's draft coach integrates all three: it provides strategic scaffolding (flagging missing IRAC elements), supports the writing process (real-time rather than post-hoc feedback), and implicitly models competent structure by highlighting departures from it.

The temporal specificity of the feedback is architecturally significant. Shute's review established that feedback is most effective when it arrives at the moment of maximum cognitive relevance — when the learner is actively engaged with the material and the error is still accessible in working memory. A writing coach that reads each paragraph as it is completed operates at precisely this temporal resolution. By contrast, feedback delivered after the entire essay is submitted requires the candidate to reconstruct their reasoning for each paragraph, reactivate the relevant rules, and re-engage with structural decisions they may no longer remember making. The real-time approach eliminates that reconstruction overhead.

For bar examination essays specifically, the stakes of structural deficiency are severe. MEE and MPT grading rubrics reward systematic rule identification and methodical fact application — not elegant prose, not creative argumentation, not exhaustive background knowledge. A candidate who identifies seven of eight testable rules but organizes them chaotically will score lower than a candidate who identifies six rules and deploys each one with structural precision. The writing coach targets this specific failure mode: it detects when a sub-issue has been omitted, when a rule statement lacks application to the operative facts, and when the analytical structure has deviated from the organizational framework examiners expect.

Transfer-appropriate processing

Exam simulation

Vrenberg delivers full-length UBE simulations — MBE, MEE, and MPT in bar-day sequence — scored at the rule level and projected against the jurisdictional passing standard.

Morris, Bransford, and Franks's 1977 paper on transfer-appropriate processing, published in the Journal of Verbal Learning and Verbal Behavior, established a principle that remains foundational in cognitive psychology: memory performance depends on the match between the cognitive operations engaged during encoding and those required at retrieval. Students who processed material in a manner congruent with the eventual test format performed significantly better than students who processed the same material in a different format, even when the latter group engaged in ostensibly deeper processing. For bar examination preparation, this means that the strongest predictor of test-day performance is not how much material was reviewed but how closely the practice conditions replicate the retrieval demands of the actual examination.

Bjork's 1994 framework of desirable difficulties extends the transfer-appropriate processing principle to training design. He demonstrated that training conditions that maximize immediate performance — blocked practice, massed repetition, constant conditions — often impair long-term retention and transfer. Conversely, conditions that impair immediate performance — interleaved practice, spaced repetition, variable conditions — produce superior long-term outcomes. A full-length mock examination is the canonical desirable difficulty: it introduces fatigue, time pressure, topic switching, and sustained concentration demands that isolated drills cannot replicate, and it forces the candidate to deploy retrieval strategies under conditions that degrade immediate fluency but build transferable competence.

Ericsson, Krampe, and Tesch-Römer's 1993 theory of deliberate practice, published in Psychological Review, identifies the specific characteristics that distinguish effective practice from mere repetition. Deliberate practice requires well-defined objectives, immediate feedback, opportunities for repetition, and cognitive demands that push the practitioner slightly beyond their current capability. A pile of disconnected drills satisfies some of these criteria; a full-length simulation satisfies all of them simultaneously. The objective is a passing score. The feedback is the scaled result and rule-level breakdown. The repetition is available through multiple distinct examinations. And the cognitive demand is precisely the demand the candidate will face on bar day.

The scoring methodology is not cosmetic. Vrenberg's mock examinations produce a scaled MBE score using the same statistical methodology that NCBE applies to the operational examination, combined with a weighted writing score derived from rule-by-rule essay evaluation. These components are projected against the candidate's target jurisdiction's passing standard — whether that is the UBE's 266 uniform passing score, California's 1390 on the 2000-point scale, or any other jurisdictional threshold. The projection includes a confidence band that reflects the statistical uncertainty inherent in any single-administration estimate. This is not a motivational score. It is a psychometric estimate designed to be calibrated, interpretable, and actionable.

The simulation fidelity extends to temporal structure. A bar examination is not simply two hundred MBE questions and six essays. It is two hundred MBE questions divided into two three-hour sessions, six MEE essays written in three hours, and two MPT tasks completed in three hours, administered over two consecutive days. The cognitive experience of question one hundred seventy-eight is categorically different from the experience of question twelve — fatigue, decision fatigue, attentional depletion, and motivational drift all compound. Vrenberg replicates this temporal architecture because the transfer-appropriate processing literature is explicit: practice that omits the contextual demands of the target performance will produce inflated competence estimates that fail to transfer.

Join the waitlist

Founding 100 - $199 locked