<- All features

Formative assessment

Essay calibration

Vrenberg returns rule-by-rule scoring, a model answer calibrated to minimum competence, and structural diagnostics — within seconds, not days.

Research to product

The case for formative assessment

Black and Wiliam's 1998 review article "Inside the Black Box" in Phi Delta Kappan synthesized over two hundred and fifty studies and concluded that formative assessment — assessment whose primary purpose is to inform the next instructional decision rather than to assign a summative grade — is among the most effective interventions in education. The effect sizes they reported ranged from 0.4 to 0.7 standard deviations, which places formative assessment among the largest-impact educational strategies ever documented. The critical condition is that the feedback must reach the learner while the performance is still cognitively accessible. A grade returned three weeks after submission is summative by default, regardless of the rubric that produced it.

Four properties of effective feedback

Shute's 2008 comprehensive review in the Review of Educational Research refined the conditions under which formative feedback produces its effects. She identified four critical properties: feedback must be nonevaluative (focused on the work, not the person), supportive (constructive rather than punitive), timely (delivered close to the performance), and specific (identifying precisely what needs to change). Generic comments like "needs more analysis" satisfy none of these criteria. Vrenberg's essay grading system is architected around Shute's framework: each essay receives a rule-by-rule breakdown that names which doctrinal rules were present, which were absent, which were incorrectly stated, and how each rule should have been deployed in the specific fact pattern.

Why speed matters cognitively

The temporal dimension of feedback is not a convenience factor — it is a cognitive one. Butler and Roediger's 2008 research demonstrated that feedback delivered immediately after a retrieval attempt produces different learning outcomes than delayed feedback, particularly when the initial response was incorrect. Immediate corrective feedback prevents the incorrect response from consolidating into long-term memory, whereas delayed feedback must compete with an already-encoded error. For essay writing specifically, this means that a candidate who receives feedback within minutes of completing a Contracts essay can rewrite the analysis while the fact pattern, their reasoning process, and their structural decisions are all still active in working memory. A candidate who waits weeks cannot.

The reference level

Sadler's 1989 theory of formative assessment introduced the concept of the "reference level" — the standard against which the learner's work is compared. Without a concrete reference level, feedback becomes uninterpretable: the learner knows they fell short but not of what. Vrenberg addresses this directly by returning a model answer with every graded essay. The model answer is not aspirational — it is calibrated to the minimum-passing threshold, showing the candidate precisely what a competent response looks like for that specific prompt. The gap between the candidate's submission and the model answer constitutes the actionable feedback, and the rule-by-rule overlay makes the gap navigable.

Task-level and process-level scoring

Hattie and Timperley's feedback typology further supports the design. Their distinction between task-level and process-level feedback maps directly to essay scoring: task-level feedback tells the candidate which rules were missed (the "what"), while process-level feedback identifies structural deficiencies in the IRAC framework, misapplication of rules to facts, or failure to address counterarguments (the "how"). Vrenberg provides both simultaneously — the rule checklist addresses the task level, and the structural diagnostics address the process level — because Hattie's evidence shows that the combination produces effects neither level achieves independently.

There is a body of firm evidence that formative assessment is an essential component of classroom work and that its development can raise standards of achievement.

Black & Wiliam, Phi Delta Kappan (1998)

Rule-by-rule scoring

Every doctrinal rule the prompt tests is scored individually — present, absent, misstated, or misapplied.

Minimum-competence model

A passing-threshold model answer accompanies every submission, calibrated to what examiners actually require.

Structural diagnostics

IRAC framework analysis identifies organizational deficiencies, not just substantive gaps.

Seconds, not weeks

Feedback arrives while the fact pattern and the candidate's reasoning are still cognitively active.

Ready when we open.

Join the waitlist

Founding 100 - $199 locked - every feature included