Research Foundation

Three Meta-Analyses

RBI is not derived from best-guess principles. It emerges from systematic synthesis of the empirical literature on rubric use.

Meta-Analysis 1

Rubrics, Self-Efficacy & Self-Regulation

Camargo, Parra-Martínez, Chang, Maeda, & Traynor (2024) — Educational Psychology Review

1,777 records screened. After removal of two influential outliers, nine primary studies (n = 2,793 students) yielded 16 and 15 effect sizes respectively.

1,777
Records screened
2,793
Students
9 & 9
Primary studies
Self-Efficacy g = 0.39
Students showed increased task confidence
Self-Regulation g = 1.00
Strong effect on goal-setting, planning, monitoring

The self-regulation effect (g = 1.00) is large by any conventional benchmark. Studies with the highest effects combined rubric use with formative feedback and explicit classroom training.


Meta-Analysis 2

Rubrics & Academic Performance

Camargo & Parra-Martínez (under review)

39 primary studies — 80 effect sizes — N = 8,968 students. 83% of estimates positive. Range: −1.65 to 2.18.

39
Primary studies
8,968
Students (N)
80
Effect sizes
83%
Positive estimates

Mean Effect Size

ĝ = 0.47

SE = 0.09, 95% CI [0.28, 0.66]
t = 4.91, p < .001

Prediction interval: [−0.72, 1.66]

Academic Performance g = 0.47

Moderator Analysis

When do rubrics work best?

No single moderator reached statistical significance, meaning the effect of rubrics is remarkably consistent. However, directional patterns point to actionable design and implementation factors.

Implementation

Co-creation

Rubrics co-created with students showed the highest effects (g = 1.13 vs. 0.45), though not statistically significant given sample sizes.

1.13

Co-creation

0.45

Without

Implementation

Feedback integration

Studies combining rubric use with formative feedback outperformed rubric-only conditions (0.54 vs. 0.34).

0.54

With feedback

0.34

Without

Design

Rubric type

Analytic rubrics (0.44) outperformed holistic rubrics (0.36). RBI exclusively recommends analytic design.

0.44

Analytic

0.36

Holistic

Qualitative Findings

What the rubrics themselves revealed

A content analysis of 20 rubrics from the primary studies — categorized by effect size magnitude — identified specific linguistic and design features associated with large vs. negative effects.

g: −1.65 to −0.20

Negative effects

Vague adjectives, task-focused rather than learning-focused, complex tasks without scaffolding.

g: −0.21 to 0.20

Negligible effects

Non-parallel descriptions, inconsistent verbs, unknown content, ambiguous quantifiers.

g: 0.21 to 0.79

Small–medium effects

Action verbs, quantifiable metrics, clear adjectives, alignment with task and objectives.

g: 0.80 to 2.18

Large effects

Explicitly parallel descriptions, multi-source feedback, familiar tasks, socialized rubric language.

Presenting This Work

At NCME AIME — October 7, 2026

Rubricator, the AI-assisted rubric design tool built on this framework, is being presented as an Innovation Demonstration at the NCME AI in Measurement and Education (AIME) Conference in Pittsburgh, PA.

Innovation Demonstration

Rubricator: An AI Tool for Research-Aligned Rubric Design and Implementation