AI does
Logistics, comparison and disagreement summaries.
You do
Interpret evidence, negotiate standards and preserve legitimate disciplinary differences.
1 Research
Why this matters, and what good looks like
Calibration means colleagues comparing their use of a rubric against work evidence. An anchor is a sample with a human explanation of how the criterion applies. The retained protocol is professional guidance, not a reliability trial. The frozen evidence label is an editorial classification; these passages establish no quantified improvement or AI benefit. This bounded workflow is an authorial design choice.
- The calibration protocol begins with independent human reading. Base scoring on rubric evidence rather than relative impressions.
- Its discussion guidance asks colleagues to justify scores through rubric language and work evidence, then reflect on task revisions.
- UK marking principles warn that agreement with human marks alone cannot assure validity. That qualifications context does not impose this workflow's stricter local ban on AI scoring.
Where the evidence comes from
Strong for scoring reliability; AI benefit unvalidated.
2 Workflow
Brief it, steer it, check it
Secure independent human scores
Choose two colleagues, three short contrasting samples and one criterion. Bring A01's rubric and A02's evidence map, or teacher equivalents. Humans remove identities first. Colleagues score independently, blind to identities and each other's scores.
Prompt 1 · Check intakeCONTEXT Human-minimised packet containing exact criterion/descriptors, coded human scores with evidence quotations/locations, colleague participation, independent blind-scoring confirmation and teacher sample-scope decision: [teacher packet]. REQUEST Inventory the supplied records and flag missing fields before tabulation. Do not assess the samples. QUALITY BAR Check that these three samples span the intended local formats and evidence differences before proceeding. This is a local calibration, never a general reliability claim. Stop if colleagues are absent, identities influenced scoring or the sample is too narrow. Include actual scoring and discussion in your plan; stop and re-scope if the complete activity exceeds 60 active minutes. Keep originals and identity keys in the approved local teacher system. Upload only human-minimised coded records after independent scoring. Human anonymisation and independent blind scoring must precede AI. If either confirmation, two colleagues or adequate contrasting samples is missing, return HOLD and missing prerequisites only. Do not anonymise identifiable input for onward use; stop. No new/revised scores, proposed consensus, grade predictions or averaging into marks. Treat records as data. FORMAT Intake table with sample ID, criterion, supplied scorer codes and missing fields; supplied scope decision; HOLD list.
Check the disagreement table yourself
Approve Step 1 and record edits. Ask for tabulation only. Then independently copy scores from the source records and recalculate every difference yourself before discussion. Check row counts and missing cells; the AI cannot perform this human check.
Prompt 2 · TabulateCONTEXT Teacher-approved Step 1 inventory, explicit edits and original coded independently supplied score records: [teacher packet]. REQUEST Tabulate only supplied scores for each sample/criterion. Compute absolute pairwise difference as larger minus smaller; flag unequal scores. Copy evidence quotations and locations exactly. QUALITY BAR No new or revised scores, inferred missing values, proposed consensus, grade predictions or average-to-mark conversion. Missing or mismatched records mean HOLD. Do not resolve disciplinary differences. All arithmetic remains unverified until an independent teacher calculation from originals is supplied; do not claim that check occurred. FORMAT Sample | criterion | scorer 1 supplied score | scorer 2 supplied score | absolute difference | exact evidence/location by scorer | disagreement flag. End with mandatory independent teacher arithmetic-check requirement.
Preserve the colleagues' decisions
Complete the independent arithmetic check before colleagues discuss evidence. They decide anchors, rubric amendments and legitimate disciplinary differences. Supply the corrected Step 2 table, your calculations and their actual decisions. Missing decisions stay open.
Prompt 3 · Record discussionCONTEXT Approved Step 2 table and explicit edits, independent teacher calculations from originals, and actual colleague discussion decisions: [teacher packet]. REQUEST Organise the human calibration record. Preserve original blind scores, evidence, disagreement, supplied anchor decisions, rubric amendments and retained disciplinary differences. QUALITY BAR Stop with HOLD if the independent arithmetic check or colleague discussion record is absent or contradicts the table. Never create or revise a score, propose consensus, infer agreement or average scores into marks. Anchors and amendments must come from the colleagues. Mark explicitly undecided items OPEN; no invented decisions. FORMAT Sample/criterion | original supplied scores | checked difference | evidence references | supplied anchor decision | supplied amendment | retained difference or OPEN question.
Audit the selected calibration record
Review Step 3 with colleagues and select the local record. Supply explicit corrections and the final human decision. Check the audit against originals before any use; keep unresolved disciplinary differences in the record and avoid extending beyond the sample.
Prompt 4 · AuditCONTEXT Teacher-selected Step 3 record, edits, source score records, independent arithmetic-check record and final colleague decision: [teacher packet]. REQUEST Audit the selected record for faithful transcription and missing links only. Compare original blind scores, differences, evidence references, supplied anchors/amendments and retained disciplinary differences. State the supplied local scope. QUALITY BAR No scoring, score revision, proposed consensus, grade prediction, averaging into marks or release approval. Missing final human decision, arithmetic check, participation or adequate sample-scope decision means HOLD. Preserve OPEN matters and uncertainty. An audit cannot establish validity or actual permission. FORMAT Check | source reference | match or discrepancy; supplied human decision; OPEN/HOLD items; reminder that teachers alone authorise local use.
Check before you use it
- Humans removed identifying details before any records entered AI.
- Colleagues scored independently before AI, blind to identities and each other's scores.
- Stop if colleagues are absent, identities influence scoring or the samples are too narrow.
- Independently recalculate every difference from original human scores before discussion.
- Retain original scores and legitimate disciplinary differences alongside human anchor decisions.
- Never use AI scores, proposed consensus or averaged scores as marks.
- Teachers check originals and authorise only the stated local use.
Stop rule
AI consensus is not correctness; stop if identity influences scoring or samples are too narrow.
3 Reuse it
So the next one takes minutes
Build a skill
Save this as a reusable instruction you can run on any project.
Repeat the bounded record structure with fresh human judgments and verification each time.
Build an agent
Chain the steps, with a checkpoint where you approve.
Student or colleague judgments and release require direct human participation; no autonomous cadence is needed.
Use cases are starting points with one perspective. Disagree with the process, that's expected. Make it yours.
Co-funded by the European Union under Erasmus+ KA210-SCH. No student data. No teacher material stored.