AI does
Evidence-map drafting and coverage checking.
You do
Decide what counts as evidence, how it will be collected and how consequential judgments will be made.
1 Research
Why this matters, and what good looks like
An evidence plan connects each intended outcome with work that can show learning. Valid evidence supports the intended judgment: a polished group product alone cannot establish each student's understanding. Formative assessment informs next teaching; summative assessment supports an end-point judgment. These sources combine professional guidance, research-informed feedback guidance and qualifications policy discussion. They do not test this AI workflow. The rule that AI never decides grades is this method's safety boundary, not a claimed statutory classroom ban.
- PBLWorks teaching practices include formative and summative assessment, with self/peer assessment of team and individual work. Plan both levels explicitly.
- PBLWorks evaluation guidance warns that end-product focus can hide learning during the process. Include evidence of development.
- EEF feedback guidance describes benefits from well-used feedback and harm or wasted time from poor feedback. Plan how students will act on it.
- UK marking principles say agreement with human marks alone cannot assure validity. That qualifications context does not directly regulate ordinary classroom drafting.
Where the evidence comes from
- Gold Standard: Teaching Practices | PBLWorks
- PBLWorks Evaluation Within Project Based Learning
- EEF Teacher Feedback to Improve Pupil Learning
- UK principles of AI use in marking
Strong for formative and summative assessment design; AI benefit unvalidated.
2 Workflow
Brief it, steer it, check it
Fix outcomes and evidence boundaries
Use optional D01 outcomes, D04 design and D05 milestones, or equivalent approved teacher documents. Limit this pass to three outcomes and three checkpoints. Keep learner records private. If active planning exceeds 60 minutes, stop and narrow scope.
Prompt 1 · InventoryCONTEXT Approved outcomes and what would count as evidence for each: [outcomes]. Approved project design and milestone windows, with source versions: [design and milestones]. Teacher decisions on evidence collection, weighting or non-graded use, accessible response routes and accommodations expressed without individual profiles: [assessment decisions]. Existing assessment plan to improve, or none: [existing plan]. REQUEST Inventory outcome coverage, possible individual and process evidence, checkpoint opportunities and unresolved teacher decisions. Identify criteria that cannot be observed or repeat the same learning judgment. QUALITY BAR Use supplied materials only. HOLD with questions only if approved outcomes or acceptable-evidence decisions are absent. Do not infer learner needs, weights, permissions or accommodations. Keep private rosters, support and safeguarding facts outside AI, including pseudonymous profiles. Do not grade or infer achievement. Flag polish, compliance, language fluency or AI output being used as substitutes for learning. FORMAT Return a table: outcome and source, acceptable evidence, individual/process coverage, checkpoint window, missing or duplicate criterion, teacher decision. Separate supplied decisions from proposals and HOLD items.
Draft the evidence and feedback map
Approve or correct the inventory outside AI. Decide what evidence can support each judgment and what students can submit accessibly. Keep weighting and accommodations under your control. AI drafts a map from those decisions.
Prompt 2 · MapCONTEXT Exact step 1 inventory checked and approved by the teacher: [approved inventory]. Explicit teacher corrections, selected outcomes, evidence decisions, approved milestone windows, weighting and non-identifying access decisions: [mapping decisions]. REQUEST Draft a candidate assessment-and-evidence map. For each outcome include individual evidence, process evidence, purpose of the check, collection method and responsible role, feedback focus, student action and a recheck inside supplied windows. Include self/peer assessment where approved and distinguish it from teacher judgment. QUALITY BAR HOLD missing approval, evidence validity decisions or checkpoint capacity. Label the map PROPOSAL. Do not set weights, accommodations, grades or attainment judgments. Group work alone cannot establish individual understanding. A revision should reveal a learning decision, not compliance with feedback. Exclude presentation polish and fluency unless an explicitly supplied outcome validly requires them; never use them as proxies for subject learning. No student work or identities are needed for this planning task. FORMAT Return rows with outcome/source, individual evidence, process evidence, checkpoint/purpose, collector/method, feedback, student response/recheck, and teacher-approved weighting/access decision. Add gaps, duplication risks and collection workload for teacher review.
Challenge the teacher-selected plan
Choose and edit the map. Check collection workload with relevant colleagues and how students will respond to feedback. Supply a selected version for challenge, keeping all learner-specific decisions outside AI.
Prompt 3 · ChallengeCONTEXT Exact step 2 map selected and edited by the teacher: [selected map]. Teacher decision addendum with approved source versions, edits, collection capacity, weighting, accessible routes and unresolved decisions: [challenge decisions]. REQUEST Check whether each outcome has evidence that can support its intended judgment, including individual understanding and learning during the process. Locate duplicate or unobservable criteria, feedback with no response opportunity and collection overload. Identify any reliance on polish, compliance, fluency or AI output. QUALITY BAR HOLD if the teacher selection or evidence boundaries are missing. Cite exact rows and supplied sources; do not evaluate students. Propose questions or repairs for teacher review, without changing weights, accommodations, outcomes or checkpoint dates. Self/peer comments do not prove individual mastery. Treat missing provenance or stale decisions as HOLD. No AI grades or automatic approvals. FORMAT Return issue table: exact row, source evidence, risk to the intended judgment, proposed repair or question, teacher decision needed. Also list source-matched coverage and remaining human checks.
Audit the final assessment plan
Resolve issues yourself and record your decisions. Compare the revised plan with approved outcomes and milestones. Verify evidence validity and access with responsible people, then retain your final plan locally. AI's comparison is not approval.
Prompt 4 · Final auditCONTEXT Exact step 3 challenge output reviewed by the teacher: [reviewed challenge]. Teacher-selected final assessment-and-evidence plan: [final plan]. Explicit teacher edits, approved outcome/evidence and milestone source versions, weighting, access decisions and unresolved issues: [final decisions]. REQUEST Audit source alignment and resolution of challenge issues. Check individual and process evidence for each outcome, collection feasibility, feedback response and recheck opportunities, self/peer roles and unchanged teacher decisions. QUALITY BAR HOLD unsupported, missing or stale decisions. Reject any plan rewarding polish, compliance, language fluency or AI output instead of learning. AI never decides grades. Do not rewrite or approve the plan, infer learner characteristics, or claim that evidence validity has been independently verified. Distinguish supplied teacher decisions from proposals and teacher-reported verified facts. Human review remains necessary before use. FORMAT Return blockers with exact plan line, source evidence and required human action; source-matched coverage; and remaining human checks. If no blockers are found in the comparison, say so without declaring the plan ready for use.
Check before you use it
- Does each approved outcome have evidence that supports the intended learning judgment?
- Can you distinguish each student's understanding from group production and presentation polish?
- Does process evidence reveal learning decisions rather than compliance, fluency or AI output?
- Are collection, feedback response and recheck feasible within approved milestone windows?
- Have you retained decisions on weighting, accommodations and every consequential judgment outside AI?
- Have you removed private learner data and resolved every HOLD before human release?
Stop rule
Reject plans that reward polish, compliance, language fluency or AI output instead of learning; AI never decides grades.
3 Reuse it
So the next one takes minutes
Build a skill
Save this as a reusable instruction you can run on any project.
The evidence-mapping and source-checking sequence can repeat with fresh teacher decisions and bounded outcomes. Each plan still requires human review.
Build an agent
Chain the steps, with a checkpoint where you approve.
A bounded assessment-design pass needs direct teacher judgments. Ongoing automated evidence collection or grading is outside this method.
Use cases are starting points with one perspective. Disagree with the process, that's expected. Make it yours.
Co-funded by the European Union under Erasmus+ KA210-SCH. No student data. No teacher material stored.