DP-012

GenAI Scaffolding for Criterion-Referenced Assessment Design

Criterion referenced assessment documents must align outcomes, criteria, descriptors, evidence and institutional requirements while remaining comprehensible to students and markers. Treat rubric development as iterative educational design: ground the work in authorised curriculum and policy, map the elements, draft descriptors and exemplars, test consistency with worked cases, and require educator moderation and approval. GenAI can support comparison and drafting, but disciplinary judgement determines the standards. Evidence from use should inform later revision without converting student data into ungoverned monitoring.

Source manuscript

When this helps

Editorial synthesis grounded in the source pattern. Educators creating or revising criterion-referenced assessment materials under curriculum, policy, moderation and student-comprehension requirements.

Academics often struggle to create or revise criterion referenced assessment documents that are aligned, consistent and readable. Misaligned learning outcomes, inconsistent language across criteria and grade levels, unclear descriptors and the absence of exemplars frequently result in rubrics that are verbose, difficult to interpret and insufficiently transparent for students and tutors.

Pattern response

Editorial synthesis grounded in the source response. Treat criterion-referenced assessment design as iterative educational alignment. Ground the work in authorised curriculum and policy, map outcomes to criteria and evidence, draft readable descriptors and exemplars, and test consistency with worked cases. Require educator moderation and approval. Collect proportionate evidence from use to inform revision while keeping disciplinary standards and consequential decisions under educator authority.

  • Aligned outcomes, criteria, descriptors, evidence expectations and readable exemplars.
  • A consistency record based on worked examples and moderation rather than surface wording alone.
  • Educator approval and a governed route for evaluation and later revision.

Workflow at a glance

DP-012 | Assessment Scaffolding Portrait workflow diagram for GPT Based Scaffolding for Criterion Referenced Assessment Design DP-012 | Assessment Scaffolding
Canonical workflow · DP-012GenAI Scaffolding for Criterion-Referenced Assessment Design
  1. 01

    Ingest policy and learning outcomes

    An accountable academic initiates, interprets, or approves this stage; automation remains bounded by the listed modules.

    MOD-003 · MOD-020
  2. 02

    Map criteria and outcomes

    An accountable academic initiates, interprets, or approves this stage; automation remains bounded by the listed modules.

    MOD-019 · MOD-027
  3. 03

    Generate descriptors and exemplars

    An accountable academic initiates, interprets, or approves this stage; automation remains bounded by the listed modules.

    MOD-028 · MOD-008
  4. 04

    Check alignment and consistency

    An accountable academic initiates, interprets, or approves this stage; automation remains bounded by the listed modules.

    MOD-019 · MOD-042
  5. 05

    Review and publish the rubric

    An accountable academic initiates, interprets, or approves this stage; automation remains bounded by the listed modules.

    MOD-044 · MOD-047
  6. 06

    Collect use evidence

    An accountable academic initiates, interprets, or approves this stage; automation remains bounded by the listed modules.

    MOD-041 · MOD-043

Human checkpoints

People define purpose and boundaries, supply or authorise source material, inspect intermediate representations, resolve ambiguity, approve consequential outputs, correct errors, and remain accountable for scholarly, pedagogical, legal, or organisational decisions. A generated recommendation or draft is not an approval decision.

Risks and misuse

Source-stated improvement concerns

  • Maintain expert academic oversight to validate clarity, disciplinary accuracy and implicit assumptions in generated outputs.
  • Integrate systems with learning management platforms to automate formatting, uploading and alignment with learning outcomes.
  • Develop real time validation tools to ensure consistency across units and assessment cycles.
  • Enhance dialogic interaction with generative systems to refine intent, outcomes and assessment focus, improving precision and pedagogical alignment.

Inferred workflow risks

  • Inferred: duplicate ingestion.
  • Inferred: unsupported formats.
  • Inferred: lost provenance.
  • Inferred: partial imports.
  • Inferred: missing rule.
  • Inferred: wrong interpretation.
  • Inferred: stale constraint.
  • Inferred: false pass.

Use this pattern

Begin with the recurring problem and the authorised inputs. Follow the workflow in order, keep intermediate outputs inspectable and retain the named human decisions.

  • Minimum viable implementation: a documented human procedure using the ordered modules: Source and Record Ingestion → Constraint Checking → Criteria and Outcome Alignment → Outline and Scaffold Generation → Draft Generation → Plain-Language Explanation → Criteria and Outcome Alignment → Output Quality Evaluation → Human Review and Approval → Document Rendering → Feedback Capture → Outcome and Impact Monitoring.
  • Robust implementation: add explicit schemas, source identifiers, access controls, logging, exception queues, independent evaluation, backups, and named approval owners.
  • Low-code implementation: use forms and a workflow orchestrator to connect bounded services, with approval gates before external communication or state changes.
  • Local or privacy-preserving implementation: keep sensitive artefacts in controlled storage and prefer local extraction, transcription, search, or model execution where capability and governance permit.
  • Speculative implementation: more autonomous coordination may be explored only with constrained tools, stop conditions, audit logs, and human authority; it is not implied by the source pattern.

Reusable modules

View in Atlas →

Technical and provenance detail

Data and information flow

Inputs may include files, records, messages, source metadata, candidate output, constraint set, outcomes, criteria, descriptors, activities or evidence, brief, source material, required structure, sources, outline, style constraints, complex source, audience profile, required terms, output, source evidence, candidate output or action, evidence, review criteria, structured content, style template, output constraints, output or experience, respondent, feedback prompt, success indicators, usage or outcome data, baseline. The stage sequence transforms, analyses, enriches, retrieves, generates, coordinates, stores, or outputs information according to each linked module contract. Outputs may include ingested source set, ingestion log, pass/fail findings, exceptions, unresolved constraints, alignment map, gaps, inconsistencies, outline, scaffold, open questions, draft, source annotations, uncertainties, audience-appropriate explanation, defined terms, evaluation findings, score or decision, required revisions, approval decision, corrections, rationale, escalation, rendered document, accessibility and build report, feedback record, correction, rating, monitoring report, alerts, improvement decisions. Source identity, permission state, uncertainty, retention, and human decisions should travel with records rather than being discarded between stages.

Source provenance

  • Pattern ID: DP-012
  • Original title: Pattern 15: GPT Based Scaffolding for Criterion Referenced Assessment Design
  • Source location: SRC-001:P0281–P0305; page number unknown.
  • Source status: explicit.
  • Extraction notes: Canonical ID follows manuscript order. Legacy numbers are not used as publication identifiers; any numbered source heading is retained only for provenance and source fidelity.
  • Editorial interventions: The public title, summary, context, response and intended outcomes are concise editorial syntheses grounded in the source pattern. Original source prose is preserved below; workflows and modules remain separated and provenance-labelled.

Open questions

  • Which elements of this pattern require empirical or practice-based evaluation in the intended setting?
  • Which implementation examples remain current, authorised and proportionate to the setting at the time of use?
  • How should the extended evaluation stage be calibrated so that it improves practice without creating disproportionate monitoring or workload?