DP-040

Academic Integrity Stress-Test

Academic integrity review often occurs after an assessment is released or focuses narrowly on detecting prohibited output. This encourages reactive policing and controls that may weaken learning, accessibility or trust without addressing why the task is vulnerable. Conduct a comprehensive pre release stress test of the assessment design, instructions, criteria, permitted assistance, policy alignment and evidence of learning. Use bounded GenAI trials with synthetic or public inputs to examine what current tools can produce, alongside plausible collaboration, outsourcing, ambiguity and false positive scenarios. Identify whether a weakness lies in the task, evidence model, communication, support or control, then make proportionate revisions. Do not use detector scores as proof, generate allegations about students or turn the review into a reusable evasion guide. Educators and authorised specialists retain responsibility for approval.

Workbook-derived reconstruction

When this helps

Workbook-grounded and author-clarified; proposed reconstruction. Before an assessment is released, educators need to review how its design, evidence, instructions, permitted assistance, policy, marking and support operate together. Current GenAI, collaboration and online resources create both legitimate uses and vulnerabilities, while proposed controls may create false-positive, access and privacy harms.

Proposed and author-clarified. Integrity review often occurs after release or focuses narrowly on detecting prohibited output. This encourages reactive policing and may leave weak learning evidence, ambiguous rules, policy conflicts, unrealistic marking, inaccessible controls and false-positive risks unexamined. Tool capability can be mistaken for evidence about a student.

Pattern response

Proposed and author-clarified. Conduct a comprehensive pre-release stress test of the assessment’s outcomes, evidence, instructions, criteria, permitted assistance, policy alignment, marking, moderation, accessibility and feasibility. Use bounded GenAI trials with synthetic or public inputs and compare plausible misuse with equally serious ambiguity and false-positive scenarios. Trace each finding to a proportionate revision of task, evidence, communication, support or judgement. Re-test and require accountable approval. Keep individual misconduct detection and adjudication outside the pattern.

  • A documented threat and false-positive model for the assessment.
  • Identified vulnerabilities in task design, evidence and communication.
  • Proportionate revisions that strengthen learning evidence and clarify permitted assistance.
  • An integrity review record that preserves due process and assessor authority.
  • Bounded GenAI trial records that document assessment capability risks without becoming evidence about students.
  • A teacher- and governance-approved release version with consistent student and marker guidance.

Workflow at a glance

DP-040 | Integrity Testing Portrait workflow diagram for Academic Integrity Stress-Test DP-040 | Integrity Testing
Canonical workflow · DP-040Academic Integrity Stress-Test
  1. 01

    Establish the assessment contract

    Educators define learning and evidence; authorised policy, integrity, accessibility and moderation contributors interpret requirements within their expertise.

    MOD-001 · MOD-020
  2. 02

    Run bounded design stress tests

    Educators interpret whether trial output threatens the evidence model; integrity and accessibility specialists test harm scenarios; no generated suspicion is attached to a student.

    MOD-021 · MOD-022
  3. 03

    Develop proportionate revisions

    Educators choose responses that preserve validity and learning; relevant specialists review policy, access, privacy, marking and moderation consequences.

    MOD-009 · MOD-019
  4. 04

    Re-test and approve release

    Accountable educators approve the assessment; integrity, policy, accessibility and moderation roles confirm relevant requirements without transferring academic judgement to the system.

    MOD-042 · MOD-044

Human checkpoints

Educators and assessment designers own validity, evidence, learning and release approval. Integrity and policy specialists advise on rules and process; accessibility, privacy, moderation and student perspectives challenge harms and ambiguity. GenAI systems produce bounded trial outputs and scenarios. They cannot determine misconduct, identify an offender, establish proof or recommend sanction.

Risks and misuse

Reconstructed risks

  • The exercise generates actionable cheating strategies without adequate containment.
  • Unreliable detection claims are embedded in policy or marking decisions.
  • Revisions increase surveillance, workload or inequity more than they improve validity.
  • Legitimate accessibility or language support is mistaken for misconduct.
  • A technical adversarial mindset displaces educational purpose and student trust.
  • Trial outputs or prompts are circulated in ways that create a reusable evasion guide.
  • A pre-release design finding is later misrepresented as evidence in an individual misconduct case.

Use this pattern

Begin with the recurring problem and the authorised inputs. Follow the workflow in order, keep intermediate outputs inspectable and retain the named human decisions.

  • Minimum viable implementation: review the complete draft against a short integrity and fairness checklist, run several bounded GenAI trials with public or synthetic inputs and revise the task before teacher approval.
  • Robust implementation: add a versioned assessment contract, recorded trial conditions, vulnerability and false-positive register, policy and accessibility review, traceable revisions, re-testing and formal release approval.
  • Low-code implementation: connect a controlled assessment store, approved GenAI trial workspace and structured review register while preventing trial findings from entering student case systems.
  • Local or privacy-preserving implementation: keep unreleased assessments and sensitive trial details in controlled storage, use no student data and process locally where exposure or policy risk warrants it.
  • Speculative implementation: maintain a reviewed suite of synthetic stress tests for recurring assessment types and flag designs for renewed review when tools or rules change, without automating release decisions.

Reusable modules

View in Atlas →

Technical and provenance detail

Data and information flow

Proposed and author-clarified. The complete pre-release assessment package and current rules enter a bounded review. Synthetic or public trial inputs remain separate from student records, and tool conditions travel with every result. Each finding links the affected learning evidence, vulnerability or false-positive risk to a proposed revision, trade-off, residual risk and accountable decision. The approved release version remains distinguishable from trial drafts. No generated suspicion, detector score or design finding becomes an allegation about a student.

Source provenance

  • Pattern ID: DP-040
  • Original workbook row(s): 70.
  • Original label(s): Academic Integrity Stress-Test.
  • Inherited source category/theme: GRADING AND FEEDBACK.
  • Source status: workbook-derived source; the source supplies concise labels rather than a complete pattern narrative.
  • Extraction notes: Original labels, row relationships, categories, themes, rationales, and confidence assessments are preserved below.
  • Editorial interventions: The problem statement, response, forces, risks, workflow, and module mapping are documented reconstructions that retain explicit provenance labels and remain open to author and practice-based validation.
  • Author clarification (2026-08-20): The pattern is a comprehensive pre-release assessment-design review. It checks academic integrity, curriculum alignment, institutional rules, permitted assistance, accessibility, feasibility, evidence quality and false-positive risk before students receive the task. Bounded trials may give the assessment to current GenAI tools using synthetic or public inputs to identify design vulnerabilities. Findings produce revisions rather than allegations; the pattern does not adjudicate individual cases and detector scores are not treated as proof.

Open questions

  • How can misuse scenarios be documented without becoming a reusable evasion guide?
  • Which fairness checks should be mandatory before adopting an integrity control?
  • What evidence demonstrates that a redesign improves assessment validity rather than only apparent security?
  • How frequently should tool trials be refreshed without making assessment design a continuous technical contest?