EdTechLab
Back to Lab Notes
Technical report · Lab Note 010 26 min read

Assessment After Generative AI: Authenticated Authorship, Process Evidence and On-Screen Exams

Generative AI can complete much unsupervised written work to a high standard, and no detector can prove that it did. This report sets out what UK regulators and sector bodies now expect, what the research says about detection, and how to engineer the evidence that still establishes authorship: supervised moments, conversations about the work, light process evidence and secure on-screen exams.

Who it is for

Assessment and academic integrity leads, learning technologists, exams officers and digital teams in universities, colleges and schools, and product teams building assessment tools.

Author
EdTechLab team
Scope
Schools and qualifications in England; UK higher education; secure on-screen delivery
Status as of
5 October 2026

Key findings

  1. Ofqual has not yet decided how to regulate on-screen exams. The exams regulator for England closed its consultation on 5 March 2026 and is still analysing responses. It proposed a default of no on-screen assessment, with up to two new on-screen specifications per exam board as exceptions, none in subjects with over 100,000 entries, no student-owned devices, and separate paper and screen specifications.[1][2]
  2. Coursework, not the exam hall, carries the risk. Ofqual says supervised exams remain largely protected, and the Curriculum and Assessment Review recommended no expansion of written coursework. Proven AI cases are still few: 100 in summer 2025, 2.0% of student malpractice cases.[8][11][17]
  3. None of the UK bodies we reviewed treats a detector score as proof, and the research agrees. Ofqual, the Joint Council for Qualifications (JCQ), Jisc and the Office of the Independent Adjudicator (OIA) treat detection as one source of evidence at most. In one test of 14 tools none reached 80% accuracy; in another, mean accuracy fell from 39.5% to 22.1% after simple machine edits.[10][14][23][25][29][31]
  4. Students already use AI in assessed work. In the Higher Education Policy Institute (HEPI) 2026 survey of 1,054 UK undergraduates, 94% used generative AI to help with assessed work and 12% had put AI-generated text straight into it.[27]
  5. Labels do not secure an assessment; structure does. Rules that students can simply ignore change little. The revised AI Assessment Scale ties no-AI tasks to controlled environments, and the University of Sydney's policy separates secure and open lanes.[28][34][35]
  6. Process evidence should be light. JCQ and the OIA point to drafts, checkpoints and short conversations. Typing patterns processed to identify someone are biometric data, and the ICO, the UK data protection regulator, requires a data protection impact assessment (DPIA) before any biometric recognition system is used.[14][25][39][46]
  7. For on-screen exams, AI must be switched off in three places. JCQ's instructions since 1 September 2026 say AI can come from the application, the operating system and the network, and that centres must test their approach before the exam.[15]

Dates to plan around

5 March 2026
Ofqual's on-screen consultation closed. It promised responses and next steps in 2026; none had appeared by 5 October.[1][2]
1 September 2026
JCQ's 2026 exam instructions took effect, with new rules for turning off AI on exam devices.[15]
Later in 2026
The Office for Students (OfS) and Advance HE expect to publish their research on AI in higher education.[20]
10 December 2026 (provisional)
Ofqual's malpractice statistics for summer 2026, including AI cases.[11]

What generative AI changed

Assessment assumes that work submitted under a student's name shows what that student can do. Generative AI removes that assumption for anything done without supervision. In a blind test at a UK university, answers written entirely by GPT-4 were entered into five undergraduate psychology exams taken at home: 94% went undetected, and on average they scored half a grade boundary higher than real students' work.[32] HEPI's 2026 survey found near-universal use for assessed work, with some students afraid of being falsely accused.[27]

The useful question is therefore less whether a student cheated than whether a result can be trusted. A 2024 paper argues that validity matters more than cheating, and that anti-cheating technology can make validity worse, for example by creating problems for inclusion.[36] Regulation in England already takes this view: OfS condition B4 requires each assessment to be valid and reliable, and its guidance says assessments that let students gain marks for work that is not their own would likely be of concern.[18]

Schools and qualifications: Ofqual and JCQ

Ofqual updated its approach to artificial intelligence on 16 July 2026. It applies its existing outcomes-based rules, has told exam boards that "AI may not be used as a sole marker", and judges that supervised exams remain largely protected while non-exam assessment (NEA) faces greater pressure.[8] In March 2026 it asked exam boards for immediate steps against AI misuse, including more rigorous candidate and teacher authentication of work.[9] Its April 2026 advice note asks boards to judge each assessment's exposure by task, output, length and supervision; to use detection tools as sources of evidence, not sole determinants; and to consider removing or replacing an assessment where no effective mitigation exists, while checking that mitigations do not damage what is assessed.[10]

Detected cases remain few. In summer 2025, AI-related plagiarism accounted for 100 proven cases: 75.0% of student plagiarism cases, but 2.0% of all 5,025 student malpractice cases.[11] Ofqual reports significant concern among teachers about the potential scale of misuse.[8]

For centres, the rules sit in JCQ's guidance AI Use in Assessments; revision two of 30 April 2025 was still current on 5 October 2026. Students must submit their own work, acknowledge any AI use with the tool's name and date, and keep a non-editable copy of prompts and outputs. Content reproduced from AI earns no credit where the student has not independently met the criteria, and AI must not be the sole marker. Suggested prevention includes completing portions of work under direct supervision, reviewing intermediate stages, short verbal discussions and restricting AI on centre devices. Concerns raised before the student signs the authentication declaration stay in the centre; after it, they go to the awarding body.[14]

The Curriculum and Assessment Review's final report (November 2025) agreed: exams should remain the principal form of assessment, written coursework should not expand, and NEA should be used only where it is the only valid method. It also asked the Department for Education (DfE) and Ofqual to seek to cut GCSE exam time by at least 10%.[17]

On-screen exams: what Ofqual proposed, and where it stands

Ofqual's rules do not currently say whether exams are taken on paper or on screen, and five GCSE and A level specifications already include on-screen components. Its consultation, open from 11 December 2025 to 5 March 2026, proposed a controlled approach:[1]

  • Default: no on-screen assessment, except where integral to the subject (computer science, British Sign Language, music technology), in existing specifications, and as a reasonable adjustment.
  • Limit: up to two new on-screen specifications per exam board, subject to accreditation: at most eight across the four boards.
  • High-entry subjects: none with over 100,000 national entries, which at consultation meant 13 GCSEs and A level maths.
  • Separation: paper and on-screen versions as separate specifications with substantially different questions; Ofqual's press release called them separate qualifications.[2]
  • Devices: no student-owned devices except as reasonable adjustments; centre-owned devices may also be used for teaching.
  • Platforms: expectations on usability, familiarisation, accessibility, security, infrastructure and support, with no single mandated platform; devices tested before assessments and protected by measures such as locked-down browsers.[1]

Four research reports underpin the proposals.[3] A PA Consulting study for Ofqual and the DfE found that wide-scale on-screen assessment would need considerable planning, gradual implementation and often extra investment.[4] A 2023 survey of 516 students and 500 parents found worries about technical failure, device fairness and security.[5] A review found screen reading often more demanding but no consistent pattern of mode effects,[6] and a companion framework locates item-level mode effects in four stages: information presentation, thinking required, response mechanism and appraisal strategy.[7]

Status on 5 October 2026: no decision. The consultation page says Ofqual is analysing feedback; Ofqual said it would publish responses and next steps in 2026, with detailed rules to follow.[1][2]

Higher education: OfS, QAA, Jisc and the OIA

OfS. Condition B4 defines "assessed effectively" to include assessment designed to minimise opportunities for misconduct and help detect it, and its guidance expects records of assessed work to be kept for five years after a course ends.[18] The OfS says its regulation requires no particular scale or form of AI use, while recognising risks to integrity and to the credibility of assessment.[19]

QAA. The Quality Assurance Agency's 2023 advice said AI outputs cannot reliably be detected, and suggested reducing assessment volume and using more synoptic and authentic tasks.[21] Its November 2025 toolkit, from universities led by Kingston University, recommends prioritising high-stakes, heavily weighted assessments; a no-AI category only where assessment is controlled or dialogic; and care over detection's false positives and ease of evasion.[22] The Russell Group's 2023 principles commit its members to adapt assessment for ethical AI use and to uphold academic integrity.[26]

Jisc. Jisc's National Centre for AI says the mood has "shifted decisively away from automated AI detection". Decisions should never rest solely on detection, by tool or person; staff must not run student work through detectors found online; and anyone using a detector should be able to show a DPIA for it. At scale, a 1% false positive rate across 480,000 assessments would mean about 4,800 false flags a year.[23]

OIA. The OIA, the higher education complaints scheme for England and Wales, says the provider must prove an AI misconduct case; students should see all evidence, including detection reports; detectors' limits must be weighed against other evidence; a viva should focus on the content of the work; and assumptions about AI use should be checked for bias against disabled students and those whose first language is not English.[25]

Abroad, the University of Sydney announced a "two-lane" approach from Semester 2 2025: secure, in-person assessment of learning outcomes, and open assessment that allows all relevant tools.[28]

What the evidence says about AI detection

A 2023 study ran 54 test documents through 14 detectors, including two commercial systems: none reached 80% accuracy, the tools leaned towards classifying AI text as human, and obfuscation made results worse.[29] A 2024 study of six detectors (805 tests) found mean accuracy of 39.5% on unmodified AI text, 67% on human-written controls, and 22.1% after simple machine edits; its authors conclude detectors cannot currently be recommended for deciding integrity cases.[31] An opinion article reported a small experiment in which seven detectors wrongly flagged, on average, 61.3% of essays by non-native English writers as AI-generated; after ChatGPT enriched the essays' vocabulary, the rate fell to 11.6%.[30]

Figure 1

Surface edits swing detector results

The same texts, scored as written and again after automated surface edits.

  1. AI-written texts correctly identified Mean of six detectors, 2024 · 805 tests · unmodified, then after edits such as added spelling errors
  2. Human-written essays wrongly flagged as AI Average of seven detectors, 2023 · 91 essays by non-native English writers · as written, then after ChatGPT enriched the vocabulary

Percentage of texts (each row states its own measure)

The rows measure different things in different studies and are not comparable with each other. Both show detector outputs moving with surface features of a text: disguised AI text was caught less often, and human essays polished by ChatGPT were flagged less often. The second row comes from an opinion article reporting a small experiment. Sources: [31][30]
Show Figure 1 as a table
Detector results before and after surface edits, from two studies
MeasureAs writtenAfter surface edits
AI-written texts correctly identified (six detectors, 2024)39.5%22.1%
Essays by non-native English writers flagged as AI (seven detectors, 2023)61.3%11.6%

Jisc notes that the best paid tools now report false positive rates of around 1 to 2%, but such figures date quickly.[23] The conclusion matches the regulators': a detector score can prompt a conversation, but it cannot decide a case.[10][25]

Design frameworks: scales, lanes and structural change

A common first response was to tell students which AI uses were allowed, through traffic lights or scales such as the five-level AI Assessment Scale (AIAS) of 2024.[33] A 2025 paper calls these changes discursive: they rely on instructions students can ignore, creating an "enforcement illusion". It argues for structural changes to the mechanics of the task itself.[35] The AIAS, revised in December 2025, now runs No AI, AI Planning, AI Collaboration, Full AI and AI Exploration. Level 1 requires a controlled environment and must still accommodate assistive technology; relabelling an unchanged task does not improve validity; and the scale is presented as complementary to two-lane models, which address security.[34]

The working rule: for each programme, decide which outcomes must be evidenced in a secure lane, under supervision or in conversation, and design the rest as an open lane in which AI use is assumed and, where useful, assessed.

Assessment approaches compared on validity for authorship, security, delivery effort and accessibility
ApproachValidity for authorshipSecurityDelivery effortAccessibility
Take-home essay or examCannot confirm who did the work[32]Low: rests on declarations[21]LowHigh: usual tools
Coursework with checkpointsShows development over time[14]Moderate: drafts can be circumvented[34]Moderate: review at each stageDemands for proof can disadvantage some[34]
Open task with AI by designValid where AI use is part of the outcome[21]None by design: assume AI use[34]Redesign, then ongoing review[21]Needs fair access to AI tools[21]
Supervised in-class taskShows unaided performance[34]Strong if devices have AI off[14]Moderate: class timeGood with word processors[16]
Viva or interactive oralTests understanding and authorship[21][25]Strong deterrent[21]High: staff time, small groups[21]Stressful for some; adjust for speech and hearing[21]
Invigilated on-screen examDepends on design; mode effects possible[6]Good with managed devices and lockdown[15][21]High: devices, testing, support[1][15]Built-in supports if designed and tested[1]
Remote proctored examAs on-screen, in an uncontrolled roomLimited: other devices within reach[21]Moderate to high: reviewing flags[13]Privacy concerns; biometrics need an alternative[13][46]
Handwritten invigilated examSamples outcomes; narrow competencies[21]Strong[21]High demands on estate[21]Poor for some; QAA calls it regressive[21]

Qualitative summary of the cited sources. Cells without a citation are the EdTechLab team's judgement.

Process evidence without surveillance

Process evidence shows how work came to be: outlines, drafts, version history, checkpoint feedback and, where AI is permitted, a short prompt log or reflection.[14][34] JCQ suggests checking that a submission continues earlier stages, and the OIA expects students to be asked for notes or drafts and told what any file metadata relied on is thought to show.[14][25] Its limits are real: version histories can be circumvented, and demands to prove authorship can disadvantage students without paid tools.[34] Process evidence supports a conversation; it does not replace one.

Why keystroke logging is high-risk. Some tools record every keypress. Studies of large-scale writing assessments show such logs can separate original from copied text with over 89% accuracy in operational data, but keystroke analysis is also used to authenticate users,[39] models trained on raw logs have been used to measure stable personal characteristics,[40] and detectors built on them need fairness testing.[38] Under UK GDPR, the way someone types, processed to identify them, is special category biometric data. The ICO says explicit consent is likely to be the appropriate condition, people who decline must be offered a suitable alternative, and a DPIA must come before any biometric recognition system.[46] Tracking behaviour online is also on the ICO's list of processing that needs a DPIA,[47] yet the ICO's June 2026 audit report found that over 40% of the edtech providers audited had no DPIA for their product.[48]

A proportionate design collects the least that answers the question: dated milestones rather than keystroke streams, kept for a set period and visible to the student. That is what data minimisation asks: data that is adequate, relevant and limited to what is necessary.[47]

Oral and in-person checks at scale

A short conversation about a piece of work is the most direct test of authorship. QAA describes structured orals by two or more examiners with clear rubrics, and mini-vivas in which small groups discuss their submissions, which both authenticate and assess. Orals take considerable staff time, can be stressful, and must treat students with speech impairments or different accents fairly.[21] A survey-based study of scaffolded tasks ending in interactive orals reported that they helped prevent misconduct.[37] In practice:

  • Scheduling. Book short slots and rotate question variants, since early candidates can pass questions on.[21]
  • Rubrics. Mark against a published rubric focused on the content of the work, allowing for how long ago it was done.[25]
  • Recording. If recordings join the assessment record, plan for the OfS expectation of five years' retention.[18] A plain recording is not biometric data; analysing a voice to identify someone would make it so.[46]
  • Adjustments. Orals fall under the same duty to make reasonable adjustments as any assessment.[45]

Secure on-screen delivery

Managed devices or bring your own. For GCSEs and A levels, Ofqual proposes to rule out students' own devices except as reasonable adjustments.[1] Universities decide for themselves. Safe Exam Browser (SEB), an open-source lockdown browser, supports unmanaged laptops through encrypted per-exam settings that the exam system can verify,[43] though QAA notes remote candidates can still reach other devices.[21]

Lockdown and network isolation. SEB pairs a kiosk application with a browser without an address bar, filters URLs against an allow-list, detects virtual machines and pins server certificates. With "Use Browser & Config Keys" on, it sends a hash of the exam's Config Key with each request so the exam system can check its configuration. Its September 2026 release for macOS and iOS fixed a certificate validation flaw, a reminder to keep clients current.[43] In Moodle 5.2, a quiz can "Require the use of Safe Exam Browser", list allowed browser exam keys, filter URLs and set a quit password.[44] JCQ's 2026 instructions add that AI can come from the application, the operating system and the network, so centres "must not assume that turning one of these off is enough", must test their approach, and must prevent unauthorised communication with other computer users.[15]

Resilience. Failure is not hypothetical: in a poll for Ofqual, 27% of schools reported a cyber incident in 2025 to 2026, and Ofqual notes the uncertainty for students when coursework or marks are lost.[12] Ofqual's definition of on-screen assessment includes offline and hybrid delivery needing no internet connection during the exam.[1] Moodle's quiz saves changed responses after a delay, 60 seconds by default, and its SEB settings warn that reloading a page offline can break offline caching.[44] JCQ requires centres to test the software before exam day and keep at least one replacement computer, and its failure procedures should let a candidate continue elsewhere, or later, without losing working time.[15]

A secure on-screen exam session, step by step

  1. Weeks beforePrepare and rehearseCentre-owned devices at the minimum specification; delivery software and lockdown client installed and tested; a practice session on the same platform; an alternative site planned.
  2. Exam set-upLock the sessionRequire the lockdown browser, check its Config Key or allowed browser exam keys, allow-list only the exam hosts, set a quit password, and turn AI off in the application, the operating system and the network.
  3. ArrivalIdentify and seatCheck identity, issue log-in details at the start with different passwords per session, and keep a seating plan that identifies each device, with screens 1.25 m apart or divided.
  4. DuringSave and superviseResponses autosave; at least one invigilator per 20 candidates; a control-centre workstation watched by staff; per-candidate time overrides for extra time.
  5. On failureRecover without lost timeMove the candidate to a spare workstation or a later slot. The invigilator controls the restart, resets timing and restores previous responses where feasible, or switches to paper.
  6. CloseSubmit and lock downConfirm every candidate's work is submitted, remove access to user areas between sessions, and report anything that may have affected integrity.
Sources: Ofqual's consultation, JCQ's 2026 instructions for on-screen examinations, Safe Exam Browser documentation and the Moodle 5.2 source code.[1][15][43][44]

Learning management systems. Vendor-neutrally, the requirement is a quiz engine that can require and verify a lockdown client, autosave, and grant individual time extensions. Moodle's core quiz does all three, with SEB support built in and user overrides for time limits.[44] SEB lists compatible modes in other systems too, including Canvas through third-party tools.[43]

Item banks and portability with QTI 3.0

Separate paper and screen specifications with substantially different questions mean larger item banks,[1] and item-level mode effects mean each item should record how it behaves on screen.[7] The 1EdTech Question and Test Interoperability (QTI) 3.0 specification moves items, tests and results between authoring tools, item banks and delivery systems, and adds HTML5 support, a shared presentation vocabulary, adaptive testing and integrated accessibility.[41]

Accessibility travels with the item. Through Personal Needs and Preferences (PNP) 3.0 profiles, a delivery system can add tools, change settings such as time limits, or inform invigilators. Alternative content sits in QTI catalogs inside the item (qti-catalog-info, holding qti-card elements whose support attribute names a need such as spoken, braille or sign-language).[41] A companion specification, still a candidate final release, defines a format for reporting item statistics.[42]

Accessibility and access arrangements

Every route to authenticated work must meet the Equality Act 2010. Section 20 sets the duty to make reasonable adjustments, including accessible information; section 91 applies it to further and higher education institutions; section 96 applies it to qualifications bodies, though the regulator may specify adjustments that should not be made, having regard to the qualification's reliability and public confidence.[45] JCQ's access arrangements allow a word processor where it is the candidate's normal way of working, with spelling and grammar checks off, and a computer reader even in papers that test reading, provided the computer holds no other software that could help.[16] Ofqual would keep on-screen delivery available as a reasonable adjustment, and says it helps disabled students only if designed for accessibility, tested with a range of users and backed by familiarisation.[1]

The risk runs both ways. QAA warns that relying on handwritten exams would reverse progress on accessibility,[21] and a Jisc community discussion found that no-AI rules can strip students of assistive tools they already use, some funded through the Disabled Students' Allowance.[24] Write the assistive-technology exception into the rule, test lockdown configurations with screen readers, and grant extra time through per-candidate overrides.[34][44]

Remote proctoring and biometrics

Ofqual's 2023 study of remote invigilation found that live supervision lets staff intervene, while record-and-review costs less and scales more easily but finds problems only afterwards; room scans, lockdown tools and thresholds for flagging footage all need judgement. The research predates generative AI and is not an endorsement of wide use in high-stakes assessment.[13] In 2026 Ofqual restated the high bar for remote evidence to be attributable to each learner, and its concern about relying on AI alone for remote invigilation.[8]

We found no ICO guidance specific to online proctoring. Its biometric guidance applies whenever a system matches faces or voices to identify candidates: the data is special category, consent needs a genuine alternative, power imbalances need care, and a DPIA must come first.[46] Its DPIA list cites eye tracking and remote working under tracking,[47] and its June 2026 edtech audit report asked providers whether collecting children's biometric information is really necessary and proportionate.[48] For most UK settings, that makes in-person supervision the default and remote proctoring an exception with a real alternative.

Implementation checklist

  • Map each programme outcome to a secure lane (supervised or oral) or an open lane (AI assumed)
  • Change the task, not just its label: mechanics, checkpoints and marking criteria
  • Publish assessment-specific AI rules with an assistive-technology exception
  • Never decide a case on a detector score; if one is used, choose it institutionally, cover it with a DPIA and share the report
  • Collect dated milestones, not keystroke streams; set retention and show students what is kept
  • No keystroke or biometric capture without a DPIA, a necessity test and a genuine alternative
  • Run short, structured orals with a published rubric, rotating questions and recorded decisions
  • Use centre-owned, managed devices for high-stakes exams, and rehearse on the same platform
  • Turn AI off in the application, operating system and network, and test before exam day
  • Require a verified lockdown client, allow-list exam hosts and keep clients patched
  • Autosave, keep a spare device, document restarts without lost time, and hold a paper fallback
  • Store items in QTI 3.0 with PNP supports where portability matters
  • Treat remote proctoring as an exception, with a DPIA and an in-person alternative

Where EdTechLab stands

We do not provide exam delivery, proctoring or AI detection. Our platforms sit on the evidence side of assessment. In intle, answer keys are validated and questions checked for common writing flaws on every generation, and intle offers per-question difficulty and discrimination analysis for hosted sessions, on higher-tier plans. EngagedLab produces per-objective evidence reports, with cohort figures suppressed below 5 learners.

When a client project needs secure or authenticated assessment, we start from what the assessment must show and what the regulator requires: mapping outcomes to secure and open lanes, completing a DPIA before collecting process data, preferring milestones and short structured conversations to keystroke or biometric capture, and testing the whole delivery stack, including assistive technology, before any live session.

Limits of this report

  • This is general information, not legal advice. Regulatory and legal points describe the cited documents as they stood on 5 October 2026.
  • Ofqual's on-screen measures are proposals; its decisions may differ.
  • We did not review regulators in Scotland, Wales or Northern Ireland, or awarding organisations' platform rules.
  • Detection studies test particular tool and model versions, which change quickly; several are small, and one is an opinion article.
  • HEPI's survey is a sponsored survey of 1,054 undergraduates, not a census.
  • Keystroke research comes from laboratory studies and large-scale writing tests, not coursework.
  • We did not test Safe Exam Browser, Moodle or any proctoring product for this report.

References

Accessed 5 and 6 October 2026.

  1. Ofqual. Regulating on-screen assessment (consultation, 11 December 2025 to 5 March 2026; status page and consultation document). Source
  2. Ofqual. Ofqual launches consultation to protect standards in on-screen exams (press release, 11 December 2025). Source
  3. Ofqual. On-screen assessment: the evidence base for Ofqual's consultation (11 December 2025). Source
  4. PA Consulting for Ofqual and the Department for Education. On-screen assessment research study (December 2025). Source
  5. El Masri Y, Stratton T. On-screen assessments in sessional high-stakes qualifications in England: opportunities and risks in the eyes of students and parents. Ofqual research report 25/7279, 2025. Source
  6. Dodds HJ, El Masri Y, Stratton T. On-screen assessment and mode effects: a review of the effects of on-screen assessment on levels of engagement, cognitive demands and performance. Ofqual research report 25/7280, 2025. Source
  7. Sweiry E. Making sense of mode effects: a framework for anticipating performance differences in equivalent paper-based and digital test items. Ofqual, 2025. Source
  8. Ofqual. Ofqual's approach to regulating the use of artificial intelligence in the qualifications sector (updated 16 July 2026). Source
  9. Ofqual. Letter to awarding organisation chief executives about malpractice (3 March 2026). Source
  10. Ofqual. Artificial intelligence malpractice and assessment: advice note (27 April 2026). Source
  11. Ofqual. Malpractice in GCSE, AS and A level: summer 2025 exam series (official statistics, 11 December 2025), and release announcement for the summer 2026 series (provisional, 10 December 2026). Source
  12. Ofqual. Schools making strides in cyber security as recovery times improve (press release, 1 October 2026; Teacher Tapp poll of 13 July 2026). Source
  13. Ofqual. Remote invigilation in vocational and technical qualifications (18 July 2023). Source
  14. Joint Council for Qualifications. AI Use in Assessments: Your role in protecting the integrity of qualifications, revision two (30 April 2025). Source
  15. Joint Council for Qualifications. Instructions for conducting examinations, effective from 1 September 2026 (paragraphs 14.24 to 14.25 and section 32). Source
  16. Joint Council for Qualifications. Access Arrangements and Reasonable Adjustments (updated March 2026, applying in 2026/27). Source
  17. Department for Education. Curriculum and Assessment Review Final Report: Building a world-class curriculum for all (independent report, 5 November 2025). Source
  18. Office for Students. Regulatory framework for higher education in England, Condition B4: Assessment and awards, with guidance. Source
  19. Office for Students. Embracing innovation in higher education: our approach to artificial intelligence (blog, 5 June 2025). Source
  20. Office for Students. OfS collaborates with Advance HE to conduct research into how universities and colleges are using artificial intelligence (27 May 2026). Source
  21. Quality Assurance Agency for Higher Education. Reconsidering assessment for the ChatGPT era: QAA advice on developing sustainable assessment strategies (31 July 2023). Source
  22. Quality Assurance Agency for Higher Education. An evidence-based toolkit on leveraging generative AI to support the graduates of the future (Collaborative Enhancement Project, 21 November 2025). Source
  23. Jisc National Centre for AI. AI Detection and assessment: an update for 2025 (blog, 24 June 2025). Source
  24. Jisc National Centre for AI. AI in assessment: Considering assistive technology users (community discussion write-up, 3 December 2024). Source
  25. Office of the Independent Adjudicator for Higher Education. Casework note: complaints relating to AI and academic misconduct (15 July 2025). Source
  26. Russell Group. Principles on the use of generative AI tools in education (3 July 2023). Source
  27. Higher Education Policy Institute. Student Generative AI Survey 2026 (HEPI Report 199, 12 March 2026; sponsored by Kortext). Source
  28. University of Sydney. University of Sydney's AI assessment policy: protecting integrity and empowering students (27 November 2024). Source
  29. Weber-Wulff D, Anohina-Naumeca A, Bjelobaba S, Foltýnek T, Guerrero-Dib J, Popoola O, Šigut P, Waddington L. Testing of detection tools for AI-generated text. International Journal for Educational Integrity, 2023, 19(1), 26. DOI
  30. Liang W, Yuksekgonul M, Mao Y, Wu E, Zou J. GPT detectors are biased against non-native English writers. Patterns, 2023, 4(7), 100779 (opinion article). DOI
  31. Perkins M, Roe J, Vu B, Postma D, Hickerson D, McGaughran J, Khuat H. Simple techniques to bypass GenAI text detectors: implications for inclusive education. International Journal of Educational Technology in Higher Education, 2024, 21(1), 53. DOI
  32. Scarfe P, Watcham K, Clarke A, Roesch E. A real-world test of artificial intelligence infiltration of a university examinations system: A "Turing Test" case study. PLOS ONE, 2024, 19(6), e0305354. DOI
  33. Perkins M, Furze L, Roe J, MacVaugh J. The Artificial Intelligence Assessment Scale (AIAS): A Framework for Ethical Integration of Generative AI in Educational Assessment. Journal of University Teaching and Learning Practice, 2024, 21(06). DOI
  34. Perkins M, Roe J, Furze L. Reimagining the Artificial Intelligence Assessment Scale: A refined framework for educational assessment. Journal of University Teaching and Learning Practice, 2025, 22(7), e1707. DOI
  35. Corbin T, Dawson P, Liu D. Talk is cheap: why structural assessment changes are needed for a time of GenAI. Assessment & Evaluation in Higher Education, 2025, 50(7), 1087–1097. DOI
  36. Dawson P, Bearman M, Dollinger M, Boud D. Validity matters more than cheating. Assessment & Evaluation in Higher Education, 2024, 49(7), 1005–1016. DOI
  37. Sotiriadou P, Logan D, Daly A, Guest R. The role of authentic assessment to preserve academic integrity and promote skill development and employability. Studies in Higher Education, 2020, 45(11), 2132–2148. DOI
  38. Jiang Y, Zhang M, Hao J, Deane P, Li C. Using Keystroke Behavior Patterns to Detect Nonauthentic Texts in Writing Assessments: Evaluating the Fairness of Predictive Models. Journal of Educational Measurement, 2024, 61(4), 571–594. DOI
  39. Deane P, Zhang M, Hao J, Li C. Using Keystroke Dynamics to Detect Nonoriginal Text. Journal of Educational Measurement, 2026, 63(1), e12431. DOI
  40. Zhang M, Deane P, Hoang A, Guo H, Li C. Applications and Modeling of Keystroke Logs in Writing Assessments. Educational Measurement: Issues and Practice, 2025, 44(2), 5–19. DOI
  41. 1EdTech. Question & Test Interoperability (QTI) 3.0 Overview (Final Release, 1 May 2022) and QTI 3.0.1 Best Practice and Implementation Guide (1 October 2024). Source
  42. 1EdTech. QTI Usage Data & Item Statistics Specification v3.0 (Candidate Final Public, 2 February 2022). Source
  43. Safe Exam Browser. About: overview; developer documentation, Config Key; and news, SEB 3.7.1 for macOS and iOS (9 September 2026). Source
  44. Moodle source code, MOODLE_502_STABLE: the Safe Exam Browser quiz access rule, quiz settings and the quiz override form. Source
  45. Equality Act 2010, sections 20, 91 and 96. Source
  46. ICO. Biometric data guidance: Biometric recognition (under review following the Data (Use and Access) Act). Source
  47. ICO. Examples of processing 'likely to result in high risk' (DPIA guidance) and Principle (c): Data minimisation. Source
  48. ICO. Edtech examined: key findings from our audits (June 2026). Source

Next step

Rethinking assessment for generative AI?

We can review an assessment plan against this checklist, or scope the evidence and delivery systems with you, starting with a DPIA.

Related

Core finding

Assume unsupervised work may involve AI. Secure a few well-chosen moments through supervision or conversation, design the rest to use AI openly, and treat detectors and surveillance data as weak evidence, never proof.

In this report