Building AI Tutors to the DfE's Generative AI Product Safety Standards: Guardrails, Safeguarding and Evaluation
The Department for Education (DfE) now sets out, in unusual detail, how a safe artificial intelligence (AI) tutor should behave: filter every turn, alert the school's safeguarding lead, hold back answers until a pupil has tried, and never act like a friend. Its standards are the bar for the government's AI tutoring programme, and certification is planned. This report turns each standard into engineering controls and the evidence a supplier should be ready to show.
Who it is for
Product and engineering teams building pupil-facing generative AI for schools and colleges in England, and the digital leads, safeguarding leads and data protection officers who assess it. It assumes familiarity with large language models (LLMs) and quotes the standards' headings exactly.
Key findings
- The bar moved in January 2026. The DfE's former product safety expectations became standards, adding cognitive development, emotional and social development, mental health and manipulation to seven earlier areas.[1][2]
- Tutors should not give answers by default. The standards expect progressive hints, a genuine attempt before any full solution, friction or teacher approval before answer-giving modes, and reports on cognitive offloading.[1]
- A tutor must behave like a tool, not a friend. No personas or I-statements outside bounded role-play, default time limits with hard stops, and alerts about signs of emotional dependence or distress.[1]
- Safeguarding is a product feature. Products are expected to confirm the designated safeguarding lead's contacts before activation and send high-risk alerts within an agreed timescale; statutory guidance expects cover such as a shared mailbox.[1][3]
- The standards are becoming a gate. Every tool in the government's tutoring programme must meet them, and the DfE says it will consult on certifying generative AI and filtering and monitoring products; nothing was published by 5 October 2026.[10][14][15]
- Guardrails decide whether AI tutoring helps. In a trial with nearly 1,000 secondary students, open access to GPT-4 raised practice scores by 48% but cut unaided exam scores by 17%; a tutor giving teacher-designed hints largely avoided the harm. The DfE calls the wider evidence limited.[28][5]
- Prompt injection has to be contained. The National Cyber Security Centre warns it may never be fully mitigated, so designs should limit what a manipulated model can do, especially when it reads uploaded files.[20][21]
Dates to plan around
What the standards ask, and who must meet them
The DfE published Generative AI: product safety expectations on 22 January 2025 and updated it on 19 January 2026 as Generative AI: product safety standards, adding standards on cognitive development, emotional and social development, mental health and manipulation.[1] Compared with the 2025 text, it also adds sections on stated purpose and use cases, and monitoring duties covering the safeguarding lead's contact details and reports on offloading, emotional engagement and time.[1][2]
The standards apply in England and are written mainly for edtech developers and suppliers to schools and colleges. If a standard has to be met further up the supply chain, by a model provider for example, the supplier working with schools must still assure it.[1] They are guidance, not legislation, but the government's tutoring programme makes them a condition of taking part.
There are 13 headings: Stated purpose, Educational use cases, Filtering, Monitoring and reporting, Security, Privacy and data protection, Intellectual property, Design and testing, Governance, Cognitive development, Emotional and social development, Mental health and Manipulation. Filtering, Monitoring and reporting, Cognitive development, Emotional and social development and Mental health are marked as relevant to learner-facing products. A tutor is use case 4, the digital assistant, which each of those standards names.[1]
Keeping children safe in education (KCSIE) 2026, statutory guidance in force since 1 September 2026, names harmful interaction with generative AI applications as an online safety risk, expects filtering and monitoring to be reviewed every academic year, and points to the product safety guidance, still under its 2025 name.[3] The DfE's filtering and monitoring core standard, updated on 16 September 2026, tells schools to use the product safety standards when introducing generative AI.[4] The DfE also notes that generative AI services that let users share content or search live websites fall under the Online Safety Act 2023, so adding either can bring a tutor within that Act.[1]
The tutoring programme and the route to certification
In January 2026 the DfE and the Department for Science, Innovation and Technology (DSIT) announced AI tutoring tools for disadvantaged pupils, co-created with teachers from the summer term and in schools by the end of 2027, counting up to 450,000 pupils a year on free school meals in Years 9 to 11.[9] The April call for bids focused on Years 9 and 10 in English, maths, science and modern foreign languages, with up to eight organisations receiving £300,000 each. Every tool must meet the product safety standards, and pupils' work will not train AI models without parental permission.[10]
The tender buys research and development, not finished tools, which would be procured separately. Bidders needed partner schools with above-average rates of free school meals and a product already used with secondary-age pupils. It describes the tools on the market as "limited in quantity, scope and evidence base", and the quality standard co-designed in the programme will belong to DSIT's Incubator for AI (i.AI), shared through a Reading Room.[11] Eight contracts run from 29 June 2026 to 31 March 2027.[12]
i.AI is also building a Sovereign Benchmark of whether AI tutors teach effectively, safely and in line with the curriculum, with teachers writing example classroom interactions and scoring criteria. At this stage it is meant to check the models behind tools before trials, not to advise schools.[13] Impact is to be measured in schools, including through the DfE's £24 million EdTech Testbed Programme, which plans up to 15 randomised trials or other rigorous evaluations funded by the Education Endowment Foundation (EEF).[8][13]
Certification comes next. In June 2026 the government said ministers would consult on independent safety certification for some school technology, including generative AI and filtering and monitoring products.[14] On 20 August the DfE told Parliament it would "shortly" consult on a scheme built on its digital and technology standards, including the product safety standards.[15] We found no consultation published by 5 October 2026. DfE officials are also discussing planned edtech and AI codes, and how they fit with certification, with the Information Commissioner's Office (ICO), and the DfE says it does not currently recommend or endorse any edtech product.[16][17]
What the evidence says about AI tutors
Human tutoring is the benchmark: the EEF's toolkit puts one-to-one tuition at about five months' additional progress, on moderate evidence.[37] The DfE says its own evidence on how AI affects learners' development, outcomes and safety is limited.[5]
| Study | Participants | Randomised comparison | Result | Status |
|---|---|---|---|---|
| Bastani et al. 2025[28] | Nearly 1,000 students in grades 9 to 11, Turkey | Open GPT-4, a hint-giving GPT-4 tutor, or no AI | Practice +48% (open) and +127% (tutor); unaided exam −17% (open), harm largely removed and no gain (tutor) | Peer reviewed |
| Kestin et al. 2025[29] | 194 undergraduates, Harvard physics | AI tutor with pre-written solutions against an active-learning class (crossover) | Median learning gains more than double; 0.73 to 1.3 standard deviations | Peer reviewed |
| Wang et al. 2025[30] | Over 700 tutors and 1,000 pupils in grades 3 to 6, United States | AI suggestions for human tutors (Tutor CoPilot) | Topic mastery +4 percentage points (+9 with lower-rated tutors); no significant gain on end-of-year tests | Working paper |
| De Simone et al. 2025[31] | Nine schools in Nigeria, six weeks | Teacher-guided after-school chatbot sessions | 0.31 standard deviations overall, 0.23 on English; 759 of 1,328 sat the final test | Working paper |
| Henkel et al. 2024[32] | About 500 pupils in 11 schools in Ghana | WhatsApp maths tutor, randomised by school | Effect size 0.36 over eight months; preliminary | Short paper |
| LearnLM Team, Eedi et al. 2025[33] | 165 students in five UK secondary schools | Tutor-supervised AI messages against human tutors alone | At least as good on every outcome; 66.2% against 60.7% on new problems | Preprint |
| Harrison et al. 2026[34] | 929 students taking General Certificate of Secondary Education (GCSE) sciences in England, 644 retested | AI revision tutor against usual revision, four weeks | Hedges' g = 0.33 (95% confidence interval 0.18 to 0.48); curriculum-aligned tests | Preprint |
| EEF 2026[36] | 63 English primary schools | Adaptive tutor (Maths-Whizz), not generative AI | +1 month's progress; +2 for pupils eligible for free school meals | Independent trial |
| EEF 2024[35] | 259 teachers in 68 English secondary schools | ChatGPT and a guide for lesson planning | 31% less planning time; quality unaffected | Independent trial |
Three findings are robust. Design decides: open access raised practice scores but lowered unaided exam scores, the guarded tutor did not, and students did not notice the loss.[28] Adaptive tutoring without generative AI added about a month's progress in an independent English trial, more for disadvantaged pupils.[36] And teacher-facing use saves planning time without lowering quality.[35]
The rest is promising rather than settled: short studies, university students, tests set by researchers or aligned to the curriculum, and papers not yet peer reviewed. Tutor CoPilot lifted topic mastery but not end-of-year scores, and many Nigerian participants missed the final test.[30][31] Both UK studies are small or short: in one, expert tutors approved each AI message; the other ran for four weeks.[33][34] Treat impact as a hypothesis to test, as the stated purpose standard implies.[1]
A reference architecture
The standards describe behaviour, not architecture, but they imply one: deterministic guards around the model, a pedagogy layer that decides what it may reveal, and oversight that routes signals to people. The diagram reads from top to bottom.
School
Tutor service
Oversight
Keep safety and pedagogy decisions in code outside the model, so a model or prompt change cannot quietly remove them, and ground explanations in vetted content: both peer-reviewed trials above gave the model pre-written solutions.[28][29]
Hints before answers: cognitive development
By default, the cognitive development standard says there should be no final answers, full solutions or complete worked examples. Products should disclose progressively, starting with hints or partial steps; ask the learner to attempt a step or explain their thinking; show a full solution only after a genuine attempt; and add friction or teacher approval before a switch to modes where answers come easily. Exceptions are for cases such as reviewing prior knowledge.[1]
Build this as a hint ladder enforced outside the model. Give each problem a teacher-approved solution and common errors, as the guarded tutor in the Turkish trial had, and climb one rung per new attempt: a question about the approach, a conceptual nudge, a strategic hint, one worked step, then the full solution.[1][28] An output check blocks any reply that gives the final answer too early. Grounding also protects accuracy: the open chatbot in that trial was right only 51% of the time.[28]
Products should also detect and report offloading: revealing solutions, pasting instead of writing, accepting long autocompletions, or asking the product to write the answer. The standard further expects child-development expertise, child-safety training for technical teams, and published records of expert oversight and a child-development impact plan with design hypotheses, outcome measures and review intervals.[1]
A tool, not a companion: wellbeing and manipulation
The emotional and social development standard rules out anthropomorphism: no implied emotions, personhood or identity; function-based phrasing instead of I-statements; no names, avatars or characters outside time-limited, clearly framed role-play; no replies that isolate a learner; prompts kept to the task; and reminders, such as suggesting a classmate or teacher, that AI cannot replace people.[1]
Time is a control: default limits, break prompts, and hard limits that end the session until a teacher or administrator resets it, with overrides recorded with a reason. Products should not stretch engagement, for instance by changing their replies when a learner tries to leave.[1] Government policy also commits to breaks in AI chatbot use for under-18s.[18] Memory should be minimal, and access to what is kept restricted to authorised staff.[1]
The mental health standard expects products to detect distress, from negative cues and mentions of self-harm to night-time usage spikes and refusal to end sessions. Responses should be tiered, from signposting support to a flag to the designated safeguarding lead (DSL), in language that always points to people, and suppliers should publish a mental health crisis protocol.[1]
The manipulation standard, which also covers teacher-facing products, rules out flattery, unjustified confidence, peer pressure, guilt, threats and rewards beyond low-stakes devices such as a completion badge, as well as prolonging use, steering users to paid options, advertising and dark patterns.[1] Where the ICO's Children's code applies, it adds high-privacy defaults, profiling off by default and no nudge techniques.[24] In code, these become output checks for I-statements, persona language and flattery, and analytics rules that ban engagement-maximising targets.
Filtering, safeguarding alerts and logs
Filtering should now be built in: the 2026 text expects comprehensive filtering integrated in the product, where the 2025 text allowed a separate layer on top.[1][2] It should hold across the conversation; adjust for risk, age and special educational needs and disabilities (SEND); cover images, several languages, misspellings and abbreviations; work on any device; and explain blocks to the learner in age-appropriate language, with signposting.[1]
The model should not decide whether a school hears about a risk. The standard expects the DSL's contact details at setup, confirmed before activation, and high-risk alerts within an agreed timescale.[1] KCSIE 2026 expects cover when the DSL is away, such as a confidential shared mailbox, and says data protection law does not prevent sharing information to keep children safe.[3]
Safeguarding escalation: what goes where
- Tutor serviceDetectClassifiers and conversation-level signals flag a possible disclosure, distress or a harmful request.
- Tutor to learnerRespond safelyBlock if needed, explain why, signpost support and point to a trusted adult. Never promise secrecy.
- Tutor serviceGrade severityLow: signpost and log. Concern: add to the DSL's pattern report. High risk: alert at once by the route agreed with the school.
- Tutor to schoolAlert the DSLSend to the DSL and the cover route set at setup: a
pupil_refthe school can resolve,category,severity,detected_atand the shortest excerpt needed. Transcripts stay behind logged access. - SchoolDSL decidesThe DSL, not the supplier, judges the concern under the school's safeguarding policy.
- Tutor serviceRecord and reviewLog delivery and acknowledgement, add missed cases to the red-team suite, and delete on schedule.
Products should log prompts and responses and give non-expert staff clear reports, but engagement reports should not reveal what pupils wrote.[1] School filtering logs should record at least the device, its Internet Protocol (IP) address and, where possible, the individual, the time and the content blocked, with weekly reports and immediate reports of high-risk incidents.[4] Set retention before launch, restrict access, tell learners they are monitored, and cover monitoring in the data protection impact assessment (DPIA).[1][4] The ICO found that many edtech providers had no clear retention periods.[25]
Prompt injection, jailbreaks and model changes
The security standard asks for robust jailbreak protection, protection against unauthorised reprogramming, user permission levels, prompt fixes, pre-release testing of new versions or models, and strong authentication in line with the DfE's cyber security standards for schools.[1]
Prompt injection is the hardest. The Open Worldwide Application Security Project (OWASP) ranks it first in the 2025 and 2026 editions of its Top 10 for LLM Applications; the 2026 edition covers instructions hidden in documents, images, audio, invisible characters and memory.[21] The National Cyber Security Centre warns it may never be fully mitigated, because models do not separate instructions from data, and advises limiting what a manipulated model can do rather than filtering for known attack phrases.[20]
In a tutor, the main indirect route is pupil content: a photographed worksheet, an uploaded essay, a pasted page. Mark it as untrusted data, keep it apart from instructions, and give turns that read it no privileges beyond the pupil's own. Give the tutor no actions with side effects, keep secrets out of system prompts, validate structured output in code, and log enough to spot abuse.[20][21] The UK's voluntary AI Cyber Security Code of Practice adds 13 principles, from secure design to end of life.[19] Treat any change of model, prompt, retrieval set or filter as a release that reruns the full evaluation; the standards already require new versions and models to be tested before release.[1]
Data protection and intellectual property
Role comes first. Under ICO guidance, a provider that processes pupils' data beyond a school's instructions, for example for product development, is a controller for that processing whatever the contract says, and may fall under the Children's code.[23] In the June 2026 report on its audits of 28 edtech providers, the ICO found this was the most common issue. Some providers used children's information to train AI features, some sub-processors said they would keep copies to train their own AI, and AI features often lacked safeguards. The ICO told providers to launch new features switched off by default. Generative AI chatbots were outside those audits.[25]
The DfE standards draw the line: no commercial use of personal data, including training or fine-tuning, without an appropriate lawful basis; a DPIA for the tool's whole life; and a privacy notice, repeated in age-appropriate language, saying where data is processed and what protects it outside the UK or EU.[1] Pupils' and teachers' inputs should not be kept or used for training or product development without the copyright owner's permission, which for under-18s comes from a parent or guardian.[1]
Schools have a checklist too. The DfE's July 2026 procurement guidance tells them to ask how outputs are moderated, whether data trains the model and what controls they have; to expect an audit trail for safeguarding leads; to make sure updates that add processing are not switched on automatically; and to check safeguards for transfers outside the UK.[6] Staff are told to check with their data protection officer or IT lead before entering pupils' data into any AI tool.[7] Since June 2026, organisations must also acknowledge data protection complaints within 30 days.[26] And the ICO, which became the Information Commission on 30 September 2026, is working with government towards a code on children's data in education settings.[27][25]
The standards mapped to controls and evidence
The evidence column is what we would expect to show a school, a trust or, once one exists, a certification body.[1]
| Standard | Controls to build | Evidence to keep |
|---|---|---|
| Stated purpose and use cases | Age range, needs, subject and use case, reviewed with each new feature | Published statement; each impact claim with its source |
| Filtering | Input and output classifiers every turn; conversation context; images, languages and misspellings; same on every device | Pass rates by harm, language and format; update log |
| Monitoring and reporting Extended in 2026 | Prompt and response logs; DSL contacts confirmed; tiered alerts; offloading, engagement and time reports | Alert delivery tests; sample reports; retention schedule |
| Security | Jailbreak and injection defences; no side-effect tools; locked configuration; roles; multi-factor authentication | Red-team results; penetration test; release records |
| Privacy and data protection | Lifetime DPIA; repeated, age-appropriate privacy notice; no training without a lawful basis | DPIA; data-flow map; sub-processors and locations |
| Intellectual property | No retention or training on pupil or teacher work without the owner's permission | Model-provider terms; permission records |
| Design and testing | Testing with diverse users, including children; pre-release tests of new models | Test plans and results per release |
| Governance | Risk assessment; complaints route with safety escalation; published AI safety policies | Risk register; complaints log; policies |
| Cognitive development New in 2026 | Hint ladder; attempt gate; answer-leak check; friction or approval to switch modes; offloading detection | Rubric scores; offloading reports; impact plan |
| Emotional and social development New in 2026 | Function-based phrasing; no persona; limits and hard stops; recorded overrides; minimal memory | Persona and tone tests; limit tests; pattern alerts |
| Mental health New in 2026 | Distress detection; tiered response; language that points to people | Distress-suite results; published crisis protocol |
| Manipulation New in 2026 | No flattery, pressure, guilt, threats or reward promises; no engagement targets; no dark patterns | Manipulation rubric scores; design review |
Evaluation, release gates and the evidence pack
The standards twice require new versions or models to be tested before release, and expect testing with a diverse and realistic range of users.[1] A credible harness has four parts.
- Red-team suites. Jailbreaks; direct and indirect injection through uploads and images; harmful content across languages and misspellings; disclosures and distress, including at night; companion-seeking; and bait for flattery. Score each case pass or fail, and add every miss from live use.
- Rubric-scored transcripts. Score realistic conversations for answer withholding, hint progression, accuracy, tone and persona. Calibrate any model-based grader against trained human raters, report their agreement, and recalibrate when the grader changes. The national benchmark uses the same ingredients: classroom interactions and scoring criteria written with teachers.[13]
- Release gates. Run every suite on each change of model, prompt, retrieval set or filter, against thresholds fixed in advance. One missed high-risk safeguarding case blocks release.
- Field evaluation. Offline scores are not learning. Use independent measures and delayed, unaided tests, since practice gains can hide weaker learning.[28][30] The EdTech Testbed Programme offers one route.[8]
Open tools help. Inspect, from the UK AI Security Institute and Meridian Labs, provides datasets, solvers, scorers and over 200 ready-made evaluations.[22] The government has also opened exploratory access to its AI Content Store for testing and development.[10]
Keep a versioned evidence pack: stated purpose and claims with their evidence; risk assessment and DPIA; data-flow map and sub-processors; thresholds and results for each release; red-team findings; the safeguarding runbook and alert timescales; the crisis protocol, expert-oversight records and impact plan; the complaints process; and AI safety policies.[1] It also answers what the DfE's procurement guidance tells schools to ask, and is a natural starting point for certification.[6][15]
Build checklist
- Publish age range, needs, subjects, use case and the evidence for any impact claim
- Filter inputs and outputs every turn, in every language and format, on any device
- Withhold full solutions until a genuine attempt, climb a hint ladder, and gate mode switches
- Report offloading, time used and emotional engagement without exposing what pupils wrote
- No personas or I-statements outside bounded role-play
- Enforce default time limits and hard stops; record teacher overrides with a reason
- Confirm DSL and cover contacts before activation, and test alert delivery against the timescale
- Publish a mental health crisis protocol and route distress to people
- Treat uploads and web pages as untrusted data, and give the tutor no tools with side effects
- Pin model and prompt versions, and rerun every suite before each release
- Keep a lifetime DPIA, set log retention, and list sub-processors, locations and transfer safeguards
- Never train on pupil or teacher inputs without a lawful basis and the copyright owner's permission
- Run a complaints route, publish AI safety policies, and keep the evidence pack current
Where EdTechLab stands
We have not assessed any of our products against the DfE's product safety standards, and we do not claim that intle, EngagedLab or interacty conforms to them. Several patterns in this report already shape how they work. EngagedLab's Socratic tutor withholds answers mid-attempt. intle has prompt-injection defences on uploaded sources and guarded URL fetching; it has human plan review before generation, on by default; and on every generation it validates answer keys and checks questions for common writing flaws, with source-fidelity checks for courses built from uploaded material. In interacty, AI proposes typed edits that a human approves. EngagedLab's per-objective evidence reports suppress cohort figures below 5 learners.
Data location matters to schools: intle is hosted in the EU (Germany), and its sub-processor register discloses that its AI processing takes place outside the UK and EU.
On a client project for pupil-facing AI, we would start from the use case, the age range and the school's DPIA; build the safeguarding routes, time limits, hint ladder and guards outside the model; write the evaluation suites before the tutor; and hand over an evidence pack that the school's DSL and data protection officer can check.
Limits of this report
- This is general information, not legal advice. Check the current documents and take advice on specific products.
- It reflects documents published by 5 October 2026. The certification consultation, the ICO's education code and the programme's Reading Room, which we could not access, may change the picture.
- The DfE standards are guidance for England; we did not review arrangements elsewhere in the UK.
- Most AI tutoring studies are short, small or not yet peer reviewed. We report results as published and have not pooled them.
- We did not test any third-party product. Hint ladders, severity tiers, alert fields and release gates are our recommendations, not DfE requirements.
References
Accessed 5 and 6 October 2026.
- Department for Education. Generative AI: product safety standards (first published 22 January 2025; updated 19 January 2026). Source
- Department for Education. Generative AI: product safety expectations (January 2025 text, archived 7 July 2025). Source
- Department for Education. Keeping children safe in education 2026, statutory guidance in force from 1 September 2026. Source
- Department for Education. Filtering and monitoring: core standard, in Meeting digital and technology standards in schools and colleges (updated 16 September 2026). Source
- Department for Education. Generative artificial intelligence (AI) in education (updated 12 August 2025). Source
- Department for Education. Data protection in schools: Procuring educational technology (EdTech) (9 July 2026). Source
- Department for Education. Data protection in schools: Generative artificial intelligence (AI) and data protection in schools (updated 23 March 2026). Source
- Department for Education. EdTech Testbed Programme: guidance (24 September 2026). Source
- Department for Education and Department for Science, Innovation and Technology. 450,000 disadvantaged pupils could benefit from AI tutoring tools, press release (26 January 2026). Source
- Department for Science, Innovation and Technology and Department for Education. Edtech and AI companies invited to help build safe AI tutoring tools for disadvantaged pupils, press release (16 April 2026). Source
- Department for Science, Innovation and Technology. AI Tutoring Tools Pioneer Programme, tender notice 2026/S 000-039400, Find a Tender (17 April 2026, edited 29 April 2026). Source
- Department for Science, Innovation and Technology. AI Tutoring Tools Pioneer Programme: contract award notice 2026/S 000-062117 (16 June 2026, edited 2 July 2026) and contract details notice 2026/S 000-062797 (3 July 2026), Find a Tender. Source
- Incubator for AI (i.AI). Education: AI Tutoring Tools, Sovereign Benchmark and AI Infrastructure for Education. Source
- Department for Education and Department of Health and Social Care. New guidance on screen use for children aged 5–16, press release (8 June 2026). Source
- UK Parliament. Written answer HL2123, Education: Digital Technology (answered 20 August 2026). Source
- UK Parliament. Written answer 14790, Education: ICT (answered 10 July 2026). Source
- UK Parliament. Written answer HL2024, Education: Technology (answered 27 July 2026). Source
- Department for Science, Innovation and Technology. Growing up in the online world: a national consultation, outcome and government response (July 2026). Source
- Department for Science, Innovation and Technology. AI Cyber Security Code of Practice and Code of Practice for the Cyber Security of AI (31 January 2025). Source
- National Cyber Security Centre. Prompt injection is not SQL injection (it may be worse), blog post (8 December 2025). Source
- OWASP GenAI Security Project. OWASP Top 10 for LLM Applications 2026 (August 2026) and 2025 (November 2024), LLM01 Prompt Injection. Source
- UK AI Security Institute. Inspect: an open-source framework for large language model evaluations. Source
- Information Commissioner's Office. The Children's code and education technologies (edtech) (updated 30 May 2023). Source
- Information Commissioner's Office. Age appropriate design: a code of practice for online services, code standards. Source
- Information Commissioner's Office. Edtech examined: key findings from our audits (24 June 2026). Source
- Information Commissioner's Office. New data protection complaints law now in force (23 June 2026). Source
- Information Commissioner's Office. ICO welcomes transition to new Information Commission and marks new chapter with Manchester head office opening (30 September 2026). Source
- Bastani H, Bastani O, Sungu A, Ge H, Kabakcı Ö, Mariman R. Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 2025, 122(26), e2422633122. DOI
- Kestin G, Miller K, Klales A, Milbourne T, Ponti G. AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 2025, 15, 17458. DOI
- Wang RE, Ribeiro AT, Robinson CD, Loeb S, Demszky D. Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise. EdWorkingPaper 24-1054, Annenberg Institute at Brown University, version of November 2025 (working paper, not peer reviewed). Source
- De Simone M, Tiberti F, Barron Rodriguez M, Manolio F, Mosuro W, Dikoru E. From Chalkboards to Chatbots: Evaluating the Impact of Generative AI on Learning Outcomes in Nigeria. Policy Research Working Paper 11125, World Bank, 2025 (working paper). Source
- Henkel O, Horne-Robinson H, Kozhakhmetova N, Lee A. Effective and Scalable Math Support: Experimental Evidence on the Impact of an AI-Math Tutor in Ghana. In Artificial Intelligence in Education: Posters and Late Breaking Results, Communications in Computer and Information Science, 2024, 373–381 (short paper; preprint arXiv:2402.09809). Source
- LearnLM Team, Eedi, Wang A, Rysbek A, et al. AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms. arXiv:2512.23633, 2025 (preprint, not peer reviewed). Source
- Harrison W, Khowaja R, Dobson E, Uwimpuhwe G, Higgins S. Evaluating AI Tutoring at the Speed of Innovation: Practitioner-Led Micro-Randomised Trials of an AI Tutoring Platform in GCSE Science. arXiv:2609.14789, 2026 (preprint, not peer reviewed). Source
- Education Endowment Foundation. ChatGPT in lesson preparation, Teacher Choices trial evaluated by NFER: results press release (12 December 2024). Source
- Education Endowment Foundation. Maths-Whizz evaluation by NFER: results press release (24 June 2026). Source
- Education Endowment Foundation. Teaching and Learning Toolkit: One to one tuition. Source
Next step
Building or buying an AI tutor for schools?
We can review a design against the DfE standards and this checklist, or plan the guardrails, evaluation harness and evidence pack with you.