Evidence by Design: Making EdTech Ready for the DfE-EEF EdTech Testbed
Between 2026 and 2030, the Department for Education (DfE) and the Education Endowment Foundation (EEF) will test edtech in more than 1,000 schools and colleges. Products that can show who used them, how much, in which version and with what result will be far easier to evaluate. This report sets out what the programme asks, how the EEF judges evidence, and how to build evaluation into a product from the start.
Who it is for
Edtech product teams, school and college digital leads, and university research and innovation teams preparing a product or pilot for independent evaluation. It assumes no training in statistics.
Key findings
- Two stages, four years. Up to 100 rapid evaluations by an Ipsos-led consortium, then up to 15 EEF trials or other rigorous evaluations funded with £7 million, in more than 1,000 schools and colleges. Total investment is £24 million.[1]
- Supplier rules are not yet published. Recruitment runs in phases from the 2026 to 2027 academic year; until DfE publishes supplier criteria, EEF practice is the best guide to Stage 2.[1][9]
- The gap is coherence, not volume. In DfE's Evidence Board pilot, relatively few of the 12 portfolios reviewed linked need, design, implementation and outcomes, and no awards were made.[7]
- Rigour is rated. Five EEF padlocks need randomisation, a minimum detectable effect of 0.2 or less and attrition of 10% or less; the weakest of the three sets the starting rating.[11]
- Usage data explains results. Maths-Whizz pupils averaged 32 minutes a week against a recommended 45 to 60, and the EEF expects compliance to be defined in advance from the logic model.[10][22]
- Effects are small, so samples are large. The mean effect across 141 large education trials was 0.06 standard deviations, and DfE's scoping report suggested 50 to 70 schools or colleges per trial arm.[6][38]
- Research use changes roles. A provider using pupils' data for research beyond a school's instructions acts as a controller. Since 5 February 2026, research processing needs safeguards such as pseudonymisation and must not drive decisions about individual pupils.[29][30]
Dates to plan around
- 24 September 2026
- Testbed guidance published and Stage 1 consortium named.[1]
- 2026 to 2027 academic year
- Phased recruitment of schools, colleges and suppliers begins.[1]
- Autumn 2026
- The EEF trial of Aila, Oak National Academy's lesson-planning assistant, is due to report.[27]
- Spring 2027
- The EEF trial of FFT Tutoring with the Lightning Squad, an online reading programme, is due to report.[28]
What the Testbed is and what it asks of suppliers
The EdTech Testbed Programme, a DfE and EEF partnership running from 2026 to 2030, replaces DfE's EdTech Impact Testbed pilot.[2] The EEF announced its role in May 2026.[3] The programme targets three priorities: learner outcomes, workload, and the inclusion of all learners, including those with special educational needs and disabilities (SEND). It will test products and practices, including generative artificial intelligence (AI) tools, in more than 1,000 school and college settings, with total investment of £24 million over four years.[1] DfE's January 2026 announcement described an additional £23 million.[5]
Stage 1 is up to 100 rapid evaluations to find products that show promise, run by a consortium led by Ipsos with University College London, EdTech Impact, Ecorys UK, the Chartered College of Teaching and King's College London, with the EEF as evidence partner. In Stage 2 the EEF will invest £7 million in up to 15 randomised controlled trials (RCTs) or other rigorous evaluations of the most promising tools.[1] It describes those alternatives as quasi-experimental designs (QEDs).[4] Research by RAND Europe will shape the scope, and results will appear in an online evidence hub.[1]
Organisations that registered for the pilot will be contacted by the Stage 1 consortium.[1] The guidance does not yet say how suppliers will be chosen, what data they must share, whether schools will pay or which methods Stage 1 will use. Until it does, EEF practice is the clearest guide.
| Question | EdTech Testbed | EEF trials (published practice) |
|---|---|---|
| Who takes part | Published Schools, colleges and suppliers, in phases from 2026/27[1] | Programmes delivered in more than 20 settings in England, with capacity for 50 in a year[9] |
| Supplier eligibility | Not published | Fully developed, needing no further development[9] |
| Implementation support | Not published | Standalone platforms without implementation support are not usually treated as programmes[9] |
| Data access | Not published | Compliance measures agreed with the evaluator; pupil data linked to national records and archived[10][18] |
| Costs to schools | Not published | Schools usually contribute, and cost to schools is weighed[9] |
| Outcomes | Published Learner outcomes, workload and inclusion[1] | Usually attainment, with one primary outcome[10] |
| Duration | Published 2026 to 2030[2] | Usually two to three years per trial[9] |
What came before: the pilot, the scoping report and the Evidence Board
DfE's scoping report for the pilot, by the Open Innovation Team and published in December 2025, compared four testbed models from a 2019 Nesta review (co-design, test and learn, evidence hub and edtech network) and found that none meets every objective alone. Because a testing window of 6 to 12 weeks is too short to measure attainment or inclusion reliably, it recommended starting with workload, especially marking and feedback tools.[6]
Its options foreshadow the two stages. A pilot evaluation, typically with around 25 schools or colleges and no control group, tests feasibility and early promise within a term, against success criteria set in advance, such as minimum usage thresholds. An impact evaluation needs an RCT or a matched comparison design; for an RCT, the rule of thumb was 50 to 70 schools and colleges per arm, so about £1 million could test just one tool. The report also warned that edtech, especially AI, changes faster than research cycles.[6]
DfE's EdTech Evidence Board pilot, delivered by the Chartered College of Teaching, tested independent review of product evidence. Its September 2026 report found the criteria workable, but relatively few portfolios gave a coherent account linking educational need, design, implementation and outcomes; evidence was often fragmented or poorly matched to the claims made. Documentation of accessibility, data protection, cybersecurity, safety and interoperability varied widely, and no awards were made. The report advised providers to make a logic model the lynchpin of their evidence strategy, and suggested building data gathering for evaluation into the product.[7]
DfE's June 2026 market assessment counted 1,123 UK edtech companies. It found reliable evidence on how edtech is used in schools still limited, with schools lacking evaluation frameworks and consistent usage data from suppliers, and decisions often resting on word of mouth.[8] The EEF judges the effectiveness of edtech interventions mixed, and smaller and less certain for disadvantaged pupils.[4]
How the EEF judges evidence
EEF evaluations are built to be comparable. Most combine an RCT with an implementation and process evaluation (IPE) and a cost evaluation, follow the sequence below, and are published whatever they find.[9][10]
An EEF trial, step by step
- Developer and evaluatorDescribe the programmeA logic model, an intervention description and an agreed definition of compliance.
- EvaluatorPublish and registerA protocol published before the evaluation starts, and registration on the International Standard Randomised Controlled Trial Number (ISRCTN) registry or the Open Science Framework.
- Developer and schoolsRecruit and agreeMemoranda of understanding, privacy notices and data agreements, then pre-tests.
- EvaluatorRandomiseAllocation revealed only after pre-tests; a statistical analysis plan (SAP) submitted three months later.
- Developer and schoolsDeliverA fixed programme, with use recorded and usual practice surveyed in every arm.
- Evaluator and EEFAnalyse, report and archiveAnalysis by allocated group, a padlock rating, publication whatever the result, and archiving of the linked data.
Everyone randomised is analysed in their allocated group (intention to treat); prior attainment is controlled for; the clustering of pupils in classes and schools is modelled; and effects are reported as Hedges' g, a standardised effect size, with 95% confidence intervals and without labels of statistical significance. Intraclass correlations (ICCs) are reported, and pupils eligible for free school meals (FSM) are always analysed separately.[10] Effects on attainment become months of additional progress: 0.05 to 0.09 is one month, 0.10 to 0.18 two, 0.19 to 0.26 three, and −0.04 to 0.04 none.[15] Each impact evaluation's primary result is rated from zero to five padlocks on three criteria.[11]
| Padlocks | Design | Minimum detectable effect size | Attrition |
|---|---|---|---|
| 5 | Randomised | 0.20 or less | 0–10% |
| 4 | Allows for some unobserved differences, such as regression discontinuity or difference-in-differences | 0.21–0.29 | 11–20% |
| 3 | Allows for all relevant observed differences, such as matching | 0.30–0.39 | 21–30% |
| 2 | Allows for some relevant differences | 0.40–0.49 | 31–40% |
| 1 | Allows for no relevant differences | 0.50–0.59 | 41–50% |
| 0 | No comparison group | 0.60 or more | Over 50% |
The weakest criterion sets the starting rating, and threats such as poor fidelity or missing data can lower it. Pilots are not rated.[11] The IPE explains results: the EEF lists twelve dimensions, including fidelity, adaptation, dosage, reach, responsiveness and programme differentiation, and every trial records usual practice in all arms before and after delivery.[12] Pilots instead ask whether a programme is feasible, shows promise and is ready for a trial, against indicators agreed in advance, such as more than 70% of pupils attending all sessions.[13]
For a trial, the EEF wants prior evidence of promise and a fully developed programme, already delivered in more than 20 settings in England and able to reach 50 in a year. Trials usually take two to three years.[9] Most link pupils to the National Pupil Database (NPD), and the EEF archives the linked data in the Office for National Statistics Secure Research Service (ONS SRS), managed by FFT Education, to track long-term effects.[18]
What recent edtech trials show
Three of these reports appeared between June and September 2026. Results are modest, and the reasons matter to product teams.
| Programme | What was tested | Scale | Result | Lesson |
|---|---|---|---|---|
| Lexia Core5 Reading, second trial (2026)[21] | Adaptive literacy software for struggling Year 2 readers | 224 schools | +2 months (effect size 0.15), 4 padlocks; heavier users tended to progress more | Record dosage against the recommended minimum |
| Maths-Whizz (2026)[22] | Online tutor, recommended for 45 to 60 minutes a week | 63 schools | +1 month (0.08), 4 padlocks; average use 32 minutes a week | Measure actual use |
| DreamBox Reading Plus, first trial (2026)[23] | Adaptive reading, 90 minutes a week in Year 5 | 124 schools, 5,406 pupils | No additional months, 4 padlocks | Know the comparison: many control schools ran other reading interventions |
| Stop and Think, second trial (2025)[24] | Whole-class quiz game in 12-minute sessions | 173 schools, 14,645 pupils | Maths: no additional months; science: +2; 3 padlocks | Attrition of 22.48% cost two padlocks |
| GraphoGame Rime (2018)[25] | Phonics game for pupils with low phonics screening scores | 15 schools, 398 pupils | −1 month, 5 padlocks; found highly engaging | Engagement is not impact |
| ChatGPT in lesson preparation (2024)[26] | Teachers planning Key Stage 3 science with ChatGPT and a guide | 68 schools, 259 teachers | 31% less planning time, which the EEF rates as high security; quality not affected in a blind review | Workload can be measured in a short trial |
Use can fall short of design, and heavier use can go with more progress;[21][22] a trial compares a product with whatever control schools actually do, which may be another intervention;[23][25] and losing pupils lowers the security of an otherwise strong trial.[24] Lab Note 004 reviews the wider evidence on gamification and children's learning.
Choosing a design: units, power and comparison groups
Who or what is randomised
Randomising classes, teachers or year groups within schools risks spillover, as staff and pupils share what they learn; randomising whole schools reduces that risk but needs a much larger sample.[6] The cost of clustering is measured by the ICC, the share of variation in outcomes that lies between schools or classes. In completed English three-level trials reviewed for the EEF, school-level ICCs for attainment ranged from 0.02 to 0.27, and class-level clustering was strongest in secondary maths, where pupils are often set by attainment.[39] A good pre-test reduces the sample needed.[10][39]
Power and the smallest effect worth detecting
Trials are sized so that the smallest effect worth finding, the minimum detectable effect size, would be detected with 80% probability (power) at a 5% significance level, using assumed ICCs and pre-test correlations,[14] and the EEF expects most to detect an effect of 0.2 or whatever would be cost-effective.[11] Real effects are often smaller: across 141 large trials commissioned by the EEF and the US National Center for Educational Evaluation and Regional Assistance, the mean effect was 0.06 standard deviations, with confidence intervals averaging 0.30 wide, so many results were uninformative.[38] Plan for a modest effect, and never promise a large one.
Attrition, contamination and business as usual
Attrition is counted at pupil level from randomisation to analysis,[11] and control schools are the hardest to keep.[17] Contamination, when the control group receives some of the intervention, dilutes the measured effect; Torgerson calculated that allocating whole clusters only reduces the total sample needed when contamination exceeds about 30%.[40] The risk is acute for consumer AI tools that control teachers can simply use.[6] The EEF's Aila trial asked both arms to avoid Oak National Academy's main website and the control arm to avoid Aila,[27] and every EEF trial measures business as usual rather than assuming it.[12]
Waitlists and year-group designs
Offering control schools the product afterwards can ease recruitment,[6] and the EEF's protocol template lists waitlist controls.[14] The Lightning Squad trial randomises each school to deliver the programme in Year 3 or Year 4, so every school implements it.[28]
A/B tests inside the product
Experiments embedded in a learning platform give strong grounds for causal claims,[42] but education adds constraints: a class may need one condition, and a learner a consistent version. The open-source UpGrade platform assigns conditions by class, school or district and resolves conflicts between group and individual consistency.[45] An A/B test shows which version works better, not whether the product beats business as usual, which is the Stage 2 question.
Quasi-experimental designs
A matched comparison design compares users with similar schools drawn from DfE administrative data, but cannot rule out unobserved differences and is limited to outcomes in those data.[6] Regression discontinuity, which compares those just either side of an eligibility cut-off, and difference-in-differences, which compares changes over time, can reach four padlocks; matching can reach three.[11] An interrupted time series compares outcomes before and after a launch; it needs an impact model set in advance and must handle autocorrelation and seasonal patterns,[41] which in schools follow the academic year.
The two stages ask different questions, and a product should serve both.[6][9][10][11][13][14]
| Feature | Rapid evaluation (Stage 1 style) | Randomised trial (Stage 2 style) |
|---|---|---|
| Question | Can it be delivered, and does it show promise? | Did it cause a difference compared with business as usual? |
| Typical scale | Around 25 schools or colleges | 50 to 70 per arm, as a rule of thumb |
| Comparison | Usually none; a comparison group only if justified | Random allocation to a control or waitlist arm |
| Duration | A term, with 6 to 12 weeks of use | Usually a year of delivery; two to three years in all |
| Set in advance | Success indicators, such as usage thresholds | Protocol, SAP and trial registration |
| Product data used for | Uptake and implementation | Compliance, dosage and fidelity analyses |
| Security rating | No padlocks | Zero to five padlocks |
University teams will recognise the ladder: the Office for Students (OfS) standards of evidence, used by the Centre for Transforming Access and Student Outcomes in Higher Education (TASO), separate narrative, empirical and causal evidence as Types 1 to 3.[20]
Logic models, fidelity and dosage
A trial tests a theory of change: how inputs and activities lead to outputs and outcomes. The EEF designs evaluations around the logic model so that a null result can be traced, where possible, to theory, implementation or methodology failure. Its January 2026 guidance uses a PLANIT description (purpose, learners, actors, nature, intervention and tailoring) that records adaptations.[16] Its report template recommends the 12-item TIDieR checklist, which covers what, who, how, where, when and how much, tailoring, modifications and how well.[15][44]
Compliance turns the logic model into numbers. Developer and evaluator agree in advance which inputs make a pupil, class or school compliant; where a programme combines teacher training and software, both count. Compliance can be binary or continuous with thresholds, and it feeds a complier average causal effect (CACE) estimate alongside the main result.[10] Fidelity is broader, covering quality, dosage and adaptation.[12] For software, these measures mostly come from product data, which must be reliable and auditable.
Maths-Whizz pupils averaged 32 minutes a week against a recommended 45 to 60, which the EEF reads as room for further impact;[22] Lexia pupils who exceeded the recommended minimum tended to progress more, a finding based on smaller numbers.[21] Neither conclusion is possible without trustworthy usage data.
Engineering a product for evaluation
Evidence by design means deciding before a trial what the product records, how arms are applied, which version is evaluated and how data reaches the evaluator. Lab Note 001 explains what researchers lose when analytics is an afterthought; this section turns that into a specification.
An event schema the evaluator can use
Our recommended minimum fields are below. The requirements behind them come from EEF analysis and archiving guidance, and from pseudonymisation guidance by the ICO, the UK data protection regulator.[10][18][33]
| Field | Example | Why the evaluator needs it |
|---|---|---|
event_id | A random unique ID | Removes duplicates when devices retry |
occurred_at, received_at | 2026-11-03T10:42:07Z | Orders events and exposes clock errors |
learner_key, class_key, school_key | Keyed tokens | Models clustering without names or pupil numbers |
trial_id, arm, unit | T1, treatment, school | Analyses each pupil in the allocated arm |
app_version, content_version | 3.2.0, unit 4 v7 | Ties outcomes to the version evaluated |
event_type, activity_id | activity_completed | Maps use onto the logic model |
active_ms, visibility | 41000, visible | Measures dosage as active time |
score, max_score, attempt | 7, 10, 2 | Proximal outcomes for the IPE, never the primary outcome |
research_status | active, withdrawn | Keeps withdrawn pupils out of research extracts |
Session time, from login to logout, counts tabs left open. Active time counts only intervals of interaction while the page is visible, which browsers report, with idle gaps capped.[48] The choice matters: in two learning analytics datasets, different estimates of time on task changed the findings, and the method is rarely reported.[43] Log raw timestamps, define the active-time rule in the SAP and let the evaluator recompute it.
Trial arms and a stable version
Apply arms with feature flags at the unit the evaluator randomised. Many flag tools split variants by a targeting key, which the OpenFeature specification defines as the subject of a flag evaluation.[47] In a school-randomised trial that key must be the school, or the product will quietly randomise pupils within schools. EEF trials are designed and randomised by the independent evaluator,[9][14] so import that allocation rather than generating your own, and decide what happens afterwards: UpGrade lets researchers keep, revert or switch conditions when an experiment ends.[45]
Then hold the product still. The EEF evaluates programmes needing no further development[9] and expects adaptations to be recorded.[16] For a product that ships weekly, that means a trial release: features frozen, fixes logged, content versioned and, for generative AI, the model version pinned and logged, because AI tools change faster than research cycles.[6]
Linking to school records
Rosters come from the school's management information system (MIS). DfE treats the unique pupil number (UPN) as personal data and a "blind number" not to be used for purposes unrelated to education; a supplier commissioned by a school may hold it as processor, and researchers may receive it only for work initiated by a school, local authority, DfE or another prescribed government department.[36] Never use it as a product identifier. For NPD matching, evaluators collect names and dates of birth plus a UPN, school identifier or postcode;[10] matched data is analysed and archived inside the ONS SRS,[18] and outputs from DfE data leave it, as a rule, only if counts are 10 or more.[37]
A pipeline that keeps research apart
Pseudonymise at collection. The ICO advises against unkeyed hashes and outdated algorithms such as MD5 and SHA-1, describes keyed hashing and random tokens, and requires the key or mapping table to be kept separately; pseudonymised data stays personal data for whoever holds that key.[33] Keep the research store one-way: since 5 February 2026, research processing needs safeguards including data minimisation measures such as pseudonymisation, and fails them if used for measures or decisions about a particular person, outside approved medical research.[29] Apply small-number suppression to every dashboard, using the strictest threshold among the sources combined.[37]
School and product
Research pipeline
Outputs
Data protection for trials
Roles come first. A supplier acting only on a school's instructions is likely to be its processor; one using pupils' data for research that is not the core service the school bought, or for its own product development, acts as a controller and is likely to fall within the Children's code, whatever the contract says.[30] In EEF trials, evaluators are controllers, delivery teams become joint controllers when they decide with evaluators what data is collected, and the EEF is controller for the archive.[19] In June 2026 the ICO reported audits of 28 edtech providers. Common issues included providers misjudging whether they were controllers or processors, especially where children's data fed product development or analytics, along with thin contracts, incomplete data flow maps and gaps in data protection impact assessments (DPIAs).[31]
Each controller needs a lawful basis. The ICO considers public task unlikely to suit an edtech provider acting as controller,[30] and the EEF archive relies on legitimate interests, with a right to withdraw.[19] The ICO requires a DPIA for data matching, which outcome linkage is.[32] The EEF treats a data sharing agreement between controllers as good practice, and a data processing agreement as a legal requirement where a delivery team acts as processor.[17][35]
The Data (Use and Access) Act 2025 changed the research rules from 5 February 2026. Scientific research explicitly includes commercial research, technological development and applied research; consent can cover an area of research; further processing for research under the new safeguards is compatible with the original purpose but still needs a lawful basis; and new Articles 84A to 84D replace Article 89 of the UK General Data Protection Regulation (UK GDPR).[29] Older documents, including the EEF's data protection statement, still cite Article 89,[19] and the ICO's guidance on research, DPIAs, pseudonymisation and data sharing is under review.[32][33][34][35]
What an evaluation-ready evidence pack contains
Drawing on the Evidence Board criteria, EEF readiness criteria and the sections above, the pack a supplier gives an evaluator and a school before a trial should contain:[7][9][16]
- A logic model and a PLANIT or TIDieR description, showing how the product differs from usual practice
- The research behind the design, and what user research has shown
- Delivery history, typical usage and the implementation support schools receive
- An agreed compliance definition and dosage measure, including the active-time rule
- An event dictionary and sample export with pseudonymous pupil, class and school keys
- A plan for applying the evaluator's allocation at the unit of randomisation
- A trial release plan covering frozen features, logged fixes and pinned AI models
- A data flow map, roles, lawful bases, a DPIA and the necessary agreements
- Rules for withdrawal, retention and small-number suppression
- Evidence on accessibility, data protection, cybersecurity, safety and interoperability
- A claims register linking each impact claim to its evidence
- Costs to schools, staff time and technical requirements
Where EdTechLab stands
We have not run or taken part in an EEF trial or a Testbed evaluation. What our products do today is narrower. EngagedLab produces per-objective evidence reports, with cohort figures suppressed below 5 learners. Its SCORM (Sharable Content Object Reference Model) packages are self-contained and make no external calls, with opt-in evidence return using hashed learning management system (LMS) learner IDs. intle provides per-question difficulty and discrimination analysis for hosted sessions, on higher-tier plans, and its exported packages send no learner data back. The Experience API (xAPI) is not a live feature of our products.
On WorkReady Finance, which we developed for UWE Bristol Business School, UWE Bristol is the data controller and EdTechLab the processor. On a client project aimed at evaluation, we would start from the logic model and the evaluator's analysis plan, agree the event dictionary and compliance definition before launch, apply the evaluator's allocation at the right unit, fix a trial release, pseudonymise at collection, and settle roles, the DPIA and agreements before any pupil data moves. Lab Note 002 explains how we evaluate digital infrastructure for education research.
Limits of this report
- This is general information, not legal advice.
- Programme details reflect publications on 5 October 2026; supplier criteria, Stage 1 methods and costs to schools were unpublished.
- Trial results come from EEF project pages; we did not re-analyse them.
- The ICC figures come from a few English trials and vary by phase and subject.
- The engineering advice is ours, not an EEF or DfE requirement.
- Testbeds also shape what counts as evidence, by making classroom practice visible and actionable in particular ways.[46]
References
Accessed 5 and 6 October 2026.
- Department for Education. EdTech Testbed Programme: guidance (24 September 2026). Source
- Department for Education. EdTech Testbed Programme, publication page (24 September 2026). Source
- Education Endowment Foundation. New partnership with the Department for Education to build the evidence base on the impact of EdTech tools in schools and colleges, press release (7 May 2026). Source
- Education Endowment Foundation. Research Agenda theme: EdTech. Source
- Department for Education. Education Secretary speech at Bett UK Conference (21 January 2026). Source
- Department for Education (research by the Open Innovation Team). Approaches to developing the education technology (EdTech) impact testbed: scoping report (December 2025). Source
- Chedzey K, Evans S, Lindroos Cermakova A. Piloting the EdTech Evidence Board: report on key findings from phase one and phase two. Chartered College of Teaching for the Department for Education (September 2026). Source
- Bolibruchova M, Hugill J, Rayman D, Yitbarek E (PUBLIC). Assessment of the education technology market in England. Department for Education research report RR1638 (June 2026). Source
- Education Endowment Foundation. Funding FAQ. Source
- Education Endowment Foundation. Statistical analysis guidance for EEF evaluations (October 2022). Source
- Education Endowment Foundation. Classification of the security of findings from EEF evaluations, version 2.0 (July 2019). Source
- Education Endowment Foundation. Implementation and process evaluation guidance for EEF evaluations (updated August 2022). Source
- Education Endowment Foundation. Guidance for EEF pilot evaluations (October 2023). Source
- Education Endowment Foundation. Protocol, study plan and SAP templates, including the 2025 trial protocol template. Source
- Education Endowment Foundation. Reporting templates, including the 2025 evaluation report template. Source
- Education Endowment Foundation. Guidance on Theory of Change and logic model development for EEF-funded evaluations (January 2026). Source
- Education Endowment Foundation. Recruitment and Retention Guidance (May 2026 update). Source
- Education Endowment Foundation. Archiving evaluation data from EEF projects (July 2025). Source
- Education Endowment Foundation. Data protection statement regarding EEF evaluations (13 May 2021). Source
- Centre for Transforming Access and Student Outcomes in Higher Education (TASO). How we rate evidence. Source
- Education Endowment Foundation. Lexia Core5 Reading (second trial), evaluation report published September 2026. Source
- Education Endowment Foundation. Maths-Whizz Intelligent Tutoring Programme: trial, evaluation report published June 2026. Source
- Education Endowment Foundation. DreamBox Reading Plus (first trial), evaluation report published September 2026. Source
- Education Endowment Foundation. Stop and Think: Learning Counterintuitive Concepts, second trial, evaluation report published March 2025. Source
- Education Endowment Foundation. GraphoGame Rime: trial, evaluation report published May 2018. Source
- Education Endowment Foundation. ChatGPT in lesson preparation: Teacher Choices trial, evaluation report published December 2024. Source
- Education Endowment Foundation. Lesson planning using AI lesson assistant, Aila: Teacher Choices trial. Source
- Education Endowment Foundation. FFT Tutoring with the Lightning Squad: trial. Source
- Data (Use and Access) Act 2025, sections 67, 68, 71 and 86, in force from 5 February 2026 (S.I. 2026/82). Source
- Information Commissioner's Office. The Children's code and education technologies (edtech) (updated 30 May 2023). Source
- Information Commissioner's Office. Edtech examined and statement on the report (24 June 2026). Source
- Information Commissioner's Office. Data protection impact assessments: examples of processing likely to result in high risk. Source
- Information Commissioner's Office. Pseudonymisation. Source
- Information Commissioner's Office. The research provisions. Source
- Information Commissioner's Office. Data sharing: a code of practice. Source
- Department for Education. Unique pupil numbers (UPNs): a guide for schools and local authorities, version 1.2 (June 2019). Source
- Department for Education. Statistical Disclosure Control Policy for DfE data: Office for National Statistics Secure Research Service (April 2026). Source
- Lortie-Forgues H, Inglis M. Rigorous Large-Scale Educational RCTs Are Often Uninformative: Should We Be Concerned? Educational Researcher, 2019, 48(3), 158–166. DOI
- Demack S. Does the classroom level matter in the design of educational trials? A theoretical & empirical review. EEF Research Paper No. 003 (May 2019). Source
- Torgerson DJ. Contamination in trials: is cluster randomisation the answer? BMJ, 2001, 322(7282), 355–357. DOI
- Lopez Bernal J, Cummins S, Gasparrini A. Interrupted time series regression for the evaluation of public health interventions: a tutorial. International Journal of Epidemiology, 2016, dyw098. DOI
- Motz BA, Carvalho PF, De Leeuw JR, Goldstone RL. Embedding Experiments: Staking Causal Inference in Authentic Educational Contexts. Journal of Learning Analytics, 2018, 5(2). DOI
- Kovanovic V, Gašević D, Dawson S, Joksimovic S, Baker R. Does Time-on-task Estimation Matter? Implications on Validity of Learning Analytics Findings. Journal of Learning Analytics, 2016, 2(3), 81–110. DOI
- Hoffmann TC, Glasziou PP, Boutron I, et al. Better reporting of interventions: template for intervention description and replication (TIDieR) checklist and guide. BMJ, 2014, 348, g1687. DOI
- Ritter S, Murphy A, Fancsali SE, Fitkariwala V, Patel N, Lomas JD. UpGrade: An Open Source Tool to Support A/B Testing in Educational Software. Workshop paper, Learning at Scale 2020 (not a journal article). Source
- Decuypere M, Hartong S. Edtechcraft: investigating the global emergence of edtech testbeds. Learning, Media and Technology, 2026, 1–17. DOI
- OpenFeature. Specification, section 3: Evaluation Context. Source
- WHATWG. HTML Living Standard, section 6.2: Page visibility. Source
Next step
Preparing a product or pilot for evaluation?
We can review your instrumentation, data flows and evidence pack against this report, or design and build them with you.