Evidence & limits

What is known, what is not, and what nobody should claim

A professional development provider that admits uncertainty is more use to you than one that promises transformation. This page states the evidence position as we understand it, including where it is thin, because your decisions should be made on what is actually known.

Last reviewed: August 2026

What the evidence supports

Students are already using these tools

Every survey of secondary and older primary pupils points the same way: a large majority have used a chatbot for schoolwork, and most teachers underestimate how often. The exact percentages vary by study and by age group, and the surveys rely on self-reporting, but the direction is not in doubt. Policy written on the assumption of non-use is policy written for a school that does not exist.

The tools fabricate, fluently

Invented references, misattributed quotations and confident arithmetic errors are documented across every major model, including the most recent ones. The rate has fallen with each generation and has not reached zero. Any classroom use that treats output as authoritative rather than as a draft to be checked is built on a false premise, and this is unlikely to change soon.

Drafting assistance saves teachers time

Where teachers use the tools for first drafts of planning, differentiation and routine feedback, and check the output properly, time savings are real and consistently reported. The savings concentrate on drafting tasks. They do not extend to judgement: assessing a child, writing a reference, or anything safeguarding-adjacent remains human work with human accountability.

Task design beats policing

Schools that redesign assessment so the AI’s answer becomes material to be marked, corrected and argued with report fewer misuse disputes than schools that rely on detection and prohibition. This is practitioner evidence rather than controlled trial data, and we are honest about that, but it is consistent, and it matches what teachers tell us term after term.

The honest half

Where the evidence is thin

These are the claims you will hear from vendors and conference speakers that the research does not currently support. If we ever make one of them in a session, your staff should challenge us on it.

Detectors do not work well enough to accuse anyone

Independent testing consistently finds meaningful false-positive rates: human writing flagged as machine-generated, with the burden falling hardest on students writing in a second language and on the most formal, careful writers. Vendors publish accuracy figures measured against the very models they test on, which flatter the tools. No detector result, on its own, is a defensible basis for a disciplinary decision, and any provider telling you otherwise is selling something.

There is no good longitudinal attainment data

Nobody yet knows what regular AI use does to attainment over years. The tools have not existed in classrooms long enough for cohort studies, and the early studies mix enthusiastic volunteers, short timescales and outcome measures chosen by people with something to prove. Anyone claiming proven attainment gains, in either direction, is ahead of the evidence.

The effect on weaker writers is genuinely unclear

One reasonable hypothesis says drafting support frees weak writers to focus on ideas. An equally reasonable one says it removes the practice that makes writers strong. Small studies point both ways. This is why the programme treats the tools as material to be criticised rather than as a ghostwriter: that position stays defensible whichever way the evidence eventually lands.

The tools will change again before your policy review

Every capability statement on this page has a shelf life. The models improve, the prices change, the terms of service are rewritten. A policy that names products expires; a policy built on principles, and reviewed annually, does not. That is why session six drafts principles and the annual contract includes a review.

How we keep this page honest

We review this page termly against the published research and against what we see in schools. Where a claim rests on practitioner evidence rather than trials, we say so, as above. Where we have changed our minds, the change is noted here rather than quietly edited away.

If you have read a study that contradicts something on this page, send it to enquiries@skylighteducation.com. The best corrections we have received came from sceptical deputy heads, and long may that continue.

What this means for a buying decision

Buy training that builds judgement and produces documents, because those hold their value as the tools change. Do not buy a course about a named product, a detector subscription, or anything whose pitch requires the evidence to be stronger than it is. That test rules out much of the market, and we are content to be measured by it.

Read the page for leaders or book a call.