“Which of our disadvantaged pupils are furthest behind in reading, and is the gap closing?”
Imagine asking an AI assistant that question and getting a confident answer in seconds. Some people would think “Wow, that’s amazing!”, and others would be more cautious. AI can be wrong, and it can hallucinate. “How can I trust this?" and "How do you know?” might be better questions to ask.
Schools have no shortage of confidently presented information. What teachers and leaders need is information they can trust, understand and act on. Our mission is to get the right information to the right person at the right time, and over two posts I want to show some of how we go about that with Juniper Intelligence. This first post is about the data we test it on. The second is about the maths underneath the answers.
A quick word on where things stand before I go any further. Assessment is the part of Juniper Intelligence that schools are using today. Attendance, behaviour, safeguarding and the rest are being built and tested in exactly the same way, but they are still in our lab rather than in your school. I will use examples from all of those areas in these posts because the method is the same throughout. So when you read an attendance example, please read it as a look at our workbench, not a description of something you can switch on tomorrow.
Synthetic data is made-up data designed to behave like the real thing. It is not a copy of a school’s records with the names changed. There is no real child behind it.
For Juniper Intelligence we generated complete fictional schools: pupils, staff, forms, attendance marks, behaviour incidents, assessment results, SEND records, interventions and workforce information. We currently use seven different settings, including a small rural primary, a large inner-city primary, a market-town secondary and a coastal secondary serving a community with high levels of disadvantage.
These schools need to be believable. A thousand randomly generated pupils might fill a database, but they would not behave like a school.
So we anchored each setting in published Department for Education data. The DfE’s school and pupil characteristics release gives us national and local patterns for measures including Free School Meals (FSM), English as an Additional Language (EAL), ethnicity and class size. We also use official releases on absence, special educational needs, attainment, suspensions, workforce and EHC plan timescales.
Those figures are a starting point, not a template. The national FSM figure was 25.7% in January 2025. It would be absurd to give every fictional school exactly that proportion. We adjust the baseline for phase, region, community and level of disadvantage so that each setting is recognisable. A fictional rural North Yorkshire primary should not have the same pupil profile as a fictional primary in Tower Hamlets, or the same attendance pattern as a coastal secondary.
The shape of the data matters as much as the average. The DfE’s 2025 Key Stage 2 figures show 47% of disadvantaged pupils meeting the expected standard in reading, writing and maths combined, compared with 69% of other pupils. Its 2025 EHC plan release reports that 46.4% of new plans were issued within 20 weeks during 2024. Patterns like those shape the attainment gaps and the mix of on-time and late plans in our test schools. They are controlled test conditions, not predictions about any real school or child.
The data also has to hang together. Results across phonics, the multiplication tables check, Key Stage 2 and in-school tracking are related to one another rather than a fresh roll of the dice each time. Attendance has a realistic persistent-absence tail. Names and likely home languages fit the fictional communities, using census data.
We include the ordinary mess that anyone who has worked with a real school MIS will recognise: a pupil awaiting a Unique Pupil Number, a mid-year arrival, a blank contact email, siblings sharing an address, phone numbers in three different formats. Clean data is useful for a sales demo. A school’s reality is usually uglier, and we want to be tested against the reality.
A plausible fictional school gives us a stage. It does not give us a plot.
Before Juniper Intelligence can be trusted to surface a meaningful pattern, the analytical tools underneath it have to prove they can find one we know is there. So we deliberately plant clues in the records, the kind of thing a talented data analyst ought to spot given time.
Some examples of what we have planted:
There are hundreds more. These are not labels pasted onto random records. We change the underlying attendance marks, incident participants, dates, measures and notes so the clue is really in the data. The clues also fit the setting. A low-absence school might contain a conspicuous cluster of unauthorised term-time holiday marks. A secondary might have a behaviour hotspot in one maths set. A high-mobility, high-EAL setting might have a group of recent arrivals with incomplete contact details. Between them, the Juniper Education team has hundreds of years of frontline experience in schools, and that experience is what we have distilled into the clues we plant.
Every planted clue is recorded in a separate ground-truth manifest. It says what was planted, where, what signal should be visible and which records are involved. Think of it as the answer sheet.
Juniper Intelligence never sees it. It sees the fictional school’s register, attendance, behaviour, assessment and other permitted records, just as it would see the authorised parts of a real school’s data. The answer sheet sits outside the set of tools available to it. Only afterwards does a separate test harness compare what the analytical tools found with what the manifest says was there.
We sometimes call this the air gap. The system being tested cannot read the answers.
The harness only gives credit when two things are true; the expected signal was recovered, and the supporting evidence resolves to the real records behind it. A plausible-sounding answer with no evidence does not pass.
This tests the analytical layer on its own terms, without relying on an AI to happen upon the right question. In a real conversation, Juniper Intelligence then uses those same tools, choosing the appropriate checks and explaining what they return.
That is the part I want to cover next. In next weeks' part 2, I’ll explain why we insist on calculations first and AI second, how every finding carries a reference back to the records it came from, and what the AI is actually for, once the numbers are in.