AI Research & Development

How We Built Juniper Intelligence - Part 2: Calculations First, AI Second

George Gardiner
George Gardiner Sep 10, 2026, 9:00:02 AM 3 min read

Introduction

In Part 1, I described the fictional schools we built to test Juniper Intelligence, the clues we plant in their data, and the mark scheme the system is never allowed to see. This post is about what happens when a real question comes in, and why the answer is grounded in arithmetic before any AI gets involved.

The same caveat as last time applies. Assessment is what schools are using today. The attendance, behaviour and safeguarding examples below come from the same design and the same testing, but those areas are still with us, rather than with you.

Calculations First, AI Second

This is the central design decision in Juniper Intelligence. We don't hand thousands of raw rows to a large language model and ask it to “spot something interesting”. Language models are probabilistic. They produce likely language, so their wording can vary and a plausible answer can still be wrong. They are excellent with words, but far less dependable at counting, applying statutory rules consistently, or resisting the temptation to turn a weak pattern into a confident claim.

Instead, Juniper Intelligence has a set of named, read-only analytical tools. These are reliable, tested calculators designed around the questions schools actually ask.
They can:

  • Compare attendance or attainment between cohorts
  • Find concentrations by pupil, form, teaching group, year group, location or day
  • Identify a change in a trend and the point at which it shifted
  • Find linked absence across a household, or an unusual cluster on one date
  • Compare an intervention’s recorded starting point with its exit measure

There are many more, and they all share one property; the same records and the same question produce the same result every time. The methods are ordinary, inspectable statistics and rules implemented in software, never a generative model’s best guess.

“Deterministic” doesn't mean pretending uncertainty has gone away. Some comparisons use statistical tests, and those tests describe uncertainty. The difference is that the calculation follows a fixed, reviewable method. Juniper Intelligence is given the cohort size, the size of the gap and the relevant caveats. It's not invited to manufacture a conclusion. Very small groups are suppressed, modest samples are labelled as indicative, and a broad scan adjusts for the risk of finding apparent patterns simply because it tried lots of comparisons.

We make a similar distinction for written records. For an auditable check, a fixed, versioned watchlist can find exact words or phrases and account for negation, so “no concern about bruising” isn't treated the same way as “concern about bruising”. An optional meaning-based search can help locate potentially relevant passages, but it returns the original words and their source. It doesn't generate a quotation, and it never uses a model’s summary as evidence.

Show Your Working

When one of those tools returns a finding, it also returns a tamper-evident evidence reference. That reference points to the precise pupils, attendance values, incidents, plans or text passages used in the calculation.

If the underlying data changes, the old reference stops working and the analysis has to be run again. If an evidence reference has been made up, it will not resolve. For a quotation, the system checks the exact passage against the source record rather than trusting that the AI remembered it correctly.

Permissions are applied before analysis, too. Juniper Intelligence can only use the information the member of staff is entitled to see. If a role can't access SEND data, that data is never placed in their analytical view, and Juniper Intelligence has to say “I cannot see SEND data” rather than wrongly implying that no SEND need exists.

This is a much stronger foundation than a line under every answer saying, in effect, “AI can make mistakes”. We still say that professional judgement matters, and it absolutely does, but we also design the system so that its factual claims can be challenged and checked.

So, What Does the AI Actually Do?

It does the part AI is good at. Once the analytical tools have returned grounded findings, Juniper Intelligence can explain them in clear language, connect related information and suggest sensible next questions. It might say that an attainment gap is worth exploring alongside SEND or mobility, suggest looking at the pupils behind a hotspot, or help a leader turn several findings into an agenda for a pastoral meeting.

Those are suggestions, not decisions. A pattern doesn't explain its own cause, and a data point never knows the whole child. Teachers, pastoral staff, SENCOs, DSLs and leaders bring the context, relationships and professional judgement that software cannot.

Trust Is Earned

Schools make decisions that matter and an interesting answer isn't enough. It needs to be based on the right records, calculated consistently, honest about its limits and open to professional challenge.

Our synthetic schools let us test all of that before anything reaches a real school. We know the schools are fictional, we know which clues we planted and we keep the mark scheme away from the system under test. Then we check whether the analytical tools can find the clues and show the records that prove it. Assessment has been through that process and is in schools now. Attendance and the other areas are going through it as I write, and we will talk about them properly when they are ready for you to use.

Our aim is to get the right information, to the right person, at the right time, with a clear route back to the evidence, while leaving the important judgement where it belongs.

Don't forget to share this post!

George Gardiner
George Gardiner
Chief Technology Officer at Juniper Education