Due diligence · Education
Enthusiasm is not evidence
Almost every educational intervention shows a positive effect in its first study, conducted by its inventor, on volunteers who knew they were being watched. The interesting question is what survives the second study.
Why education research is unusually easy to get wrong
Education is the discipline where the measurement instrument is a child who wants to please you. Novelty raises engagement; attention from an adult raises performance; teachers who volunteer for a pilot are not the teachers who will be assigned to it later. These are not fringe effects. They are frequently larger than the intervention being tested.
Compounding this, a good deal of the field's most-cited work has not replicated cleanly. Learning-styles matching — the idea that teaching to a student's visual or auditory preference improves learning — remains widely believed by teachers and has repeatedly failed controlled test. Meta-analytic summaries of "what works" are themselves contested on the grounds that averaging heterogeneous studies of varying quality produces a number with no clear referent. A person doing honest diligence has to hold both facts at once: the evidence base is real, and it is weaker than its presentation.
The questions to ask of any educational claim
- Who ran the study, and who paid for it? Vendor-run evaluations of vendor products belong in a separate category from independent replication.
- What was the comparison group actually doing? “Better than nothing” and “better than a good teacher with the same time” are different claims.
- Was the outcome pre-registered, or chosen after the data came in? An unregistered outcome measure is a licence to find something.
- Is the effect size meaningful, or merely statistically detectable? A reliable effect on a twenty-minute quiz is not a claim about a life.
- How long did the effect last? Most measured gains are tested within days. Durability is the whole point and the least measured property.
- Does it replicate outside the inventor's setting, with ordinary staff who did not choose it?
- What is the sample? Twelve motivated children in a university lab is a hypothesis, not a programme.
- What happened to the students it did not work for? Averages hide the children an intervention harmed.
Named cases
Discipline — the standard we borrow
Bloom's two-sigma problem (1984)
Benjamin Bloom reported that one-to-one tutoring with mastery learning moved average students roughly two standard deviations above conventionally taught peers, and framed it explicitly as a problem: find a group method that reaches the same result.
Read carefully, this is a diligence document, not a promise. Bloom named the effect, named the cost, and named the unsolved part. Subsequent scholarship has questioned whether two sigma replicates at that magnitude — which is exactly the sort of correction a claim ought to invite.
Failure — belief without evidence
Learning styles
The proposition that instruction matched to a student's preferred modality improves learning has been tested and has failed the test that matters — the interaction between style and instruction — while remaining one of the most widely held beliefs among practising teachers.
The lesson is not that teachers are credulous. It is that an idea which is intuitive, humane, and flattering to individual difference can survive for decades without support. Our own method is intuitive, humane, and flattering. That is a warning, not a credential.
Contested — read the fine print
Socratic Mind, Georgia Tech
Georgia Tech researchers have developed an AI-based Socratic oral-assessment tool and published early results. It is a university-level assessment instrument, studied with undergraduates.
We cite it as evidence that Socratic questioning can be operationalised with AI. We do not cite it as evidence about five-year-olds, and there is no partnership or affiliation between that work and this Academy.
Open question — our own weakest point
Transfer
Music training reliably improves performance on tasks close to music. Whether it transfers to distant domains — general reasoning, academic attainment — is genuinely disputed, with several well-conducted analyses finding little to no far transfer.
Our curriculum is music-first. If far transfer is the justification, the justification is shaky. We therefore justify music on grounds that do not depend on transfer: it structures time, it trains sustained attention on a task with an audible standard of correctness, and children will do it willingly for longer than they will do almost anything else.
The model we are most often confused with
Live — the closest comparison, and the sharpest distinction
Teacher-free AI schools
A network of private AI-first schools is expanding through 2026 into Oklahoma, Massachusetts, Illinois and elsewhere. Core academics — maths, reading, language, science — are delivered by adaptive software in a compressed two-hour morning window, with the remainder of the day given to workshops and life skills. The adults in the room are described as "guides" rather than certified teachers, and act as motivators and coaches rather than planning lessons or grading. Reported tuition runs roughly $30,000 to $55,000 a year.
It is a commercialized enterprise: a for-profit operation whose growth depends on opening more campuses and enrolling more families, not on publishing outcomes. That is not a criticism; it is the incentive structure. A business that scales by expansion has a structural reason to move before the evidence does.
At roughly $55,000 per year, these schools are also out of reach for the average family. That tuition is not a public option; it is a private purchase for families who can absorb the cost. That does not make it wrong; it makes the question of who gets access inseparable from the question of whether the model works. If the evidence for teacher-free AI education is gathered only from families who can pay $55,000, we will know what it does for that population and nothing else.
There is a sharper question underneath the price. If the academic gains these schools advertise come from the adaptive software — and almost any family can now put that same software in front of a child at home for a fraction of the cost — then what the $50,000 tuition is actually purchasing is not the education but the room, the supervision, and the schedule. That is a child-care service. There is nothing disreputable about child care; it is necessary work and honest work. But it should be named for what it is. A family paying fifty thousand dollars a year is not buying AI instruction that is unavailable at home — they are buying an adult presence and a building, and the AI is the part they could have had for free. The diligence question is whether that distinction is being made honestly to the families paying for it.
This is not a distant comparison. An Alpha campus is launching in Nashville in Fall 2026, K–8, with published tuition of $50,000 a year — the same metropolitan area in which this Academy intends to operate. Its public page advertises that students "crush their academics in 2 hours," that classes rank in the top 1% nationally, and that children learn twice as fast, alongside parent and student testimonial videos. Those are outcome claims, presented as marketing, without linked evidence a parent could check. We note them as advertised claims, not as findings.
We use the word guide too, and we mean something close to the opposite. In this model the software instructs and the adult supports morale. In ours the adult conducts the questioning and the AI is an instrument the adult puts in front of the student for a stated purpose and then puts away. If the human can be removed from the teaching without changing the method, it was never a Socratic method.
We are not aware of independent, peer-reviewed outcome evidence for the teacher-free model at the ages we intend to serve. That is not an accusation; it is the state of the record, and the same absence applies to us. The difference we are asking to be judged on is whether we publish our results when we have them.
Why it matters here
Speed, scale, and the missing protocol
These schools are opening faster than anyone can evaluate them, at price points that select for families able to absorb the risk, in a regulatory environment that has not settled what an AI-delivered curriculum for a young child even is. Lawmakers and educators in the affected communities have raised concerns; those concerns are, so far, ahead of the evidence in both directions.
This is the concrete form of the pace problem: the technology arrived, the model scaled, and the protocol that would tell anyone whether it works has not been written. We are not in a position to slow that down. We are in a position to publish ours first.
The information environment parents actually meet
When a parent types a question about AI and young children into a search engine, the answers that come back are mostly forums — open, anonymous, uncredentialled threads where anyone can post and the most confident reply often has the least behind it. A question like “How is AI changing children's education in 2026?” on Quora collects dozens of replies ranging from earnest personal anecdote to marketing copy, with no way for a reader to sort evidence from opinion.
We link it here for what it is and nothing more: active public discussion, a real-time record of what families are hearing and asking, not a source that can be qualified or quantified. It is valuable as a window into the discourse environment. It is not a finding about what AI does to children, and treating it as one would be precisely the error this page is about.
The diligence point is simple: the conversation a parent meets is louder and faster than the evidence that supports it. Our job is not to join the noise but to say, plainly, what we know, what we do not, and how we intend to narrow the gap.
The public is against it, uses it anyway, and lacks the training to judge it
A paradox shapes the current moment: the general public is largely hostile to AI in education while using it constantly in daily life, and almost no one — parent, teacher, or critic — has been trained to evaluate what the tools do to learning. Opposition without literacy is not caution; it is a posture. Use without literacy is not adoption; it is exposure. Both are happening at once, and neither tells us whether the tool helps a child think.
In August 2026, Google for Education announced a partnership with ISTE+ASCD to provide free Gemini training to all six million U.S. educators. The announcement itself acknowledged the gap: “too many teachers tell us they're being asked to navigate AI without the training to use it effectively.” If six million teachers need training, the literacy deficit is structural, not individual.
One diligence question follows from who provides the training. Google is a vendor of the tools being taught. The modules teach educators to use Gemini and NotebookLM in classrooms — Google's products, in Google's pedagogical frame. That does not make the training worthless; it makes it interested. Corporate programs such as this are well-intended, and we hope they live up to the very best standards. Vendor-provided literacy is better than no literacy, and it is not the same as independent literacy. A teacher trained by a vendor to use that vendor's tool has learned what the vendor believes the tool is for — which may be everything a teacher needs, if the vendor holds itself to the standards it asks of the educators it trains.
The diligence we can hold ourselves to is narrower but more honest: we do not need six million teachers to be wrong for our method to matter. We need to say what we are doing, why, and what would prove us wrong — and to let parents and educators who have no vendor stake judge whether the answer holds up.
How we intend to be tested
Commitment — binding on us
Commitment — binding on us
Hypothesis — untested in our setting
Commitment — binding on us
See also transparency, the pilot, and the plan.
Sources
- 1.Bloom, B. S., “The 2 Sigma Problem: The Search for Methods of Group Instruction as Effective as One-to-One Tutoring,” Educational Researcher 13(6), 1984
- 2.Pashler, McDaniel, Rohrer & Bjork, “Learning Styles: Concepts and Evidence,” Psychological Science in the Public Interest 9(3), 2008
- 3.Georgia Tech, Socratic Mind — AI-supported Socratic oral assessment (project page) — university-level assessment research; no affiliation with this Academy
- 4.Sala & Gobet, “Cognitive and academic benefits of music training with children: A multilevel meta-analysis,” Memory & Cognition, 2020 — finds little evidence of far transfer; cited here against our own position
- 5.Kraus & Chandrasekaran, “Music training for the development of auditory skills,” Nature Reviews Neuroscience 11, 2010
- 6.Alpha School — programme description and 2026 campus expansion — primary source; operator's own description, no affiliation with this Academy
- 7.Alpha Nashville — campus page, tuition and programme claims (Fall 2026) — primary source; operator's own advertising, cited for its claims, not as evidence
- 8.FOX 2, “New private school opening with AI being the teachers”
- 9.The Union Democrat, report on AI-first private schools, guides, and tuition levels
- 10.Quora, “How is AI changing children's education in 2026?” — open discussion thread — cited as active public discussion, not as evidence; not qualifiable or quantifiable
- 11.Google for Education, “Our commitment to make AI training available to all 6 million U.S. educators,” August 2026 — vendor-published initiative; cited for the literacy gap it acknowledges, not as endorsement of the tools