What Fills the Space Where the Phone Was?

School started back a couple of weeks ago, and the phone argument came back with it — except this time there is real evidence on the table. About two-thirds of states have passed something restricting phones in schools. Many districts went to locking pouches. And we now have a study large enough to say what actually happened.

It is not the result either side was hoping for.

What the numbers say

A team of economists and psychologists looked at tens of thousands of middle and high schools, comparing schools that adopted strict phone policies against schools that didn't. They looked at test scores, attendance, discipline, and surveys of students and teachers.

The policies worked at the thing they literally do. Teachers reported that in-class personal phone use fell from roughly six in ten students to roughly one in ten.

Then, beyond that, mostly nothing. Attendance didn't move. Test scores, on average, didn't move. Students didn't report paying more attention in class. They didn't report less online bullying.

Two things did move. Disciplinary incidents jumped in the first year of a pouch policy — suspensions up roughly sixteen percent — and then the effect faded. And well-being followed a shape rather than a line: it fell sharply the year the pouches arrived, then recovered, and two years on it was positive. The authors are careful about that rebound and so am I. The early dip comes mostly from schools that adopted late, the recovery from schools that adopted early, and with only three years of data a real time path can't be cleanly separated from a difference between cohorts. Averaged across all the years, the well-being estimate is still negative.

The test scores hide one more thing, and it points the wrong way. In middle schools they got slightly worse — combined scores down a fraction, small but statistically real, with math leaning the same direction more weakly. High school math, meanwhile, showed a modest gain.

The authors offer two explanations for that split and flag both as speculative. The first is the one that matters most for anyone who works with kids: younger students have less developed impulse control, so when the phone goes away they are more likely to substitute toward other disruptive behavior. The second is plainer — phone use fell further in high schools than in middle schools, so middle schools may have paid the cost of enforcing the rule without collecting as much of the benefit.

The public argument is about whether the policy worked. The researchers asked a better question.

In setting up their analysis, they say plainly that the academic consequences of taking phones away depend on what students substitute toward. Not on the removal. On the substitution.

That is the whole thing. That is the finding.

Nobody designed the substitution

We ran a national experiment in removing something. In almost none of those schools was the replacement designed. Whatever filled the twelve minutes before the bell, the walk between classes, the empty stretch at lunch — it assembled itself out of whatever happened to be lying around.

I see a small version of this every season on a robotics team. The kids who spend the most time building elaborate fidget toys out of spare parts are, reliably, the kids for whom the competition hasn't become a reason yet — nothing in it has caught them, so nothing in it is pulling their attention. The fidget toy isn't the problem. It's the readout. It's what a good pair of hands does when nothing has claimed them — and if I take it away, all I've made is a kid with nothing in their hands.

That is roughly what the middle school numbers describe, and exactly what the suspension spike is: a year of kids finding something else to do with their hands and their attention, a good deal of it landing them in the office. An undesigned substitution produces mostly nothing, plus a small penalty landing on the kids with the least self-regulation to fall back on. The eleven-year-olds. The ones for whom the phone was a lid on something, and removing the lid without building anything underneath just meant the something came out sideways.

That is not an argument against the policies. It is an argument that the policy was never the intervention — and so at least insufficient. It was the clearing of a space. And a space is not a plan.

Why we reach for the object

There is a pattern underneath this that shows up far outside phones, and once you see it you see it everywhere.

When something is going wrong, we reach for the object. Ban it, confiscate it, block it. It feels like action because it is action — visible, announceable, enforceable by lunchtime.

What we reach for much less often is the design: the arrangement of time, attention, relationship and purpose that made the object attractive in the first place.

A phone in a school is absolutely competing with algebra. Anyone who has taught a lesson while a kid steals glances at a screen under the desk knows the cost, and knows it is real. That is the fight the pouch was built to win — and it won it. In-class use fell from roughly six students in ten to roughly one. That is not a small victory; it is the policy doing exactly what it was designed to do.

And then the test scores didn't move.

So the attention came back and the learning didn't. Getting a kid to look up is not the same as giving them a reason to lean in. And outside the classroom, the stretch of unstructured time in which a thirteen-year-old has nothing to belong to is still there. It is just emptier.

Anyone who has monitored group work knows the shape of this. You circulate, you catch the drift, you redirect — and you are always doing it after the fact, because that is what monitoring is. The redirect works. It buys back four minutes. What it cannot do is answer what those four minutes are for, and that question had to be settled before the period started, in the design of the work itself. I got very good at the redirect over twenty-two years. I was much slower to admit that every one I performed cleanly was evidence I hadn't built something compelling enough to make it unnecessary.

And I want to be fair to teachers here, because the constraint usually isn't skill or will. Building the compelling thing takes discretionary time and room to spend it, and both have been getting scarcer for a long while. When the week is fully spoken for, the redirect isn't chosen; it's the only move that fits.

The experiment where someone designed it

Here is why I'm confident this is the right read, rather than just a nice-sounding one.

A research team ran the experiment the phone studies couldn't. Nearly a thousand high school students, four sessions covering about fifteen percent of a semester's math. Three groups. One got a standard chatbot interface. One got the same underlying model behind guardrails — give hints, never give the answer, ask the student what they've tried first — with every problem's prompt carrying a worked solution and the common mistakes to watch for, supplied by math teachers from the school. One got textbooks and notes, as usual.

Both AI groups did much better on the practice problems — the standard interface by about half again, the guardrailed one by more than double. Then the laptops closed and everyone took the same exam alone.

The students who'd had the standard chatbot scored about seventeen percent below the students who'd never had it at all. The students who'd had the guardrailed version scored the same as the control group. No harm.

Same model. Same students. Same problems. The only variable was the design, and the design was the difference between a seventeen-point deficit and none.

Notice what that design ran on. Not the company's defaults, not a district policy. For every problem, the prompt carried a worked solution, the mistakes students actually make on it, and the hint to give for each — written by math teachers from that school, working from what they knew about where kids get stuck. The variable that eliminated the harm was teacher knowledge, applied to a tool — which is exactly the thing we have been steadily taking away.

The researchers open their paper with an analogy I did not expect to find there. They point to aviation: regulators have advised pilots to hand-fly more, precisely so the skill is still there on the day the automation isn't. Autopilot is superb until the moment you need to actually fly the plane. That is their framing, not mine, and it is the clearest statement of the problem I have read from anyone.

A second team, working in university physics, built an AI tutor on the same principle: one step at a time, no full solutions, make them try first. The design worked there too. Two research groups, different subjects, different countries, different students, converging independently on the same guardrail. That is about as close to a design rule as education research gets.

And now we're doing it again

In January, Denver Public Schools blocked student access to ChatGPT across a district of eighty-nine thousand students.

I want to be precise about this, because the district's reasoning was more careful than the headlines suggested. The stated concerns were specific: a new twenty-person group chat feature, the platform's announced move to permit adult content, and the fact that the company had no data privacy agreement with the district. Those are real, concrete, and squarely a district's job. The district also did not ban AI — it named Google Gemini as its approved tool and pointed students at it.

The district substituted. It just substituted at the level of vendor. Swap one chatbot for another and you have changed which company holds the data — which matters, and I'm not being glib about it — but you have not changed what happens when a fourteen-year-old asks a machine to write the paragraph. The guardrail in that experiment wasn't a procurement decision. It ran on what teachers knew about the exact problems their students were working — the solutions, the common wrong turns, the hint to give for each.

That is available. It costs a conversation with the people who teach the subject. It is nobody's default.

Except AI is not phones

There's a reason to be more careful here rather than less.

In that same math study, the researchers checked how often the ungated chatbot actually got the problems right. It was correct about half the time. Roughly four in ten answers used the wrong method outright.

Here's the part that stopped me. Those wrong methods left no trace on the exam. If students had been reading the bad reasoning, you'd see it surface later as bad reasoning. It didn't. The authors' conclusion is that the students weren't reading it at all — they were copying it. Half the thinking in front of them was wrong, and nobody noticed, because nobody was looking.

And they couldn't feel it. Asked afterward, the students who'd been quietly harmed didn't think they'd learned less or done worse. There is no internal signal. That is the difference between this and a calculator. You can tell when a calculator gives you a ridiculous answer, because you learned place value somewhere else. The capacity to check the machine came from outside the machine. But you cannot develop judgment by outsourcing judgment. The competence you'd need to evaluate the output is the competence the tool just performed for you.

I should be honest that the study doesn't prove this for writing. The authors say so themselves — they could measure math because math has objective answers, and that isn't available in writing. The extension is my argument, not their finding. But it's the extension I'd bet on, because long division was a proxy for understanding and the essay never was. The essay is the training mechanism. You learn to think by writing the thing.

What the alternative actually asks

I'm not arguing against restriction. Restriction buys time, and time is worth buying.

But the evidence now says plainly that removal alone buys you very little, and costs the youngest kids something. The variable that decides the outcome is the one almost nobody schedules: what goes in the space.

Years ago I ran a renewable energy project that straddled spring break. For the first few days I kept redirecting a student — I'll call her Sam — whose group was researching wind turbines. She was content to let others lead and to take whatever was left over.

Over the break she went to Baja. Then, in a park, a ranger was explaining how a warming climate is pulling migratory birds out of sync with the food they arrive expecting to find — species like the Western tanager — and while he was talking, a tanager landed on her shoulder.

Something in that moment claimed her. She came back with a question she had not had before: how do you weigh birds killed by turbines against birds lost to a warming climate? That is a real question. It has numbers on both sides and no comfortable answer, and for the rest of the project it was hers. The birds had become a reason to care about the turbines.

I want to be careful about what I'm claiming. I didn't arrange the tanager. What the project did was be about the real world, rich and complex — and consequential. Design doesn't guarantee the moment for any particular kid. It decides whether there's anywhere for the moment to go.

Here is the practical version of all this, the part I'd hand to any teacher or coach: you can't predict what will grab which kid, so open more doors. The same unit can be entered through the physics, the money, the ethics, the birds, or the sheer pleasure of figuring the thing out. One entry point is a bet on a single kind of kid. Five is a design.

The phones will keep being an argument. AI is already a bigger one. We'll be tempted to settle it the same way, with a block announced on a Tuesday that moves the behavior somewhere we can't see it.

We just ran the experiment. The finding was that removal, by itself, buys almost nothing — and that where somebody sat down and designed what came next, the harm disappeared.

Something always fills the space. We can choose what.

Previous
Previous

Walk Nearby First

Next
Next

What Has to Exist Before a Goal Can Be Real?