A couple of days ago, Stanford University released a study that may be surprising to many. In a blind, head-to-head test, law professors were shown answers to student questions—some written by AI, some written by fellow professors—and asked which were better, without knowing the source. The AI answers won roughly three out of four matchups. More striking still: the professors flagged the human-written answers as potentially misleading or harmful about three times as often as the AI ones. And this was law—a domain the authors deliberately chose because it rewards judgment and the weighing of competing arguments, not because it has tidy right answers.
Here’s what I claim:
In many cases, AI can produce feedback on student work that is genuinely better—better at finding the real problems in a paper, better at suggesting how to fix them. In these cases, professors not only may use it, they probably should.
Let me be clear. I’m not defending lazy automation or rubber-stamping whatever the model spits out. I’m defending the use of a tool that, under the right conditions, outperforms what I can do alone in the time I have. This is the key. Many professors teach hundreds of students. In these cases, AI can do better than what those professors can do in the time they have. If this is not true now, it will be true very soon as AI improves.
AI can give good feedback
The Stanford result is one data point, but it doesn’t stand alone. These systems are now doing things at the frontier of human knowledge. An internal OpenAI model recently produced a proof of an 80-year-old open problem in geometry that working mathematicians verified—with Fields Medalist Tim Gowers noting no previous machine-generated proof had come close. (Earlier, looser claims about AI cracking unsolved problems were rightly walked back, so I’m pointing only at the verified case.) Seting aside the hype, this is just common sense: a system that can contribute at the edge of mathematical research is not going to be defeated by a stack of History 101 papers on the causes of the French Revolution.
That asymmetry is the whole point. If the hard cases are within reach, the ordinary cases—the freshman essay, the problem set, the close reading of an assigned text—are squarely in range. And grading is, for the most part, ordinary cases.
I practice what I’m preaching here. Here is what this looks like in my own teaching. In a philosophy course of sixty students, I’ve used AI to generate roughly thousand-word comment sets on student work. I don’t know many professors that are writing 1,000 word comments on 5-page papers in classes that have over 60 students. On the whole, the comments are great. I read every one. I cut the bad sections—and there are bad ones—and I add my own where I can do better. I read the paper, the model reads it too; I am the editor and the final authority. It works best when the assignment is tightly tethered to a text I can put in front of the model, and worst when the prompt is wide open. That’s a real limitation that I can talk about on a different post. But within its zone, the output is feedback I’m proud to hand back, delivered at a scale and speed I could never match by hand. (I grant that when I first did this, it was an emotionally difficult moment. But reason had to win out.)
So the question is no longer can it. The question is should we. Let me consider some objections.
Objection 1: “Students pay for a professor’s comments, not a machine’s.”
Is that true? Is it written anywhere—in a syllabus, a catalog, a contract—that the marginal notes must be composed, word by word, by a human holding a credential? I’ve never seen such a clause. What students are owed is good feedback that a qualified professional stands behind. That’s the actual promise.
Consider your doctor. When you walk into a clinic, are you paying for the personal, unaided opinion of one human being—even when that opinion is worse than what the tools can produce? Or are you paying for the best available care, delivered by a professional who is accountable for it? Modern medicine runs on statistics, imaging, diagnostic algorithms, and increasingly on machine and AI systems, with the physician in the loop: interpreting, overriding, taking responsibility. We don’t feel cheated that a radiologist used software. We’d feel cheated if they refused to. We don’t say “Please, do not use that computer to detect tumors, use your eyes. I pay for your eyes!”
That’s exactly the structure I’m describing. The professor stays in the loop—reading, correcting, vouching. What you’re paying for is a professional’s guarantee that the comments are right, not a guarantee that a human typed every keystroke. The credential certifies the judgment, not the typing or the concatenation of words.
Objection 2: “If students can’t use AI, professors can’t either. Double standard!”
To be honest, I don’t have patience for this objection. It treats teaching as a symmetrical contest of equals in which whatever is forbidden to one side must be forbidden to the other. But students and teachers don’t have the same role, and almost nothing about education is symmetrical.
We never apply this rule anywhere else. A professor may use an answer key; the student may not. A math instructor uses a calculator—or a script—to check computations the students were required to do by hand. No one cries foul, because the student’s task is to learn to compute and the teacher’s task is to evaluate accurately. Different jobs, different tools.
Return to the doctor for a moment. If your physician tells you that exercise is good for you, do you demand that the physician also exercise before you’ll accept the advice? It would be nice. It is not required. The advice is good or bad on its own terms because of the science; the doctor’s personal habits are beside the point. Likewise, the prohibition on student AI use exists to protect the student’s learning. The professor’s use of AI exists to improve the evaluation. Conflating the two is an adversarial reflex, not an argument.
Objection 3: “Real comments track the student’s whole trajectory—what they struggled with last time.”
Some people assume AI can’t overcome this worry. But it’s easily answered. Tailoring feedback to a student’s history is precisely the kind of thing these systems do well when you give them the history. Put the student’s earlier work, your earlier comments, the recurring errors, into the context, and the model will connect this week’s paper to last month’s mistake more reliably and more patiently than a tired professor grading the fiftieth essay at midnight. The trajectory objection assumes the AI is starting fresh each time. It doesn’t have to. You decide what it remembers.
Objection 4: The Human Relationship
The strongest objection isn’t about contracts or fairness or personalization. It’s this: grading is one of the few sustained, individual points of contact between a professor and a student. Sit with someone’s writing long enough and you come to know their mind. Outsource that, and you may quietly sever a thread that mattered—the slow, accumulating relationship that is, for many students, the real education.
And this premise isn’t sentimental. It’s supported by evidence. Gallup’s survey of more than 30,000 graduates found that those who recalled a professor who cared about them as a person, made them excited about learning, and mentored them toward their goals had roughly double the odds of being engaged at work and thriving in life years later—an effect that outweighed whether or not they attended a regional school. Decades of research in the Pascarella and Astin tradition point the same way: quality faculty contact predicts gains in critical thinking, persistence, and intellectual development. And it turns out to be the relationship that matters, not the procedural contact. The thing that does the most good is also the thing we are scarcest with. That is precisely the resource I’m trying to protect.
I’m talking about real human mentorship.
I think it’s the heart of the matter. But notice it’s not an argument that the comments are worse. It’s an argument about relationship, and that points to the solution rather than against it.
Here is my answer: spend the time you save on the part a machine cannot do. The hours I no longer pour into composing comments from scratch don’t go away—they convert. Into office hours that aren’t rushed. Into actually learning students’ names and what they care about. Into the conversation after class that turns a course into a turning point. The relationship wasn’t built by the marginal notes anyways. It’s built by presence. If AI handles the mechanical labor of high-quality feedback so that I can be more human with my students, not less, then we haven’t lost the relationship. We’ve finally got time for it. And a big bonus: professors rarely enjoy grading, they prefer real human mentorship.
That’s the bargain I’m proposing. Not professors replaced by machines, but professors freed by them—on the condition, always, that the feedback is better, and that a real professional remains in the loop to make sure it is.
But won’t professors just dial it in?
My whole case rests on professors converting the saved hours into mentorship. But nothing forces them to. Some will simply pocket the time. Let the model write the comments, never read them, skip the office hours, coast. If AI makes it possible to do the human part of teaching better, it also makes it possible to skip the human part entirely and hide behind competent-looking output. The lazy professor is now a more dangerous figure than before, because the laziness is harder to see.
I think this will happen. What I don’t think is that it will last—and the reason is the market, working through what we measure.
Consider what a professor’s perceived value currently rests on. For the teaching part of the job (I will write another post about AI and research), a large share of it is tied up in exactly the things AI is about to make cheap and abundant: producing clear, competent feedback; delivering a lecture that could just as easily be a recording; answering the routine question. When a machine can do those at a quality floor that everyone can reach, they stop distinguishing anyone. Everyone is the same. And what’s left—the only thing left—is the part that can’t be commoditized: the relationship, the mentorship, the judgment, the human presence that we already know does the real pedagogical work.
So the basis on which we judge teaching has to shift, and it will. Right now teaching evaluations are blunt instruments, often measuring (at best) little more than clarity, workload, and likability. That made a rough kind of sense when those were the scarce goods. But when the routine teaching tasks are essentially free, students choosing courses, departments reviewing faculty, and institutions competing for enrollment will all start rewarding the thing that’s actually scarce—did this professor mentor me, did they know me, did they change where I was going. The professor who outsources both the comments and the connection will, in that world, have nothing left to offer that a cheaper tool can’t already provide. The dial-it-in strategy stops being invisible and starts being the whole story.
I won’t oversell this. Markets in higher education are slow, tenure has great benefits but it also insulates people, and evaluation regimes are notoriously sticky and easy to game. The correction will take years, not semesters. But the direction is not in doubt: once the routine work is free, you get judged on the irreplaceable work. The professors who saw that early will look, in hindsight, like the ones who understood what the job was always actually for.
What this means for the future of higher ed
If I’m right—and if this is the direction we’re inevitably headed—then AI will, ironically, make room for more human connection and more fulfilling college experiences, not less.
But I have no illusions about how some institutions will respond. There will be administrators who seize on this new productivity and squander it, either by hiring fewer professors or by packing more students into each one’s load. That is exactly the wrong move, for every reason I’ve given here. The way to really benefit from a productivity gain is to give the individual student a better education—real mentorship, a real relationship with a professor who has the time to offer it. Spending the windfall on bigger classes and thinner faculty throws away the most important thing.
The right move is the opposite: use the gains to hire more professors, precisely because each one can now do so much more. 1 My guess is that some colleges will take the first route and others the second—and that choice, more than any ranking or endowment, is what will separate the great colleges from the mediocre ones or the ones that will be left in the dust.
1When the automated teller machine arrived in the 1970s, the obvious prediction was that a machine built to do the teller's job would eliminate tellers. The opposite happened. The economist James Bessen has documented that ATMs cut the staff needed to run a branch from about twenty to about thirteen, which made branches cheap enough to open in far greater numbers—urban branches grew by roughly 40%—so total teller employment actually rose, even as each branch employed fewer people. Crucially, the work changed: freed from routine cash handling, tellers shifted toward relationship banking and the more complex, human side of the job. The same pattern showed up a century earlier in textiles, where power looms automated nearly all the labor of weaving yet the number of weavers kept climbing for decades, because cheaper cloth made people buy so much more of it. The lesson is not that automation always expands a workforce—when productivity gains are sudden and total, machines can simply substitute for people (mobile banking, not the ATM, is what finally shrank the teller ranks). The lesson is that whether a technology multiplies a profession or hollows it out depends on, among other things, how its gains are spent.




This post was written by AI…
As a former English teacher, I have to say….Unless people have experienced what it’s like to trudge through 80+ student essays (or worse, research papers! 😂) and try to write quality comments throughout each one, I’m not sure they can appreciate how truly revolutionary this might be!
They also might not appreciate how much easier it is to read the essays than it is to write all the feedback. It’s totally believable to me that you are able to read the papers and edit/approve all the comments in a high-quality way.
I’m sure there might be pitfalls, as you say, but I’m much more excited about the possibilities! ✨🩵