Last week, the University of Chicago announced it would provide Claude Enterprise access to its entire community, and Mark Levin, a chemistry professor there, responded in the Maroon with an essay arguing that there are “vanishingly few good places” for AI in the classroom. Faculty who use these tools to interface with students, he writes, “should be embarrassed”. I use such tools for productivity and I know many other professors who do so as well. Should we be embarrassed?
Let me start with what’s not in dispute. Students come to university to learn to think — to read carefully, write clearly, build and test arguments — and AI poses a genuine threat to that. A student who outsources the struggle never builds the capacities the struggle exists to build. On this, Levin and I are in complete agreement. Although I think students need to learn to work with AI, I also support meaningful restrictions (I have even argued that wealthy AI companies should pay for some of the disruption they have caused in higher education—by building massive in-person testing centers around the world).
But Levin’s argument is not about students using AI. His target is faculty: professors using AI as a tool in their own work — writing exams and problem sets, grading, responding to email, anything that interfaces with students. And that, I want to argue, is a different question entirely. Here I disagree, and I think the disagreement matters, because conflating the two questions — what novices should be allowed to skip, and what experts should be allowed to delegate — is a serious confusion running through the current AI and education debate.
Below, I list Levin’s central arguments. They are all wrong. And I add that the reason why I bother writing this piece is because I believe that judicious use of AI tools can help bring a better educational experience for students. So I see attempts by professors like Levin to shame other faculty for using AI, as ultimately, a step backwards in higher education.
1. The quality argument: worse outputs don’t settle anything
Levin’s first move is empirical: LLMs are bad at organic chemistry, students are paying for an artisanal product, and a professor who uses Claude for exams, grading, or email is rolling the Oscar Mayer Wienermobile into a Michelin-starred restaurant (this is his metaphor, presumably equating the University of Chicago to a fine Michelin restaurant, while the scientific technology of Artificial Intelligence is equated to a Wienermobile).
Let’s grant the empirical premise entirely, and not just for organic chemistry. Suppose that for every task — writing a test, drafting comments, answering a routine email — the AI-assisted output is somewhat worse than what the professor would have produced unassisted. Does it follow that professors shouldn’t use it?
No, and here’s why: teaching is not one task. It’s a portfolio of tasks competing for a fixed budget of time and attention, and the right question is never “is the AI version of task X worse?” but “what is the best allocation of my hours across the whole portfolio?”
Suppose writing a problem set from scratch takes me five hours, and an AI-assisted version takes twenty minutes and is, let’s say, 5% worse. Those recovered hours don’t vanish. They can go into office hours, into close mentoring, into actually reading the senior theses, into the research that is — let’s be honest — a large part of why students chose a “Michelin” university in the first place (again, his metaphor). A slightly worse problem set plus five hours of individualized attention may be a strictly better educational product than the artisanal problem set and a professor with no time to talk.
Maybe it isn’t! That’s an empirical question about trade-offs, and it will vary by course, by professor, by task. But Levin’s argument needs the trade-off to never be worth making, and he gives no reason to think that. In fact, he doesn’t even consider that his claim is really one about trade offs.
Notice that his own restaurant analogy makes my point. In a Michelin-starred kitchen, the head chef does not make the salad. The head chef delegates the salad — to someone faster and, yes, probably slightly worse at salads than the chef would be — precisely so that the chef’s attention goes where it matters most. Delegation of lower-stakes tasks in the service of excellence at higher-stakes ones is not a betrayal of the Michelin ethos. It is the Michelin ethos. Nobody walks out of Alinea demanding a refund because Grant Achatz didn’t personally dress the greens.
As for what students “rightly expect”: I suspect that if you actually asked students whether they’d prefer (a) hand-crafted problem sets and a professor with few office hours, or (b) good-enough problem sets made by AI and a professor with time to mentor them closely and answer questions about that problem set, a great many would take (b). Their expectation is for an excellent education, not for made-from-scratch problem sets. Levin must assume students would choose (a), but I seriously doubt that.
2. “If Claude is doing your job, why do students need you?”
This is Levin’s rhetorical centerpiece, and I confess I don’t quite get what the argument is supposed to be. Read literally, it suggests that professors who use LLMs are revealing themselves to be replaceable, and so — what? Should they stop, in order to conceal this fact and keep their jobs? That can’t be right; “professors should provide worse service so that professors remain employed” is not a principle anyone wants to defend out loud, and it certainly isn’t an argument about what’s good for students.
But the deeper problem is that the premise is a strawman. Nobody — literally nobody in this debate — thinks Claude is currently doing the professor’s job. The claim is that Claude is a tool the professor uses while doing the professor’s job. The physician who consults an AI diagnostic system and then exercises clinical judgment about its output is not being replaced by the system; she is using it to be a better physician. The radiologist who uses AI-assisted image screening still reads the scans. The human in the loop is not a vestigial organ.
“If the stethoscope is doing your job, why do patients need you?” has the same form and the same force, which is to say none. Tools that extend a professional’s reach are not evidence that the professional is dispensable. If anything, the opposite: the value of expert judgment goes up when there’s powerful but fallible machinery whose outputs need evaluating. Which brings us to the third argument.
3. The siren song: faculty are not students
Levin’s most interesting argument concedes, arguendo, that the models might get good. Even so, he says, faculty must abstain, because faculty set an intellectual example: “If we cannot bring ourselves to engage with our own material, why should they?” For professor Levin, LLMs are Homeric sirens that atrophy the cognitive muscles needed to use them well. He quotes Berkeley Law’s AI policy and asks how we’d react if the university hired robots to lift weights for student-athletes.
This argument depends on an equivalence between faculty and students that doesn’t hold. The whole case for restricting student AI use — and I think it’s often a good case — is that students are in the process of building cognitive capacities they don’t yet have. The struggle is the point. The Berkeley policy Levin quotes says exactly this: the concern is that students develop the skills needed to deploy and critically assess the technology.
But faculty are not in that position. A professor of organic chemistry has a Ph.D. in the subject. She has engaged with the material for decades; she has not merely mastered the curriculum but extended the field beyond it. When she uses an LLM to draft a problem set, she is not declaring the material unworthy of engagement — she has already engaged with it, at a depth no undergraduate will reach in four years, and her expertise is precisely what lets her evaluate and correct the tool’s output. The relationship between an expert and a tool is categorically different from the relationship between a novice and a shortcut. I think college students are well aware of this difference between them and their instructors.
The weightlifting analogy gives the game away. Robots lifting weights for athletes would indeed be absurd — the athlete’s muscles are the point. But the coach uses all manner of machinery the athletes don’t: film analysis, biomechanical modeling, load-management software. Nobody thinks the coach sets a bad example by using tools, because everyone understands that the coach’s job is different from the athlete’s job.
There’s a residual worry here worth taking seriously — that even experts can deskill with overreliance, the way pilots’ manual flying can erode on autopilot. Fine: that’s a reason for experts to stay deliberate about which tasks they delegate and to keep their hand in. It is not a reason for blanket abstinence, any more than autopilot’s existence is a reason pilots should hand-fly the Atlantic.
4. The pricing argument: speculation plus a loaded metaphor
Finally, Levin argues that LLMs are sold below cost, that prices will eventually rise to reflect true compute costs, and that universities risk getting “hooked” — he reaches for the drug-dealer-free-samples analogy — before the bill comes due.
Note first that this is an argument from a price nobody knows. The University hasn’t disclosed the terms of its deal; the future unit economics of inference are genuinely uncertain (and compute costs per token have been falling, not rising, for any fixed capability level). Building a case for present abstinence on speculation about future prices is weak sauce.
But set that aside and look at the metaphor doing the work. What does “hooked” mean here? If it means addiction — a compulsion that persists against one’s own interests, with withdrawal costs that trap you — then we need an actual mechanism, and Levin offers only the image of atrophied cognitive muscles, which is the previous argument wearing a trench coat. If faculty are experts whose skills predate the tool (see point 3), the lock-in story doesn’t get started: a professor who wrote her own problem sets for fifteen years before 2023 can write them again if the price triples.
If, on the other hand, “hooked” just means customers came to like a product introduced at a low price, then the sinister framing dissolves into a description of ordinary commerce. Introductory pricing followed by market-rate pricing is how streaming services, cloud computing, and for that matter academic journal bundles have always worked. We can dislike it, and universities should certainly negotiate with eyes open and avoid building anything mission-critical atop a single vendor. But “this product might cost more later” is an argument for prudent procurement, not for refusing to discover what the tool is good for. Universities make exactly this kind of bet — with library subscriptions, with lab equipment, with software site licenses — every year.
What’s actually at stake
Levin ends by calling for “a reinvestment in the human experience of teaching and learning,” and on that I agree with him completely. Where we differ is on how we get there. He thinks the human experience is protected by faculty refusing the tools. I think the human experience — the office hours, the mentoring, the seminar where a student’s idea gets taken seriously by an expert who has time to take it seriously — is exactly what the tools, used judiciously, can free professors to do.




"The weightlifting analogy gives the game away. Robots lifting weights for athletes would indeed be absurd — the athlete’s muscles are the point. But the coach uses all manner of machinery the athletes don’t: film analysis, biomechanical modeling, load-management software. Nobody thinks the coach sets a bad example by using tools, because everyone understands that the coach’s job is different from the athlete’s job."
It's even more accurate than that. If the coach wanted to move those weights from one room to another, no one would bat an eye at them using a forklift, even if the coach would also personally use those weights for fitness.
As you lay them out, each Levin’s arguments become *more* convincing, not less.
You present false data here: LLMs present a false sense of efficiency and tasks actually take 20% longer to complete.
Additionally, LLMs indeed create atrophy in your brain’s abilities.
You are a very smart person and a great writer. It’d be sad to see your skills deteriorate in some kind of induced LLM degeneracy.