Rethinking Assessment in the Age of AI
What if AI did not create the problem with traditional assignments, but made an old problem harder to ignore?
For a very long time, higher education has operated on a remarkably simple bargain. We ask students to learn something, then ask them to produce something—a paper, an essay, a discussion post, a problem set—and use the finished product as evidence that the learning occurred.
The arrangement is so familiar that it is easy to forget how much inference it requires.
An instructor who reads a thoughtful research paper does not actually observe the student researching, weighing competing evidence, abandoning a weak argument, discovering a better one, or gradually developing an understanding of the subject. The instructor sees the artifact left behind by those activities. From that artifact, we infer the intellectual process that produced it.
For most of modern higher education, that distinction was manageable enough that we rarely had to dwell on it. Then generative AI arrived.
Suddenly, a student can produce polished prose, summarize difficult material, formulate an argument, reorganize ideas, generate citations, or answer a discussion question with assistance from a machine that can perform many of the visible tasks we once associated with learning. The immediate response has understandably focused on the technology: How much AI use should be permitted? How can inappropriate use be detected? What counts as cheating? Should students disclose prompts? Should instructors redesign assignments to make them “AI-proof”?
These are legitimate questions. But they may not be the most interesting one.
What if AI has not merely created a new assessment problem? What if it has exposed an old one?
The Artifact Is Not the Learning
Consider what a traditional assignment actually tells us.
A student submits an excellent essay explaining why the Wright brothers succeeded where earlier experimenters had failed. We reasonably conclude that the student understands what made the Wrights’ achievement possible. But strictly speaking, the essay demonstrates that an excellent essay now exists. Our judgment about the student’s understanding depends upon assumptions about the process by which it came into existence.
Those assumptions have never been perfectly secure. Students have always received help from tutors, parents, friends, editors, study groups, textbooks, sample essays, spell-checkers, search engines, and countless other resources. They have also always been capable of producing a good paper with shallow understanding—or a mediocre paper despite understanding the material quite well.
Generative AI changes the scale of that problem because it weakens the relationship between the sophistication of the artifact and the sophistication of the intellectual process required of the student to produce it. Experimental evidence has begun to document that divergence. In a 2025 study of 117 university students, Yizhou Fan and colleagues found that students using ChatGPT improved their essay performance but did not demonstrate significantly greater gains in knowledge or transfer.[1]
If the purpose of education were simply to produce competent artifacts, AI would often be an excellent solution. In many professional settings, it already is. But education has a different problem. We care not only whether the answer can be produced, but whether the learner has developed the knowledge and judgment that make the answer meaningful.
The artifact and the learning were never the same thing. AI simply makes it much harder to pretend that they are.
Policing the Product
One response is to restore confidence in the artifact by controlling how it is produced: prohibit or restrict AI, require drafts, compare writing styles, use detection software, move some assignments into supervised classrooms, or ask students to document their process.
Some of these approaches are entirely appropriate in particular circumstances. There are skills students genuinely need to demonstrate independently, just as pilots must sometimes demonstrate that they can perform a maneuver without relying on automation.
But an escalating contest between increasingly capable technology and increasingly elaborate restrictions can only take us so far. It makes the instructor responsible not only for evaluating learning, but for reconstructing the provenance of every artifact submitted as evidence of it.
Early in the semester, a student emailed to tell me that he had inadvertently submitted an AI-generated template he had used to help organize his essay rather than the completed essay he intended to submit. The interesting pedagogical problem was not simply that AI had been involved. It was that the document sitting in the submission inbox told me remarkably little about what the student actually knew, what work the student had done, or even what the student had intended to submit. To discover that, I had to talk to the student.
There is another possibility.
Instead of asking how we can make the finished product trustworthy again, we can ask whether more of the evidence of learning can be moved into the learning process itself.
That question became the organizing principle behind an experiment I have been conducting in the asynchronous aviation history course that I am currently teaching. I offered my college students two pathways through the same historical content. One follows a conventional model: students read, watch lectures, participate in discussions, write historical reflections, and complete other familiar academic assignments. The other pathway embeds readings, primary sources, images, historical analysis, interactive exercises, and knowledge checks into a guided learning sequence.
The important distinction is not digital versus traditional, or easy versus rigorous. It is where the evidence of learning resides.
In the conventional model, students do much of the learning independently and then produce an artifact from which I infer what they learned. Even a recorded lecture marked complete tells me little about what happened while the student watched it. In the guided interactive model, more of the intellectual work becomes visible while the student is actually encountering the material. A student may have to interpret historical evidence before moving forward, commit to an answer before seeing an explanation, or reconsider an initial judgment in light of new information. This does not tell me everything about how deeply a student has learned, but it does give me something a finished paper often cannot: traces of the student’s thinking while that thinking is still developing.
The assignment is no longer merely what comes after the learning. More of the learning process becomes the assignment.
A Different Question About AI
When more of the evidence of learning resides within a deliberately designed learning experience, the question changes. AI no longer has to be treated primarily as a tool whose use must be restricted for the student’s work to count as evidence of learning.
It can become one of many tools in the learner’s environment, while the course design itself carries more of the burden of making learning observable.
Recent experimental research suggests that how AI participates in that learning process may matter. In a 2026 randomized experiment, Zara Contractor and Germán Reyes found that students given access to generative AI demonstrated learning gains that persisted when the technology was removed. Those later gains were stronger among students who used AI to augment their learning—for example, to explain concepts—than among those who used it to automate their work.[2]
This does not make traditional forms of assessment obsolete, nor should it. Research papers, essays, presentations, examinations, and independent projects can require students to synthesize ideas in ways that guided instruction cannot. Nor does interactive instructional design magically solve every problem created by generative AI.
The point is more modest—and perhaps more consequential.
For generations, higher education could often afford to conflate the evidence of learning with the artifact produced after learning. Generative AI has made that inference less comfortable. That may ultimately be useful.
The most productive response to AI in education may therefore be neither enthusiastic adoption nor increasingly sophisticated prohibition. It may be to revisit a more fundamental question: What evidence would actually persuade us that learning occurred?
Once we ask that question seriously, the possibilities become larger than AI policy. We can design courses in which students make more decisions, encounter more evidence, receive more feedback, reveal more misconceptions, and practice more of the intellectual work we actually care about—all during the process of learning rather than only after it.
Perhaps the arrival of generative AI is forcing higher education to examine something we should have examined more closely all along.
We have spent a great deal of time asking whether students really produced the work.
We might spend more time asking whether the work really produced the learning.
Notes:
[1] Yizhou Fan et al., “Beware of Metacognitive Laziness: Effects of Generative Artificial Intelligence on Learning Motivation, Processes, and Performance,” British Journal of Educational Technology 56, no. 2 (2025): 489–530,https://doi.org/10.1111/bjet.13544.
[2] Zara Contractor and Germán Reyes, “Experimental Evidence on the Learning Impact of Generative AI,” IZA Discussion Paper No. 18792 (July 2026).