“Why can’t students just use ChatGPT to mark their work?” It’s a reasonable question, and one we’ve been asked more than once since launching the AI marking in Practical Writing. If a general-purpose AI can write a sonnet, plan a wedding and diagnose a rash, surely it can mark an intermediate learner’s essay?
Yes and no. It can have a go. Whether it does the job well is a different question, and worth exploring.
Here’s the analogy I keep coming back to. If you had a friend who had read every book in every library in the world, they could tell you a great deal about almost anything. Medicine, plumbing, tax law, the offside rule. But you wouldn’t want them performing your surgery, and you wouldn’t want them wiring your house. Being extraordinarily well-read is not the same thing as being trained to do a specific job to a specific standard.
That’s the difference between ChatGPT and the AI in Practical Writing: it’s the same underlying technology, but with a very different level of preparation for the specific task.
1. The generalist doesn’t know what to look for
Ask ChatGPT to mark my essay and it will produce something – but with no criteria, no level, and no information about the task, the response is essentially its own choice of what to notice and how strictly to judge. Feedback varies in emphasis from one session to the next. Grades, if you ask for them, can vary too. There’s no consistent measure being applied, because you haven’t given it one.
The AI in Practical Writing was built on those criteria. It marks against CEFR descriptors – for task achievement, organisation, vocabulary and grammar – and applies them consistently, script after script.
2. The generalist doesn’t know the task
Marking writing, of course, isn’t just about the grammar, the vocabulary and the way the ideas are organised. It’s also about communicative competence. A for-and-against essay and a descriptive one call for different things. A formal email and a friendly one call for different things. The CEFR ‘can do’ statements for a task at B1 aren’t going to be the same as the ones at C1.
The AI in Practical Writing has been given criteria for every task in the program – what the student is being asked to do, what skills the task is designed to develop, what success looks like at each level. It isn’t guessing at the target.
3. The generalist hasn’t been tested
It’s one thing to build an AI marker; it’s quite another to build one whose grades you can trust. We calibrated Practical Writing’s marker against experienced human examiners, who graded the same scripts independently. The Practical Writing AI’s scores sat squarely within the range of the human scores, and were statistically indistinguishable.
You can’t say that about ChatGPT, because nobody has run the equivalent test with it. And if they did, they couldn’t repeat it – next month’s ChatGPT won’t be quite the same as this month’s.
In short
There’s a lot a bright generalist can do. But marking student writing is a specialist job, and students deserve feedback from something that has been built, and checked, for that job. That’s what you’ll find in Practical Writing.
