Where marking time goes

Break down the marking of one essay or exam paper and a teacher's time falls into four parts: reading and understanding each script, marking each item against the rubric, writing comments, and tallying and recording. Marking against the rubric and writing repetitive comments take up the most, and they are the most mechanical. The same mistakes recur across a class, and a teacher often writes the same comment for the thirtieth time. These are the two parts that AI can genuinely take over.

Where AI marking really saves time

  • Item-by-item first marking: the system gives a first mark for each item against the rubric, so the teacher moves from marking from scratch to reviewing and correcting.
  • Repetitive comments: the system generates the explanations for common mistakes, and the teacher only handles the parts that need a personal touch.
  • Tallying and analysis: the class's item-by-item performance and the spread of common mistakes are collated automatically. Doing this by hand takes the longest, yet it is the information with the most teaching value.
  • Following up on several rounds of revision: the system gives the first mark on each resubmission, so one draft with one round of feedback becomes several drafts with several rounds, without adding to the teacher's load.

The parts you cannot, and should not, save

The final decision has to sit with the teacher. Every mark that affects a student should be confirmed by a teacher. In a well-designed system the review is quick: each mark comes with the candidate's own words attached, so the teacher can see at a glance whether to agree. A poorly designed one gives only a total with no reasons, which forces the teacher to mark it all again, and the time goes up rather than down.

Two other things cannot be handed to AI: judging where an individual student stands, where the AI sees this one script and the teacher sees a year of this student's progress, and high-stakes assessment, the final call on internal exams and banding assessments. The right division of labour is AI for the volume, teachers for the judgement.

The right way to run a school pilot

Run the pilot at the scale of one panel over one term. Pick a panel with a heavy marking load that is willing to try, and Chinese or English writing is a common starting point, set a clear scope such as Form 4 argumentative essays, and keep normal marking standards through the pilot. Avoid both extremes. Roll it out across the whole school at once and training and support fall behind. Let one or two teachers dabble and there is too little data to make a purchasing decision.

During the pilot, do two things in parallel: collect teacher feedback every two weeks, where a ten-minute meeting is enough, and record usage data. By the end of the term the panel has enough evidence to decide whether to expand, adjust or drop it.

How to measure the time saved

The method is simple, but it has to start before the trial. Ask the teachers taking part to record a baseline first: how long, on average, it takes to mark one assignment of a set type, timing three to five scripts and taking the average. After a month of use, measure again the same way. The difference between the two figures, multiplied by the year's marking volume, is the quantitative result you can put in the report, for example, "Chinese essay marking fell from about 15 minutes per script to about 6, freeing up roughly X hours across the year for lesson preparation and open-lesson planning."

Record the changes in quality at the same time: how many times students revise each essay, and how many days they wait for feedback. "Students receive item-by-item comments within two days" is as persuasive in a report and in talking to parents as the time a teacher saves.