THEAccounting EducatorEvidence, ideas and practice for accounting teachers
Two related paper-based tasks sit side by side, one completed with a translucent guide and the other completed independently without it.
AI & Technology

What can students do after AI is switched off?

Polished AI-assisted work may show what a student can produce, but not necessarily what they have learned. A meta-analysis in programming education offers accounting educators a useful assessment distinction: evaluate performance with AI separately from competence demonstrated without it.

By The Accounting Educator · Published 16 September 2026 · 7 minute read
Listen to this articleNarrated edition

A student submits a polished analysis of an unfamiliar transaction. The explanation is well structured, the calculations reconcile and the conclusion sounds confident. Generative AI was permitted, so the work may demonstrate productive use of the tool. But does it also show that the student can identify the accounting issue, evaluate the output and reach a defensible conclusion independently?

Those are different questions, and assessment design can easily blur them.

A 2026 meta-analysis by Mian Wu and Fan Ouyang offers a useful way to separate them. The authors examined 35 studies in programming education, published between 2022 and 2025, which contributed 131 effect sizes. They classified outcomes into four categories: performance while using AI, performance without AI, higher-order skills such as problem-solving, and motivational or emotional outcomes.

Under the main analysis, generative AI had small-to-medium positive average effects across all four categories. Of particular interest, students who learned with AI also had a positive average result on programming assessments completed without it at the end of the intervention. The independent-performance estimate was positive and credible in the main model, but the lower end of its credible interval was close to zero. Under some assumptions about correlations among sampling errors, that lower end crossed zero marginally, and an adjustment for possible small-study bias left the positive estimate uncertain.

That main-model finding is more encouraging than a simple improvement in AI-assisted output. It suggests that using AI did not, on average, prevent students in the included programming studies from developing some independent capability. But it does not tell us how much of an individual student’s assisted performance remained when AI was removed. Few studies measured assisted and unaided performance in the same participants, and none included a delayed independent test.

The evidence comes entirely from programming education. It does not establish that generative AI improves accounting learning, nor does it validate a particular accounting assessment model. Its value for accounting educators lies in the distinction it makes visible: assisted performance and independent competence are related possibilities, not interchangeable measures.

One submission can conceal two learning questions

Suppose students complete an accounting case with access to generative AI. They must identify the relevant issues, perform calculations, propose entries and explain the financial-statement effects. The resulting submission could support several conclusions.

It might show that students can prompt an AI system effectively. It might show that they can edit and organise generated material. It might demonstrate that they can verify calculations and reject inappropriate suggestions. It might also reflect genuine accounting knowledge developed while using the system.

The finished document alone may not reveal which of these occurred.

This does not make AI-assisted work educationally worthless. Using AI well can itself be a legitimate learning outcome, particularly where students must interrogate an answer, find errors or improve an incomplete analysis. The assessment problem arises when the quality of an assisted product is treated as sufficient evidence of knowledge or judgment that students are expected to exercise without assistance.

The programming meta-analysis avoided that mistake by analysing assisted and independent outcomes separately. That distinction can also guide the assessment question in accounting.

Before designing an activity, ask what capability the assessment is intended to reveal. If students are being assessed on responsible use of AI, access to the tool is part of the task. If the intended outcome is independent preparation of an adjusting entry, AI-assisted coursework cannot by itself demonstrate that outcome. If both matter, both should be observed.

Pair the assisted task with an independent check

A practical response is to use two connected tasks rather than trying to make one submission answer every question.

For example, students might first complete an AI-permitted exercise on reporting-period adjustments. They could be required to submit the AI output, identify any weaknesses, correct the reasoning and provide their final answer. This assesses how they work with generated material, not merely whether they can obtain a plausible response.

A later unaided task could test the underlying accounting principle with different facts. Consider this classroom exercise:

  • On 1 January, an entity pays $12,000 for insurance coverage running from 1 January to 31 December and initially records the full amount as prepaid insurance.
  • The reporting date is 31 March.
  • For this exercise, insurance cost is recognised evenly over the coverage period.

Students working without AI could be asked to calculate the adjustment, prepare the entry and explain its effects. Three months of coverage have elapsed, so the adjustment is $3,000:

Debit insurance expense $3,000. Credit prepaid insurance $3,000.

Ignoring tax effects, expenses increase and assets decrease by $3,000, reducing profit and equity by the same amount.

The point is not to repeat an AI-assisted question with new numbers. The independent task should be aligned with the same intended learning outcome while requiring students to interpret the facts themselves. If the goal is to assess recognition across reporting periods, the assisted activity might concern an accrued expense while the independent check concerns a prepayment. Students must then recognise the underlying timing issue rather than reproduce the same entry pattern.

The two performances can be reviewed separately. A student who produces strong assisted work but struggles independently may need more practice with the accounting principle. A student who understands the accounting but accepts weak AI output may need support with verification and critical use. A student who succeeds in both has demonstrated two different capabilities.

This paired design also improves the information available to the lecturer. Most studies in the programming synthesis did not measure assisted and unaided outcomes in the same students, which prevented the researchers from estimating what individuals retained when support was withdrawn. A local accounting exercise that captures both results can show how the same students perform under assisted and unaided conditions, even if only for formative purposes. It cannot by itself establish what was retained from the assisted task or attribute any difference to AI use, because the two tasks may differ in difficulty or familiarity.

An immediate check is not evidence of durable learning

Timing matters. In the reviewed studies, independent programming performance was measured at the end of the intervention. No study included a delayed assessment. The positive average finding therefore concerns immediate unaided performance, not long-term retention or transfer to substantially different problems.

An accounting lecturer interested in retention could add a short unaided check one or two weeks later. It need not be another major assessment. A brief tutorial problem, an individual explanation or a low-stakes quiz may be enough to show whether students can still select and justify the method after the original activity has faded from memory.

The delayed task should match the claim the lecturer wants to make. A calculation closely resembling the original exercise tests retention of that procedure. A less familiar case tests whether students can adapt their knowledge. These are both useful, but they are not the same outcome.

The evidence does not identify a winning AI design

The meta-analysis also examined whether results varied systematically with educational context, instructional design or features of the AI system. These moderator analyses tested whether any of those conditions reliably predicted stronger or weaker results. None of the models reliably improved prediction. That does not mean the choice of tool, guidance or teaching approach makes no difference. The evidence base was small and uneven, implementation was often poorly described, and the statistical intervals remained wide enough to include educationally meaningful differences.

For accounting educators, there is no research-backed formula here for choosing a particular system or prescribing an ideal amount of AI access. Local implementation needs to be documented and evaluated. Useful details include what the system was allowed to generate, whether students received guidance on checking its output, when assistance was available and how the unaided assessment was controlled.

It is also sensible to keep outcomes separate when evaluating a pilot. Better AI-assisted submissions, stronger independent performance, improved judgment and greater confidence should not be collapsed into a single claim that learning improved. A change in one does not guarantee a change in the others.

The most useful next step may therefore be modest. Select one existing AI-permitted activity and add a short, aligned task completed without AI. If durable learning matters, repeat the check after a delay. The comparison will not settle the wider debate about generative AI, but it will provide better evidence about what your students can produce with assistance, what they can do independently and where teaching needs to intervene.

Sources and further reading