University of New South Wales Study Shows ChatGPT Increases Student Coding Performance While Reducing Knowledge Retention
Bot Mutiny |
Researchers found that computer science students using ChatGPT-4.5 scored 20% higher on coding tasks but performed significantly worse on recall tests administered 48 hours later.
Researchers at the University of New South Wales have documented a significant gap between the ability of students to complete programming tasks using AI and their ability to remember the underlying concepts. The study, first reported by arxiv.org, tracked 55 undergraduate computer science students as they completed introductory C programming tasks. The data shows that while ChatGPT-4.5 helps students produce better code in the short term, it creates a "dissociation" between immediate success and durable learning. Students using the AI tool achieved coding scores of 89%, compared to just 69% for students using conventional web search. However, this performance advantage disappeared during testing. ## Higher Scores Lower Recall
The researchers measured retention through cued recall quizzes administered immediately after the tasks and again 48 hours later. The AI-assisted group scored 41% on the immediate quiz, while the control group scored 53%. Two days later, the gap persisted. ChatGPT users scored 39% on retention, while those who used traditional search tools maintained a score of 52%. The study suggests that the ease of using generative AI prevents the "productive cognitive effort" required to move information from working memory into long-term memory. Students using AI reported a smaller increase in mental effort across their tasks. The AI essentially acted as a shortcut that bypassed the cognitive load necessary for learning. ## The Ownership Gap
The experiment also measured "psychological ownership" of the work produced. Students were asked to attribute what percentage of the submitted code they felt was truly their own. The results showed a stark divide in how students perceive their own contributions when an LLM is involved. Students in the control group attributed 81% of their code to themselves. Students using ChatGPT attributed only 45% of the work to their own efforts. This reduced sense of ownership indicates that students are aware they are acting as editors or facilitators rather than creators, even when the final output is functionally superior. ## Physiological Monitoring
To track cognitive load, the researchers used pupillary responses and heart rate variability monitoring. While the self-reported data showed lower mental effort for the AI group, the physiological tests did not provide clear trajectories due to substantial data loss during the experiment. The authors, including Christian Bergh and Jake Renzella, noted that tertiary institutions have shifted from prohibiting AI to permitting disclosed use. This study argues that current assessment practices are flawed if they rely on the final "artefact" as a proxy for a student's actual capability. ## Policy Implications
The findings suggest that current education models are poorly equipped for a post-AI environment. Because LLMs can produce functional code without the student acquiring knowledge, the researchers recommend new assessment designs. These include requiring students to explain, retrieve, and contribute to work as active participants rather than passive prompters. The study concludes that GenAI assistance reduces the opportunities for students to generate and retrieve information independently. Without these cognitive hurdles, the knowledge does not stick, regardless of how high the initial assignment score might be.
References
(2026). Your Programming Students' Cognition with ChatGPT: Higher Performance, Lower Retention, and Reduced Ownership. arxiv.org. https://doi.org/10.3390/COMPUTERS14050185/S1