{"id":21306,"date":"2026-05-24T12:00:12","date_gmt":"2026-05-24T10:00:12","guid":{"rendered":"https:\/\/thesmartcityjournal.cibeles.net\/sin-categoria\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance\/"},"modified":"2026-05-24T12:02:00","modified_gmt":"2026-05-24T10:02:00","slug":"ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance","status":"publish","type":"post","link":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance","title":{"rendered":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019"},"content":{"rendered":"<p><i><span style=\"font-weight: 400\">Top\u00a0AI\u00a0systems show\u00a0bias\u00a0towards rewarding overly complex\u00a0prose\u00a0styles and only match\u00a0human\u00a0examiners for\u00a0grade\u00a0bands around\u00a0half\u00a0the time, research finds<\/span><\/i><\/p>\n<p><span style=\"font-weight: 400\">Researchers have used top Generative AI models to grade hundreds of undergraduate essays and found that AI only matched human-awarded degree classification around half the time, with AI often failing to assess the best and worst submissions accurately.<\/span><\/p>\n<p><span style=\"font-weight: 400\">A University of Cambridge-led team of psychologists and AI experts tested three \u201cfrontier\u201d systems, including the latest versions (as of April 2026) of Claude and ChatGPT, on over 750 student essays from three UK universities submitted as part of a psychology degree.<\/span><\/p>\n<p><span style=\"font-weight: 400\">While accuracy of AI in grading the essays, from coursework to exam answers, was \u201cnot uniformly high\u201d, say researchers, it did manage to match the broad grading bands \u2013 a first, 2:1, 2:2 and so on \u2013 given out by human examiners between 35-65% of the time.<\/span><\/p>\n<p><span style=\"font-weight: 400\">However, major stumbling blocks for AI include routinely undervaluing work awarded top marks by humans, or overvaluing essays ranked among the lowest.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Unlike human examiners, all the AI systems were \u201coversensitive to linguistic features\u201d: giving out higher marks based on essay length, vocabulary range and sentence complexity, which are often unrelated to academic standards.<\/span><\/p>\n<p><a href=\"https:\/\/www.emotional-cognition.psychol.cam.ac.uk\/sites\/default\/files\/OpRaise%20Report_DIGITAL.pdf\" target=\"_blank\" rel=\"nofollow noopener\"><span style=\"font-weight: 400\">In the latest report<\/span><\/a><span style=\"font-weight: 400\">, researchers suggest that AI could be valuable for aspects of student assessment such as error detection and consistency checks \u2013 a \u201csecond pair of eyes\u201d \u2013 as well as triaging feedback for students.<\/span><\/p>\n<p><span style=\"font-weight: 400\">For example, large discrepancies between AI and human marks could help flag assignments requiring further review by a human assessor.<\/span><\/p>\n<p><span style=\"font-weight: 400\">However, the team cautions that AI alone is far too shallow and inconsistent to grade undergraduate work, and a human should always determine the final mark.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cUniversities are under huge pressure to reduce staff workload and improve efficiency, all while meeting rising student expectations, and some may start to lean on AI for assessment,\u201d said Dr Deborah Talmi, the Cambridge psychologist who leads the\u00a0<\/span><a href=\"https:\/\/science.ai.cam.ac.uk\/projects\/2025-01-30-opportunities-and-potential-risks-of-ai-in-supporting-evaluation-opraise\" target=\"_blank\" rel=\"nofollow noopener\"><span style=\"font-weight: 400\">OpRaise project<\/span><\/a><span style=\"font-weight: 400\">\u00a0behind the new report.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cAI could perhaps automate some of the labour-intensive aspects of marking, freeing academics up for direct student engagement.\u201d<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cWe find that leaning heavily on the best current AI models would see student grading that is homogenised, underestimates brilliance, and favours linguistic style over the substance of sound academic judgement,\u201d said Talmi.\u00a0\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cAssessment is not just a system for distributing marks. It is part of how educational meaning is made, so students feel seen, standards are upheld, and trust is maintained. Use of AI in assessment poses a risk to these values.\u201d\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400\">The report, \u2018<\/span><a href=\"https:\/\/www.emotional-cognition.psychol.cam.ac.uk\/sites\/default\/files\/OpRaise%20Report_DIGITAL.pdf\" target=\"_blank\" rel=\"nofollow noopener\"><span style=\"font-weight: 400\">AI in University Assessment: Evaluating the Opportunities and Risks of Automated Marking<\/span><\/a><span style=\"font-weight: 400\">\u2019, is supported by\u00a0<\/span><a href=\"https:\/\/www.ai.cam.ac.uk\/\" target=\"_blank\" rel=\"nofollow noopener\"><b>ai@cam<\/b><\/a><span style=\"font-weight: 400\">, Cambridge University&#8217;s flagship mission to develop AI for the benefit of society, and the Accelerate Programme for Scientific Discovery, made possible by a donation from Schmidt Sciences. It is launched at an event with the British Psychological Society.<\/span><\/p>\n<p><span style=\"font-weight: 400\">For the study, AI was also asked to provide student feedback, and it churned out reflections between three and eight times longer than those provided by the original assessors.<\/span><\/p>\n<p><span style=\"font-weight: 400\">However, when AI responses were kept to a word count comparable to those from humans, focus groups of staff and students found it difficult to distinguish between human and AI feedback. Once the identity of the writer was revealed, not everyone appreciated AI-generated insights.<\/span><\/p>\n<p><span style=\"font-weight: 400\">University staff and students who took part in the study told researchers that, while current assessment practices are not perfect, being graded and receiving feedback from humans is fundamental to the \u201csocial contract\u201d between academics and students.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cMany students said they would feel cheated if AI marked their work, and staff warned that relying on AI risks weakening trust, motivation, professional judgement, and the human engagement at the heart of higher education,\u201d said Dr Yael Benn, a collaborator on the project from Manchester Metropolitan University.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The study used 761 undergraduate essays in psychology submitted and marked between 2022 and 2025 from a total of 125 students from the universities of Cambridge, Manchester Metropolitan and Nottingham.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The researchers chose to focus on psychology as essays are central to degree results in the subject. \u201cAcademic psychology is an ideal testing ground for AI assessment as it values evidence synthesis and critical judgement over single correct answers,\u201d said Talmi.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Researchers tested AI systems with the same essays at different times, and found AI gave the same or similar marks each time. The different AI models were much closer to each other than to humans in their marking.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The AI managed to match the right UK degree classification band of the five available (First, 2:1, 2:2, Third, Fail) some 63% of the time for Cambridge essays, while for Nottingham it was 53% and for Manchester Metropolitan it was 35%.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Researchers suspect that the difference in AI accuracy across institutions is due to the range of grades, which was narrowest among Cambridge students, whose essays were all written in invigilated exam halls, and widest at Manchester Metropolitan, where all analysed essays were coursework. Nottingham essays were a mixture of both.<\/span><\/p>\n<p><span style=\"font-weight: 400\">This illustrates the heart of the problem when relying on AI to assess students: inconsistent performances across institutions, types of prompting, and work that sits near grading boundaries, say the report\u2019s authors, who describe AI as having a \u201ccentral tendency bias\u201d.<\/span><\/p>\n<p><span style=\"font-weight: 400\">All papers are scored out of 100, standard practice in higher education. An essay marked 75 \u2013 a solid first \u2013 by a human is, on average, scored several points lower by every AI system. While an essay marked 50 \u2013 a low 2:2 \u2013 is scored several points higher.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The range on the marking scale where AI and humans most frequently align across institutions lies in the upper-50s to low-60s, so around a low 2:1, near the centre of the grade distribution.<\/span><\/p>\n<p><span style=\"font-weight: 400\">&#8220;Human assessors judge each essay on its own argumentative and conceptual merits while AI marks are based on statistical predictions,\u201d said co-author Dr Alexandru Marcoci, from Cambridge\u2019s Institute for Technology and Humanity.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cAcross models, the same pattern emerges. The AI assigns middling marks to all submissions, resulting in particularly inaccurate marking of the best and worst essays, whereas humans are more willing to distinguish genuinely exceptional or weak work.\u201d<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cThe practical consequence of this bias is that the AI is least accurate precisely where assessment decisions matter most, at the boundaries that distinguish Firsts from Upper Seconds, or passes from fails,\u201d he added.<\/span><\/p>\n<p><i><span style=\"font-weight: 400\">Researchers tested the performance of three frontier LLMs: Claude Opus 4.6 (Anthropic), GPT-5.4 (OpenAI), and Gemini 3 Flash (Google).<\/span><\/i><\/p>\n<p><i><span style=\"font-weight: 400\">The dataset: 125 students in 3 UK universities volunteered 761 authentic long-form undergraduate psychology essays (University of Cambridge: 133, University of Nottingham: 172, Manchester Metropolitan University: 456). All essays were submissions to formal assessments between 2022-2025.<\/span><\/i><\/p>\n<p><i><span style=\"font-weight: 400\">They spanned 50 modules and 87 distinct assignments across all years of study. Assessments spanned coursework, open book at-home examinations and invigilated examinations. Essay marks, on a 0-100 scale, were moderated formal marks provided by expert human assessors who followed routine institutional processes.<\/span><\/i><\/p>\n<p><i><span style=\"font-weight: 400\">Prompt design: Rather than committing to a single prompt, the team systematically varied the prompt under three dimensions &#8211; criteria specificity, calibration intervention, and scoring strategy &#8211; to isolate each component&#8217;s influence on scoring accuracy and identify the best prompt for each model.<\/span><\/i><\/p>\n<p><i><span style=\"font-weight: 400\">At the most basic level, models were prompted by the following statement: \u201cYou are an experienced &lt;University name&gt; examiner marking &lt;degree name&gt; undergraduate assignment.\u201d<\/span><\/i><\/p>\n<p><i><span style=\"font-weight: 400\">At the other end, models were given the full marking rubric, information about the expected mark distribution, and asked to justify aspects of the evaluation prior to providing a mark.<\/span><\/i><\/p>\n<p><i><span style=\"font-weight: 400\">Best-performing prompts per model were selected on a 20% calibration subset (n = 153); the same prompt configurations were then applied to the full corpus for the analyses reported here<\/span><\/i><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Top\u00a0AI\u00a0systems show\u00a0bias\u00a0towards rewarding overly complex\u00a0prose\u00a0styles and only match\u00a0human\u00a0examiners for\u00a0grade\u00a0bands around\u00a0half\u00a0the time, research finds Researchers have used top Generative AI models to grade hundreds of undergraduate essays and found that AI only matched human-awarded degree classification around half the time, with AI often failing to assess the best and worst submissions accurately. A University of Cambridge-led\u2026<\/p>\n","protected":false},"author":2,"featured_media":21305,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_scj_primary_category_id":55,"_scj_featured":false,"_scj_featured_from":"","_scj_featured_until":"","_scj_featured_order":0,"_scj_visibility_class":"current","_scj_layout_family":"standard","_scj_article_style":"","_scj_display_overrides":[],"_scj_intro_image_id":21305,"_scj_intro_alt":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019","_scj_intro_caption":"","_scj_intro_class":"","_scj_intro_float":"","_scj_full_image_id":21305,"_scj_full_alt":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019","_scj_full_caption":"","_scj_full_class":"","_scj_full_float":"","_scj_media_type":"","_scj_media_provider":"","_scj_media_external_id":"","_scj_media_url":"","_scj_media_poster_id":0,"_scj_media_width":0,"_scj_media_height":0,"_scj_media_aspect_ratio":"","_scj_media_description":"","_scj_gallery_items":[],"_scj_related_post_ids":[],"_scj_additional_authors":[],"scj_layout_family":"standard","scj_media_provider":"","scj_media_url":"","scj_full_caption":"","footnotes":""},"categories":[55],"tags":[2408,3884,3885,3526,3883,3882],"class_list":["post-21306","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","tag-academic-research","tag-automated-grading","tag-educational-technology","tag-generative-ai","tag-higher-education","tag-university-assessment"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.4 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019 - thesmartcityjournal.com<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019 - thesmartcityjournal.com\" \/>\n<meta property=\"og:description\" content=\"Top\u00a0AI\u00a0systems show\u00a0bias\u00a0towards rewarding overly complex\u00a0prose\u00a0styles and only match\u00a0human\u00a0examiners for\u00a0grade\u00a0bands around\u00a0half\u00a0the time, research finds Researchers have used top Generative AI models to grade hundreds of undergraduate essays and found that AI only matched human-awarded degree classification around half the time, with AI often failing to assess the best and worst submissions accurately. A University of Cambridge-led\u2026\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance\" \/>\n<meta property=\"og:site_name\" content=\"thesmartcityjournal.com\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/thesmartcityjournal\" \/>\n<meta property=\"article:published_time\" content=\"2026-05-24T10:00:12+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-05-24T10:02:00+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/05\/ai-university-essay-marking-bias.jpg.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"2058\" \/>\n\t<meta property=\"og:image:height\" content=\"1158\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Redacci\u00f3n de The Smart City Journal\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@SmartCityJour_\" \/>\n<meta name=\"twitter:site\" content=\"@SmartCityJour_\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Redacci\u00f3n de The Smart City Journal\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"6 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance\"},\"author\":{\"@type\":\"Organization\",\"name\":\"Redacci\u00f3n de The Smart City Journal\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/author\\\/smart-city\"},\"headline\":\"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019\",\"datePublished\":\"2026-05-24T10:00:12+00:00\",\"dateModified\":\"2026-05-24T10:02:00+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance\"},\"wordCount\":1275,\"publisher\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/ai-university-essay-marking-bias.jpg.jpg\",\"keywords\":[\"Academic Research\",\"automated grading\",\"educational technology\",\"generative AI\",\"higher education\",\"university assessment\"],\"articleSection\":[\"Artificial Intelligence\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance\",\"name\":\"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019 - thesmartcityjournal.com\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/ai-university-essay-marking-bias.jpg.jpg\",\"datePublished\":\"2026-05-24T10:00:12+00:00\",\"dateModified\":\"2026-05-24T10:02:00+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#primaryimage\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/ai-university-essay-marking-bias.jpg.jpg\",\"contentUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/05\\\/ai-university-essay-marking-bias.jpg.jpg\",\"width\":2058,\"height\":1158,\"caption\":\"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#website\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\",\"name\":\"thesmartcityjournal.com\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":[\"Organization\",\"NewsMediaOrganization\"],\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\",\"name\":\"thesmartcityjournal.com\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/logo-smart-city-journal-1.png\",\"contentUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/logo-smart-city-journal-1.png\",\"width\":346,\"height\":91,\"caption\":\"thesmartcityjournal.com\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/logo\\\/image\\\/\"},\"sameAs\":[\"https:\\\/\\\/www.facebook.com\\\/thesmartcityjournal\",\"https:\\\/\\\/x.com\\\/SmartCityJour_\",\"https:\\\/\\\/www.instagram.com\\\/thesmartcityjournal_\\\/\",\"https:\\\/\\\/www.tiktok.com\\\/@thesmartcityjournal\",\"https:\\\/\\\/www.linkedin.com\\\/company\\\/the-smart-city-journal\\\/\",\"https:\\\/\\\/www.youtube.com\\\/channel\\\/UCIPTMzK206pVFxMNjO7SSxw\"]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/person\\\/39c9da1aeac1b5c2d179ef9e4e293cf4\",\"name\":\"Redacci\u00f3n de The Smart City Journal\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"caption\":\"Redacci\u00f3n de The Smart City Journal\"},\"description\":\"Equipo de redacci\u00f3n de The Smart City Journal.\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/author\\\/smart-city\"},{\"@type\":\"SiteNavigationElement\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#site-navigation\",\"cssSelector\":[\".cib-main-menu\"]},{\"@type\":\"WPHeader\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#header\",\"cssSelector\":[\".cib-header\"]},{\"@type\":\"WPFooter\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#footer\",\"cssSelector\":[\".cib-footer\"]}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019 - thesmartcityjournal.com","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance","og_locale":"en_US","og_type":"article","og_title":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019 - thesmartcityjournal.com","og_description":"Top\u00a0AI\u00a0systems show\u00a0bias\u00a0towards rewarding overly complex\u00a0prose\u00a0styles and only match\u00a0human\u00a0examiners for\u00a0grade\u00a0bands around\u00a0half\u00a0the time, research finds Researchers have used top Generative AI models to grade hundreds of undergraduate essays and found that AI only matched human-awarded degree classification around half the time, with AI often failing to assess the best and worst submissions accurately. A University of Cambridge-led\u2026","og_url":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance","og_site_name":"thesmartcityjournal.com","article_publisher":"https:\/\/www.facebook.com\/thesmartcityjournal","article_published_time":"2026-05-24T10:00:12+00:00","article_modified_time":"2026-05-24T10:02:00+00:00","og_image":[{"width":2058,"height":1158,"url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/05\/ai-university-essay-marking-bias.jpg.jpg","type":"image\/jpeg"}],"author":"Redacci\u00f3n de The Smart City Journal","twitter_card":"summary_large_image","twitter_creator":"@SmartCityJour_","twitter_site":"@SmartCityJour_","twitter_misc":{"Written by":"Redacci\u00f3n de The Smart City Journal","Est. reading time":"6 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#article","isPartOf":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance"},"author":{"@type":"Organization","name":"Redacci\u00f3n de The Smart City Journal","url":"https:\/\/www.thesmartcityjournal.com\/en\/author\/smart-city"},"headline":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019","datePublished":"2026-05-24T10:00:12+00:00","dateModified":"2026-05-24T10:02:00+00:00","mainEntityOfPage":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance"},"wordCount":1275,"publisher":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#primaryimage"},"thumbnailUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/05\/ai-university-essay-marking-bias.jpg.jpg","keywords":["Academic Research","automated grading","educational technology","generative AI","higher education","university assessment"],"articleSection":["Artificial Intelligence"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance","url":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance","name":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019 - thesmartcityjournal.com","isPartOf":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#website"},"primaryImageOfPage":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#primaryimage"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#primaryimage"},"thumbnailUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/05\/ai-university-essay-marking-bias.jpg.jpg","datePublished":"2026-05-24T10:00:12+00:00","dateModified":"2026-05-24T10:02:00+00:00","breadcrumb":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#primaryimage","url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/05\/ai-university-essay-marking-bias.jpg.jpg","contentUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/05\/ai-university-essay-marking-bias.jpg.jpg","width":2058,"height":1158,"caption":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019"},{"@type":"BreadcrumbList","@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/ai-not-yet-good-enough-to-mark-university-essays-rewarding-style-over-substance#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.thesmartcityjournal.com\/en"},{"@type":"ListItem","position":2,"name":"AI not yet good enough to mark university essays, rewarding \u2018style over substance\u2019"}]},{"@type":"WebSite","@id":"https:\/\/www.thesmartcityjournal.com\/en#website","url":"https:\/\/www.thesmartcityjournal.com\/en","name":"thesmartcityjournal.com","description":"","publisher":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.thesmartcityjournal.com\/en?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":["Organization","NewsMediaOrganization"],"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization","name":"thesmartcityjournal.com","url":"https:\/\/www.thesmartcityjournal.com\/en","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/logo\/image\/","url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/09\/logo-smart-city-journal-1.png","contentUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/09\/logo-smart-city-journal-1.png","width":346,"height":91,"caption":"thesmartcityjournal.com"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/thesmartcityjournal","https:\/\/x.com\/SmartCityJour_","https:\/\/www.instagram.com\/thesmartcityjournal_\/","https:\/\/www.tiktok.com\/@thesmartcityjournal","https:\/\/www.linkedin.com\/company\/the-smart-city-journal\/","https:\/\/www.youtube.com\/channel\/UCIPTMzK206pVFxMNjO7SSxw"]},{"@type":"Organization","@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/person\/39c9da1aeac1b5c2d179ef9e4e293cf4","name":"Redacci\u00f3n de The Smart City Journal","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","caption":"Redacci\u00f3n de The Smart City Journal"},"description":"Equipo de redacci\u00f3n de The Smart City Journal.","url":"https:\/\/www.thesmartcityjournal.com\/en\/author\/smart-city"},{"@type":"SiteNavigationElement","@id":"https:\/\/www.thesmartcityjournal.com\/en#site-navigation","cssSelector":[".cib-main-menu"]},{"@type":"WPHeader","@id":"https:\/\/www.thesmartcityjournal.com\/en#header","cssSelector":[".cib-header"]},{"@type":"WPFooter","@id":"https:\/\/www.thesmartcityjournal.com\/en#footer","cssSelector":[".cib-footer"]}]}},"_links":{"self":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts\/21306","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/comments?post=21306"}],"version-history":[{"count":0,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts\/21306\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/media\/21305"}],"wp:attachment":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/media?parent=21306"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/categories?post=21306"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/tags?post=21306"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}