{"id":20044,"date":"2025-07-20T10:16:01","date_gmt":"2025-07-20T08:16:01","guid":{"rendered":"https:\/\/thesmartcityjournal.cibeles.net\/sin-categoria\/evaluating-ai-language-models-just-got-more-effective-and-efficient\/"},"modified":"2025-07-20T10:17:40","modified_gmt":"2025-07-20T08:17:40","slug":"evaluating-ai-language-models-just-got-more-effective-and-efficient","status":"publish","type":"post","link":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient","title":{"rendered":"Evaluating AI language models just got more effective and efficient"},"content":{"rendered":"<p><i><span style=\"font-weight: 400\">Assessing the progress of new AI language models can be as challenging as training them. Stanford researchers offer a new approach<\/span><\/i><\/p>\n<p><span style=\"font-weight: 400\">As new versions of artificial intelligence language models roll out with increasing frequency, many do so with claims of improved performance. Demonstrating that a new model is actually better than the last, however, remains an elusive and expensive challenge for the field.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Typically, to prove their mettle and improve trust that new models are indeed better, developers subject new models to a battery of benchmark questions. Potentially hundreds of thousands of such benchmark questions are stored in question banks, and the answers must be reviewed by humans, adding time and cost to the process. Practical constraints make it impossible to ask every model every benchmark question, so developers choose a subset, introducing the risk of overestimating improvements based on softer questions. Stanford researchers have now introduced a cost-effective way to do these evaluations in a new paper published at the International Conference on Machine Learning.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cThe key observation we make is that you must also account for how hard the questions are,\u201d said Sanmi Koyejo, an assistant professor of computer science in the School of Engineering who led the research. \u201cSome models may do better or worse just by luck of the draw. We\u2019re trying to anticipate that and adjust for it to make fairer comparisons.\u201d<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cThis evaluation process can often cost as much or more than the training itself,\u201d added co-author Sang Truong, a doctoral candidate at the Stanford Artificial Intelligence Lab (SAIL). \u201cWe\u2019ve built an infrastructure that allows us to adaptively select subsets of questions based on difficulty. It levels the playing field.\u201d<\/span><\/p>\n<h2>Apples and oranges<\/h2>\n<p><span style=\"font-weight: 400\">To achieve their goal, Koyejo, Truong, and colleagues have borrowed a decades-old concept from education, known as Item Response Theory, which takes into account question difficulty when scoring test-takers. Koyejo compares it to the way standardized tests like the SAT and other kinds of adaptive testing work. Every right or wrong answer changes the question that follows.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The researchers use language models to analyze questions and score them on difficulty, reducing the costs by half and in some cases by more than 80%. That difficulty score allows the researchers to compare the relative performance of two models.<\/span><\/p>\n<p><span style=\"font-weight: 400\">To construct a large, diverse, and well-calibrated question bank in a cost-effective way, the researchers use AI\u2019s generative powers to create a question generator that can be fine-tuned to any desired level of difficulty. This helps automate the replenishing of question banks and the culling of \u201ccontaminated\u201d questions from the database.<\/span><\/p>\n<h2>Fast and fair<\/h2>\n<p><span style=\"font-weight: 400\">With better-designed questions, the authors say, others in the field can make better performance evaluations with a far smaller subset of queries. This approach is faster, fairer, and less expensive.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The new approach also works across knowledge domains \u2013 from medicine and mathematics to law. Koyejo has tested the system against 22 datasets and 172 language models and found that it can adapt easily to both new models and questions. Their approach was able to chart subtle shifts in GPT 3.5\u2019s safety over time, at first getting better and then retreating in several variations tested in 2023. Language model safety is a metric of how robust a model is to data manipulation, adversarial attacks, exploitation, and other risks.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Where once reliably evaluating language models was an expensive and inconsistent prospect, the new Item Response Theory approach puts rigorous, scalable, and adaptive evaluation within reach. For developers, this means better diagnostics and more accurate performance evaluations. For users, it means fairer and more transparent model assessments.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cAnd, for everyone else,\u201d Koyejo said. \u201cIt will mean more rapid progress and greater trust in the quickly evolving tools of artificial intelligence.\u201d<\/span><\/p>\n","protected":false},"excerpt":{"rendered":"<p>Assessing the progress of new AI language models can be as challenging as training them. Stanford researchers offer a new approach As new versions of artificial intelligence language models roll out with increasing frequency, many do so with claims of improved performance. Demonstrating that a new model is actually better than the last, however, remains\u2026<\/p>\n","protected":false},"author":2,"featured_media":20043,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_scj_primary_category_id":55,"_scj_featured":false,"_scj_featured_from":"","_scj_featured_until":"","_scj_featured_order":0,"_scj_visibility_class":"current","_scj_layout_family":"standard","_scj_article_style":"","_scj_display_overrides":[],"_scj_intro_image_id":20043,"_scj_intro_alt":"Evaluating AI language models just got more effective and efficient","_scj_intro_caption":"","_scj_intro_class":"","_scj_intro_float":"","_scj_full_image_id":20043,"_scj_full_alt":"Evaluating AI language models just got more effective and efficient","_scj_full_caption":"","_scj_full_class":"","_scj_full_float":"","_scj_media_type":"","_scj_media_provider":"","_scj_media_external_id":"","_scj_media_url":"","_scj_media_poster_id":0,"_scj_media_width":0,"_scj_media_height":0,"_scj_media_aspect_ratio":"","_scj_media_description":"","_scj_gallery_items":[],"_scj_related_post_ids":[],"_scj_additional_authors":[],"scj_layout_family":"standard","scj_media_provider":"","scj_media_url":"","scj_full_caption":"","footnotes":""},"categories":[55],"tags":[173,2740,299,2220],"class_list":["post-20044","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-artificial-intelligence","tag-artificial-intelligence","tag-language-models","tag-machine-learning","tag-stanford-research"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.4 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Evaluating AI language models just got more effective and efficient - thesmartcityjournal.com<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Evaluating AI language models just got more effective and efficient - thesmartcityjournal.com\" \/>\n<meta property=\"og:description\" content=\"Assessing the progress of new AI language models can be as challenging as training them. Stanford researchers offer a new approach As new versions of artificial intelligence language models roll out with increasing frequency, many do so with claims of improved performance. Demonstrating that a new model is actually better than the last, however, remains\u2026\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient\" \/>\n<meta property=\"og:site_name\" content=\"thesmartcityjournal.com\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/thesmartcityjournal\" \/>\n<meta property=\"article:published_time\" content=\"2025-07-20T08:16:01+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2025-07-20T08:17:40+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2025\/07\/evaluating-ai-language-models-stanford.jpg\" \/>\n\t<meta property=\"og:image:width\" content=\"1500\" \/>\n\t<meta property=\"og:image:height\" content=\"1000\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/jpeg\" \/>\n<meta name=\"author\" content=\"Redacci\u00f3n de The Smart City Journal\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@SmartCityJour_\" \/>\n<meta name=\"twitter:site\" content=\"@SmartCityJour_\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Redacci\u00f3n de The Smart City Journal\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"3 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient\"},\"author\":{\"@type\":\"Organization\",\"name\":\"Redacci\u00f3n de The Smart City Journal\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/author\\\/smart-city\"},\"headline\":\"Evaluating AI language models just got more effective and efficient\",\"datePublished\":\"2025-07-20T08:16:01+00:00\",\"dateModified\":\"2025-07-20T08:17:40+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient\"},\"wordCount\":626,\"publisher\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2025\\\/07\\\/evaluating-ai-language-models-stanford.jpg\",\"keywords\":[\"artificial intelligence\",\"language models\",\"machine learning\",\"Stanford Research\"],\"articleSection\":[\"Artificial Intelligence\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient\",\"name\":\"Evaluating AI language models just got more effective and efficient - thesmartcityjournal.com\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2025\\\/07\\\/evaluating-ai-language-models-stanford.jpg\",\"datePublished\":\"2025-07-20T08:16:01+00:00\",\"dateModified\":\"2025-07-20T08:17:40+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient#primaryimage\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2025\\\/07\\\/evaluating-ai-language-models-stanford.jpg\",\"contentUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2025\\\/07\\\/evaluating-ai-language-models-stanford.jpg\",\"width\":1500,\"height\":1000,\"caption\":\"Evaluating AI language models just got more effective and efficient\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/artificial-intelligence\\\/evaluating-ai-language-models-just-got-more-effective-and-efficient#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Evaluating AI language models just got more effective and efficient\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#website\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\",\"name\":\"thesmartcityjournal.com\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":[\"Organization\",\"NewsMediaOrganization\"],\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\",\"name\":\"thesmartcityjournal.com\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/logo-smart-city-journal-1.png\",\"contentUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/logo-smart-city-journal-1.png\",\"width\":346,\"height\":91,\"caption\":\"thesmartcityjournal.com\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/logo\\\/image\\\/\"},\"sameAs\":[\"https:\\\/\\\/www.facebook.com\\\/thesmartcityjournal\",\"https:\\\/\\\/x.com\\\/SmartCityJour_\",\"https:\\\/\\\/www.instagram.com\\\/thesmartcityjournal_\\\/\",\"https:\\\/\\\/www.tiktok.com\\\/@thesmartcityjournal\",\"https:\\\/\\\/www.linkedin.com\\\/company\\\/the-smart-city-journal\\\/\",\"https:\\\/\\\/www.youtube.com\\\/channel\\\/UCIPTMzK206pVFxMNjO7SSxw\"]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/person\\\/39c9da1aeac1b5c2d179ef9e4e293cf4\",\"name\":\"Redacci\u00f3n de The Smart City Journal\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"caption\":\"Redacci\u00f3n de The Smart City Journal\"},\"description\":\"Equipo de redacci\u00f3n de The Smart City Journal.\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/author\\\/smart-city\"},{\"@type\":\"SiteNavigationElement\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#site-navigation\",\"cssSelector\":[\".cib-main-menu\"]},{\"@type\":\"WPHeader\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#header\",\"cssSelector\":[\".cib-header\"]},{\"@type\":\"WPFooter\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#footer\",\"cssSelector\":[\".cib-footer\"]}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Evaluating AI language models just got more effective and efficient - thesmartcityjournal.com","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient","og_locale":"en_US","og_type":"article","og_title":"Evaluating AI language models just got more effective and efficient - thesmartcityjournal.com","og_description":"Assessing the progress of new AI language models can be as challenging as training them. Stanford researchers offer a new approach As new versions of artificial intelligence language models roll out with increasing frequency, many do so with claims of improved performance. Demonstrating that a new model is actually better than the last, however, remains\u2026","og_url":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient","og_site_name":"thesmartcityjournal.com","article_publisher":"https:\/\/www.facebook.com\/thesmartcityjournal","article_published_time":"2025-07-20T08:16:01+00:00","article_modified_time":"2025-07-20T08:17:40+00:00","og_image":[{"width":1500,"height":1000,"url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2025\/07\/evaluating-ai-language-models-stanford.jpg","type":"image\/jpeg"}],"author":"Redacci\u00f3n de The Smart City Journal","twitter_card":"summary_large_image","twitter_creator":"@SmartCityJour_","twitter_site":"@SmartCityJour_","twitter_misc":{"Written by":"Redacci\u00f3n de The Smart City Journal","Est. reading time":"3 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient#article","isPartOf":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient"},"author":{"@type":"Organization","name":"Redacci\u00f3n de The Smart City Journal","url":"https:\/\/www.thesmartcityjournal.com\/en\/author\/smart-city"},"headline":"Evaluating AI language models just got more effective and efficient","datePublished":"2025-07-20T08:16:01+00:00","dateModified":"2025-07-20T08:17:40+00:00","mainEntityOfPage":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient"},"wordCount":626,"publisher":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient#primaryimage"},"thumbnailUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2025\/07\/evaluating-ai-language-models-stanford.jpg","keywords":["artificial intelligence","language models","machine learning","Stanford Research"],"articleSection":["Artificial Intelligence"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient","url":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient","name":"Evaluating AI language models just got more effective and efficient - thesmartcityjournal.com","isPartOf":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#website"},"primaryImageOfPage":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient#primaryimage"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient#primaryimage"},"thumbnailUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2025\/07\/evaluating-ai-language-models-stanford.jpg","datePublished":"2025-07-20T08:16:01+00:00","dateModified":"2025-07-20T08:17:40+00:00","breadcrumb":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient#primaryimage","url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2025\/07\/evaluating-ai-language-models-stanford.jpg","contentUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2025\/07\/evaluating-ai-language-models-stanford.jpg","width":1500,"height":1000,"caption":"Evaluating AI language models just got more effective and efficient"},{"@type":"BreadcrumbList","@id":"https:\/\/www.thesmartcityjournal.com\/en\/artificial-intelligence\/evaluating-ai-language-models-just-got-more-effective-and-efficient#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.thesmartcityjournal.com\/en"},{"@type":"ListItem","position":2,"name":"Evaluating AI language models just got more effective and efficient"}]},{"@type":"WebSite","@id":"https:\/\/www.thesmartcityjournal.com\/en#website","url":"https:\/\/www.thesmartcityjournal.com\/en","name":"thesmartcityjournal.com","description":"","publisher":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.thesmartcityjournal.com\/en?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":["Organization","NewsMediaOrganization"],"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization","name":"thesmartcityjournal.com","url":"https:\/\/www.thesmartcityjournal.com\/en","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/logo\/image\/","url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/09\/logo-smart-city-journal-1.png","contentUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/09\/logo-smart-city-journal-1.png","width":346,"height":91,"caption":"thesmartcityjournal.com"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/thesmartcityjournal","https:\/\/x.com\/SmartCityJour_","https:\/\/www.instagram.com\/thesmartcityjournal_\/","https:\/\/www.tiktok.com\/@thesmartcityjournal","https:\/\/www.linkedin.com\/company\/the-smart-city-journal\/","https:\/\/www.youtube.com\/channel\/UCIPTMzK206pVFxMNjO7SSxw"]},{"@type":"Organization","@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/person\/39c9da1aeac1b5c2d179ef9e4e293cf4","name":"Redacci\u00f3n de The Smart City Journal","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","caption":"Redacci\u00f3n de The Smart City Journal"},"description":"Equipo de redacci\u00f3n de The Smart City Journal.","url":"https:\/\/www.thesmartcityjournal.com\/en\/author\/smart-city"},{"@type":"SiteNavigationElement","@id":"https:\/\/www.thesmartcityjournal.com\/en#site-navigation","cssSelector":[".cib-main-menu"]},{"@type":"WPHeader","@id":"https:\/\/www.thesmartcityjournal.com\/en#header","cssSelector":[".cib-header"]},{"@type":"WPFooter","@id":"https:\/\/www.thesmartcityjournal.com\/en#footer","cssSelector":[".cib-footer"]}]}},"_links":{"self":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts\/20044","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/comments?post=20044"}],"version-history":[{"count":0,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts\/20044\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/media\/20043"}],"wp:attachment":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/media?parent=20044"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/categories?post=20044"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/tags?post=20044"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}