{"id":20901,"date":"2026-02-22T16:36:06","date_gmt":"2026-02-22T15:36:06","guid":{"rendered":"https:\/\/thesmartcityjournal.cibeles.net\/sin-categoria\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models\/"},"modified":"2026-02-22T16:38:40","modified_gmt":"2026-02-22T15:38:40","slug":"exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models","status":"publish","type":"post","link":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models","title":{"rendered":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models"},"content":{"rendered":"<p><i><span style=\"font-weight: 400\">A new method developed at MIT could root out vulnerabilities and improve LLM safety and performance<\/span><\/i><\/p>\n<p><span style=\"font-weight: 400\">By now, ChatGPT, Claude, and other large language models have accumulated so much human knowledge that they\u2019re far from simple answer-generators; they can also express abstract concepts, such as certain tones, personalities, biases, and moods. However, it\u2019s not obvious exactly how these models represent abstract concepts to begin with from the knowledge they contain.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Now a team from MIT and the University of California San Diego has developed a way to test whether a large language model (LLM) contains hidden biases, personalities, moods, or\u00a0other abstract concepts. Their method can zero in on connections within a model that encode for a concept of interest. What\u2019s more, the method can then manipulate, or \u201csteer\u201d these connections, to strengthen or weaken the concept in any answer a model is prompted to give.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The team proved their method could quickly root out and steer more than 500 general concepts in some of the largest LLMs used today. For instance, the researchers could home in on a model\u2019s representations for personalities such as \u201csocial influencer\u201d and \u201cconspiracy theorist,\u201d and stances such as \u201cfear of marriage\u201d and \u201cfan of Boston.\u201d They could then tune these representations to enhance or minimize the concepts in any answers that a model generates.<\/span><\/p>\n<p><span style=\"font-weight: 400\">In the case of the \u201cconspiracy theorist\u201d concept, the team successfully identified a representation of this concept within one of the largest vision language models available today. When they enhanced the representation, and then prompted the model to explain the origins of the famous \u201cBlue Marble\u201d image of Earth taken from Apollo 17, the model generated an answer with the tone and perspective of a conspiracy theorist.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The team acknowledges there are risks to extracting certain concepts, which they also illustrate (and caution against). Overall, however, they see the new approach as a way to illuminate hidden concepts and potential vulnerabilities in LLMs, that could then be turned up or down to improve a model\u2019s safety or enhance its performance.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cWhat this really says about LLMs is that they have these concepts in them, but they\u2019re not all actively exposed,\u201d says\u00a0Adityanarayanan \u201cAdit\u201d Radhakrishnan, assistant professor of mathematics at MIT. \u201cWith our method, there\u2019s ways to extract these different concepts and activate them in ways that prompting cannot give you answers to.\u201d<\/span><\/p>\n<p><span style=\"font-weight: 400\">The team published their findings today in a study\u00a0<\/span><a href=\"http:\/\/doi.org\/10.1126\/science.aea6792\" target=\"_blank\" rel=\"nofollow noopener\"><span style=\"font-weight: 400\">appearing in the journal\u2002Science<\/span><\/a><span style=\"font-weight: 400\">. The study\u2019s co-authors include Radhakrishnan, Daniel Beaglehole and Mikhail Belkin of UC San Diego, and\u00a0Enric Boix-Adser\u00e0 of the University of Pennsylvania.<\/span><\/p>\n<h2>A fish in a black box<\/h2>\n<p><span style=\"font-weight: 400\">As use of OpenAI\u2019s ChatGPT, Google\u2019s Gemini, Anthropic\u2019s Claude, and other artificial intelligence assistants has exploded, scientists are racing to understand how models represent certain abstract concepts such as \u201challucination\u201d and \u201cdeception.\u201d In the context of an LLM, a hallucination is a response that is false or contains misleading information, which the model has \u201challucinated,\u201d or constructed erroneously as fact.<\/span><\/p>\n<p><span style=\"font-weight: 400\">To find out whether a concept such as \u201challucination\u201d is encoded in an LLM, scientists have often taken an approach of \u201cunsupervised learning\u201d \u2014 a type of machine learning in which algorithms broadly trawl through unlabeled representations to find patterns that might relate to a concept such as \u201challucination.\u201d But to Radhakrishnan, such an approach can be too broad and computationally expensive.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cIt\u2019s like\u00a0going fishing with a big net, trying to catch one species of fish. You\u2019re gonna get a lot of fish that you have to look through to find the right one,\u201d\u00a0he says. \u201cInstead, we\u2019re going in with bait for the right species of fish.\u201d<\/span><\/p>\n<p><span style=\"font-weight: 400\">He and his colleagues had previously developed the beginnings of a more targeted approach with a type of predictive modeling algorithm known as a recursive feature machine (RFM). An RFM is designed to directly identify features or patterns within data by leveraging a mathematical mechanism that neural networks \u2014 a broad category of AI models that includes LLMs \u2014 implicitly use to learn features.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Since the algorithm was an effective, efficient approach for capturing features in general, the team wondered whether they could use it to root out representations of concepts, in LLMs, which are by far the most widely used type of neural network and perhaps the least well-understood.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cWe wanted to apply our feature learning algorithms to LLMs to, in a targeted way, discover representations of concepts in these large and complex models,\u201d\u00a0Radhakrishnan says.<\/span><\/p>\n<h2>Converging on a concept<\/h2>\n<p><span style=\"font-weight: 400\">The team\u2019s new approach identifies any concept of interest within a LLM and \u201csteers\u201d or guides a model\u2019s response based on this concept. The researchers looked for 512 concepts within five classes: fears (such as of marriage, insects, and even buttons); experts (social influencer, medievalist); moods (boastful, detachedly amused); a preference for locations (Boston, Kuala Lumpur); and personas (Ada Lovelace, Neil deGrasse Tyson).<\/span><\/p>\n<p><span style=\"font-weight: 400\">The researchers then searched for representations of each concept in several of today\u2019s large language and vision models. They did so by training RFMs to recognize numerical patterns in an LLM that could represent a particular concept of interest.<\/span><\/p>\n<p><span style=\"font-weight: 400\">A standard large language model is, broadly, a\u00a0<\/span><a href=\"https:\/\/news.mit.edu\/2017\/explained-neural-networks-deep-learning-0414\" target=\"_blank\" rel=\"nofollow noopener\"><span style=\"font-weight: 400\">neural network<\/span><\/a><span style=\"font-weight: 400\">\u00a0that takes a natural language prompt, such as \u201cWhy is the sky blue?\u201d and divides the prompt into individual words, each of which is encoded mathematically as a list, or vector, of numbers. The model takes these vectors through a series of computational layers, creating matrices of many numbers that, throughout each layer, are used to identify other words that are most likely to be used to respond to the original prompt. Eventually, the layers converge on a set of numbers that is decoded back into text, in the form of a natural language response.<\/span><\/p>\n<p><span style=\"font-weight: 400\">The team\u2019s approach trains RFMs to recognize numerical patterns in an LLM that could be associated with a specific concept. As an example, to see whether an LLM contains any representation of a \u201cconspiracy theorist,\u201d the researchers would first train the algorithm to identify patterns among LLM representations of 100 prompts that are clearly related to conspiracies, and 100 other prompts that are not. In this way, the algorithm would learn patterns associated with the conspiracy theorist concept. Then, the researchers can mathematically modulate the activity of the conspiracy theorist concept by perturbing LLM representations with these identified patterns.\u00a0<\/span><\/p>\n<p><span style=\"font-weight: 400\">The method can be applied to search for and manipulate any general concept in an LLM. Among many examples, the researchers identified representations and manipulated an LLM to give answers in the tone and perspective of a \u201cconspiracy theorist.\u201d They also identified and enhanced the concept of \u201canti-refusal,\u201d and showed that whereas normally, a model would be programmed to refuse certain prompts, it instead answered, for instance giving instructions on how to rob a bank.<\/span><\/p>\n<p><span style=\"font-weight: 400\">Radhakrishnan says the approach can be used to quickly search for and minimize vulnerabilities in LLMs. It can also be used to enhance certain traits, personalities, moods, or preferences, such as emphasizing the concept of \u201cbrevity\u201d or \u201creasoning\u201d in any response an LLM generates. The team has made the method\u2019s underlying code publicly available.<\/span><\/p>\n<p><span style=\"font-weight: 400\">\u201cLLMs clearly have a lot of these abstract concepts stored within them, in some representation,\u201d\u00a0Radhakrishnan says<\/span><b>. \u201c<\/b><span style=\"font-weight: 400\">There are ways where, if we understand these representations well enough, we can build highly specialized LLMs that are still safe to use but really effective at certain tasks.\u201d<\/span><\/p>\n<p><span style=\"font-weight: 400\">This work was supported, in part, by the National Science Foundation, the Simons Foundation, the TILOS institute, and the U.S. Office of Naval Research.\u00a0<\/span><\/p>\n<p><i><span style=\"font-weight: 400\">Text: <\/span><\/i><b>Jennifer Chu\u00a0<\/b><span style=\"font-weight: 400\">|<\/span><b>\u00a0MIT News<\/b><\/p>\n","protected":false},"excerpt":{"rendered":"<p>A new method developed at MIT could root out vulnerabilities and improve LLM safety and performance By now, ChatGPT, Claude, and other large language models have accumulated so much human knowledge that they\u2019re far from simple answer-generators; they can also express abstract concepts, such as certain tones, personalities, biases, and moods. However, it\u2019s not obvious\u2026<\/p>\n","protected":false},"author":2,"featured_media":20900,"comment_status":"closed","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_scj_primary_category_id":130,"_scj_featured":true,"_scj_featured_from":"","_scj_featured_until":"","_scj_featured_order":60,"_scj_visibility_class":"current","_scj_layout_family":"standard","_scj_article_style":"","_scj_display_overrides":[],"_scj_intro_image_id":20900,"_scj_intro_alt":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models","_scj_intro_caption":"","_scj_intro_class":"","_scj_intro_float":"","_scj_full_image_id":20900,"_scj_full_alt":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models","_scj_full_caption":"","_scj_full_class":"","_scj_full_float":"","_scj_media_type":"","_scj_media_provider":"","_scj_media_external_id":"","_scj_media_url":"","_scj_media_poster_id":0,"_scj_media_width":0,"_scj_media_height":0,"_scj_media_aspect_ratio":"","_scj_media_description":"","_scj_gallery_items":[],"_scj_related_post_ids":[],"_scj_additional_authors":[],"scj_layout_family":"standard","scj_media_provider":"","scj_media_url":"","scj_full_caption":"","footnotes":""},"categories":[130],"tags":[3434,173,3433,299,2431,3435],"class_list":["post-20901","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-articles","tag-ai-safety","tag-artificial-intelligence","tag-large-language-models","tag-machine-learning","tag-mit","tag-neural-networks"],"yoast_head":"<!-- This site is optimized with the Yoast SEO plugin v28.4 - https:\/\/yoast.com\/product\/yoast-seo-wordpress\/ -->\n<title>Exposing biases, moods, personalities, and abstract concepts hidden in large language models - thesmartcityjournal.com<\/title>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"Exposing biases, moods, personalities, and abstract concepts hidden in large language models - thesmartcityjournal.com\" \/>\n<meta property=\"og:description\" content=\"A new method developed at MIT could root out vulnerabilities and improve LLM safety and performance By now, ChatGPT, Claude, and other large language models have accumulated so much human knowledge that they\u2019re far from simple answer-generators; they can also express abstract concepts, such as certain tones, personalities, biases, and moods. However, it\u2019s not obvious\u2026\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models\" \/>\n<meta property=\"og:site_name\" content=\"thesmartcityjournal.com\" \/>\n<meta property=\"article:publisher\" content=\"https:\/\/www.facebook.com\/thesmartcityjournal\" \/>\n<meta property=\"article:published_time\" content=\"2026-02-22T15:36:06+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2026-02-22T15:38:40+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/02\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm-1024x677.png\" \/>\n\t<meta property=\"og:image:width\" content=\"1024\" \/>\n\t<meta property=\"og:image:height\" content=\"677\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Redacci\u00f3n de The Smart City Journal\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:creator\" content=\"@SmartCityJour_\" \/>\n<meta name=\"twitter:site\" content=\"@SmartCityJour_\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Redacci\u00f3n de The Smart City Journal\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"6 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\\\/\\\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#article\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models\"},\"author\":{\"@type\":\"Organization\",\"name\":\"Redacci\u00f3n de The Smart City Journal\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/author\\\/smart-city\"},\"headline\":\"Exposing biases, moods, personalities, and abstract concepts hidden in large language models\",\"datePublished\":\"2026-02-22T15:36:06+00:00\",\"dateModified\":\"2026-02-22T15:38:40+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models\"},\"wordCount\":1278,\"publisher\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/02\\\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm.png\",\"keywords\":[\"AI safety\",\"artificial intelligence\",\"large language models\",\"machine learning\",\"MIT\",\"neural networks\"],\"articleSection\":[\"Articles\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models\",\"name\":\"Exposing biases, moods, personalities, and abstract concepts hidden in large language models - thesmartcityjournal.com\",\"isPartOf\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#primaryimage\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#primaryimage\"},\"thumbnailUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/02\\\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm.png\",\"datePublished\":\"2026-02-22T15:36:06+00:00\",\"dateModified\":\"2026-02-22T15:38:40+00:00\",\"breadcrumb\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#primaryimage\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/02\\\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm.png\",\"contentUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/02\\\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm.png\",\"width\":1428,\"height\":944,\"caption\":\"Exposing biases, moods, personalities, and abstract concepts hidden in large language models\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/articles\\\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"Exposing biases, moods, personalities, and abstract concepts hidden in large language models\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#website\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\",\"name\":\"thesmartcityjournal.com\",\"description\":\"\",\"publisher\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":[\"Organization\",\"NewsMediaOrganization\"],\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#organization\",\"name\":\"thesmartcityjournal.com\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/logo\\\/image\\\/\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/logo-smart-city-journal-1.png\",\"contentUrl\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/wp-content\\\/uploads\\\/2026\\\/09\\\/logo-smart-city-journal-1.png\",\"width\":346,\"height\":91,\"caption\":\"thesmartcityjournal.com\"},\"image\":{\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/logo\\\/image\\\/\"},\"sameAs\":[\"https:\\\/\\\/www.facebook.com\\\/thesmartcityjournal\",\"https:\\\/\\\/x.com\\\/SmartCityJour_\",\"https:\\\/\\\/www.instagram.com\\\/thesmartcityjournal_\\\/\",\"https:\\\/\\\/www.tiktok.com\\\/@thesmartcityjournal\",\"https:\\\/\\\/www.linkedin.com\\\/company\\\/the-smart-city-journal\\\/\",\"https:\\\/\\\/www.youtube.com\\\/channel\\\/UCIPTMzK206pVFxMNjO7SSxw\"]},{\"@type\":\"Organization\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#\\\/schema\\\/person\\\/39c9da1aeac1b5c2d179ef9e4e293cf4\",\"name\":\"Redacci\u00f3n de The Smart City Journal\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"url\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"contentUrl\":\"https:\\\/\\\/secure.gravatar.com\\\/avatar\\\/?s=96&d=mm&r=g\",\"caption\":\"Redacci\u00f3n de The Smart City Journal\"},\"description\":\"Equipo de redacci\u00f3n de The Smart City Journal.\",\"url\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en\\\/author\\\/smart-city\"},{\"@type\":\"SiteNavigationElement\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#site-navigation\",\"cssSelector\":[\".cib-main-menu\"]},{\"@type\":\"WPHeader\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#header\",\"cssSelector\":[\".cib-header\"]},{\"@type\":\"WPFooter\",\"@id\":\"https:\\\/\\\/www.thesmartcityjournal.com\\\/en#footer\",\"cssSelector\":[\".cib-footer\"]}]}<\/script>\n<!-- \/ Yoast SEO plugin. -->","yoast_head_json":{"title":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models - thesmartcityjournal.com","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models","og_locale":"en_US","og_type":"article","og_title":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models - thesmartcityjournal.com","og_description":"A new method developed at MIT could root out vulnerabilities and improve LLM safety and performance By now, ChatGPT, Claude, and other large language models have accumulated so much human knowledge that they\u2019re far from simple answer-generators; they can also express abstract concepts, such as certain tones, personalities, biases, and moods. However, it\u2019s not obvious\u2026","og_url":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models","og_site_name":"thesmartcityjournal.com","article_publisher":"https:\/\/www.facebook.com\/thesmartcityjournal","article_published_time":"2026-02-22T15:36:06+00:00","article_modified_time":"2026-02-22T15:38:40+00:00","og_image":[{"width":1024,"height":677,"url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/02\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm-1024x677.png","type":"image\/png"}],"author":"Redacci\u00f3n de The Smart City Journal","twitter_card":"summary_large_image","twitter_creator":"@SmartCityJour_","twitter_site":"@SmartCityJour_","twitter_misc":{"Written by":"Redacci\u00f3n de The Smart City Journal","Est. reading time":"6 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#article","isPartOf":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models"},"author":{"@type":"Organization","name":"Redacci\u00f3n de The Smart City Journal","url":"https:\/\/www.thesmartcityjournal.com\/en\/author\/smart-city"},"headline":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models","datePublished":"2026-02-22T15:36:06+00:00","dateModified":"2026-02-22T15:38:40+00:00","mainEntityOfPage":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models"},"wordCount":1278,"publisher":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#primaryimage"},"thumbnailUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/02\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm.png","keywords":["AI safety","artificial intelligence","large language models","machine learning","MIT","neural networks"],"articleSection":["Articles"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models","url":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models","name":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models - thesmartcityjournal.com","isPartOf":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#website"},"primaryImageOfPage":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#primaryimage"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#primaryimage"},"thumbnailUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/02\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm.png","datePublished":"2026-02-22T15:36:06+00:00","dateModified":"2026-02-22T15:38:40+00:00","breadcrumb":{"@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#primaryimage","url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/02\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm.png","contentUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/02\/mit-metodo-detectar-sesgos-conceptos-ocultos-llm.png","width":1428,"height":944,"caption":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models"},{"@type":"BreadcrumbList","@id":"https:\/\/www.thesmartcityjournal.com\/en\/articles\/exposing-biases-moods-personalities-and-abstract-concepts-hidden-in-large-language-models#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.thesmartcityjournal.com\/en"},{"@type":"ListItem","position":2,"name":"Exposing biases, moods, personalities, and abstract concepts hidden in large language models"}]},{"@type":"WebSite","@id":"https:\/\/www.thesmartcityjournal.com\/en#website","url":"https:\/\/www.thesmartcityjournal.com\/en","name":"thesmartcityjournal.com","description":"","publisher":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.thesmartcityjournal.com\/en?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":["Organization","NewsMediaOrganization"],"@id":"https:\/\/www.thesmartcityjournal.com\/en#organization","name":"thesmartcityjournal.com","url":"https:\/\/www.thesmartcityjournal.com\/en","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/logo\/image\/","url":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/09\/logo-smart-city-journal-1.png","contentUrl":"https:\/\/www.thesmartcityjournal.com\/wp-content\/uploads\/2026\/09\/logo-smart-city-journal-1.png","width":346,"height":91,"caption":"thesmartcityjournal.com"},"image":{"@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/logo\/image\/"},"sameAs":["https:\/\/www.facebook.com\/thesmartcityjournal","https:\/\/x.com\/SmartCityJour_","https:\/\/www.instagram.com\/thesmartcityjournal_\/","https:\/\/www.tiktok.com\/@thesmartcityjournal","https:\/\/www.linkedin.com\/company\/the-smart-city-journal\/","https:\/\/www.youtube.com\/channel\/UCIPTMzK206pVFxMNjO7SSxw"]},{"@type":"Organization","@id":"https:\/\/www.thesmartcityjournal.com\/en#\/schema\/person\/39c9da1aeac1b5c2d179ef9e4e293cf4","name":"Redacci\u00f3n de The Smart City Journal","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","url":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/?s=96&d=mm&r=g","caption":"Redacci\u00f3n de The Smart City Journal"},"description":"Equipo de redacci\u00f3n de The Smart City Journal.","url":"https:\/\/www.thesmartcityjournal.com\/en\/author\/smart-city"},{"@type":"SiteNavigationElement","@id":"https:\/\/www.thesmartcityjournal.com\/en#site-navigation","cssSelector":[".cib-main-menu"]},{"@type":"WPHeader","@id":"https:\/\/www.thesmartcityjournal.com\/en#header","cssSelector":[".cib-header"]},{"@type":"WPFooter","@id":"https:\/\/www.thesmartcityjournal.com\/en#footer","cssSelector":[".cib-footer"]}]}},"_links":{"self":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts\/20901","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/comments?post=20901"}],"version-history":[{"count":0,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/posts\/20901\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/media\/20900"}],"wp:attachment":[{"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/media?parent=20901"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/categories?post=20901"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.thesmartcityjournal.com\/en\/wp-json\/wp\/v2\/tags?post=20901"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}