News

Hidden hearing loss affects millions, but has no diagnosis or treatment options. EU-funded researchers have developed tools that could change both

Imagine struggling to follow conversations in a noisy restaurant or keep up in a busy office, yet being told by a doctor that your hearing is fine. That is the frustrating reality for people living with cochlear synaptopathy, or hidden hearing loss.

Around 34 million adults in the EU live with a disabling hearing loss, according to the Hearing Health Forum EU. But the true scale of the problem is likely much larger, because conventional hearing tests cannot detect cochlear synaptopathy at all.

There is no clinical diagnosis for the condition and no approved treatment, leaving many people with real hearing difficulties and no help available.

This hidden hearing loss can be debilitating, explained Professor Ingeborg Dhooge, who heads the Ear, Nose and Throat department at Ghent University hospital in Belgium. Dhooge was a clinical consultant on a research initiative called EarDiTech, which investigated the condition between 2022 and 2026.

The project was funded by the European Innovation Council, which helps researchers and innovators convert bold scientific ideas into real-world products.

“Everywhere you go, there is always noise,” Dhooge said. “When you go shopping, there is noise, music playing. At work, people are often placed in open plan offices. This creates a lot of background noise, which interferes with the productivity of those patients.”

Current options are limited: training people to lip read and focus more on speech sounds. But the EarDiTech researchers developed a test that can detect hidden hearing loss, and have been working on new software for hearing aids.

How hearing goes wrong

In the inner ear, sound is detected by tiny hair cells that convert vibrations into neural signals, which travel to the brain via synaptic connections between the hair cells and the auditory nerve.

Until recently, it was thought that age-related hearing loss starts with damaged hair cells.

“We now know that the first damage to the auditory structures is not the hair cells, but it’s actually these synapses that connect the hair cells to the brain,” explained Professor Sarah Verhulst, who leads the Hearing Technology lab at Ghent University.

The distinction matters, stressed Verhulst. As coordinator of the EarDiTech research, she understood the stakes: you only need one synapse per hair cell to detect a sound, but you need many working together to decode speech in a noisy environment.

The standard clinical tool for assessing hearing is the audiogram, in which patients are played tones at different pitches and volumes through headphones and asked to signal if they can hear them.

It measures hair cell performance, but says nothing about how many working synapses a patient has. Therefore, it cannot capture the difficulties those with cochlear synaptopathy face in the real world.

Counting synapses

The EarDiTech test combines electrodes placed on the forehead and earlobes with a specially designed audio stimulus, developed through years of research into modelling the auditory system.

“We built these computer models of the hair cells and the synapses,” Verhulst said. “Basically, we can simulate the auditory nerve responses or the synapses’ response in the whole cochlear model. Using this computational method, we can then see which audio stimulus is best able to fire off all these synapses at the same time.”

The test is designed to produce the strongest possible synaptic response. When someone with healthy synapses hears the stimulus, the electrodes detect a large spike in neural activity. In those with synapse loss, the response is smaller.

“By comparing the size of a patient’s response to the normal hearing response, we can estimate the degree of cochlear synaptopathy,” Verhulst said.

The researchers have shrunk the technology into a compact, portable device. Clinical trials at Ghent University have shown it can identify cochlear synaptopathy across different age groups.

The test is simple and requires no surgical procedure, so it can fit easily into everyday clinical practice, giving audiologists and ear, nose and throat specialists a tool they currently do not have.

The next step is obtaining CE marking, to certify that the device meets EU requirements for medical diagnostic equipment. Verhulst said that once they have investment to fund this, she expects it to take around 18 months to 2 years. “I think when it’s on the market, ear, nose and throat doctors and hospitals will use it,” she added.

A smarter hearing aid

Detecting hidden hearing loss is only half the challenge. The other part is finding a solution.

Standard hearing aids mainly amplify sound, which does not really help people who struggle to pick out speech from background noise.

“To some extent, hearing aids can help, but it’s difficult because with them everything becomes louder, but not necessarily clearer,” said Dhooge. “This means that patients can still struggle to pick out different sounds in noisy environments.”

To address this, the team used machine learning to train a software algorithm on two computational models: one simulating normal hearing, the other simulating the faulty sound processing of someone with cochlear synaptopathy.

The goal was an algorithm that could modify incoming sound in a way that best stimulated the synapses a patient still has. That would improve their ability to follow speech in noisy environments.

Clinical trials have shown measurable benefits in patients with cochlear synaptopathy. The team have also shown that the algorithm can run on low battery power and on processing chips commonly used in hearing aids.

In theory, this means the algorithm could run on the next generation of hearing aids, or even consumer earbuds. Verhulst envisages a future in which people will be able to tell their phone that they want to hear better and it will activate the hidden hearing loss software on the earbuds.

For now, the EarDiTech team is actively seeking hardware manufacturers, such as hearing aid makers and chipset companies, who can help them further develop their software and turn it into products patients can actually buy.

The algorithm uses each patient’s CochSyn test results to adjust sound to their specific degree of damage – something no current hearing aid can do.

Beyond the clinic

Hidden hearing loss has only been properly identified in the last decade, and there is still limited understanding of how many people are affected or how it develops with age.

“Now we have a tool that we can use in the clinic to gain a better idea of how big the problem is and how impaired people are,” Dhooge said. “We can conduct long-term studies following up patients and see how it evolves.”

That evidence base, researchers hope, will be what persuades health systems to act.

Text: By Michael Allen

This article was originally published in  the EU Research and Innovation Magazine.

The study emphasizes that tourism is evolving toward more connected, flexible and lifestyle-driven experiences

The global tourism industry is entering a new phase defined by technological innovation, service personalization and shifting traveler preferences. This is highlighted in the Travel Trends 2026 report, produced by Amadeus in collaboration with consulting firm Globetrender, which identifies six major trends expected to significantly influence how people plan, book and experience travel in the coming years.

The study emphasizes that tourism is evolving toward more connected, flexible and lifestyle-driven experiences. This transformation is fueled by rapid technological advances and by travelers seeking journeys that better reflect their personal interests, priorities and expectations. Today’s travelers increasingly value convenience, tailored itineraries and experiences that combine culture, entertainment and well-being.

One of the most notable developments identified in the report is the “Pawprint Economy,” a trend reflecting the growing role of pets in travel decisions. More travelers are choosing to include their pets in their vacations, encouraging the travel sector to expand services designed specifically for animal companions. Hotels are introducing more sophisticated pet-friendly amenities, while transportation providers and destinations are developing solutions that make traveling with pets easier and more comfortable. As a result, tourism offerings are gradually adapting to a reality where pets are no longer an afterthought but an integral part of many travelers’ plans.

Another key trend is “Travel Mixology,” which describes the blending of artificial intelligence with human expertise in travel planning. AI-powered tools can analyze large volumes of traveler data to generate recommendations tailored to individual preferences, past behaviors and interests. However, the report notes that human insight remains essential. Travel advisors and professionals bring creativity, cultural understanding and contextual knowledge that complement algorithm-based suggestions. The future of trip planning, therefore, is expected to rely on a hybrid approach where technology enhances — rather than replaces — human expertise.

The aviation sector will also play a crucial role in shaping travel patterns through the trend known as “Point-to-Point Precision.” Advances in aircraft technology, particularly the expansion of long-range narrow-body aircraft, are enabling airlines to operate more direct routes between cities that previously required connecting flights. This shift can significantly reduce travel times and open new international links for secondary and emerging destinations. As a result, travelers may gain access to a wider range of places without the inconvenience of traditional hub-and-spoke routes.

In the hospitality sector, the report highlights the emergence of “Pick ‘n’ Stays,” a trend centered on the hyper-personalization of accommodation experiences. Hotels are increasingly offering guests the ability to customize multiple aspects of their stay. From selecting room features and amenities to adjusting services based on individual lifestyle needs, travelers can shape their hotel experience according to their preferences. In this model, the hotel room evolves from a standardized space into a flexible environment designed specifically for each guest.

The influence of entertainment and media on travel choices is also gaining momentum through the trend called “Pop-Culture Pilgrimages.” Movies, television series, video games and global entertainment phenomena are increasingly inspiring travelers to visit destinations associated with their favorite stories or celebrities. Fans seek to experience real-world locations connected to fictional universes or cultural icons, creating new tourism opportunities for destinations that appear in widely consumed media content. This growing link between popular culture and tourism is reshaping how certain places gain global visibility.

Finally, the report identifies the rise of “Innovation Tourism,” a concept linked to travelers interested in exploring destinations associated with technological progress and scientific advancement. Cities known for innovation, research centers, futuristic museums and interactive science attractions are becoming appealing travel experiences in their own right. Visitors are increasingly curious about emerging technologies, sustainable urban development and experimental projects that offer a glimpse into the future.

Taken together, these six trends illustrate a broader transformation within the global tourism landscape. Modern travelers are not simply looking to move from one place to another; they are seeking experiences that reflect their lifestyles, interests and values, supported by technology that simplifies every stage of the journey.

For the tourism industry — including airlines, hotels, travel agencies and destinations — adapting to these evolving expectations will be essential. Companies that successfully integrate innovation, personalization and meaningful experiences will be better positioned to attract a new generation of travelers who are more informed, connected and selective than ever before.

According to the Amadeus report, the future of travel will be shaped by the intersection of technology, culture and human behavior. This convergence is expected to redefine not only how people travel, but also how they discover and engage with the world.

The network fosters the enhancement of cities’ digital capabilities, the ethical use of technology, and the creation of opportunities for global collaboration that accelerate digital innovation

Madrid Digital Capital takes a further step in its commitment to connect spaces for cooperation and collaboration by joining the international network WeGO | World Smart Sustainable Cities Organization, one of the leading global alliances dedicated to promoting smart and sustainable cities. 

WeGO brings together more than 160 local governments along with a wide range of technological and academic organizations to promote smart urban development, international cooperation, and the exchange of best practices and innovative technological solutions.

The network promotes the improvement of cities’ digital capabilities, the ethical use of technology and the creation of opportunities for global collaboration to accelerate digital innovation on Data, AI, IoT, 5G, Cybersecurity and CityResilience, DigitalTwin and the development of the principles of usercentricities and of HumanAdaptiveCities with a vision of Human at the core. 

With this membership, Madrid joins an international community of cities leading digital transformation, strengthening its global position, expanding access to specialized knowledge, and reinforcing strategic alliances with key regions such as Asia-Pacific. 

This membership also boosts the international projection of Madrid’s technological ecosystem and encourages participation in projects, workshops, and training programs that enrich the Madrid Digital Capital strategy. 

Joining hashtag#WeGO (https://we-gov.org/) is an opportunity to continue learning, sharing, and building together. Because the future of cities is defined by cooperation, collaboration, and connection. 

Madrid moves forward with determination in deploying a technological model that enables the development of the city’s collective intelligence, which is shared within the international context. 

Promoting alliances, sharing knowledge, and opening new pathways for international cooperation is a way to build a more cognitive and competitive city—one that is closer to the needs of people and the environment. 

Because collaboration is the key to respond to the challenges of the city.

The Global South City Competitiveness Index (GS-CCI 2025/2026) offers a Global South-focused diagnosis of competitiveness, evaluating 48 cities across 3 dimensions, 10 pillars, and 257 indicators

Cities are entering a decisive decade. Rapid technological change, shifting global demand, and deepening geo-economic fragmentation are reshaping how value is created, where investment flows, and which places succeed. Today, the SuperSymmetry Institute (SSI) released the Driving Urban Advantage in the Next Economy Report and the Global South City Competitiveness Index (GS-CCI 2025/2026), the major benchmarking tool designed explicitly for cities in the Global South.

The GS-CCI 2025/2026 introduces a multilayered urban intelligence framework that evaluates 48 global cities across three pillars – Urban Competitiveness, Business Environment, and Investment Attractiveness. Powered by 25 million datapoints and 257 indicators organized across 10 pillars and 5 economic hubs, GS-CCI 2025/2026 moves beyond outcome rankings to assess the institutional and economic capacity to act, the core of urban competitiveness in the Next Economy.

“Cities can no longer treat technology as an experiment or a side project,” said Dr. Basma AlBuhairan, Managing Director of Center for the Fourth Industrial Revolution Saudi Arabia.“It is central to how cities compete, grow, and serve their people. Those that combine technology with strong governance, coordinated planning, and capable institutions will define the Next Economy – where opportunity is inclusive, growth is sustainable, and cities lead with purpose.”

Overall Results

The Index ranks cities into five competitiveness tiers: Leading, Advanced, Developing, Emerging, and Nascent. Leading cities demonstrate strong alignment between economic hubs, business environment, and governance capacity, enabling them to convert scale into productivity, resilience, and sustained advantage. Cities in lower tiers face deeper structural constraints, particularly in governance effectiveness, business conditions, and investment readiness. The headline finding is the near-elimination of the traditional North-South gap at the top tier:

  • Leading Tier cities, including Singapore (1st), New York(2nd), London (3rd), Dubai (4th), and Beijing (5th), now perform on par with established global leaders. Dubai’s strength in Talent & Livability, for instance, now surpasses all Global North peers, demonstrating how targeted investment can create world-leading livability.
  • Cities in the Advanced through Nascent tiersshow varying levels of alignment between ambition, enabling conditions, and delivery capacity, highlighting distinct developmental trajectories.

“Global South cities are no longer catching up. They are actively reshaping the geography of global competitiveness,” said Alexey Prazdnichnykh, Project Lead. “Location still matters, but governance increasingly determines whether cities succeed or fall behind.”

Regional Patterns

Competitiveness across the Global South is marked by significant regional variation and internal disparities. 

  • East Asia leads with the highest median score (59.9), driven by scale, innovation ecosystems, and infrastructure. Beijing’s leadership in the Technology Development & Innovation pillar exemplifies this, with a score rivaling Silicon Valley.
  • MENA(50.1) shows how strategic investment and governance reform can propel cities forward. Abu Dhabi and Riyadh have leveraged sovereign investment and clear strategy to jump into the Advanced and Leading tiers.
  • South Asia andSoutheast Asia cluster near the global median (48.1 and 46.7), with wide internal gaps. Bengaluru(Advanced Tier) thrives as a deep-tech hub, while others in the region face infrastructure constraints.
  • Latin America, Sub-Saharan Africa, and Eurasiarecord lower median scores (42.6, 42.1, and 40.4, respectively), reflecting persistent challenges. Yet, Cape Town’s strong creative cluster and Santiago’s relative stability show pathways to regional leadership.
  • North America: remains the global benchmark for urban competitiveness, with leadership in innovation, finance, and city performance. Records the highest median GS-CCI score (~64) among all regions in the Index. Both cities included New York and Toronto rank in the Leading Tier, reflecting consistently strong economic foundations, institutional quality, and infrastructure maturity. 

International Engagement

“International cooperation requires a shared, data-driven understanding of global asymmetries. At the SSI, as an ecosystem of policy experts, visionaries, and business leads, we believe that open dialogue all over the globe is the cornerstone of equitable progress,” said Anastasia Kalinina, Founding Steward at SSI. “This report is considered as a tool for action. By showing the gaps between the Global South and North, we aim to rebalance the scales of opportunity and foster a new era of genuine international cooperation.”

The report reframes international cooperation as a practical competitiveness lever. Success comes from moving beyond symbolic partnerships to focused action. Nairobi’s role in the LAPSSET corridor embeds it in regional trade flows; Cape Town’s film industry leverages global airline partnerships and co-production treaties, and the Vietnam-Singapore Industrial Parks (VSIP) model has transferred world-class zone management and investor services to Ho Chi Minh City and Hanoi. 

Governance and Delivery

  • The Index establishes Future-Proof Governance as the foundation that turns ambition into results. It is the decisive differentiator between cities that strategize and those that deliver sustained competitiveness.
  • Effective governance is characterized by a city’s ability to set a clear long-term vision, coordinate action across fragmented institutions, ensure financial sustainability, and engage business as a core partner in execution. The report highlights that cities with strong governance mechanisms, like Singapore’s integrated network of specialized agencies or Dubai’s program-based budgeting that ties the D33 economic agenda directly to funding and key performance indicators, demonstrate how institutional capacity translates strategy into on-the-ground outcomes.
  • Also, the Index reveals that even cities with abundant resources or strategic advantages can be held back by fragmented, reactive, or opaque governance. This gap between planning and implementation is a primary constraint for cities in lower tiers. The report concludes that without this foundational governance capacity, efforts to improve business environments or develop economic hubs are likely to be unsustainable.

The GS-CCI 2025/2026 is designed as a strategic dialogue tool, that helps cities diagnose their unique structural strengths and binding constraints, prioritize interventions, and design evidence-based roadmaps for transformation.

About the Global South City Competitiveness Index (GS-CCI 2025/2026)

  • Scope:48 global cities with a principal focus on the Global South across three core dimensions that shape long-term competitiveness: Economic & Enabling Environment, Infrastructure & Systems, and Talent & Livability.
  • Architecture:257 indicators, 10 pillars and 50 sub-pillars, allowing for a detailed and comparable assessment of urban strengths and constraints, and 5 economic hubs – cluster strength, business dynamism, innovation capacity, investment readiness, and global connectivity.
  • Methodology:25 million datapoints combining big data with traditional and new economy metrics to capture foundational capabilities, not just outcomes. 

Has officially become the second African nation to join the Horizon Europe programme, marking a historic step forward in the country’s research and innovation ambitions

The agreement, signed during the EU–Egypt Summit in Brussels, opens new opportunities for Egyptian researchers, universities, and innovators to collaborate with their European counterparts via Horizon Europe on equal footing with EU Member States.

A major boost for Egyptian research and innovation

Under the new association agreement, Egypt gains full access to all components of the Horizon Europe programme, the European Union’s flagship initiative for research and innovation.

With a total budget of €93.5bn for 2021–2027, the programme is designed to drive solutions to global challenges – from climate change and sustainable development to technological competitiveness.

The move allows Egyptian institutions not only to participate in projects but also to lead research consortia, shaping Europe’s scientific and innovation agenda.

Participation will also help Egypt modernise its research infrastructure, promote institutional reforms, and enhance its national capacity-building efforts.

Strengthening EU–Egypt relations through Horizon Europe

The agreement was formally signed by EU Commissioner for Startups, Research, and Innovation Ekaterina Zaharieva and Egyptian Foreign Minister Badr Abdelatty, in the presence of European Commission President Ursula von der Leyen and Egyptian President Abdel Fattah al-Sisi.

This milestone reflects the deepening cooperation between Cairo and Brussels, extending beyond trade and diplomacy into the realms of science, technology, and innovation.

According to EU officials, the partnership will foster joint efforts in areas such as climate resilience, digital transformation, and sustainable development – all priorities shared by both regions.

Speaking on the partnership, President von der Leyen said: “People are at the core of the EU-Egypt partnership – supporting them to develop their talents, ideas, and skills.

“Egypt’s association to Horizon Europe will create opportunities for cutting-edge projects and innovations.

“These developments will be in vital research areas like water management, sustainable farming, and food security, bringing tangible benefits to our societies. “

Building on a longstanding partnership

Egypt’s association builds upon two decades of collaboration under the 2005 EU–Egypt Science and Technology Cooperation Agreement and its active role in the Partnership for Research and Innovation in the Mediterranean Area (PRIMA), which focuses on sustainable water management, agricultural innovation, and food security.

Following successful negotiations concluded in April 2025, Egypt’s full membership in Horizon Europe signals a new era for regional and international scientific cooperation.

It strengthens the Mediterranean’s role as a hub for cross-border innovation, linking Europe, Africa, and the Middle East through shared research goals.

With this landmark agreement, Egypt’s integration into the Horizon Europe programme not only enhances its scientific potential but also reaffirms its position as a vital bridge between continents in addressing the challenges of a rapidly changing world.

A look at the work of three experts examining the successes and failures of AI to create solutions that lead to a better world

When John McCarthy, founder of the Stanford AI Lab, coined the term “artificial intelligence” in 1955 to describe “the science and engineering of making intelligent machines,” the defining feature of this tech was its smarts. But as AI becomes more sophisticated and widespread, it’s increasingly evident intelligence can’t be the only priority – AI must be trustworthy and socially responsible too.

AI systems are trained on specific sets of data in order to learn how to behave in various scenarios. This means that any biases or errors in the training data can result in unreliable or unfair outcomes. Because AI lacks its own sense of morality or ability to fact-check its own work, humans need to take an active role in monitoring how AI systems are trained and used.

“For many applications of AI, whether it is for aerospace, medicine, or financial systems, there are unlikely but possible ‘edge cases,’” said Mykel Kochenderfer, associate professor of aeronautics and astronautics in Stanford Engineering and director of the Stanford Intelligent Systems Laboratory, discussing scenarios where problems occur under extreme or unlikely conditions. “If these edge cases are not anticipated, the system could fail, and any one failure can be catastrophic. A lack of trust in these AI systems often stands in the way of their deployment.”

Stanford researchers want to go beyond identifying edge cases and fairness gaps in AI systems; their ultimate goal is to design AI that actively improves a project’s reliability, fairness, and trustworthiness compared to doing it without AI.

“AI’s capabilities aren’t magic, they’re measurable phenomena that we can study scientifically,” said Sanmi Koyejo, assistant professor of computer science in Stanford Engineering, who leads the Stanford Trustworthy AI Research Lab. “Understanding this helps us make more informed decisions about AI deployment rather than being swayed by hype or fear. Scientific measurement, not speculation, should guide how we integrate AI into society.”

As AI’s influence grows, many Stanford researchers are working to make it more fair, cautious, and secure. The work of three faculty members offers a window into these broader efforts to build AI systems people can trust.

AI and law

“When we purchased our home, I had to sign a document that said that the ‘property shall not be used or occupied by any person of African, Japanese or Chinese or any Mongolian descent,’” said Daniel E. Ho, the William Benjamin Scott and Luna M. Scott Professor of Law at Stanford Law School.

The experience underscored for him how deeply racism is embedded in legal infrastructure – something he’s now working to help dismantle through technology. Ho directs Stanford’s RegLab, which partners with government agencies to explore how AI can improve policy, services, and legal processes.

Recently, RegLab worked with Santa Clara County, which was implementing a state law that mandated all counties identify and redact racist property deeds.

“For Santa Clara County, this meant revising about 84 million pages of records dating back to the 1800s,” Ho said. “We developed an AI system to enable the county to spot, map, and redact deed records, saving some 86,500 person hours.”

Procedural bloat – an issue highlighted by both Democrats and Republicans – is another problem RegLab addressed with AI.

“One of the odd areas of consensus right now is that government processes aren’t working terribly well,” said Ho. “We developed a Statutory Research Assistant (STARA) AI system that can identify obsolete requirements, such as reports that can needlessly consume tons of staff time.”

San Francisco City Attorney David Chiu introduced a resolution to cut over a third of these requirements based on the RegLab collaboration.

“Implemented responsibly, I think AI has tremendous potential to improve government programs and access to justice,” said Ho.

Fairness in medical AI

Koyejo’s lab has developed algorithmic methods that directly address bias in medical AI systems. When AI diagnostic tools are trained primarily on data from specific populations – such as chest X-ray datasets that underrepresent certain racial groups – they often fail to work accurately across diverse patient populations.

“My lab’s algorithms help ensure that AI systems diagnosing diseases from chest X-rays work equally well for patients of all racial backgrounds, preventing health care disparities,” said Koyejo. His team has also tackled the challenge of ‘unlearning,’ developing techniques that allow AI systems to forget harmful training data or private medical information without compromising their overall performance.

Beyond health care, Koyejo’s graduate student Sang Truong collaborated with researchers in Vietnam to fine-tune an open-source large language model specifically for Vietnamese speakers. This work exemplifies how AI fairness extends beyond bias correction to ensuring technological benefits reach underserved linguistic communities.

“These applications matter because AI failures can lead to harm and exclude entire populations from technological benefits,” said Koyejo.

AI for safer systems

Although his work exists outside medicine, Kochenderfer’s research is another area where AI failures could mean life or death.

Kochenderfer studies decision making under uncertain conditions – which translates to systems used to monitor and control in air traffic, uncrewed aircraft, and automated cars. In these applications, the AI systems must account for efficiency of travel, complexity of movement (especially at high speeds), and limitations in sensor technologies that gather real-time data about how these vehicles move and what’s around them.

“Ensuring the safety of AI systems that interact with the real world is harder than many people realize,” said Kochenderfer. “We need the system to be extremely safe, even when there is a broad spectrum of plausible behavior and noise in the sensor systems.”

His current and future projects include the book Algorithms for Validation (forthcoming, free download online) and the new course Validation of Safety Critical Systems, with free YouTube lectures by postdoctoral researcher Sydney Katz.

“I am very excited to understand to what extent language models can help monitor safety critical systems,” said Kochenderfer. “Language models seem to be able to encode a wealth of commonsense knowledge. If so, they could enhance safety when automated subsystems, such as those on aircraft, fail in unexpected ways.”

What is worthy

While these examples highlight applications in law, medicine, and engineering, any discipline where there are AI tools – creative arts, communications, security technologies, and education, to name a few – offers opportunities to create AI that not only functions better and works more efficiently but also serves society in a fair, reliable, and trustworthy way.

“These applications matter because AI failures don’t just produce poor results – they can cause real harm and systematically exclude entire populations from technological benefits,” said Koyejo.

For him, the path forward requires acknowledging AI’s limitations while working systematically to address them. “Perfect AI is neither possible nor the right goal,” said Koyejo. “Instead, we should aim for AI systems that are worthy of the trust placed in them by society.”

Photo: Daniel Ho, Mykel Kochenderfer, and Sanmi Koyejo | Courtesy Daniel Ho, Marc Schlichting, Ananya Navale

The CodeSteer system could boost large language models’ accuracy when solving complex problems, such as scheduling shipments in a supply chain

Large language models (LLMs) excel at using textual reasoning to understand the context of a document and provide a logical answer about its contents. But these same LLMs often struggle to correctly answer even the simplest math problems.

Textual reasoning is usually a less-than-ideal way to deliberate over computational or algorithmic tasks. While some LLMs can generate code like Python to handle symbolic queries, the models don’t always know when to use code, or what kind of code would work best.

LLMs, it seems, may need a coach to steer them toward the best technique.

Enter CodeSteer, a smart assistant developed by MIT researchers that guides an LLM to switch between code and text generation until it correctly answers a query.

CodeSteer, itself a smaller LLM, automatically generates a series of prompts to iteratively steer a larger LLM. It reviews the model’s current and previous answers after each round and provides guidance for how it can fix or refine that solution until it deems the answer is correct.

The researchers found that augmenting a larger LLM with CodeSteer boosted its accuracy on symbolic tasks, like multiplying numbers, playing Sudoku, and stacking blocks, by more than 30 percent. It also enabled less sophisticated models to outperform more advanced models with enhanced reasoning skills.

This advance could improve the problem-solving capabilities of LLMs for complex tasks that are especially difficult to solve with textual reasoning alone, such as generating paths for robots in uncertain environments or scheduling shipments in an international supply chain.

“There is a race to develop better and better models that are capable of doing everything, but we’ve taken a complementary approach. Researchers have spent years developing effective technologies and tools to tackle problems in many domains. We want to enable LLMs to select the right tools and methods, and make use of others’ expertise to enhance their own capabilities,” says Chuchu Fan, an associate professor of aeronautics and astronautics (AeroAstro) and principal investigator in the MIT Laboratory for Information and Decision Systems (LIDS).

Fan, the senior author of the study, is joined on a paper about the work by LIDS graduate student Yongchao Chen; AeroAstro graduate student Yilun Hao; University of Illinois at Urbana-Champaign graduate student Yueying Liu; and MIT-IBM Watson AI Lab Research Scientist Yang Zhang. The research will be presented at the International Conference on Machine Learning.

An LLM “trainer”  

Ask an LLM which number is bigger, 9.11 or 9.9, and it will often give the wrong answer by using textual reasoning. But ask it to use code to answer the same question, and it can generate and execute a Python script to compare the two numbers, easily solving the problem.

Initially trained to understand and predict human language, LLMs are more likely to answer queries using text, even when code would be more effective. And while they have learned to generate code through fine-tuning, these models often generate an incorrect or less efficient version of the code.

Rather than trying to retrain a powerful LLM like GPT-4 or Claude to improve these capabilities, the MIT researchers fine-tune a smaller, lightweight LLM to guide a larger model between text and code. Fine-tuning a smaller model doesn’t change the larger LLM, so there is no risk it would undermine the larger model’s other abilities.

“We were also inspired by humans. In sports, a trainer may not be better than the star athlete on the team, but the trainer can still give helpful suggestions to guide the athlete. This steering method works for LLMs, too,” Chen says.

This trainer, CodeSteer, works in conjunction with the larger LLM. It first reviews a query and determines whether text or code is suitable for this problem, and which sort of code would be best.

Then it generates a prompt for the larger LLM, telling it to use a coding method or textual reasoning to answer the query. The larger model follows this prompt to answer the query and sends the result back to CodeSteer, which reviews it.

If the answer is not correct, CodeSteer will continue prompting the LLM to try different things that might fix the problem, such as incorporating a search algorithm or constraint into its Python code, until the answer is correct.

“We found that oftentimes, the larger LLM will try to be lazy and use a shorter, less efficient code that will not carry the correct symbolic calculation. We’ve designed CodeSteer to avoid this phenomenon,” Chen says.

A symbolic checker evaluates the code’s complexity and sends a signal to CodeSteer if it is too simple or inefficient. The researchers also incorporate a self-answer checker into CodeSteer, which prompts the LLM to generate code that calculates the answer to verify it is correct.

Tackling complex tasks

As the researchers designed CodeSteer, they couldn’t find suitable symbolic datasets to fine-tune and test the model, since many existing benchmarks don’t point out whether a certain query could be best solved with text or code.

So, they gathered a corpus of 37 complex symbolic tasks, including spatial reasoning, mathematics, order reasoning, and optimization, and built their own dataset, called SymBench. They implemented a fine-tuning approach that leverages SymBench to maximize the performance of CodeSteer.

In their experiments, CodeSteer outperformed all nine baseline methods they evaluated and boosted average accuracy from 53.3 percent to 86.4 percent. It maintains similar performance even on unseen tasks, and on a variety of LLMs.

In addition, a general-purpose model augmented with CodeSteer can achieve higher accuracy than state-of-the-art models designed to focus on complex reasoning and planning, while requiring much less computation.

“Our method uses an LLM’s own capabilities. By augmenting an LLM with the ability to smartly use coding, we can take a model that is already very strong and improve its performance even more,” Chen says.

In the future, the researchers want to streamline CodeSteer to speed up its iterative prompting process. In addition, they are studying how to effectively fine-tune a unified model with the ability to switch between textual reasoning and code generation, rather than relying on a separate assistant.

“The authors present an elegant solution to the critical challenge of tool utilization in LLMs. This simple yet impactful method enables state-of-the-art LLMs to achieve significant performance improvements without requiring direct fine-tuning,” says Jinsung Yoon, a staff research scientist at Google Cloud AI, who was not involved with this work. “This research represents a substantial contribution that promises to significantly enhance the application of LLMs to a diverse range of tasks with which they currently struggle.”

“Their success in training a smaller, specialized model to strategically guide larger, advanced models is particularly impactful,” adds Chi Wang, a senior staff scientist at Google DeepMind who was not involved with this work. “This intelligent collaboration among diverse AI ‘agents’ paves the way for more robust and versatile applications in complex real-world scenarios.”

This research is supported, in part, by the U.S. Office of Naval Research and the MIT-IBM Watson AI Lab.

Text: Adam Zewe | MIT News

Chief Executive of the Hong Kong Special Administrative Region (HKSAR), John Lee led a delegation of more than 50 business leaders and entrepreneurs from Hong Kong and Mainland China on a visit to Qatar and Kuwait.

Chief Executive of the Hong Kong Special Administrative Region (HKSAR), John Lee, hailed as a "success" his recent visit to the Middle East, where he led a delegation of more than 50 business leaders and entrepreneurs from Hong Kong and Mainland China on a visit to Qatar and Kuwait.

"This Middle East visit has elevated Hong Kong's relations with Qatar and Kuwait to a new level, bringing more business opportunities to Hong Kong," Mr Lee said.

Summing up the trip, the Chief Executive said the delegation had achieved three key objectives: to strengthen government-to-government relations; to find new areas of collaboration; and to make friends and extend networks in the region.

"We share a common commitment to deepening bilateral co-operation in trade, investment and cultural exchanges," Mr Lee said, noting his roundtable discussion with senior Kuwaiti officials hosted by the Acting Prime Minister of Kuwait His Excellency Sheikh Fahad Yousuf Saud Al-Sabah.

The Chief Executive highlighted six particularly successful areas of the trip:

First, strengthening relations with the governments of Qatar and Kuwait, and building consensus for collaboration.

Second, reaching a total of 59 memoranda of understanding (MOUs) and agreements (35 in Qatar and 24 in Kuwait), laying a diversified foundation for multifaceted co-operation.

Third, leveraging Hong Kong's strengths under "one country, two systems" as a "super connector" and "super value-adder", bridging global opportunities and linking the Mainland and the world.

Fourth, bolstering ties between Hong Kong and the Gulf Cooperation Council (GCC) member states. The Chief Executive, together with his previous Middle East mission to Saudi Arabia and the United Arab Emirates in 2023, has visited four of the six GCC member states. "The HKSAR Government is now actively exploring a free trade agreement with the GCC to further access this vital market," Mr Lee said.

Fifth, deepening mutual understanding and strengthening business networks and connections by promoting the strengths and opportunities of Hong Kong and Mainland China to partners in Qatar and Kuwait.

Sixth, advancing cultural exchanges and people-to-people connections with GCC countries.

"Middle East countries are seeking diversification of risks and looking for opportunities in (Mainland) China and the HKSAR in order to join the tide of the global economic shift towards the East," Mr Lee said.

"In this, Hong Kong has boundless opportunities."

Expansion of an open access database of critical materials characterisation data is well underway and an open call has been launched for the development of innovative open science projects

Momentum Transfer, a materials characterisation platform based in Mannheim, Germany, has successfully completed the first phase of its open Materials Scattering Network database. Developed as part of the EU-funded OSCARS project, the database will help to speed up materials research and development by offering open access to high-quality materials characterisation data.

This action is expected to play an important role in democratising access to crucial data for the international research community. Now included are 500 high-quality diffraction patterns contributed by researchers from around the world. These contributions currently cover a wide range of materials, including minerals, metals, metal oxides and organic compounds. Next, Momentum Transfer plans to further expand the database to host 20 000 high-fidelity datasets across diverse material classes.

“The response from the scientific community has been extraordinary,” states Momentum Transfer CEO Sebastian Winkler in a ‘Metal Powder Technology’ news item. “This level of engagement demonstrates the critical need for accessible, high-quality materials characterisation data to drive scientific discovery and industrial innovation.”

Bernd Hinrichsen, Chief Science Officer at Momentum Transfer, remarks on the “unparalleled insights into both crystalline and amorphous materials” made possible by combining high-resolution powder X-ray diffraction and total scattering measurements. “By making this data openly available, we aim to significantly advance global materials science research,” he adds.

One of the institutions benefiting from access to the platform is Arizona State University in the United States. Christina Birkel, a contributing researcher from the university, explains: “Access to synchrotron-quality data through Momentum Transfer’s platform has accelerated our research capabilities tremendously. This has enabled a significant big-picture study on functional materials that would not have been possible otherwise. Besides, the data are excellent training and learning opportunities for students who are working with diffraction techniques and solid-state compounds. The OSCARS database is set to become an invaluable resource for the entire materials science community.”

Under the expanded initiative, free X-ray diffraction and total scattering measurements will be provided for 10 000 samples under ambient conditions at the European Synchrotron Radiation Facility. Researchers can contribute up to five samples. Participants in the initiative will enjoy free access to synchrotron-quality measurements, complete data ownership during research and publication, and opportunities for global collaboration. Their research will also be integrated into the OSCARS database post-publication.

Calling for open science projects

The OSCARS project has now also launched its second open call for open science projects and services. The aim is to support researchers in carrying out FAIR – findability, accessibility, interoperability and reusability – data analysis and providing a series of new scientific results and services. Projects can be proposed in the field of any of the following five science clusters: humanities and social sciences; life sciences; environmental sciences; photon and neutron science; and astronomy, nuclear and particle physics.

Successful proposals will each receive between EUR 100 000 and EUR 250 000 in a lump sum, and will have up to 24 months to complete their project. OSCARS (O.S.C.A.R.S. - Open Science Clusters’ Action for Research and Society) has also made a video available online to help interested researchers prepare for the open call. The application deadline is 14 May 2025.

We use cookies

We use cookies on our website. Some of them are essential for the operation of the site, while others help us to improve this site and the user experience (tracking cookies). You can decide for yourself whether you want to allow cookies or not. Please note that if you reject them, you may not be able to use all the functionalities of the site.