By employing a refined questioning strategy in a modified version of the Battleship game, researchers have shown that AI language models can enhance their ability to ask informative questions. This improvement boosts performance in the game and suggests more effective applications in complex fields like medical diagnosis and scientific research.
The Research Setting
In 2026, excitement around artificial intelligence agents reaches new heights as these semi-autonomous systems find applications across sectors from customer service to software development. However, challenges arise when these models must operate in uncertain environments, where asking the right questions becomes essential. Researchers from MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL) and Harvard University’s School of Engineering and Applied Sciences (SEAS) tackled this challenge by adapting the classic game of Battleship. Their experiment, called “Collaborative Battleship,” involved one player acting as the captain, asking questions about hidden ship locations, while a second player, the spotter, provided real-time answers.
Insights from Human Players
To establish a benchmark, the researchers had over 40 human participants play the game, recording their questions and answers to create the “BattleshipQA” dataset. This dataset served as a crucial reference when the team tested advanced language models, including GPT-5 and Llama-4-Scout, against human players. The findings indicated that while the top language models could outperform humans in terms of the number of game turns taken, smaller models struggled significantly.
The Monte Carlo Approach
One key challenge was the models' difficulty in generating effective questions. To tackle this, the research team applied a Monte Carlo inference strategy to the language models, enabling them to evaluate and weigh the likelihood of various responses. This method transformed how the models approached questioning, leading to a marked increase in their ability to deduce ship locations based on human responses.
The results were particularly notable for Llama 4 Scout, which initially won against humans only 8% of the time. After enhancements to its questioning strategy, its win rate climbed to 82%. This improvement came at a fraction of the operational cost compared to leading models like GPT-5.
Enhancing Answer Accuracy
In addition to refining their questioning techniques, researchers aimed to improve the accuracy of the models’ answers. The implementation of auto-formalization strategies allowed models to convert questions into programming commands that provided clear instructions for verifying answers. For instance, a question like “Is there a ship in column one that spans two rows?” would become a directive to check that specific area on the board. This approach resulted in a significant average accuracy increase of 15% across the tested models.
The Implications for AI Development
According to Gabriel Grand, a lead researcher and MIT PhD student, “Today’s language models are primarily optimized to answer complex queries, but it’s less clear whether they learn to ask good questions for themselves.” The findings highlight that the ability to ask informative questions is linked to a model's capability to predict and simulate scenarios accurately. By equipping agents with a ‘world model,’ the researchers demonstrated that these systems can inquire more effectively and make discoveries more efficiently.
The implications of this research extend beyond gaming. As AI agents improve their questioning abilities, they could enhance performance in high-stakes environments such as healthcare and scientific research, where navigating uncertainty is crucial. This evolution in AI questioning capabilities may lead to more sophisticated interactions between humans and machines, ultimately resulting in better outcomes across various domains.
The stories that move AI & crypto markets — before the market reacts.
Free. 7am ET. Five stories. 62,400 readers.

