Yale Researchers Ask: How Many LLMs Does It Take to Make a Decision?

A woman using artificial intelligence
Photo Credit: Getty Images

A Yale School of Medicine research team has introduced a new benchmark, MedicalAgentsBench, to evaluate how different large language model (LLM) approaches perform on complex medical questions.

 

“With MedicalAgentsBench, we’re trying to evaluate how these externalized platforms, which require a lot of interaction between the LLM agents to make a decision, compare to a more complex model that internalizes all the discussion in its training,” lead researcher Mark Gerstein, PhD, said. “Is it better to have something that acts like a committee of experts, or is it better to have this incredibly well-trained super oracle that was trained with the knowledge of all of the experts?”

 

The study, published in Cell Patterns, compares two broad AI methods used in clinical reasoning: externalized agents, where multiple specialized models interact and reach a consensus, and internalized reasoning, where a single model works through a problem step by step.

 

Researchers said the benchmark was designed to address limits in existing medical AI tests, which can be too easy for newer models and may not separate reasoning from memorization.

 

“As the training of models gets more and more complicated and uses more data, we have this problem where the system just memorizes the answers,” Dr. Gerstein said. “We wanted to come up with a more sophisticated benchmark that dealt with this ceiling effect.”

Building MedicalAgentsBench

To build MedicalAgentsBench, the team drew from eight medical datasets and created more than 800 complex questions intended to require multistep thinking.

 

The researchers found neither approach consistently outperformed the other, but said the two methods may work well together. In testing, adding externalized agents to an internalized reasoning model improved performance.

 

“We’re trying to find a better recipe to help people build their own clinical support system that can answer questions from humans,” researcher Yanjun (Daniel) Shao said.

 

Read more of the latest AI in Eye Care news here

Author

Leave a Reply

Your email address will not be published. Required fields are marked *