Studying steps in FARSCAPE process

Assessing Scientific Claims: Agent-Based AI System Answers Tough Questions

09.03.2026

Advances in materials science and drug discovery increasingly depend on robotic laboratories that synthesize and test candidate compounds without human intervention. Artificial intelligence models can now propose custom materials and potential drug molecules far faster than robotic labs can physically test them. The resulting pools of candidate materials dramatically outstrip the availability of experimental testing.
 

Deciding which candidates deserve scarce laboratory time is itself a hard problem. It requires weighing scientific plausibility against practical constraints — likely impact, expected effectiveness, and eventual manufacturability — considerations that are difficult to apply consistently across thousands of candidates.
 

Studying drug discovery candidates
FARSCAPE could help narrow down which candidate materials and drug designs are worth pursuing. (iStockphoto)

 

An artificial intelligence (AI) program developed at the Georgia Tech Research Institute (GTRI) could provide a new way to narrow that hypothesis pool down to a rate that robotic laboratories can actually test, and more broadly, to evaluate a wide range of scientific and technical claims that would otherwise overwhelm human assessment.
 

Known as FARSCAPE, the program was developed to support a research project organized by the U.S. Defense Advanced Research Projects Agency (DARPA) to evaluate a broad range of feasibility questions. By breaking down challenging questions into smaller components that can be evaluated by a team of independent computer agents, FARSCAPE uses a “chain of thought” approach to provide answers in the form of probabilities. It then helps humans check the rationales for its assessment.
 

“We’re trying to capture the kind of high-level patterns that will allow AI systems to explore and evaluate in a way that is very natural for humans,” explained Clayton Kerce, a GTRI principal research scientist who leads development of FARSCAPE. “Humans are very good at making connections and intuitively addressing important issues that computational systems now aren’t so good at doing.”
 

While certain innate human abilities provide an advantage over machines, humans can’t possibly know as much as today’s large language model AI systems do now. By combining human and machine strengths, Kerce and his team can identify the right hypothetical materials to test in available laboratory time. 
 

FARSCAPE is an acronym for Formal and Associative Reasoning for Scientific Claim Assessment and Proposition Evaluation.
 

Putting Agents to Work
 

FARSCAPE works like this:
 

First, a contested scientific or technical claim is entered into the FARSCAPE system. The question may be too risky for generic chatbots to tackle or too important to wait for review by an expert panel of humans. A claim might involve the potential efficacy of a proposed new drug molecule, for instance.
 

FARSCAPE breaks the claim into a number of smaller sub-claims that can be evaluated independently — an approach similar to what a human expert might do. The sub-claims are then assigned to individual computerized agents who work separately, but in parallel, on tasks such as searching scientific literature, testing the claim against established physics or other scientific principles, and building a formal reasoning chain. As many as 16 agents can work simultaneously, accessing a volume of knowledge beyond what any human could possess. 
 

Steps FARSCAPE follows
Chart shows the steps FARSCAPE follows to evaluate scientific and technical claims, in this case involving lithium ion battery technology.

 

“The agents search wide, not just deep,” Kerce said. “They’re able to explore pathways no individual person would be able to explore.”
 

Information retrieved by the agents is then evaluated by FARSCAPE’s “Main Agent,” whose job is to analyze what each individual agent has found, critique and reconcile contradictions between what the different agents retrieved, recheck evidence, and synthesize a judgment of feasibility and confidence in the ultimate conclusion. A human analyst oversees this process to inspect, challenge, and redirect the software agents, if necessary.
 

The system’s output includes a score ranging from “highly feasible” to “highly infeasible.” The system explains its reasoning, lists any questions it couldn’t resolve, and shows key information sources, allowing humans to independently review the outcome.
 

FARSCAPE can use as many as 16 agents
FARSCAPE provides a new way to evaluate a wide range of scientific and technical claims, using as many as 16 agents simultaneously. (Credit: Sean McNeil, GTRI)

 

FARSCAPE can evaluate difficult questions across many disciplines, ranging from quantum computing and materials science to electronic design and drug discovery, Kerce said.
 

Expanding FARSCAPE’s Applications
 

Many technical fields advance rapidly, so today’s feasibility answer can become outdated quickly. Using its Living Deep Research (LIDR) extension, FARSCAPE can be set to re-run each feasibility analysis periodically — daily, weekly, or monthly. Rather than repeating the entire analysis, the system can tap a specific agent to evaluate an issue in a quickly developing scientific discipline or an area with a recent scientific advance relevant to the original big-picture question.
 

From experience with phones and personal computers, users may expect instant answers. But studying millions of journal papers and reviewing scientific principles takes time — minutes, hours, days, even weeks, Kerce noted. 
 

GTRI’s team can access a broad range of AI hosts to run FARSCAPE, each with different strengths. The researchers choose the right machines based on the question they want to evaluate, taking into account time constraints and the cost of using the AI service. 
 

What is Agentic AI?
 

Large language models (LLMs), familiarly known as chatbots, help write papers, answer questions, and take on other tasks by tapping the large amount of information used to train them. But many other types of AI can be applied to scientific needs.
 

Studying steps in FARSCAPE process
FARSCAPE provides a new way to evaluate a wide range of scientific and technical claims. The chart shows the step-by-step process. (Credit: Sean McNeil, GTRI)

 

FARSCAPE’s agentic AI uses autonomous or semi-autonomous software agents to interact with tools such as databases to pursue goals assigned by humans — in this example, evaluating large numbers of hypothetical materials that may or may not be feasible to produce or test.
 

“Large language models have captured the public imagination, but there has been a tremendous amount of progress across many different fields, and these disciplines end up interrelated and intermingled,” Kerce said. “In FARSCAPE, agentic AI finds the pieces of the puzzle that need to be fit together to create a pathway showing the feasibility of the claim — or alternatively, a pathway to refute the claim.”
 

One challenge of the process, he added, is that some of the puzzle pieces unearthed by the agents may turn out to be irrelevant. It’s up to the AI to figure out which pieces to use.
 

Getting to FARSCAPE and Beyond
 

For many years, GTRI has been working on artificial intelligence, including a specialty called computational symbolic reasoning, which brings a discipline known as deep learning together with AI logic. That background provides a natural fit for the needs of DARPA’s SciFy initiative. 
 

The 32-month SciFy is divided into three technical sprints that began with evaluating a materials science claim. Competitors, including GTRI, were asked to analyze more than 200 Ph.D.-level materials problems over a 48-hour time period. GTRI was a top performer as SciFy advanced into future AI and quantum-focused sprints.
 

“GTRI understands how to understand things,” Kerce said. “What’s important is how to organize information and knowledge to push forward on a high-level integrated understanding using agentic AI architectures. We design systems to explore and evaluate in ways that are very natural for humans.”
 

Because of GTRI’s performance, the team was asked to help DARPA develop an evaluation tool known as Tech Council in a Box. When complete, the tool is expected to use FARSCAPE techniques to help DARPA’s staff screen technical proposals submitted to the agency.
 

“Tech Council in a Box uses the same reasoning engine, pointed at a different job,” Kerce said. “Think of it as putting an expert technical review council on call. It will pull in outside context, flag potential technical risks and prior work, and raise questions a program manager would ask. It will present information to a human user who can then think critically about it.”
 

In addition to Kerce, the project includes GTRI Research Engineer Blair Johnson, Professor Faramarz Fekri from Georgia Tech’s School of Electrical and Computer Engineering, and Ph.D. Student Siheng Xiong, from Georgia Tech’s Machine Learning program.

About GTRI: The Georgia Tech Research Institute (GTRI) is the nonprofit, applied research division of the Georgia Institute of Technology (Georgia Tech). Founded in 1934 as the Engineering Experiment Station, GTRI has grown to more than 3,000 employees, supporting eight laboratories in over 20 locations around the country and performing more than $1 billion of problem-solving research annually for government and industry. GTRI's renowned researchers combine science, engineering, economics, policy, and technical expertise to solve complex problems for the U.S. federal government, state, and industry.

Writer: John Toon (john.toon@gtri.gatech.edu)

GTRI Communications
Georgia Tech Research Institute
Atlanta, Georgia USA
 

Newsletter

Sign up for monthly updates on GTRI’s research, activity, and more.

Related News

News stories
GTRI researchers are supporting the Quantum Benchmarking Initiative, a project of the Defense Advanced Research Projects Agency (DARPA) to evaluate approaches being pursued by a number of quantum computing companies.
News stories
An optical principle discovered more than a century ago may soon find new applications in such areas as monitoring atmospheric turbulence, tracking airborne objects, and mapping the environment.
News stories
Chuck Eassa, Washington Field Office Manager of the Georgia Tech Research Institute (GTRI), focuses on strategic engagement and business intelligence.