Insider Brief
- AI tools like GPT4-Turbo and Elicit show potential in helping researchers extract data for systematic reviews, though they struggle with identifying nuanced contextual information compared to human reviewers.
- The study found Elicit outperformed GPT4-Turbo, providing higher quality responses, but both AI tools demonstrated inconsistency in accuracy without human oversight.
- Researchers suggest a collaborative workflow where AI performs initial extractions, with human reviewers verifying data, which could streamline evidence synthesis without sacrificing accuracy.
As researchers face an overwhelming number of published studies, the potential of artificial intelligence (AI) to streamline systematic reviews has become a growing interest in academia. A new preprint study on Research Square examines the performance of AI in extracting qualitative data from peer-reviewed documents, evaluating AI models’ capabilities to assist in this rigorous and resource-intensive process.
Researchers tested AI’s potential to perform evidence syntheses by comparing extractions done by AI systems with those conducted by human experts. Using two AI tools, GPT4-Turbo and Elicit, the study found mixed results: while AI performed similarly to humans in certain areas, it struggled significantly in discerning the presence of relevant contextual data.
The findings point to AI’s value as a supporting tool in evidence synthesis but underscore the need for human oversight to ensure accuracy and relevance.
Rising Demand for Systematic Reviews and Evidence Synthesis
Systematic reviews have become critical for informing evidence-based decision-making in fields like environmental management and policy. By examining and synthesizing data from multiple studies, these reviews offer comprehensive insights that can guide public policy, scientific research and practical interventions. However, with over five million scientific articles published each year, researchers face challenges in covering all relevant data, especially given the extensive time and labor involved in systematic reviews.
The study notes that AI’s ability to streamline systematic reviews could provide a solution to these challenges. AI tools may help researchers sift through large volumes of literature more quickly, potentially increasing both the efficiency and reliability of systematic reviews.
Methods: Comparing AI and Human Extractions
The researchers conducted the study by setting up an experiment in which two AI tools — GPT4-Turbo, a popular large language model (LLM), and Elicit, a specialized data extraction platform — were tasked with extracting information from a series of academic articles. Human reviewers also analyzed the same set of documents, focusing on a systematic review of benefits and barriers in community-based fisheries management in Pacific island nations.
The review team posed 11 contextual questions, such as the geographic focus of studies and specific management practices. Both human and AI reviewers answered these questions for 33 articles, and a separate team evaluated the quality and accuracy of the responses.
AI’s Strengths and Limitations
The study found that AI’s ability to identify relevant contextual data was limited, especially in comparison to human reviewers. The tools struggled with certain contextual extractions, with inconsistent agreement between AI and human reviewers, particularly when distinguishing the presence or absence of relevant information. GPT4-Turbo’s false positive rate — cases where it incorrectly indicated relevant data — reached 15%, while Elicit’s response rate was nearly 100%, though it generated many false positives.
Despite these limitations, AI tools demonstrated their utility in specific contexts. Elicit, for instance, outperformed GPT4-Turbo in extracting nuanced information, producing results that were more in line with human assessments. However, GPT4-Turbo’s results improved when it ran multiple extractions for each question, reducing its error rates.
“While the AI tools we tested (GPT4-Turbo and Elicit) were not reliable in discerning the presence or absence of contextual data, at least one of the AI tools consistently returned responses that were on par with human reviewers,” the researchers wrote. “These results highlight the utility of AI tools in the extraction phase of evidence synthesis for supporting human-led reviews and underscore the ongoing need for human oversight.”
This iterative process reflects one way AI can become a more dependable partner in research, though it still falls short of human reliability.
Where AI Fits in Knowledge Production
The findings suggest that AI can indeed support human researchers in the extraction phase of evidence synthesis, especially for structured information retrieval. The researchers emphasize that AI’s most promising role lies in supplementing human review processes rather than replacing them entirely.
One potential benefit of using AI in systematic reviews is its ability to handle vast amounts of information and to do so with reproducibility. AI models can apply the same logic and parameters consistently across multiple documents, potentially minimizing human error and increasing transparency. However, the researchers caution that AI’s tendency to miss or misinterpret nuanced details necessitates human oversight to ensure accuracy.
Limitations and Future Directions
The study acknowledges several limitations. First, the findings are specific to the topic of community-based fisheries management and may not be generalizable to all fields or topics. Second, the AI models tested were not trained specifically for evidence synthesis, which may have affected their performance in extracting nuanced contextual data.
To address these limitations, future research might focus on developing specialized AI models explicitly tailored for systematic reviews. Additionally, refining workflows that combine AI with human reviewers could enhance the quality of evidence syntheses. The study suggests that collaborative workflows—where AI performs initial data extractions that are then verified by human experts—could improve efficiency without compromising accuracy.
The researchers also highlight the proprietary nature of AI models as a potential barrier. Since tools like GPT4-Turbo and Elicit operate within closed systems, the models and methods they employ are not always transparent, raising questions about the replicability of their outputs.