Breaking
Personal Health

FDA funds study on LLMs evaluating AI radiology reports

FDA funds study on LLMs evaluating AI radiology reports - ai radiology reports
Chaudhari stressed that the findings should provide a framework for future AI regulations, ensuring the approach remains adaptable beyond Cognita’s own products.

Cognita Imaging has won a $1.29 million agreement with the FDA to examine whether large language models can accurately assess AI-generated radiology reports. The 18-month study, which began June 22, targets a major hurdle in AI oversight: evaluating systems that produce open-ended text rather than simple diagnostic indicators.

Current FDA-approved radiology AI tools handle specific tasks—such as identifying blood clots in CT scans or lung collapses in X-rays—making their assessment relatively simple. However, newer platforms, including those developed by Cognita, generate complete radiology reports containing hundreds of potential findings. This complexity renders traditional radiologist reviews impractical for large-scale use.

The project will employ a panel of multiple LLMs to evaluate AI-generated reports against a dataset of 1 million patient exams collected from various U.S. healthcare settings. Performance will be measured across patient demographics, clinical environments, and imaging equipment types, including rare conditions. When the LLMs disagree on significant clinical issues, human radiologists will intervene to determine whether errors originate from the original AI, the LLM panel, or the baseline report.

Read Also: Healthcare AI adoption outpaces governance controls

Akshay Chaudhari, Cognita’s co-founder and a Stanford radiology associate professor, stated the initiative will quantify two primary risks in generative AI: hallucinations, plausible but incorrect information, and omissions of critical details. The aim is to assess whether these errors hold clinical significance and how frequently they appear.

The final deliverables will include a full report for the FDA, along with open-source software code, guidelines for constructing LLM review panels, and comparisons between large and smaller validation datasets. Chaudhari stressed that the findings should provide a framework for future AI regulations, ensuring the approach remains adaptable beyond Cognita’s own products.

The FDA’s increasing focus on AI regulation is evident in other recent actions. Earlier this year, Cognita received a breakthrough device designation for a vision-language model designed to interpret chest X-rays and produce draft radiology reports for physician review. This designation, granted separately from the research contract, reflects the agency’s commitment to accelerating approvals for tools addressing critical clinical needs. Chaudhari noted that while the study plan for this model remains under development, the FDA’s partnerships have clarified regulatory pathways.

Read Also: Medicare urged to crack down on suppliers

A separate FDA discussion paper on generative AI highlights broader efforts to establish standardized evaluation methods. The shift toward LLM-based review systems could reduce dependence on radiologist-led panels, which are slow and struggle to keep up with AI advancements. Yet challenges remain: if LLMs introduce biases or inaccuracies, the system’s reliability depends on human oversight at key decision points.

Founded in 2024 and acquired by Radiology Partners, the largest U.S. radiology group, just one year later, Cognita has drawn on insights from radiologists in both academic and community practices. Real-world pressures, such as radiologist shortages and mounting backlog cases, highlight why AI tools capable of automating report generation could address pressing workforce gaps. Chaudhari framed the project’s success as dependent on proving that LLM juries can consistently match or improve upon traditional review methods without creating new risks.

clinical diagnosis fda llms
Seraphina Wentworth

Leave a Reply

Your email address will not be published. Required fields are marked *