Independent directory Public ClinicalTrials.gov records United States
ClinicalTrials.gov record
Active, not recruiting No phase listed Observational Accepts healthy volunteers

OpenEvidence Safety and Comparative Efficacy of Four LLM's in Clinical Practice

ClinicalTrials.gov ID: NCT07199231

Public ClinicalTrials.gov record NCT07199231. Field values are reproduced from the official study page; the official ClinicalTrials.gov record remains the source of truth for eligibility, enrollment, and contact information.

ClinicalTrials.gov public records Last synced Sep 2, 2026, 9:21 PM EDT

Data is sourced from official ClinicalTrials.gov public API records. Always review the official ClinicalTrials.gov record for the latest information.

Official title

A Comparative Performance Evaluation of Four Publicly Available Large Language Models Against Gold Standard Medical References

Brief summary

Reproduced verbatim from the official ClinicalTrials.gov record. Not medical advice.

OpenEvidence is an online tool that aggregates and synthesizes data from peer-reviewed medical studies, then producing a response to a user's questions using generative AI. While it is in use by a number of clinicians (including residents) today, there is little to no published data on whether the tool's outputs are accurate and whether this information appropriately informs clinical decision making. Similarly, a number of clinicians are turning to other large language models (LLM's) to assist in decision making when providing clinical care. While there have been a number of studies published on the accuracy of these LLM's responses to medical boards questions or clinical vignettes, there have been few studies to date examining their performance in a real world clinical setting, and even fewer comparing this performance. In this study, investigators have two goals: 1. To determine whether the use of the AI tool "OpenEvidence" leads to clinically appropriate decisions when utilized by family medicine, internal medicine, and psychiatry residents in the course of clinical practice. 2. To determine how the output of the OpenEvidence tool compares with three other commonly-used, publicly-available large language models (OpenAI's ChatGPT, Anthropic's Claude, and Google's Gemini) in answering common questions that residents have in the course of clinical practice. To accomplish study goal #1, investigators have enlisted residents in the above specialties to use the OpenEvidence tool in the course of clinical practice. In order to mitigate any safety risks, the residents will also use a typical reference tool for their question, which is referred to as the "Gold Standard" tool. These tools include PubMed and UpToDate. The residents will: 1. State their clinical question. 2. Query OpenEvidence, capturing their prompt and the OpenEvidence output for data analysis. All residents will undergo training in prompt engineering at the start of the study. 3. State their clinical conclusion based on the OpenEvidence data. 4. Query the Gold Standard Resource. 5. State their final clinical conclusion. 6. Answer a question on whether their clinical conclusion was modified by the Gold Standard reference. 7. Answer a question on whether they had any clinical safety concerns on the output from OpenEvidence. Attending physician Subject Matter Experts (SMEs) matched by specialty with at least 5 years of post-training clinical experience will then evaluate the residents' responses. 5 years was chosen based the book "Outliers" by Malcolm Gladwell, in which he asserts that 10,000 hours of focused practice is needed to achieve expertise in a field. SMEs will be asked to evaluate the residents' initial clinical questions and their conclusions based only on OpenEvidence. They will be asked to rate the clinical appropriateness of those conclusions on a scale of 1-10. For questions where the SME's rate the clinical appropriateness of the residents' conclusions poorly (\< 5/10), they will be asked to review the OpenEvidence output and answer an additional question as to whether the output was incorrect or the resident misinterpreted the output from the tool. To accomplish goal #2, the initial prompt entered by the residents into OpenEvidence will be copied by the research team into ChatGPT, Gemini, and Claude. The outputs from each tool (including OpenEvidence) will be surfaced to SMEs, who will be asked to rate each output based on accuracy, completeness, and bias. Likert scales will be used for these ratings. SMEs will also be asked an open-ended question to identify any patient safety issues from any of the outputs.

Study identification

NCT ID
NCT07199231
Recruitment status
Active, not recruiting
Study type
Observational
Phase
Not listed
Lead sponsor
Cambridge Health Alliance
Other
Enrollment
20 participants

Conditions and interventions

Eligibility (public fields only)

Age range
Not listed
Sex
All
Healthy volunteers
Accepts healthy volunteers

This page does not interpret eligibility. Detailed inclusion and exclusion criteria are on the official ClinicalTrials.gov record.

Study timeline

Start date
Sep 30, 2025
Primary completion
Jul 29, 2026
Completion
Sep 29, 2026
Last update posted
Aug 17, 2026

2025 – 2026

United States locations

U.S. sites
1
U.S. states
1
U.S. cities
1
Facility City State ZIP Site status
Cambridge Health Alliance Cambridge Massachusetts 02193

Site contact phone numbers, emails, and investigator names are intentionally not displayed here. Open the official ClinicalTrials.gov record for site contact information.

About this trial record page

What this page shows
Public field values for ClinicalTrials.gov record NCT07199231, including study identification, conditions, interventions, eligibility (age, sex, healthy volunteer), timeline, and U.S. site list.
What this page does not do
No medical advice, eligibility judgments, treatment recommendations, study quality scoring, or AI-generated medical summaries. No site contact phone numbers, emails, or investigator names.
Where the data comes from
Sourced from the official ClinicalTrials.gov public API. The official record is the source of truth.
Last refresh
Last update posted Aug 17, 2026 · Synced Sep 2, 2026

Related: full search, browse by condition, browse by drug or therapy, browse by sponsor, browse by U.S. city.

Open the official record

The complete protocol, eligibility criteria, and contact information for NCT07199231 live on ClinicalTrials.gov.

View official ClinicalTrials.gov record →