Communications Director, Connecticut Hospital Association
110 Barnes Road, Wallingford, CT
rall@chime.org, 203-265-7611
STAT News – Wednesday, July 29, 2026
By Katie Palmer
Hundreds of thousands of U.S. doctors use clinical large language models, pitched by companies like OpenEvidence, Doximity, and UpToDate as an antidote to the dangers of hallucination-prone generalist models from Big Tech. Yet few studies have pitted them against each other — and this summer brought a high-profile head-to-head.
Researchers from NYU Langone Health had tested general and clinical models, including OpenEvidence and UpToDate Expert AI, on three sets of clinical questions. The findings, published in Nature Medicine in June: The clinical AI performed worse than the general models.
The results rang out like a gunshot. “I’ve never seen a single paper trigger the kind of reactions this one has in the health AI community,” wrote Kaiser Permanente’s vice president of AI and emerging technologies on LinkedIn. The paper’s findings, like all science, are subject to interpretation and debate — but many online reactions treated them more like a clear victory for general frontier models.
OpenEvidence, which describes its product as an “AI copilot” that helps doctors “make high-stakes decisions at the point of care,” went on defense. On X and LinkedIn, the company criticized the study for an “obvious conflict of interest and inattention to basic data contamination concerns.”
The stakes are high for patients — and the nascent industry. OpenEvidence, which claims to be the dominant LLM for doctors, has rocketed to a $12 billion valuation in five years since its founding. The company and its competitor Doximity, valued at nearly $4 billion, have been battling it out in court over allegations of stolen trade secrets, poached employees, and defamation through coordinated paid social media campaigns.
As the competition mounts, an academic paper isn’t just part of an evolving body of research; it can be a corporate threat. When scientists independently analyze clinical AI, their results can be weaponized in skirmishes between companies. And in some cases, they’re setting off a clash between scientists and the companies they study.
“Some of the tension that’s going on right now in the industry is not even between companies,” said Daniel Nadler, OpenEvidence’s founder and CEO. “It’s between academia and doctors.” Nadler sees academics trying to reassert a position of power over the clinical tools that get used, just as AI gives individual physicians more agency.
Researchers say the technology, which like many other clinical resources is not regulated by the Food and Drug Administration, is influencing clinical decision-making at enormous scale before its risks and benefits are well-understood. They want to fill the gap.
“As researchers who are nonaffiliated, we are the only ones who are able to say, well, actually this can produce harm,” said Ethan Goh, executive director of AI Research and Science Evaluation network, or ARISE, which interrogated that question in a preprint published this month. “Because why would someone who’s producing it want to highlight this when there’s so much at stake in terms of investment and valuation?”
But the research that gets done depends on the level of access companies are willing to offer researchers, and the funding available to understand the tools’ impact on real patient outcomes. As the evidence base evolves, how can doctors and health systems make decisions about which of these systems to trust — and when?
The benchmark back-and-forth
Unlike general chatbots like ChatGPT, AI tools targeted at doctors usually ground their responses in clinical text: guidelines, peer-reviewed research, expert summaries. But models using the technique, called retrieval-augmented generation, can grab the wrong references, said Eric Oermann, who directs the Health AI Research Lab at NYU Langone Health.
“All these tools that we use in medicine are pitching themselves like, we use RAG, we have all this great content that makes things better,” said Oermann, senior author on the Nature Medicine paper. “And we’re like, does it really?”
So Oermann and his colleagues threw a battery of tests at both general and clinical LLMs: 500 multiple-choice questions testing medical knowledge, 500 conversation-based questions released by OpenAI, and 100 real-life questions that NYU physicians had asked their internal version of GPT.
It took forever. While tech companies behind frontier models give researchers access to an API that makes it easier to test hundreds of queries, clinical AI companies typically don’t. “We had to manually sit there and enter every single query and get every single response,” Oermann said.
The results came out June 12: On all three tests, the clinical LLMs, OpenEvidence and UpToDate’s Expert AI, fell behind AI from Google, Anthropic, and OpenAl. On the real-world queries from NYU, OpenEvidence tied with one extra comparator — Google’s AI-generated search results.
It wasn’t a perfect experiment. Medical AI researchers have spent the last several years developing standardized tests for clinical tools — what the field calls benchmarks — like the multiple-choice MedQA and OpenAI’s HealthBench. But there’s a well-known problem: When a new benchmark is published in full, a model can easily incorporate the answer key, essentially cheating on the test.
The authors acknowledged those limitations. Careful readers dissected them online. And OpenEvidence made them the crux of its public criticism.
“These bad-faith things are comparing a model that has all the test answers and questions in the training data to one that has made the active choice not to,” Nadler said. “If we do that, besides misleading physicians, we’re sort of condoning the premise, and we kind of reject the premise of the entire thing.”
Benchmarks developed by an AI company, like HealthBench, also can be biased in their model’s favor, Nadler said — a point reiterated by Peter Bonis, chief medical officer at Wolters Kluwer Health, which owns UpToDate. While physicians helped develop the questions, “the scoring in HealthBench is then automated by LLMs,” he said, “and they tend to reward family resemblance.”
The last test used in the study — the 100 real-world queries from NYU docs — weren’t public, so they couldn’t have been used to train any of the models. But they were flawed too, OpenEvidence wrote, because the authors weren’t transparent about their selection process, and NYU doctors were developing their own clinical LLM.
“We should not accept that any model gets deployed without proper vetting and study and care and attention,” Bonis said. “But to go off to say that one performs better than the other, or generalist performs better, I think is overstating the degree of investigation that was conducted in that study.”
Three days after the NYU paper came out, OpenEvidence reached out to another set of researchers: Would they be interested in doing another head-to-head? Less than two weeks after that, the researchers had published their own preprint, which has not been peer-reviewed: Given more than 600 questions posed by OpenEvidence users and provided directly by the company, physicians rated OpenEvidence higher than three general LLMs on five measures, including accuracy, source quality, and verifiability.
This is the trouble with studying large language models: They’re finicky. Their clean, well-formatted text looks so reliable, so confident. But over and over, researchers have shown that they’re remarkably sensitive to small tweaks in the questions or prompts they’re given. Put those responses in the hands of subjective physician evaluators, and the signals can get even muddier.
“One potential reason” for the studies’ opposing results, said lead author Jean Feng, an associate professor of epidemiology and biostatistics at the University of California, San Francisco, “is the choice of questions and how the evaluation scheme was designed.” For example, NYU physicians tended to ask GPT short, concise questions, Oermann said; the ones posed to OpenEvidence in Feng’s study were far longer.
“Unsurprisingly,” Oermann said, “depending on how you use the models, you get different results.”
Clinical AI starts playing ball
Meanwhile, another research group was grappling with that problem.
“The main and fair critique is that benchmarks test different things,” said Goh, of the ARISE network. “It’s not like a doctor will use AI for just one thing.” So over the last few years, the group has developed a suite of benchmarks to measure different clinical capabilities.
Their latest test focused on what they considered a crucial unanswered question. “Is this stuff even safe? Let alone is it good — is it even just not dangerous?” said Jonathan Chen, a hospitalist and director for AI education at Stanford University.
Like NYU, the researchers were running into problems testing that benchmark, which it dubbed NOHARM. They had assembled more than 1,000 clinical questions, drawn from real consultations sent to specialists, and built a harm rating system for the laundry list of plausible responses. But they still didn’t have API access for the biggest-name clinical models.
“We reached out to them to try to partner with them, but it is their choice to offer an API, right?” Chen said. “Not everybody was playing ball.”
The researchers resorted to manual testing. They fed 100 benchmark questions, one by one, to major clinical LLMs, and shared companies’ individual results with them.
As their work continued, the Nature Medicine study came out. Amid the hubbub, ARISE published a version of its NOHARM work as a preprint, along with a subset of sample benchmark cases, its grading rubric, and code. The goal was to show the holdouts that “we’re not trying to trick or catch anybody,” Chen said. It worked.
The latest analysis, published as an updated preprint on July 13, includes API-based harm scores for four clinical LLMs — from OpenEvidence, Doximity, Amboss, and Glass Health.
“This is the very first time anyone’s gotten the clinical AIs to make that more transparent,” said Doximity CEO Jeff Tangney, “which, credit to Stanford and the ARISE team for doing that and getting to that.”
UpToDate’s Expert AI does not appear in the paper. “I think it would be really great if there was a set of benchmarks that were credible and were valid and truly did reflect how these things performed in the real world,” said Bonis, who pointed to an internal validation framework the company described in May. But “there is a reason to be a little bit cautious.” A Wolters Kluwer Health spokesperson said the company carefully evaluates research requests for API access: “Since research approaches vary widely, we want to collaborate with researchers to make progress on initiatives that can best measure the effectiveness of AI systems in the messy real world of clinical care.”
For the companies that opened up, transparency paid off. Every tested model produced some severe, harmful errors — most often, by leaving out critical recommendations for the clinical cases presented. But the four clinical LLMs showed less potential for harm: 2.9% to 5.4% of their cases introduced severe errors, significantly lower than the general LLMs’ 8.9% to 24.6%.
The test looked good for clinical AI, but not any particular company: None of the differences between the clinical models’ API results were big enough to be statistically significant. But the industry couldn’t resist capitalizing on the study.
Doximity crowed that it outranked OpenEvidence. OpenEvidence promoted that it was the most popular LLM when a separate trial in the study gave physicians the option to use any tool — and no one chose Doximity.
“All the social media drama is about company A vs. company B, and to us as the authors, it’s just not the point,” Chen said. “But this is a very high-stakes multibillion-dollar industry. So people got to say what they got to say, I guess.”
What really matters for doctors and health systems
Scientists acknowledge that no one test will ever tell the whole story about clinical AI performance. And day-to-day, the choice of when and how to use large language models in care — by individual doctors, and increasingly health system leaders — is shaped by more than research.
“I don’t see the benchmarks really playing a role among doctors,” said OpenEvidence’s Nadler. “Academics like talking about them.”
The more important question, companies, doctors, and researchers agree, is how these tools impact patients.
Anytime a new technology is introduced into medicine, its impact on those outcomes depends on how it’s used by physicians. Are they using the tool safely and correctly, on the right patients, in the right situations?
In the case of AI, that often comes down to trust. An accurate model inspires confidence, but it can still occasionally be wrong — potentially leading a physician astray, if they rely on it outright. Conversely, a good model can leave gains on the table if a physician doesn’t trust its output.
In a trial within the ARISE study that randomized doctors to use OpenAI’s GPT-5.4, conventional resources, or a tool of their choice, physicians using the AI performed better on the safety benchmark — but not as much better as they could have, because they frequently ignored the AI’s good advice.
“It’s not about model A vs. model B,” Chen said. “It’s how do you get humans and computers to do the combination you want?”
Travis Zack, OpenEvidence chief medical officer, sees doctors using the tool to surface vetted medical literature that can inform their decisions, not dictate them, as some benchmarks can suggest. But doctors know that just because research is peer-reviewed doesn’t mean it’s definitive. So this month, OpenEvidence launched an automated, LLM-powered system called EvidenceGrade that will help clinicians gauge their trust in an answer, and its citations. It will give a letter grade to an individual response based on factors such as study design and strength of the results.
OpenEvidence says it’s building on a pattern it sees in its users: A doctor will start by asking questions they know the answer to, use the responses to gauge their trust, and then start to branch out. “Because they went to medical school, because they’re experts, because they’re doctors, they’re able to calibrate the accuracy of these systems,” Nadler said.
Doctors also know an abstract summary or LLM-generated response can leave out critical context. “When you’re a specialist using this day-to-day, you see it gets stuff wrong,” Doximity’s Tangney said. Last December, Doximity announced its response to the trust question: a system called PeerCheck that calls on physician specialists to review a subset of AI-generated answers for accuracy, the strength of evidence they draw on, and potential bias.
“It’s less about what’s going on with academic benchmarks,” Tangney said, “and more about, hey, is this putting me and my license at risk — particularly if these tools are being deployed by a health system that’s using it to effectively upskill people.”
The question of liability is more pressing as developers make a concerted push to get hospitals to adopt their platforms. OpenEvidence notably got off the ground by allowing physicians free individual access and making money off the ads it served to those users. But more recently, it has started inking deals with hospitals. Doximity, whose original products were also ad-supported and targeted at doctors, similarly turned to enterprise over time; it says that 150 health systems now use its AI tools.
“It’s just a matter of time,” Tangney said, until health systems have to address legal concerns about outcomes associated with clinical AI.
The future of clinical LLM research
Design features can help physicians decide when and how much to trust an individual AI response, but they still can’t show whether they leave patients better or worse off.
“We need to know how safe these things are, when they should be relied upon,” Nadler said. But instead of “hand-wavy” benchmarks, he said, “I think we need to know those things to FDA-grade research levels.”
The question is how to design that real-world research — and how to fund it. While a yearslong randomized controlled trial is the gold standard to study patient outcomes, “the technology moves faster than RCTs can finish,” Chen said.
One option, Oermann said, would be large-scale studies that track patient outcomes for doctors with and without access to LLMs. Nadler similarly advocated for “quasi-natural experiments” that compare health systems that have adopted different tools due to chance. “Go look at readmissions rates, go look at medical error, go look at malpractice claims, go look at all-cause mortality,” Nadler said. “We would throw our hand up way up high and say, ‘Study the hell out of us, poke it.’”
That kind of study wouldn’t be impossible, Goh said, but there would be many confounding factors to control for. Another hard question is who would be able to drive that kind of study. Many researchers pointed to the collaboration between Penda Health, a Kenyan health system, and OpenAI, as a strong example of real-world, outcomes-based research. But industry-backed research will always come with conflict-of-interest concerns.
Until researchers, health systems, and industry can figure out how to execute more outcomes-based research, clinical LLM developers said they’d be open to providing API access to some groups. Nadler said OpenEvidence will provide access on a “case by case” basis, and “only to the extent that it allows us to be very vocal about giving them notes about how to do this properly” — no multiple-choice benchmarks, no tests that some LLMs have obviously memorized the answers to.
“There are independent benchmarks that are truly closed, that aren’t these open-book benchmarks that everyone can train to, and I do think there’s real value to them today,” said Tangney, who added Doximity will be “selective” about API access. “I also think there’s real value to longer-term trials too.”
In the meantime, Goh and Chen think it’s important to continue their work on benchmarks — despite the fact that they “don’t want to be the benchmark people,” Chen said. “Benchmarks are not the end answer, but they’re at least something that you can automate, streamline. They can keep pace a little bit, so you’re at least getting a signal,” he said. “Otherwise we’re unleashing things with just a guess and a vibe and a heuristic.”
