DAILY NEWS CLIP: August 20, 2026

How health systems are embracing chatbots to query and summarize patient records


STAT News – Thursday, August 20, 2026
By Katie Palmer

The pathologists were stumped. Six people had tried to identify a patient’s cancer based on a recent lymph node biopsy. They’d stained the cells 70 times to try to draw out more distinguishing features — but still, nobody had an answer.

At Stanford, a physician called onto the case was trying out a new tool called ChatEHR, one of several large language model-powered tools being deployed by health systems to summarize patients’ often-extensive medical records. It got a question: Did the patient have any history of skin lesions? After some back-and-forth, from the depths of the patient’s history, ChatEHR delivered an answer: In a different health system, the patient had previously been diagnosed with sarcomatoid squamous cell carcinoma.

It “completely explained the findings in the lymph node,” wrote the happy doctor in their feedback for the chatbot. “If that doesn’t prove the value of ChatEHR, I don’t know what does!”

This was the kind of needle in a haystack doctors hoped to find when health systems first started experimenting with generative AI tools like ChatEHR to search and synthesize patients’ health records. Clinicians often struggle to find the information they need to care for patients, because modern electronic health records have gotten so bloated. Today, a number of health systems are moving toward broad implementation of chatbots for EHRs, both homegrown and vendor-built. And it turns out that solving diagnostic mysteries is the least of their selling points.

“People like talking about this complex ‘zebra case’ that was only solved with technology with some heroic effort, but that’s not the majority,” said Nigam Shah, chief data scientist for Stanford Health Care. Much more of the feedback he’s gotten about ChatEHR comes from physicians who are using the tool to save time as they prepare to see patients every day. “This is real impact,” Shah said. “This overworked guy can now do 25 consults a day.”

As health systems increasingly turn to AI for chart summaries, they know no chatbot will be free of errors. Now, the question is what kind of mistakes summarization tools will make, what health systems will do to catch them — and how many errors are worth accepting in the name of efficiency and the odd zebra-case breakthrough. 

“We’re all worried about making sure that AI is trustworthy and safe,” said Jeffrey Ferranti, chief digital officer at Duke Health, which rolled out its chatbot Scout in July. But AI has reached an “interesting pivot point,” he said. In a randomized controlled trial recently published as a preprint, the health system found that Scout reduced task time by 37.6%. “The value being created, the things being caught, the nuances of the record that were coming through, were so important clinically and so impactful that expanding it seemed like the right thing to do.”

Errors of commission vs. omission

Just two years ago, AI researchers and health system leaders were nervous about using generative AI for something as high-stakes as accurately summarizing a patient’s history. “I don’t think anybody that I know of would trust [generative AI] to consume a patient’s medical record and create an accurate summary,” Rob Bart, chief medical information officer at University of Pittsburgh Medical Center, told STAT at the time.

Today, though, helping clinicians prepare for visits — whether in the minutes before a patient arrives at the ER, or the evening before a complex cancer patient arrives for a follow-up — is emerging as a dominant application of AI chart summarization tools.

“I think there is a growing comfort with its abilities to accurately summarize and accurately represent the patient’s health condition,” Bart said. UPMC, with a network of 40 hospitals across Pennsylvania, has implemented features of its ambient scribe Abridge that provide pre-visit summaries and allow clinicians to ask questions about patients’ medical records while drawing from the medical literature — Abridge just made those tools broadly available. Clinician chatbot OpenEvidence can be used similarly to answer questions about patients as it directly integrates with more health systems’ EHRs.

EHR vendor Epic says it is delivering 50 million inpatient and outpatient summaries every month to health system customers that have turned on that feature of its generative AI tool Art.

But it expects more systems to move from those “batched, pre-generated, one-size-fits-all” summaries “into dynamic, on demand” queries for its chatbot feature Ask Art, said vice president of research and development Garrett Adams. Four systems are piloting Ask Art, whose conversational interface mirrors homegrown EHR-summary systems from Stanford, Duke, Penn Medicine, and Children’s Hospital of Philadelphia — Stanford and Duke have broadly deployed the systems and made them available to thousands of users.

Behind the uptick in adoption are rapidly-advancing large language models, which have reduced early concerns about frequent hallucinations. Models have also gotten cheaper, making it more realistic for health systems to search through hundreds of pages of patient records and implement open-ended chatbot interfaces. Bart in particular used to be worried that a model would disregard important clinical information, but that hasn’t borne out, he said: “It understands that an echocardiogram of heart function might be more important than the next hemoglobin, for example.”

If hallucinating false information was the top concern for health systems in the earlier days of generative AI, today the focus is on what critical information summaries might leave out or misinterpret.

“If someone flags there’s an issue, that’s easy to go verify,” said Eugene Gitelman, executive director for transformation platforms at University of Pennsylvania Health System, whose EHR chatbot Chart Hero has been released to about 250 users (Penn is also piloting Epic’s Ask Art). The biggest challenge by far, according to Gitelman, is verifying whether a summary omitted something important from the patient record. 

In a proof-of-concept study of a 2024 implementation of Epic’s AI-generated summaries at UC San Diego Health, physicians identified omissions in 46 of 208 summaries, compared to only five hallucinations. During the first four months of Stanford’s ChatEHR rollout, every generation had about 0.73 hallucinations, 2.33 unsupported claims, and 1.6 inaccuracies, researchers shared in a recent comment in Nature Medicine.

“You might not want to omit your step of reviewing the chart, because there could be additional information that you might find,” said Nicolas Kahl, a clinical informaticist and emergency physician who led the UCSD study.

Users also have to watch out for how the models handle errors or conflicting information within the medical record itself. “‘Rule out pneumonia,’ ‘doesn’t have pneumonia,’ ‘evidence of pneumonia’ can appear in an ED doctor’s note, a hospitalist’s note, a radiologist’s note,” said Shaun Miller, chief health informatics officer at Cedars-Sinai in Los Angeles, which has given Epic’s inpatient summaries to all clinicians. Reconciling those positive and negative phrases in a summary has “always been looked at as the highest risk,” he said.

Trust, but verify

Early adopters point out that AI summaries shouldn’t be held to a standard of perfection. Human physicians aren’t perfect. Neither are their documents.

Sometimes, AI-based summaries can even help identify and correct discrepancies in the underlying patient record: Scout, for example, has flagged typos in the shorthand used to document how many times a patient has been pregnant and given birth. And when Art notes a discrepancy, Epic’s Adams said, “we’re able to either temporally take the latest thing where the error doesn’t exist, or flag that there’s something not quite right about this and fix it.”

Knowing those trade-offs exist, and that AI-based chart summaries will always include some errors, the next step is for health systems to build features and monitoring systems that make it easier to catch them. They’re learning as they go.

One approach to managing model performance is to put guardrails around their uses. There are functionally no limits to the questions a user can submit to a chatbot, and the task of “summarization” spans hundreds of distinct use cases. But some health systems, including Duke and Stanford, have built a menu of pre-specified tasks, automations, or prompts, which are easier to individually test and validate than the chatbot’s performance as a whole.

That removes some of the valuable versatility of a chatbot, though — so for those open-ended platforms, most AI-based summaries include feedback mechanisms for users to report when they see something off. Epic’s summaries, for example, include a thumbs-up or thumbs-down, and an option for the user to write in a free text response. “It’s hard to expect the whole clinical workforce to continually be providing feedback,” said Miller, so Cedars-Sinai has designated “super users” that are “tasked to remain vigilant.”

AI-based summaries almost always build in citations that take a user back to the original documentation, and training materials remind users that it’s always on them to verify the AI’s output. “It’s not AI operating on its own, it’s AI teeing things up for review and decision-making by physicians,” Duke’s Ferranti said. “Part of that decision-making is clicking these links and double-checking things in the record.”

When Stanford rolls out citations in ChatEHR, Shah said, they will install a logging system to see how often users actually click through. “I bet you lunch that the first month people click on it,” he said, “and then they’ll stop clicking on it.” That’s the commonly reported challenge of automation bias: As users get more comfortable with a new technology and find it to be generally reliable, they stop applying their own checks and balances.

That tendency to become less vigilant over time is part of why many health systems are moving to build at least semiautomated monitoring systems.

“By the nature of these being probabilistic, you can have drift, there’s model changes, there’s lots of things that sort of happen in there” that necessitate ongoing monitoring, said Nishit Patel, chief medical informatics officer at Tampa General, which in March became the first health system to pilot Epic’s Ask Art conversational tool, now having released it to about 50 users. “There are opportunities to use AI to monitor AI, with humans in the loop.”

That’s the approach Stanford took when it built VeriFact, an LLM-as-judge system that can check whether an AI-generated summary lines up with the underlying medical record. It has also implemented a system called CARE, which allows them to “set the level of hallucinations we’re accepting to let out from the generation,” Shah said. “The monitoring has to be computational.”

Combining chart summarization and model-based monitoring will be expensive — presenting health systems with challenging choices as their use expands. When you multiply the technology “across a million-plus encounters” with no additional revenue, Patel said, “how do you absorb those costs?” Stanford’s Shah is starting to contemplate hosting its own models so they don’t have to pay per-usage costs.

Those questions around cost will become more pressing as clinicians’ performance concerns are outweighed by the perceived benefits of AI-generated chart summaries. Published research has found that despite summary errors, users consider them acceptable to use in clinical practice. “We want to review every piece of information in the electronic health record, but it’s simply not possible if we want to be there to serve our patients,” said UCSD’s Kahl, the emergency physician. Stanford has found that 76% of users who try ChatEHR keep using it.

The next step is researching whether AI might even systematically improve the quality of care. Tampa General’s Patel recalled a cardiologist who shared how Ask Art helped him quickly surface details of a patient’s nearly decade-old stent that changed his treatment plan. “If I save that cardiologist 10 minutes of their time, that’s wonderful,” he said. “And we made patient care better because we identified things that helped in the planning of the next intervention for that patient.”

Access this article at its original source.

Digital Millennium Copyright Act Designated Agent Contact Information:

Communications Director, Connecticut Hospital Association
110 Barnes Road, Wallingford, CT
rall@chime.org, 203-265-7611