Skip To Content

NEWS

Seoul National University Hospital Develops Chest X-ray Report Encoder Based on Large Language Models

Hit : 7 Date : 2026-09-14

- Trained on Approximately 1.6 Million Paired Chest X-ray Images and Reports... Maintains Accuracy Even with Abbreviations and Formats Varying by Hospital

- Recorded a GREEN Score of 0.618... Actual Reports Can Be Traced, Raising Expectations for Integration with PACS (Picture Archiving and Communication System)

 

external_image

[Figure 1] Principle and example of LLM-based chest X-ray report retrieval Left) While existing BERT models classify reports with the same meaning as different sentences. LLM groups sentences that make sense together for recognition. (Right) Images and reports are each converted using encoders, placed in the same space, and the report closest to the image is selected from the candidates.


A research team from the Department of Radiology at Seoul National University Hospital has developed a retrieval technology to address the issue of varying X-ray report styles across different hospitals. The AI model, based on large language models (LLM), finds past reports that match chest X-ray images and is characterized by its ability to maintain consistent accuracy even with different abbreviations and narrative styles used by hospitals and doctors.

On the August 26th, a research team led by Professors Chang-Min Park and Dong-Heon Lee of the Department of Radiology at Seoul National University Hospital (with first author Researcher Han-Bin Ko) announced the development of a bidirectional chest X-ray imagetext retrieval AI model by integrating and analyzing approximately 1.6 million large-scale imagetext pairs, including globally available data and actual clinical data from Seoul National University Hospital.

Chest X-ray reports frequently use abbreviations that vary by institution, such as "BLLF" for bilateral lower lung field and "PTX" for pneumothorax. Formatting is also inconsistent, as some hospitals record findings in detail while others leave only a brief conclusion.

As such, the format and abbreviations of reports differ from hospital to hospital. Consequently, existing AI systems using BERT text encoders*a type of AI language modelstruggled to accurately connect images and sentences. There were also limitations, such as performance stagnating or even declining as training data was increased.

*Encoder: A device that converts text, images, etc., into numbers that a computer can understand.

To address this problem, the research team converted a large language model into a sentence encoder instead of BERT. They changed the existing unidirectional structure, which reads sentences only from front to back, into a bidirectional structure that reads the surrounding context as well. They created a retrieval-type encoder that finds existing reports, rather than a generative model that writes new reports. Next, the training data was increased to a total of approximately 1.1 million cases by adding AI-generated variant sentences to the original text, and then an encoder (LLM2VEC4CXR) was completed that groups texts with the same meaning closely together through masked token prediction (MTP) and supervised contrastive learning.

The research team combined this encoder with a vision encoder to create the "LLM2CLIP4CXR" model. This method converts both images and reports into numbers, places them in the same space, and compares them. Subsequently, the model was trained on a total of approximately 1.6 million imagereport pairs by adding de-identified data from Seoul National University Hospital and public datasets (MIMIC-CXR, CheXpert-plus, and PadChest), and its performance was verified using an internal validation dataset (MIMIC-CXR) and an external validation set (Open-I).

 

그림입니다.  원본 그림의 이름: [Figure 2].jpg  원본 그림의 크기: 가로 1180pixel, 세로 1050pixel  프로그램 이름 : BandiView 7.23 (quality: 100)

[Figure 2] Performance comparison of LLM2VEC4CXR compared to existing encoders

 

그림입니다.  원본 그림의 이름: [Figure 3].jpg  원본 그림의 크기: 가로 790pixel, 세로 69pixel

[Figure 3] Ranking evaluation by thoracic radiology specialists for retrieval method (LLM) retrieval method (BERT) generation method (MAIRA-2) using 72 randomly selected samples out of 200 from Open-I data: Rankings were calculated based on only three representative models, and a lower average rank indicates a higher clinical agreement with the correct answer report.

 

The analysis results showed that on the internal validation dataset (MIMIC-CXR), the GREEN score (where a value closer to 1 indicates greater accuracy) representing clinical accuracy was 0.308, the highest among the comparison models. On the external validation set (Open-I), the GREEN score was recorded at 0.618, and the rate of finding the correct answer within the top-1 ranking was 4.9%, exceeding the highest value (2.9%) among retrieval-based comparison models; thus, the model outperformed the comparison models in both retrieval accuracy and clinical indicators.

Notably, accuracy was maintained or improved even when the model was additionally trained on reports containing many abbreviations and written briefly with a focus on conclusions. In contrast, the accuracy of existing BERT-based models declined as the number of reports increased.

Its excellence was also confirmed in expert evaluations. As a result of evaluating 200 cases on the external validation set by three medical students and four large language models, the research team's model received the highest rating. Among the 72 cases separately reviewed by thoracic radiologists, 76% selected the results of this model as their top choice.

The research team predicted that this technology could be utilized in future PACS/RIS environments as a "retrieval layer" that finds past examinations or similar cases based on semantics. They cited the ability to trace the basis of results and allow specialists to verify them directly as an advantage, given that the method involves finding and presenting actual existing images and reports.

Professor Chang-Min Park (Department of Radiology) emphasized, "The performance of medical AI depends more on how deeply it understands context and clinical knowledge than on the amount of data," adding, "This retrieval method has the advantage of allowing the tracing of the basis of results."

Professor Dong-Heon Lee (Department of Radiology) stated, "This study has opened the way to utilizing unstructured reports in their original form," and added, "We plan to expand the scope to include chest CT scans and pursue multi-institutional follow-up verification to confirm its usefulness in actual clinical practice."

Meanwhile, this study was published in the international academic journal IEEE Journal of Biomedical and Health Informatics (JBHI), and the developed model was released for research purposes on the global AI platforms Hugging Face and GitHub.

 

Public model link

- https://huggingface.co/lukeingawesome/llm2vec4cxr

- https://github.com/lukeingawesome/llm2vec4cxr

 

그림입니다.  원본 그림의 이름: [사진 왼쪽부터] 서울대병원 영상의학과 박창민, 이동헌 교수.jpg  원본 그림의 크기: 가로 1304pixel, 세로 914pixel

[From left] Professors Chang-Min Park and Dong-Heon Lee of the Department of Radiology, Seoul National University Hospital

전체 메뉴

전체 검색

전체 검색