๐Ÿ“š Full Archive

Radiologist-Specific Report Style Adaptation
โ† Dashboard ๐Ÿ“„ 121 papers ยท Generated 2026-08-23 00:10

1Assessing diagnostic performance of multimodal AI and human experts in oral and maxillofacial radiography: a comparative analysis of ChatGPT, Grok, and MANUS.

2026-12Annals of medicineโญ Q1DOI 10.1080/07853890.2026.2664903
BACKGROUND

Artificial intelligence (AI), particularly large language models (LLMs), is increasingly applied to radiographic interpretation in healthcare. In dentistry, radiographic imaging is essential for diagnosis and treatment planning, yet remains subject to variability and human error. AI may enhance diagnostic accuracy and consistency.

OBJECTIVE

To evaluate and compare the diagnostic accuracy, consistency, and interpretive performance of multimodal AI models-ChatGPT, Grok, and MANUS-with expert radiologists in dental radiograph interpretation.

METHODS

A total of 120 anonymised radiographs (40 OPGs, 40 periapical, 40 CT slices) were selected from validated academic sources. Two board-certified oral and maxillofacial radiologists established gold standard diagnoses. Each image was independently assessed by the three AI models under standardised prompting. Diagnostic accuracy and intra-model consistency were analysed using descriptive statistics, Cohen's kappa, McNemar's test, and logistic regression.

RESULTS

In the first assessment, MANUS and ChatGPT achieved 92.5% accuracy (111/120), while Grok reached 88.3% (106/120). Performance improved in the second round: MANUS 95.0%, ChatGPT 93.3%, and Grok 90.8%, compared with 96.7% for radiologists. ChatGPT showed the highest reproducibility (ฮบ = 0.937), whereas MANUS demonstrated the highest overall accuracy. Strong agreement was observed between ChatGPT and MANUS, with greater variability in Grok. No significant systematic bias was detected between AI outputs and radiologist benchmarks.

CONCLUSION

The evaluated LLMs demonstrated diagnostic performance comparable to expert radiologists. MANUS excelled in accuracy and ChatGPT in reproducibility, supporting their potential as adjunct tools in dental radiology, while maintaining the need for expert oversight.Clinical trial number: Not applicable.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์น˜๊ณผ ๋ฐฉ์‚ฌ์„  ์˜์ƒ ํŒ๋…์—์„œ ChatGPT, Grok, MANUS ์„ธ ๊ฐ€์ง€ ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ AI ๋ชจ๋ธ์˜ ์ง„๋‹จ ์ •ํ™•๋„์™€ ์ผ๊ด€์„ฑ์„ ๊ตฌ๊ฐ•์•…์•ˆ๋ฉด๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜์™€ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. OPG, ์น˜๊ทผ๋‹จ ๋ฐฉ์‚ฌ์„ ์‚ฌ์ง„, CT ์Šฌ๋ผ์ด์Šค๋ฅผ ํฌํ•จํ•œ ์ด 120๊ฐœ์˜ ์ต๋ช…ํ™”๋œ ์˜์ƒ์„ ๋Œ€์ƒ์œผ๋กœ ํ‘œ์ค€ํ™”๋œ ํ”„๋กฌํ”„ํŠธ ์กฐ๊ฑด ํ•˜์— ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, MANUS๊ฐ€ ์ตœ๊ณ  ์ •ํ™•๋„(95.0%), ChatGPT๊ฐ€ ์ตœ๊ณ  ์žฌํ˜„์„ฑ(ฮบ = 0.937)์„ ๊ธฐ๋กํ•˜์˜€์œผ๋ฉฐ, ์ „๋ฌธ์˜(96.7%)์™€ ๋น„๊ตํ•ด ํ†ต๊ณ„์ ์œผ๋กœ ์œ ์˜ํ•œ ์ฒด๊ณ„์  ํŽธํ–ฅ์€ ๊ด€์ฐฐ๋˜์ง€ ์•Š์•˜๋‹ค. ์ด ๊ฒฐ๊ณผ๋Š” LLM ๊ธฐ๋ฐ˜ AI ๋ชจ๋ธ์ด ์น˜๊ณผ ๋ฐฉ์‚ฌ์„  ํŒ๋…์—์„œ ์ „๋ฌธ์˜ ์ˆ˜์ค€์— ๊ทผ์ ‘ํ•œ ์ง„๋‹จ ์„ฑ๋Šฅ์„ ๋ฐœํœ˜ํ•˜๋ฉฐ ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ์˜ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ์ง€๋‹ˆ๋‚˜, ์—ฌ์ „ํžˆ ์ „๋ฌธ๊ฐ€ ๊ฐ๋…์ด ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-05-10 00:00View โ†—

2Error Detection and Correction in Chinese Radiology Reports Using Large Language Models: Real-World Clinical Validation Study.

2026-08Journal of medical Internet researchโญ Q1DOI 10.2196/94689
BACKGROUND

Large language models (LLMs) show promise in automatically detecting errors in radiology reports, but their performance remains insufficiently validated in large-scale, real-world clinical datasets.

OBJECTIVE

This study aimed to systematically evaluate the performance of LLMs in detecting and correcting errors in Chinese radiology reports derived from authentic clinical data.

METHODS

A large-scale dataset of 4480 Chinese radiology reports with modification records containing real clinical practice-generated errors was retrospectively collected between January 2023 and June 2024 at a single institution. After exclusions, 1363 reports containing 1551 errors were included. The dataset covers various anatomical parts of the body from different imaging modalities and was randomly divided into a test set (n=1263) and an internal validation set (n=100). Additionally, 100 error-free reports were added to the internal validation set. An additional 200 English-language reports from the Medical Information Mart for Intensive Care (MIMIC-III) were used for external validation. Eight human readers and 8 widely adopted LLMs, enhanced by prompt engineering, were tasked with error detection. Overall and subgroup detection performance and reading time were evaluated. Correction suggestions from the 2 best-performing LLMs were reviewed by a senior radiologist.

RESULTS

On the test set, DeepSeek-R1 achieved the highest overall detection rate at 89% (95% CI 87%-90%), significantly better than the other 7 models (P=.001-.007). On the internal validation set, DeepSeek-R1 and Claude-3.5-Sonnet achieved detection rates of 83% (100/120; 95% CI 76%-89%) and 80% (96/120; 95% CI 72%-86%), respectively. DeepSeek-R1 showed performance comparable to radiologists (83%, 95% CI 76%-89% vs 80%, 95% CI 72%-86% for junior radiologists and 78%, 95% CI 70%-85% for senior radiologists; P=.39 and P=.19, respectively) and significantly better performance than that of nonradiologists and nonphysicians (83%, 95% CI 76%-89% vs 66%, 95% CI 57%-74% and 38%, 95% CI 30%-47%; P<.001, respectively). DeepSeek-R1 showed a false-positive rate comparable to radiologists (DeepSeek-R1 vs senior radiologists and junior radiologists, 3% vs 0% and 1%; P=.25 and P=.61, respectively) and a significantly lower rate than nonradiologists and nonphysicians (3% vs 13% and 17%; P=.02 and P=.002, respectively). On the external validation set, DeepSeek-R1 and Claude-3.5-Sonnet achieved detection rates of 94% (95% CI 89%-97%) and 93% (95% CI 88%-97%), respectively. The correction accuracy of DeepSeek-R1 and Claude-3.5-Sonnet was 95% and 91%, respectively.

CONCLUSION

Enhanced LLMs, particularly DeepSeek-R1, demonstrated robust performance in error detection and correction within real-world Chinese radiology reports, supporting their clinical use for automated quality assurance and integration into workflows to improve reporting accuracy and efficiency.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์‹ค์ œ ์ž„์ƒ ํ™˜๊ฒฝ์—์„œ ์ค‘๊ตญ์–ด ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ์˜ ์˜ค๋ฅ˜๋ฅผ ์ž๋™์œผ๋กœ ํƒ์ง€ํ•˜๊ณ  ๊ต์ •ํ•˜๋Š” ์„ฑ๋Šฅ์„ ์ฒด๊ณ„์ ์œผ๋กœ ํ‰๊ฐ€ํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ๋‹จ์ผ ๊ธฐ๊ด€์—์„œ 2023๋…„ 1์›”๋ถ€ํ„ฐ 2024๋…„ 6์›”๊นŒ์ง€ ์ˆ˜์ง‘๋œ 1,551๊ฑด์˜ ์˜ค๋ฅ˜๋ฅผ ํฌํ•จํ•œ 1,363๊ฐœ์˜ ์‹ค์ œ ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ, ํ”„๋กฌํ”„ํŠธ ์—”์ง€๋‹ˆ์–ด๋ง์œผ๋กœ ๊ฐ•ํ™”๋œ 8์ข…์˜ LLM๊ณผ ์ธ๊ฐ„ ํŒ๋…์ž์˜ ์˜ค๋ฅ˜ ํƒ์ง€ ์„ฑ๋Šฅ์„ ๋น„๊ต ๋ถ„์„ํ•˜์˜€๋‹ค. DeepSeek-R1์ด ์ „์ฒด ์˜ค๋ฅ˜ ํƒ์ง€์œจ 89%, ๊ต์ • ์ •ํ™•๋„ 95%๋กœ ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜ ์ˆ˜์ค€์— ํ•„์ ํ•˜๋Š” ํƒ์ง€์œจ๊ณผ ๋‚ฎ์€ ์œ„์–‘์„ฑ๋ฅ ์„ ๋‹ฌ์„ฑํ•˜์—ฌ ์ž„์ƒ ํ’ˆ์งˆ ๋ณด์ฆ ์ž๋™ํ™” ๋„๊ตฌ๋กœ์„œ์˜ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-08-23 00:00View โ†—

3The consensus-based CINEX guideline for reporting clinical information extraction studies.

2026-08Journal of the American Medical Informatics Association : JAMIAโญ Q1DOI 10.1093/jamia/ocag136
OBJECTIVE

Information extraction (IE) from clinical texts has advanced rapidly with recent advances in natural language processing, particularly the advent of large language models (LLMs). However, inconsistent and incomplete reporting of methodologies limits reproducibility, comparability, and clinical translation. We aimed to develop a consensus-based reporting guideline tailored to clinical IE studies.

METHODS

We developed the Clinical Information Extraction Reporting Guideline (CINEX) through a multi-phase process. The initiative was prospectively registered on the EQUATOR Network as a reporting guideline under development, and a detailed Delphi study protocol was published in advance. A scoping review informed an initial set of items, which was refined through a 3-round electronic Delphi study with 20 international experts, followed by a final consensus meeting. Items were iteratively refined based on predefined inclusion criteria and expert feedback.

RESULTS

The final CINEX guideline comprises 29 checklist items grouped into 5 domains: information model, architecture, data, annotation, and outcomes. The 3 eDelphi rounds included 20, 15, and 12 experts, respectively. Two items were added after round one. Consensus for inclusion was reached for a 21 and additional 7 items after rounds 2 and 3, respectively.

CONCLUSION

CINEX provides a structured framework to improve transparency, reproducibility, and interpretability in clinical IE research. By standardizing reporting of key methodological components, such as data provenance, annotation processes, and evaluation strategies, it facilitates meaningful comparison across studies and supports safer clinical implementation.

CONCLUSION

CINEX complements existing AI reporting standards by addressing domain-specific challenges in clinical IE.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์ž์—ฐ์–ด ์ฒ˜๋ฆฌ ๋ฐ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์˜ ๋ฐœ์ „์œผ๋กœ ๊ธ‰์†ํžˆ ์„ฑ์žฅํ•œ ์ž„์ƒ ์ •๋ณด ์ถ”์ถœ(Information Extraction, IE) ์—ฐ๊ตฌ ๋ถ„์•ผ์—์„œ ๋ฐฉ๋ฒ•๋ก  ๋ณด๊ณ ์˜ ๋น„์ผ๊ด€์„ฑ ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๊ธฐ ์œ„ํ•ด ํ•ฉ์˜ ๊ธฐ๋ฐ˜ ๋ณด๊ณ  ์ง€์นจ์ธ CINEX๋ฅผ ๊ฐœ๋ฐœํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ๊ฐœ๋ฐœ ๋ฐฉ๋ฒ•์œผ๋กœ๋Š” ๋ฒ”์œ„ ๋ฌธํ—Œ๊ณ ์ฐฐ์„ ํ†ตํ•ด ์ดˆ๊ธฐ ํ•ญ๋ชฉ์„ ๋„์ถœํ•œ ํ›„, ๊ตญ์ œ ์ „๋ฌธ๊ฐ€ 20๋ช…์„ ๋Œ€์ƒ์œผ๋กœ 3๋ผ์šด๋“œ์˜ ์ „์ž ๋ธํŒŒ์ด(eDelphi) ์—ฐ๊ตฌ์™€ ์ตœ์ข… ํ•ฉ์˜ ํšŒ์˜๋ฅผ ์‹œํ–‰ํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, ์ •๋ณด ๋ชจ๋ธยท์•„ํ‚คํ…์ฒ˜ยท๋ฐ์ดํ„ฐยท์–ด๋…ธํ…Œ์ด์…˜ยท๊ฒฐ๊ณผ ํ‰๊ฐ€์˜ 5๊ฐœ ์˜์—ญ์œผ๋กœ ๊ตฌ์„ฑ๋œ ์ด 29๊ฐœ ์ฒดํฌ๋ฆฌ์ŠคํŠธ ํ•ญ๋ชฉ์ด ๋„์ถœ๋˜์—ˆ์œผ๋ฉฐ, ์ด๋Š” ์ž„์ƒ IE ์—ฐ๊ตฌ์˜ ํˆฌ๋ช…์„ฑ, ์žฌํ˜„์„ฑ ๋ฐ ์ž„์ƒ ์ ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ํ–ฅ์ƒ์‹œํ‚ค๋Š” ํ‘œ์ค€ํ™”๋œ ๋ณด๊ณ  ์ฒด๊ณ„๋ฅผ ์ œ๊ณตํ•œ๋‹ค.
Added: 2026-08-16 00:00View โ†—

4AI-powered radiology report simplification in Arabic: A prospective evaluation of patient-perceived understandability and clinical safety.

2026-08International journal of medical informaticsโญ Q1DOI 10.1016/j.ijmedinf.2026.106637
BACKGROUND

Radiology reports are written for clinicians, leaving the majority of patients unable to understand their own imaging findings. Large language models (LLMs) offer a means to simplify reports for patient-facing use; however, most implementations rely on cloud-based platforms that raise data privacy and regulatory concerns. Arabic-speaking populations - over 400 million people worldwide - remain substantially underserved in medical AI research, with no prospectively evaluated AI Arabic radiology report simplification system previously reported.

OBJECTIVE

This study aimed to conduct a preliminary prospective feasibility and safety evaluation of patient-perceived understandability outcomes and radiologist-assessed clinical safety of an on-site, institutionally governed AI system that generates simplified radiology reports in plain English and Arabic for Arabic-speaking outpatients.

METHODS

In this prospective single-center observational study conducted at King Faisal Specialist Hospital and Research Centre - Jeddah (KFSHRC-J) in January 2026, 98 adult outpatients (mean age 50.0ย ยฑย 14.9ย years; 50 men, 48 women; IRB #2251488) reviewed three versions of their own radiology report in randomized order: a traditional radiologist report, an AI-simplified English report, and an AI-simplified Arabic report. The system used a two-stage pipeline: Qwen3-14B-FP8 with structured prompt engineering for English simplification (Stage 1), and a fine-tuned Hala-1.2B model for Arabic translation (Stage 2; 8,319 training samples). Patient-perceived understandability was measured on a 5-point Likert scale (1ย =ย Very Easy; 5ย =ย Very Difficult). Three board-certified radiologists independently assessed clinical safety using a standardized seven-domain rubric; inter-rater agreement was quantified using ICC(2,k).

RESULTS

AI-simplified Arabic reports produced a median perceived understandability score of 1 (IQR 1-2) versus 5 (IQR 3-5) for traditional reports (pย <ย 0.001; rank-biserial rย =ย 0.97), with 94.9% (93/98) of patients preferring the Arabic simplified version. Safety evaluation demonstrated 96.9% (95/98) of reports safe for patient release, with a hallucination rate of 3.1% (nย =ย 3; 1 unsafe, 2 safe by consensus). Inter-rater reliability was moderate by conventional thresholds (ICC 0.48-0.55) due to a ceiling effect from uniformly high safety scores (mean fidelity 4.8/5.0; 90.8% highest rating; SDย =ย 0.3), with 91.2% absolute agreement on binary safety classification and Fleiss' kappaย =ย 0.71 for safety.

CONCLUSION

In this preliminary prospective feasibility evaluation, a locally deployed, clinician- and AI team-built system significantly improved patient-perceived understandability of radiology reports for Arabic-speaking patients, with a radiologist-assessed safety profile that is promising but requires further validation before broad clinical deployment. The hallucination rate of 3.1% (1.0% unsafe) is lower than published pooled benchmarks, though direct comparisons are limited by methodological heterogeneity across studies (pooled 7.2% across 38 simplification studies; 1.12% for fine-tuned models). To our knowledge, this is among the first prospectively evaluated AI systems for Arabic radiology report simplification, highlighting the potential of on-site AI deployment as a safe, privacy-preserving approach for multilingual radiology communication, warranting further investigation of its contribution to health equity.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์•„๋ž์–ด๊ถŒ ์™ธ๋ž˜ ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ, ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ๋ฅผ ํ‰์ดํ•œ ์˜์–ด ๋ฐ ์•„๋ž์–ด๋กœ ์ž๋™ ๊ฐ„์†Œํ™”ํ•˜๋Š” ์›๋‚ด ๊ตฌ์ถ•ํ˜• AI ์‹œ์Šคํ…œ์˜ ์‹คํ˜„ ๊ฐ€๋Šฅ์„ฑ๊ณผ ์ž„์ƒ ์•ˆ์ „์„ฑ์„ ์ „ํ–ฅ์ ์œผ๋กœ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. Qwen3-14B-FP8 ๋ชจ๋ธ ๊ธฐ๋ฐ˜ ์˜์–ด ๋‹จ์ˆœํ™” ํ›„ ๋ฏธ์„ธ ์กฐ์ •๋œ Hala-1.2B ๋ชจ๋ธ๋กœ ์•„๋ž์–ด ๋ฒˆ์—ญํ•˜๋Š” 2๋‹จ๊ณ„ ํŒŒ์ดํ”„๋ผ์ธ์„ ์‚ฌ์šฉํ•˜์—ฌ, 98๋ช…์˜ ์„ฑ์ธ ์™ธ๋ž˜ ํ™˜์ž์—๊ฒŒ ๊ธฐ์กด ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ, AI ๊ฐ„์†Œํ™” ์˜์–ด ๋ณด๊ณ ์„œ, AI ๊ฐ„์†Œํ™” ์•„๋ž์–ด ๋ณด๊ณ ์„œ๋ฅผ ๋ฌด์ž‘์œ„ ์ˆœ์„œ๋กœ ์ œ์‹œํ•˜์˜€๋‹ค. AI ๊ฐ„์†Œํ™” ์•„๋ž์–ด ๋ณด๊ณ ์„œ์˜ ํ™˜์ž ์ธ์ง€ ์ดํ•ด๋„ ์ค‘์•™๊ฐ’์€ 1์ (๋งค์šฐ ์‰ฌ์›€)์œผ๋กœ ๊ธฐ์กด ๋ณด๊ณ ์„œ์˜ 5์ (๋งค์šฐ ์–ด๋ ค์›€) ๋Œ€๋น„ ์œ ์˜ํ•˜๊ฒŒ ์šฐ์ˆ˜ํ•˜์˜€์œผ๋ฉฐ(p < 0.001), ์ „๋ฌธ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ ํ‰๊ฐ€์—์„œ 96.9%์˜ ๋ณด๊ณ ์„œ๊ฐ€ ํ™˜์ž ์ œ๊ณต์— ์•ˆ์ „ํ•œ ๊ฒƒ์œผ๋กœ ํŒ์ •๋˜์–ด ์•„๋ž์–ด๊ถŒ ํ™˜์ž๋ฅผ ์œ„ํ•œ ํ”„๋ผ์ด๋ฒ„์‹œ ๋ณดํ˜ธํ˜• ํ˜„์ง€ AI ๋ฐฐํฌ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-08-09 00:00View โ†—

5Automatic Patient Eligibility for Photon-counting CT Using Discriminative and Generative AI Models in Neuroradiology.

2026-08Academic radiologyโญ Q1DOI 10.1016/j.acra.2026.07.026

RATIONALE AND

OBJECTIVE

Photon -Counting CT (PCCT) provides higher resolution, reduced dose, and valuable spectral data, but generates a large number of images per study, underlining radiologist and Picture Archiving and Communication Systemsย (PACS) capacity. While no guidelines defining when PCCT should be preferred over conventional Energy Integrating Detector (EID) CT in neuroradiology, and manual decision-making being impractical, Natural Language Processing (NLP)-based large language models (LLMs) may help automate routing of CT requests to the most appropriate CT technology.

METHODS

We conducted a retrospective study using Spanish-language neuroradiology CT requests retrieved from the Radiology Information System (RIS) (January 2012-October 2025). A random sample of 800 requests was independently labeled by two neuroradiologists as Basic Protocol (BP); feasible on EID-CT or basic PCCT) or Advanced Protocol (AP); requiring full-advanced PCCT), considering the gold standard an expert consensus (Cohen's k = 0.74). For discriminative modeling, transformer classifiers (Spanish bidirectional encoder representations from transformers (BERT), biomedical RoBERTa, Llama-3.1-8B, and Mistral-7B) were fine-tuned for binary BP/AP classification and evaluated with stratified five-fold cross-validation. For generative modeling, ChatGPT 5.2 and Gemini 3 were applied in a zero-shot setting using structured prompting and post-processed into binary outputs. Performance was assessed using accuracy, precision, recall, specificity, F1-score, and AUC.

RESULTS

Discriminative transformers achieved the best performance for the classification, with RoBERTa leading (accuracy 0.925, precision 0.906, recall 0.967, F1 0.935, AUC 0.920) and the lowest total error, without a significant difference versus Spanish BERT. In contrast, RoBERTa outperformed Llama, Mistral, Gemini, and ChatGPT with statistically significant improvements, while the generative LLMs yielded the poorest overall performance.

CONCLUSION

Discriminative transformer models, particularly RoBERTa, enabled highly accurate automatic identification of CT requests requiring AP to be performed at PCCT, substantially outperforming general-purpose generative LLMs. These findings support this exploratory proof-of-concept study of AI-driven eligibility routing to optimize PCCT utilization and streamline neuroradiology workflows.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์‹ ๊ฒฝ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์—ญ์—์„œ ๊ด‘์ž๊ณ„์ˆ˜ CT(PCCT)์™€ ๊ธฐ์กด ์—๋„ˆ์ง€ํ†ตํ•ฉ๊ฒ€์ถœ๊ธฐ CT ์ค‘ ์ ํ•ฉํ•œ ์žฅ๋น„๋ฅผ ์ž๋™์œผ๋กœ ์„ ํƒํ•˜๊ธฐ ์œ„ํ•ด ์ž์—ฐ์–ด์ฒ˜๋ฆฌ ๊ธฐ๋ฐ˜ AI ๋ชจ๋ธ์˜ ์œ ์šฉ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ŠคํŽ˜์ธ์–ด๋กœ ์ž‘์„ฑ๋œ 800๊ฑด์˜ CT ์ดฌ์˜ ์˜๋ขฐ์„œ๋ฅผ ์ „๋ฌธ์˜ ํ•ฉ์˜๋ฅผ ๊ธฐ์ค€์œผ๋กœ ๊ธฐ๋ณธ ํ”„๋กœํ† ์ฝœ ๋ฐ ๊ณ ๊ธ‰ ํ”„๋กœํ† ์ฝœ๋กœ ์ด์ง„ ๋ถ„๋ฅ˜ํ•˜์—ฌ, ํŒ๋ณ„ํ˜• ํŠธ๋žœ์Šคํฌ๋จธ(Spanish BERT, biomedical RoBERTa, Llama-3.1-8B, Mistral-7B) ๋ฏธ์„ธ์กฐ์ • ๋ชจ๋ธ๊ณผ ์ƒ์„ฑํ˜• ๋Œ€๊ทœ๋ชจ ์–ธ์–ด๋ชจ๋ธ(ChatGPT, Gemini) ์ œ๋กœ์ƒท ๋ฐฉ์‹์˜ ์„ฑ๋Šฅ์„ ๋น„๊ตํ•˜์˜€๋‹ค. ํŒ๋ณ„ํ˜• ํŠธ๋žœ์Šคํฌ๋จธ, ํŠนํžˆ RoBERTa๊ฐ€ ์ •ํ™•๋„ 0.925, AUC 0.920์œผ๋กœ ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ๋ฒ”์šฉ ์ƒ์„ฑํ˜• ๋ชจ๋ธ๋“ค์„ ํ†ต๊ณ„์ ์œผ๋กœ ์œ ์˜ํ•˜๊ฒŒ ์ƒํšŒํ•˜์—ฌ AI ๊ธฐ๋ฐ˜ PCCT ๋Œ€์ƒ ํ™˜์ž ์ž๋™ ์„ ๋ณ„ ๋ฐ ์‹ ๊ฒฝ๋ฐฉ์‚ฌ์„ ๊ณผ ์›Œํฌํ”Œ๋กœ์šฐ ์ตœ์ ํ™”์— ๋Œ€ํ•œ ๊ฐ€๋Šฅ์„ฑ์„ ์ œ์‹œํ•˜์˜€๋‹ค.
Added: 2026-08-09 00:00View โ†—

6Improved Readability and Translational Instability in LLM-Generated Radiology Reports.

2026-08Academic radiologyโญ Q1DOI 10.1016/j.acra.2026.07.040
BACKGROUND

Large language models (LLMs) show promise for converting complex radiology reports into patient-centric language, but inherent output instability may limit clinical application.

OBJECTIVE

To quantitatively assess the translational accuracy, error rates, and instability of various LLMs when generating patient-centric radiology reports, and evaluate demographic influences on report readability.

METHODS

This retrospective study evaluated 320 de-identified radiology reports processed by three LLMs using a two-stage (baseline and optimized) prompt engineering strategy. Two senior radiologists evaluated medical accuracy, completeness, and recommendation suitability. Readability was evaluated by 16 non-medical participants stratified by age and education.

RESULTS

Professional radiological evaluation revealed that all tested models exhibited inherent instability, omitted information, and tended to generate risk-averse, generalized clinical recommendations. To address these limitations, optimized structured prompts significantly reduced model output variance and improved translational accuracy, with particularly prominent effects observed in DeepSeek-R1 and ChatGPT-4.0. Overall, large language models significantly enhanced the readability of radiology reports (P < 0.05), with DeepSeek-R1 achieving the best performance. However, patients' self-reported comprehension of the reports was affected by demographic characteristics.

CONCLUSION

Large language models can effectively improve the readability of radiology reports, yet all such models inherently suffer from output instability and information omission. Optimized structured prompting can substantially reduce the variability of model outputs and improve the accuracy of medical text translation. Nevertheless, LLMs should currently be strictly confined to human-supervised auxiliary tools rather than applied as standalone clinical solutions.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ๋ณต์žกํ•œ ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ๋ฅผ ํ™˜์ž ์ค‘์‹ฌ์˜ ์–ธ์–ด๋กœ ๋ณ€ํ™˜ํ•  ๋•Œ์˜ ๋ฒˆ์—ญ ์ •ํ™•๋„, ์˜ค๋ฅ˜์œจ ๋ฐ ์ถœ๋ ฅ ๋ถˆ์•ˆ์ •์„ฑ์„ ์ •๋Ÿ‰์ ์œผ๋กœ ํ‰๊ฐ€ํ•˜๊ณ , ์ธ๊ตฌํ†ต๊ณ„ํ•™์  ์š”์ธ์ด ๊ฐ€๋…์„ฑ์— ๋ฏธ์น˜๋Š” ์˜ํ–ฅ์„ ๋ถ„์„ํ•˜์˜€๋‹ค. 320๊ฑด์˜ ๋น„์‹๋ณ„ํ™”๋œ ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ ์„ธ ๊ฐ€์ง€ LLM(DeepSeek-R1, ChatGPT-4.0 ํฌํ•จ)์— 2๋‹จ๊ณ„ ํ”„๋กฌํ”„ํŠธ ์—”์ง€๋‹ˆ์–ด๋ง ์ „๋žต์„ ์ ์šฉํ•˜์˜€์œผ๋ฉฐ, 2๋ช…์˜ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜์™€ 16๋ช…์˜ ๋น„์ „๋ฌธ ํ‰๊ฐ€์ž๊ฐ€ ๊ฐ๊ฐ ์˜ํ•™์  ์ •ํ™•์„ฑ๊ณผ ๊ฐ€๋…์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ๋ชจ๋“  ๋ชจ๋ธ์—์„œ ๊ณ ์œ ํ•œ ์ถœ๋ ฅ ๋ถˆ์•ˆ์ •์„ฑ๊ณผ ์ •๋ณด ๋ˆ„๋ฝ์ด ํ™•์ธ๋˜์—ˆ์œผ๋‚˜, ์ตœ์ ํ™”๋œ ๊ตฌ์กฐ์  ํ”„๋กฌํ”„ํŠธ ์ ์šฉ ์‹œ ์ถœ๋ ฅ ๋ณ€๋™์„ฑ์ด ์œ ์˜ํ•˜๊ฒŒ ๊ฐ์†Œํ•˜๊ณ  ๋ฒˆ์—ญ ์ •ํ™•๋„๊ฐ€ ํ–ฅ์ƒ๋˜์—ˆ์œผ๋ฉฐ, LLM์€ ์ „๋ฐ˜์ ์œผ๋กœ ๋ณด๊ณ ์„œ ๊ฐ€๋…์„ฑ์„ ์œ ์˜ํ•˜๊ฒŒ ๊ฐœ์„ ํ•˜์˜€์ง€๋งŒ(P < 0.05) ํ˜„์žฌ๋กœ์„œ๋Š” ์ธ๊ฐ„ ๊ฐ๋… ํ•˜์˜ ๋ณด์กฐ ๋„๊ตฌ๋กœ ํ™œ์šฉ์„ ์ œํ•œํ•ด์•ผ ํ•œ๋‹ค๊ณ  ๊ฒฐ๋ก ์ง€์—ˆ๋‹ค.
Added: 2026-08-09 00:00View โ†—

7Large Language Models for Classifying Usual Interstitial Pneumonia from Radiology Reports: Native Reasoning Versus Structured Prompting.

2026-08Journal of imaging informatics in medicineโญ Q1DOI 10.1007/s10278-026-02178-6

Extracting disease labels from radiology reports is essential for developing deep learning-based diagnostic models and enabling large-scale retrospective clinical research. Classification of usual interstitial pneumonia (UIP) patterns from high-resolution computed tomography (HRCT) reports according to Fleischner Society guidelines is a particularly demanding task, requiring synthesis of spatial distribution, fibrotic features, and exclusion criteria. As open-source large language models (LLMs) are released at an accelerating pace with steadily improving general benchmarks, a practical question arises: Do these improvements translate to better performance on complex, real-world clinical classification, and does the optimal prompting strategy differ across model architectures? While prior studies have evaluated LLMs for radiology report labeling, none have compared how prompting strategies interact with the native reasoning capabilities of newer model architectures. We evaluated 10 open-source LLMs from three architecture families (Llama, Qwen, Gemma) spanning 8 to 405 billion parameters, each tested with three prompting strategies on 270 HRCT reports classified by expert consensus of two senior thoracic radiologists. Four models with native reasoning ("thinking") capability were additionally tested in thinking mode. The best configuration achieved a Cohen's kappa (ฮบ) of 0.70 and 82% four-class accuracy. Structured reasoning prompting improved all Llama models but degraded all models with native reasoning capability (Qwen 3.5 and Gemma 4), revealing an architecture-dependent interaction. Thinking mode hurt performance on criteria-based prompts and never yielded the best configuration. Larger models did not consistently outperform smaller ones: Llama 3.1 405B offered no advantage over Llama 3.3 70B, and the Qwen 397B model underperformed the dense Qwen 27B. These findings demonstrate that newer model generations with improved general benchmarks and larger parameter counts do not guarantee better performance on specialized medical classification tasks and that prompt design must be matched to model architecture.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” Fleischner Society ๊ฐ€์ด๋“œ๋ผ์ธ์— ๋”ฐ๋ฅธ ํ†ต์ƒ์„ฑ ๊ฐ„์งˆ์„ฑ ํ๋ ด(UIP) ํŒจํ„ด ๋ถ„๋ฅ˜๋ฅผ ์œ„ํ•ด 10์ข…์˜ ์˜คํ”ˆ์†Œ์Šค ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ(LLM; Llama, Qwen, Gemma ๊ณ„์—ด, 8~405B ํŒŒ๋ผ๋ฏธํ„ฐ)์„ ์„ธ ๊ฐ€์ง€ ํ”„๋กฌํ”„ํŒ… ์ „๋žต๊ณผ ์กฐํ•ฉํ•˜์—ฌ, ํ‰๋ถ€ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜ 2์ธ์˜ ํ•ฉ์˜๋กœ ๋ถ„๋ฅ˜๋œ HRCT ๋ณด๊ณ ์„œ 270๊ฑด์„ ๊ธฐ์ค€์œผ๋กœ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ตœ์  ๋ชจ๋ธ ์กฐํ•ฉ์€ Cohen's kappa 0.70, 4๋ถ„๋ฅ˜ ์ •ํ™•๋„ 82%๋ฅผ ๋‹ฌ์„ฑํ•˜์˜€์œผ๋ฉฐ, ๊ตฌ์กฐํ™”๋œ ์ถ”๋ก  ํ”„๋กฌํ”„ํŒ…์€ Llama ๊ณ„์—ด ๋ชจ๋ธ์˜ ์„ฑ๋Šฅ์„ ํ–ฅ์ƒ์‹œํ‚จ ๋ฐ˜๋ฉด ์ž์ฒด ์ถ”๋ก (native reasoning) ๊ธฐ๋Šฅ์„ ๊ฐ–์ถ˜ ๋ชจ๋ธ(Qwen, Gemma)์—์„œ๋Š” ์˜คํžˆ๋ ค ์„ฑ๋Šฅ์„ ์ €ํ•˜์‹œ์ผœ ๋ชจ๋ธ ์•„ํ‚คํ…์ฒ˜์— ๋”ฐ๋ฅธ ํ”„๋กฌํ”„ํŒ… ์ „๋žต์˜ ์ƒํ˜ธ์ž‘์šฉ์ด ํ™•์ธ๋˜์—ˆ๋‹ค. ๋˜ํ•œ ํŒŒ๋ผ๋ฏธํ„ฐ ์ˆ˜๊ฐ€ ๋งŽ์€ ๋Œ€ํ˜• ๋ชจ๋ธ์ด ์†Œํ˜• ๋ชจ๋ธ๋ณด๋‹ค ์ผ๊ด€๋˜๊ฒŒ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์ด์ง€ ์•Š์•„, ๋ฒ”์šฉ ๋ฒค์น˜๋งˆํฌ ๊ฐœ์„ ์ด๋‚˜ ๋ชจ๋ธ ๊ทœ๋ชจ ํ™•๋Œ€๊ฐ€ ์ „๋ฌธ์  ์˜ํ•™ ๋ถ„๋ฅ˜ ๊ณผ์ œ์—์„œ์˜ ์„ฑ๋Šฅ ํ–ฅ์ƒ์„ ๋ณด์žฅํ•˜์ง€ ์•Š์œผ๋ฉฐ ํ”„๋กฌํ”„ํŠธ ์„ค๊ณ„๋Š” ๋ชจ๋ธ ์•„ํ‚คํ…์ฒ˜์— ๋งž๊ฒŒ ์ตœ์ ํ™”๋˜์–ด์•ผ ํ•จ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-08-09 00:00View โ†—

8Radiologically Relevant Clinical History Summarization with Large Language Models: A Multireader Performance Study.

2026-08Radiologyโญ Q1DOI 10.1148/radiol.253238

Background Clinical histories accompanying imaging orders guide protocol selection and diagnostic focus. However, they are often incomplete, potentially compromising diagnostic accuracy and workflow efficiency. Purpose To evaluate whether large language models (LLMs) can improve the clinical utility of provided imaging indications by leveraging clinical notes. Materials and Methods This retrospective study curated a dataset from deidentified electronic health records at the University of California San Francisco (January 2012 to August 2024), consisting of radiology reports with paired referring clinician-provided and radiologist-curated indications linked to clinical notes. The dataset was stratified across five body systems and five pathophysiologic categories to derive LLM selection and reader study internal test sets. For the reader study, 20 radiologists with 2-25 years of experience compared indications from the referring clinician, radiologist, and best-performing LLMs. Readers scored comprehensiveness, factuality, and conciseness and ranked indications for usefulness in protocoling, usefulness in interpretation, and overall ranking. Models and clinicians were compared using cumulative link mixed models with Tukey-adjusted post hoc comparisons. Results From 28โ€‰313 patients (mean age, 59 years ยฑ 20.6 [SD]; 14โ€‰912 women), 250 examinations from 247 patients were sampled for the reader study. After nine exclusions, 241 examinations were analyzed, yielding 482 reader-examination evaluations. Indications from the best-performing proprietary (Claude 3.5 Sonnet; Anthropic) and open-source (Qwen 2.5-7B Instruct; Alibaba) LLM were rated as more comprehensive (Likert rating of 5: 37.14% and 28.42%, respectively; both P < .001) and factual (68.05% and 59.75%; both P < .001) than referring clinician indications. The proprietary LLM ranked most useful in protocoling (rank 1: 40.87%; all P < .001), useful in interpretation (44.61%; all P < .001), and overall ranking (44.19%, all P < .001). Comprehensiveness (65.77% of ratings; both P < .001) most strongly influenced overall rankings. Conclusion LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation. ยฉ RSNA, 2026 Supplemental material is available for this article. See also the editorial by Yilmaz and Cardoza-Ochoa in this issue.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์ž„์ƒ ๊ธฐ๋ก์„ ํ™œ์šฉํ•˜์—ฌ ์˜์ƒ๊ฒ€์‚ฌ ์˜๋ขฐ ์‹œ ์ œ๊ณต๋˜๋Š” ์ž„์ƒ ์ •๋ณด์˜ ์œ ์šฉ์„ฑ์„ ํ–ฅ์ƒ์‹œํ‚ฌ ์ˆ˜ ์žˆ๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€์œผ๋ฉฐ, UCSF์˜ ๋น„์‹๋ณ„ํ™”๋œ ์ „์ž์˜๋ฌด๊ธฐ๋ก(2012โ€“2024๋…„)์„ ๊ธฐ๋ฐ˜์œผ๋กœ 241๊ฑด์˜ ๊ฒ€์‚ฌ์— ๋Œ€ํ•ด 20๋ช…์˜ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜๊ฐ€ ์˜๋ขฐ ์ž„์ƒ์˜, ์˜์ƒ์˜ํ•™๊ณผ ์˜์‚ฌ, ๋ฐ LLM์ด ์ƒ์„ฑํ•œ ๊ฒ€์‚ฌ ์ ์‘์ฆ์„ ํฌ๊ด„์„ฑยท์‚ฌ์‹ค์„ฑยท๊ฐ„๊ฒฐ์„ฑ ์ธก๋ฉด์—์„œ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ตœ์šฐ์ˆ˜ ๋…์  ๋ชจ๋ธ(Claude 3.5 Sonnet)๊ณผ ์˜คํ”ˆ์†Œ์Šค ๋ชจ๋ธ(Qwen 2.5-7B)์ด ์ƒ์„ฑํ•œ ์ ์‘์ฆ์€ ์ž„์ƒ์˜๊ฐ€ ์ œ๊ณตํ•œ ๊ฒƒ๋ณด๋‹ค ํฌ๊ด„์„ฑ๊ณผ ์‚ฌ์‹ค์„ฑ์—์„œ ์œ ์˜ํ•˜๊ฒŒ ๋†’์€ ํ‰๊ฐ€๋ฅผ ๋ฐ›์•˜์œผ๋ฉฐ(๋ชจ๋‘ P < .001), ๋…์  LLM์€ ํ”„๋กœํ† ์ฝœ ์„ ํƒ ๋ฐ ์˜์ƒ ํŒ๋… ์œ ์šฉ์„ฑ์—์„œ ๊ฐ€์žฅ ๋†’์€ ์ˆœ์œ„๋ฅผ ๊ธฐ๋กํ•˜์˜€๋‹ค. ์ด ๊ฒฐ๊ณผ๋Š” LLM์ด ์˜์ƒ์˜ํ•™ ์›Œํฌํ”Œ๋กœ์šฐ์—์„œ ์ž„์ƒ ์ •๋ณด์˜ ์งˆ์„ ํ–ฅ์ƒ์‹œํ‚ค๋Š” ๋ฐ ์‹ค์งˆ์ ์œผ๋กœ ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-08-09 00:00View โ†—

9AI achieves board-level performance on the Japan diagnostic radiology board examination through direct image interpretation.

2026-08Japanese journal of radiologyโญ Q1DOI 10.1007/s11604-026-01983-x
OBJECTIVE

To evaluate text-only versus vision-enabled performance of late-2025 large language models (LLMs) on the Japan Diagnostic Radiology Board Examination (JDRBE) and compare model performance with newly board-certified radiologists.

METHODS

Image-based questions from the JDRBE 2021 and 2023โ€“2025 were collected, and ground truth answers were determined by expert consensus. Four commercial multimodal LLMs were evaluated: Gemini 2.5 Pro (March 2025, baseline), Gemini 3 Pro, GPT-5.1, and Claude Opus 4.5 (all released in November 2025). Each question was answered with image input (โ€œvisionโ€) and without images (โ€œtext-onlyโ€). For the JDRBE 2025, subjective legitimacy of responses was independently rated by two radiologists using a five-point Likert scale, and low-rated responses were further analyzed by error type. Additional analyses on the JDRBE 2025 subset included image-shuffling and multi-run variability assessment (five runs). Model accuracies were also compared with those of five newly board-certified radiologists who passed the JDRBE 2025.

RESULTS

Gemini 3 Pro achieved the highest accuracy among all models, scoring 85.3% (279/327) in the vision condition and significantly outperforming its text-only accuracy (74.3%, Pโ€‰<โ€‰0.001). Gemini 2.5 Pro and Claude Opus 4.5 also improved with image input, whereas GPT-5.1 did not. For the JDRBE 2025, Gemini 3 Pro in the vision condition received the highest legitimacy ratings, and its accuracy (88%) was above the range observed in a reference group of five newly board-certified radiologists (65%โ€“83%), but hallucination was still the most common error type. Image-shuffling analysis using the 2025 subset showed no performance gain in all models, supporting reliance on visual input. Multi-run variability analysis showed high agreement across runs.

CONCLUSION

Among late-2025 commercial LLMs, Gemini 3 Pro demonstrated board-level performance on the JDRBE through direct medical image interpretation. SECONDARY ABSTRACT: The performance of vision-enabled large language models on the Japan Diagnostic Radiology Board Examination was evaluated. Among the models released in November 2025, Gemini 3 Pro demonstrated significant capabilities in direct medical image interpretation, achieving accuracy above that of a reference group of five newly board-certified radiologists. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1007/s11604-026-01983-x.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 2025๋…„ ํ›„๋ฐ˜์— ์ถœ์‹œ๋œ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ๋น„์ „ ํ™œ์„ฑํ™” ์—ฌ๋ถ€์— ๋”ฐ๋ฅธ ์„ฑ๋Šฅ ์ฐจ์ด๋ฅผ ์ผ๋ณธ ์ง„๋‹จ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜ ์ž๊ฒฉ์‹œํ—˜(JDRBE)์„ ํ†ตํ•ด ํ‰๊ฐ€ํ•˜๊ณ , ์‹ ๊ทœ ์ทจ๋“์ž์™€์˜ ๋น„๊ต ๋ถ„์„์„ ๋ชฉ์ ์œผ๋กœ ํ•˜์˜€๋‹ค. Gemini 2.5 Pro, Gemini 3 Pro, GPT-5.1, Claude Opus 4.5 ๋“ฑ 4์ข…์˜ ์ƒ์šฉ ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ LLM์— ๋Œ€ํ•ด JDRBE 2021๋…„ ๋ฐ 2023โ€“2025๋…„๋„ ์˜์ƒ ๊ธฐ๋ฐ˜ ๋ฌธ์ œ๋ฅผ ์ด๋ฏธ์ง€ ์ž…๋ ฅ ์กฐ๊ฑด(๋น„์ „)๊ณผ ํ…์ŠคํŠธ ์ „์šฉ ์กฐ๊ฑด์œผ๋กœ ๊ฐ๊ฐ ํ‰๊ฐ€ํ•˜์˜€์œผ๋ฉฐ, ์‹ ๊ทœ ํ•ฉ๊ฒฉ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜ 5์ธ์˜ ์„ฑ์ ๊ณผ ๋น„๊ตํ•˜์˜€๋‹ค. Gemini 3 Pro๊ฐ€ ๋น„์ „ ์กฐ๊ฑด์—์„œ ์ตœ๊ณ  ์ •ํ™•๋„์ธ 85.3%(2025๋…„๋„ ์‹œํ—˜ ๊ธฐ์ค€ 88%)๋ฅผ ๋‹ฌ์„ฑํ•˜์—ฌ ์‹ ๊ทœ ์ „๋ฌธ์˜๊ตฐ(65โ€“83%)์˜ ๋ฒ”์œ„๋ฅผ ์ƒํšŒํ•˜์˜€์œผ๋‚˜, ํ™˜๊ฐ(hallucination)์ด ์—ฌ์ „ํžˆ ๊ฐ€์žฅ ํ”ํ•œ ์˜ค๋ฅ˜ ์œ ํ˜•์œผ๋กœ ๋‚˜ํƒ€๋‚˜ ์ž„์ƒ ์ ์šฉ ์‹œ ์ฃผ์˜๊ฐ€ ์š”๊ตฌ๋œ๋‹ค.
Added: 2026-08-02 00:00View โ†—

10Pre-Imaging Clinical Factors Associated With Cardiac MR Image Quality Using Large Language Model-Enabled Data Extraction.

2026-08Journal of magnetic resonance imaging : JMRIโญ Q1DOI 10.1002/jmri.70336
BACKGROUND

Poor cardiac MR image quality can prompt repeat examinations and hinder clinical decision-making.

OBJECTIVE

To evaluate whether pre-imaging clinical information, extracted using a large language model (LLM), is independently associated with cardiac MR image quality. STUDY TYPE: Retrospective. POPULATION: 1006 adults undergoing clinical cardiac MR examinations. FIELD STRENGTH/SEQUENCE: 1.5โ€‰T and 3โ€‰T scanners with cine, black blood, MR angiogram, or late gadolinium enhancement protocols. ASSESSMENT: Image quality was categorized per study as excellent, slightly limited, severely limited, or nondiagnostic using institutional reporting conventions finalized by radiologists and cardiologists. A HIPAA-compliant LLM assigned image quality labels based on radiology reports through an iteratively refined prompt, with reliability confirmed by two radiologists. Labels were binarized as Good (excellent and slightly limited) versus Poor (severely limited and nondiagnostic). A repeat-imaging-adjusted image quality label was used in a sensitivity analysis. Pre-imaging clinical and patient variables were extracted from electronic health records. Associations between variables and image quality labels were investigated. STATISTICAL TESTS: Cohen's kappa (ฮบ) for label agreement. Chi-square and t-tests for univariate analysis. Variance inflation factor (VIF) and multivariable logistic regression. Significance level: pโ€‰<โ€‰0.05.

RESULTS

Binarized image quality labels showed substantial agreement with interpreters' assessments for both the primary dataset (ฮบโ€‰=โ€‰0.689) and the repeat-adjusted dataset (ฮบโ€‰=โ€‰0.879). There was no significant multicollinearity (VIFโ€‰=โ€‰1.01-1.39). Cognitive and communication impairment (OR 1.81, 95% CI [1.30-2.54], pโ€‰<โ€‰0.001) and respiratory issues (1.57 [1.14-2.17], pโ€‰=โ€‰0.006) were independently associated with poor image quality. These associations remained significant in the repeat-adjusted sensitivity analysis (cognitive and communication impairment (OR 1.75, 95% CI [1.27-2.44], pโ€‰<โ€‰0.001) and respiratory compromise (OR 1.37, 95% CI [1.04-1.82], pโ€‰=โ€‰0.027)). Other clinical variables were not independently associated after adjustment. DATA

CONCLUSION

Cognitive/communication impairment and respiratory compromise were independently associated with poor cardiac MR image quality. TECHNICAL EFFICACY: Stage 2. This study examined why some cardiac magnetic resonance scans have poor image quality. Poorโ€quality scans can limit diagnosis and may require repeat imaging. It analyzed adult examinations and used artificial intelligence (AI) to extract clinical information from the medical records available before the scan. Cognitive or communication difficulties and breathingโ€related problems were associated with worse image quality. These findings might suggest that these patients may benefit from additional preparation before scanning. More broadly, this study shows that information in the medical record can be efficiently extracted using AI to support largeโ€scale clinical research.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ํ›„ํ–ฅ์  ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ํ™œ์šฉํ•˜์—ฌ ์ „์ž์˜๋ฌด๊ธฐ๋ก์—์„œ ์ž„์ƒ ๋ณ€์ˆ˜๋ฅผ ์ถ”์ถœํ•˜๊ณ , ์ดฌ์˜ ์ „ ์ž„์ƒ ์ •๋ณด๊ฐ€ ์‹ฌ์žฅ MRI ์˜์ƒ ํ’ˆ์งˆ๊ณผ ๋…๋ฆฝ์ ์œผ๋กœ ์—ฐ๊ด€๋˜๋Š”์ง€๋ฅผ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 1,006๋ช…์˜ ์„ฑ์ธ ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ LLM์ด ๋ฐฉ์‚ฌ์„ ๊ณผ ๋ณด๊ณ ์„œ์—์„œ ์˜์ƒ ํ’ˆ์งˆ ๋ ˆ์ด๋ธ”์„ ์ถ”์ถœํ•˜๊ณ  ์ด์ง„ ๋ถ„๋ฅ˜(์–‘ํ˜ธ/๋ถˆ๋Ÿ‰)ํ•œ ๋’ค, ๋‹ค๋ณ€๋Ÿ‰ ๋กœ์ง€์Šคํ‹ฑ ํšŒ๊ท€๋ถ„์„์„ ํ†ตํ•ด ์‚ฌ์ „ ์ž„์ƒ ๋ณ€์ˆ˜์™€์˜ ์—ฐ๊ด€์„ฑ์„ ๋ถ„์„ํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, ์ธ์ง€ยท์˜์‚ฌ์†Œํ†ต ์žฅ์• (OR 1.81, 95% CI 1.30โ€“2.54)์™€ ํ˜ธํก๊ธฐ๊ณ„ ๋ฌธ์ œ(OR 1.57, 95% CI 1.14โ€“2.17)๊ฐ€ ๋ถˆ๋Ÿ‰ํ•œ ์‹ฌ์žฅ MRI ์˜์ƒ ํ’ˆ์งˆ๊ณผ ๋…๋ฆฝ์ ์œผ๋กœ ์œ ์˜ํ•˜๊ฒŒ ์—ฐ๊ด€๋˜์–ด, ํ•ด๋‹น ํ™˜์ž๊ตฐ์—์„œ ์ดฌ์˜ ์ „ ์ถ”๊ฐ€์ ์ธ ์ค€๋น„๊ฐ€ ํ•„์š”ํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-08-02 00:00View โ†—

11ARKE: An ontology-driven framework for automated mapping of local radiology procedure terms to the LOINC-RadLex playbook using large language model.

2026-08Journal of biomedical informaticsโญ Q1DOI 10.1016/j.jbi.2026.105071
OBJECTIVE

To develop an ontology-driven framework that standardizes heterogeneous local radiology procedure names by decomposing them into semantic components and aligning them to LOINC/RSNA Radiology Playbook codes using constrained large language model (LLM)-based parsing and selection.

METHODS

Radiology procedure names from two tertiary hospitals in Korea were parsed into semantic components using LLM prompting with retrieval-augmented generation, aligned with the LOINC/RSNA Radiology Playbook (version 2.80). Ontology-based similarity scoring quantify correspondence between parsed components and Playbook candidates' attributes, and retrieve the Top 10 candidates, followed by LLM-based selection within this candidate set. Performance was evaluated against direct Playbook code name-based mapping and conventional similarity metrics using a radiologist-curated gold reference.

RESULTS

A total of 3,326 local procedure names were analyzed. Ontology-based mapping approach substantially outperformed direct Playbook code name mapping across all evaluation metrics. At the candidate retrieval stage, ontology-based attribute matching achieved recall@5 of up to 0.78 (internal) and 0.89 (external), compared with 0.51 and 0.47 for direct mapping. After LLM-based selection, the ontology-based approach achieved a final selection recall@1 of up to 0.70 (internal) and 0.81 (external), exceeding direct mapping (0.48 and 0.50) more than 20 percentage points in both settings (pย <ย 0.001).

CONCLUSION

Decomposing procedure names into ontology-grounded semantic components enables robust handling of heterogeneous local terminology, while constraining LLM reasoning to structured selection tasks mitigates hallucination and preserves semantic fidelity. Ontology-driven knowledge encoding provides a scalable and reliable approach to standardizing radiology procedure names, supporting cross-institutional interoperability and secondary data use of imaging research.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ตญ๋‚ด ๋‘ ๊ฐœ 3์ฐจ ๋ณ‘์›์˜ ์ด์งˆ์ ์ธ ๋ฐฉ์‚ฌ์„  ๊ฒ€์‚ฌ ๋ช…์นญ์„ ํ‘œ์ค€ํ™”ํ•˜๊ธฐ ์œ„ํ•ด, ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)๊ณผ ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG)์„ ํ™œ์šฉํ•˜์—ฌ ์ ˆ์ฐจ๋ช…์„ ์˜๋ฏธ ๊ตฌ์„ฑ ์š”์†Œ๋กœ ๋ถ„ํ•ดํ•œ ํ›„ ์˜จํ†จ๋กœ์ง€ ๊ธฐ๋ฐ˜ ์œ ์‚ฌ๋„ ์ ์ˆ˜๋กœ LOINC/RSNA Radiology Playbook ์ฝ”๋“œ์™€ ๋งคํ•‘ํ•˜๋Š” ํ”„๋ ˆ์ž„์›Œํฌ(ARKE)๋ฅผ ๊ฐœ๋ฐœํ•˜์˜€๋‹ค. ์ด 3,326๊ฐœ์˜ ๋กœ์ปฌ ๊ฒ€์‚ฌ ๋ช…์นญ์„ ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, ์˜จํ†จ๋กœ์ง€ ๊ธฐ๋ฐ˜ ํ›„๋ณด ๊ฒ€์ƒ‰ ๋‹จ๊ณ„์—์„œ recall@5๊ฐ€ ๋‚ด๋ถ€ 0.78, ์™ธ๋ถ€ 0.89๋ฅผ ๋‹ฌ์„ฑํ•˜์˜€์œผ๋ฉฐ, ์ตœ์ข… LLM ์„ ํƒ ํ›„ recall@1์€ ๋‚ด๋ถ€ 0.70, ์™ธ๋ถ€ 0.81๋กœ ์ง์ ‘ ์ฝ”๋“œ๋ช… ๋งคํ•‘(๊ฐ๊ฐ 0.48, 0.50) ๋Œ€๋น„ 20%p ์ด์ƒ ์œ ์˜ํ•˜๊ฒŒ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€๋‹ค(p < 0.001). ์ ˆ์ฐจ๋ช…์„ ์˜จํ†จ๋กœ์ง€ ๊ธฐ๋ฐ˜ ์˜๋ฏธ ๋‹จ์œ„๋กœ ๋ถ„ํ•ดํ•จ์œผ๋กœ์จ LLM์˜ ํ™˜๊ฐ(hallucination)์„ ์–ต์ œํ•˜๊ณ  ์˜๋ฏธ์  ์ถฉ์‹ค๋„๋ฅผ ์œ ์ง€ํ•˜๋Š” ํ™•์žฅ ๊ฐ€๋Šฅํ•œ ํ‘œ์ค€ํ™” ๋ฐฉ๋ฒ•์„ ์ œ์‹œํ•˜์˜€์œผ๋ฉฐ, ์ด๋Š” ๊ธฐ๊ด€ ๊ฐ„ ์ƒํ˜ธ์šด์šฉ์„ฑ ๋ฐ ์˜์ƒ ์—ฐ๊ตฌ์˜ 2์ฐจ ๋ฐ์ดํ„ฐ ํ™œ์šฉ์„ ์ง€์›ํ•  ์ˆ˜ ์žˆ๋‹ค.
Added: 2026-06-28 00:00View โ†—

12Decreasing Administrative Effort Related to Non-Approval of Image-guidED Procedures Using Large Language Models - The DENIED-AI Pilot Study.

2026-08Academic radiologyโญ Q1DOI 10.1016/j.acra.2026.04.021

RATIONALE AND

OBJECTIVE

To evaluate whether large language models (LLMs) can generate accurate, clinically valid, and usable letters to appeal insurance denials for radiology procedures.

METHODS

This pilot study generated insurance appeal letters for a simulated clinical scenario. Four LLMsย (Claude 3.5, Nova Pro, Llama-3.1-70B, ChatGPT-4o) were used with zero-shot, few-shot, and retrieval-augmented generation (RAG) techniques. Four board-certified interventional radiologists, blinded to model and technique, scored letters for content (accuracy, personalization, references), grammar and structure (readability, tone, persuasiveness), and usability (estimated editing time, usefulness as a template). References were verified for accuracy, and outputs were carefully examined for hallucinations. Statistical analyses included ANOVA, Chi-square, and Fleiss' Kappa for interrater reliability.

RESULTS

Mean content and grammar scores were 3.9 ยฑ 0.95 and 4.3 ยฑ 0.9 (out of 5), with no significant differences by model or technique (p >.05). Reviewer agreement was poor (Fleiss' Kappa -0.18 for content, -0.085 for grammar). Hallucinations were flagged by reviewers in 16/48 assessments, significantly more often with the online model (ChatGPT-4o: 58% vs offline 25%; p =.03). Of 44 references, 80% from the offline models were fabricated compared with 29% from ChatGPT-4o (p <.001). Estimated editing time was less than 10ย min in 71% of responses, and the reviewers felt the letters would be useful as templates in 73% of cases.

CONCLUSION

LLM-generated appeal letters for insurance denials were generally well received, with high usability and adequate quality. However, fabricated references and hallucinations remain prevalent, necessitating careful human review before clinical use.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ํŒŒ์ผ๋Ÿฟ ์—ฐ๊ตฌ๋Š” ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์˜์ƒ์˜ํ•™์  ์‹œ์ˆ ์— ๋Œ€ํ•œ ๋ณดํ—˜ ๊ธ‰์—ฌ ๊ฑฐ๋ถ€์— ๋Œ€์‘ํ•˜๋Š” ์ด์˜์‹ ์ฒญ ์„œํ•œ์„ ์ž๋™ ์ƒ์„ฑํ•  ์ˆ˜ ์žˆ๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€์œผ๋ฉฐ, Claude 3.5, Nova Pro, Llama-3.1-70B, ChatGPT-4o ๋“ฑ 4์ข…์˜ LLM์„ zero-shot, few-shot, RAG ๊ธฐ๋ฒ•๊ณผ ํ•จ๊ป˜ ์ ์šฉํ•˜๊ณ  4๋ช…์˜ ์ค‘์žฌ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜๊ฐ€ ๋‚ด์šฉ, ๋ฌธ๋ฒ•, ํ™œ์šฉ๋„๋ฅผ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ƒ์„ฑ๋œ ์„œํ•œ์˜ ๋‚ด์šฉ ๋ฐ ๋ฌธ๋ฒ• ์ ์ˆ˜๋Š” ๊ฐ๊ฐ ํ‰๊ท  3.9์  ๋ฐ 4.3์ (5์  ๋งŒ์ )์œผ๋กœ ์–‘ํ˜ธํ•˜์˜€๊ณ , 71%์˜ ์‘๋‹ต์—์„œ ํŽธ์ง‘ ์†Œ์š” ์‹œ๊ฐ„์ด 10๋ถ„ ๋ฏธ๋งŒ์ด์—ˆ์œผ๋ฉฐ 73%๊ฐ€ ์‹ค์šฉ์ ์ธ ์ดˆ์•ˆ ํ…œํ”Œ๋ฆฟ์œผ๋กœ ํ‰๊ฐ€๋˜์—ˆ๋‹ค. ๊ทธ๋Ÿฌ๋‚˜ ์˜คํ”„๋ผ์ธ ๋ชจ๋ธ์—์„œ ์ƒ์„ฑ๋œ ์ฐธ๊ณ ๋ฌธํ—Œ์˜ 80%๊ฐ€ ํ—ˆ์œ„๋กœ ํ™•์ธ๋˜์—ˆ๊ณ  ํ™˜๊ฐ(hallucination) ํ˜„์ƒ๋„ ์ƒ๋‹น์ˆ˜ ๊ด€์ฐฐ๋˜์–ด, ์ž„์ƒ ์ ์šฉ ์ „ ๋ฐ˜๋“œ์‹œ ์ „๋ฌธ๊ฐ€์˜ ๋ฉด๋ฐ€ํ•œ ๊ฒ€ํ† ๊ฐ€ ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-05-10 00:00View โ†—

13Leveraging Fine-Tuned Large Language Models for Interpretable Pancreatic Cystic Lesion Feature Extraction and Risk Categorization.

2026-08AJR. American journal of roentgenologyDOI 10.2214/ajr.25.34076

BACKGROUND. Manual extraction of pancreatic cystic lesion (PCL) features from radiology reports is labor-intensive, limiting large-scale studies needed to advance PCL research. OBJECTIVE. The purpose of this study was to evaluate GPT-4o (closed source [ OpenAI]), Llama (Llama-3.1-8B-Instruct, open source), and DeepSeek (DeepSeek-R1-Distill-Llama-8B, open source) large language models (LLMs) for PCL feature extraction, without and with chain-of-thought (CoT) reasoning. METHODS. We curated a dataset of 6469 abdominal MRI or CT reports (2005-2024) that described PCLs from 5615 patients. Llama and DeepSeek were fine-tuned using quantized low-rank adaptation on GPT-4o-generated CoT labels for extracting PCL and main pancreatic duct features. Features were mapped to risk categories per institutional policy. Evaluation was performed on 285 held-out human-annotated reports from 281 patients. Model outputs for 100 cases were independently reviewed by three radiologists. Feature extraction was evaluated using exact-match accuracy, risk categorization with a macro-averaged F1 score, and radiologist-model agreement with Fleiss kappa values. Error analyses were performed to assess how and why models made mistakes. RESULTS. CoT fine-tuned LLMs showed a feature extraction accuracy of 97% (95% CI, 97-98%) for Llama, 98% (95% CI, 97-98%) for DeepSeek, and 97% (95% CI, 97-98%) for GPT-4o. Risk categorization F1 scores were 0.93 (95% CI, 0.89-0.97) for Llama, 0.94 (95% CI, 0.90-0.98) for DeepSeek, and 0.97 (95% CI, 0.93-0.99) for GPT-4o. Radiologist interreader agreement was high (ฮบ = 0.888) and showed no significant difference with the addition of Llama (ฮบ = 0.882; p > .99), DeepSeek (ฮบ = 0.893, p > .99), or GPT-4o (ฮบ = 0.897, p > .99). Across all models, object identification and clinical reasoning were the most frequent error types, accounting for 29.3-37.3% and 18.1-21.1% of total errors, respectively. CONCLUSION. LLMs show feasibility for automatically extracting PCL features from radiology reports. Fine-tuned open-source LLMs achieved performance comparable to that of GPT-4o. CoT reasoning improved accuracy and enabled interpretable error analysis. Model-assigned risk categories showed high agreement with abdominal radiologists. CLINICAL IMPACT. LLMs have the potential to enable creation of large structured registries from existing radiology reports to support population-level research on PCLs.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์—์„œ ์ทŒ์žฅ ๋‚ญ์„ฑ ๋ณ‘๋ณ€(PCL) ํŠน์ง•์„ ์ž๋™์œผ๋กœ ์ถ”์ถœํ•˜๊ณ  ์œ„ํ—˜๋„๋ฅผ ๋ถ„๋ฅ˜ํ•˜๊ธฐ ์œ„ํ•ด GPT-4o, Llama, DeepSeek ๋“ฑ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ด 5,615๋ช…์˜ ๋ณต๋ถ€ MRI/CT ๋ณด๊ณ ์„œ 6,469๊ฑด์„ ํ™œ์šฉํ•˜์—ฌ ์˜คํ”ˆ์†Œ์Šค ๋ชจ๋ธ์„ chain-of-thought(CoT) ์ถ”๋ก  ๊ธฐ๋ฐ˜์œผ๋กœ ํŒŒ์ธํŠœ๋‹ํ•˜์˜€์œผ๋ฉฐ, 285๊ฑด์˜ ์ „๋ฌธ๊ฐ€ ์ฃผ์„ ๋ณด๊ณ ์„œ๋กœ ๊ฒ€์ฆํ•˜์˜€๋‹ค. CoT ํŒŒ์ธํŠœ๋‹ ์ ์šฉ ์‹œ Llama์™€ DeepSeek ๋ชจ๋‘ ํŠน์ง• ์ถ”์ถœ ์ •ํ™•๋„ 97โ€“98%, ์œ„ํ—˜๋„ ๋ถ„๋ฅ˜ F1 ์ ์ˆ˜ 0.93โ€“0.94๋ฅผ ๋‹ฌ์„ฑํ•˜์—ฌ GPT-4o์— ํ•„์ ํ•˜๋Š” ์„ฑ๋Šฅ์„ ๋ณด์˜€๊ณ , ๋ชจ๋ธ์ด ๋ถ„๋ฅ˜ํ•œ ์œ„ํ—˜ ๋“ฑ๊ธ‰์€ ๋ณต๋ถ€์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜ ๊ฐ„ ํŒ๋… ์ผ์น˜๋„(ฮบ = 0.888)์™€ ์œ ์˜ํ•œ ์ฐจ์ด ์—†์ด ๋†’์€ ์ผ์น˜์œจ์„ ๋‚˜ํƒ€๋ƒˆ๋‹ค.
Added: 2026-05-10 00:00View โ†—

14Integrating AI Into Emergency Radiology: Promises, Pitfalls, and Practical Approaches.

2026-07Seminars in roentgenologyDOI 10.1016/j.ro.2026.151009

Emergency radiology operates in a high-acuity, time-sensitive environment where imaging is tightly integrated into real-time clinical decision-making. Growing imaging demand, increasing case complexity, and workforce constraints have intensified pressure on emergency radiologists. Artificial intelligence (AI) has emerged as a potential tool to support imaging prioritization, interpretation, and operational efficiency. However, to meaningfully advance care delivery, the role of AI must be considered beyond algorithm performance, including its implementation, reliability, and real-world clinical impact. In this narrative review, we examine the role of AI across the emergency radiology workflow through three lenses: current capabilities, limitations of the supporting evidence, and practical considerations for clinical implementation. We review applications spanning pre-image acquisition, image acquisition and reconstruction, computer-aided triage and detection, reporting, and follow-up, integrating published evidence with practical insights. Discrepancies between reported and real-world performance, the influence of human-AI interaction on clinical decision-making, and the potential for subtle errors and bias are also discussed. As national regulatory and local governance frameworks continue to evolve, including emerging challenges posed by large language models, gaps remain between reported and real-world AI performance. In emergency radiology, the true impact of AI will depend on how seamlessly and effectively these tools are integrated into existing clinical workflows. Local validation, ongoing performance monitoring, and multidisciplinary institutional oversight are essential to identify performance variability, mitigate biases, and support reliable use in a high-stakes clinical environment.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
์ด ์„œ์ˆ ์  ๋ฌธํ—Œ๊ณ ์ฐฐ์€ ์ฆ๊ฐ€ํ•˜๋Š” ์˜์ƒ ์ˆ˜์š”์™€ ์ธ๋ ฅ ๋ถ€์กฑ์œผ๋กœ ์••๋ฐ•๋ฐ›๋Š” ์‘๊ธ‰ ์˜์ƒ์˜ํ•™ ๋ถ„์•ผ์—์„œ ์ธ๊ณต์ง€๋Šฅ(AI)์˜ ์—ญํ• ์„ ์ „์ž„์ƒ ๋‹จ๊ณ„๋ถ€ํ„ฐ ํŒ๋… ๋ฐ ์ถ”์ ๊นŒ์ง€ ์ „์ฒด ์›Œํฌํ”Œ๋กœ์šฐ์— ๊ฑธ์ณ ์ฒด๊ณ„์ ์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. ์ €์ž๋“ค์€ ํ˜„์žฌ AI์˜ ์ ์šฉ ๊ฐ€๋Šฅ์„ฑ, ๊ทผ๊ฑฐ์˜ ํ•œ๊ณ„, ๊ทธ๋ฆฌ๊ณ  ์‹ค์ œ ์ž„์ƒ ๋„์ž… ์‹œ ๊ณ ๋ ค์‚ฌํ•ญ์ด๋ผ๋Š” ์„ธ ๊ฐ€์ง€ ๊ด€์ ์—์„œ ๊ธฐ์กด ๋ฌธํ—Œ๊ณผ ์‹ค๋ฌด์  ํ†ต์ฐฐ์„ ํ†ตํ•ฉํ•˜์—ฌ ๊ฒ€ํ† ํ•˜์˜€๋‹ค. ๋ณด๊ณ ๋œ ์„ฑ๋Šฅ๊ณผ ์‹ค์ œ ์ž„์ƒ ์„ฑ๋Šฅ ๊ฐ„์˜ ๊ดด๋ฆฌ, ์ธ๊ฐ„-AI ์ƒํ˜ธ์ž‘์šฉ์ด ์˜์‚ฌ๊ฒฐ์ •์— ๋ฏธ์น˜๋Š” ์˜ํ–ฅ, ํŽธํ–ฅ์˜ ์œ„ํ—˜์„ฑ์„ ๊ณ ๋ คํ•  ๋•Œ, AI์˜ ์‹ค์งˆ์  ์ž„์ƒ ํšจ๊ณผ๋Š” ๊ธฐ๊ด€ ๋‹จ์œ„์˜ ๋กœ์ปฌ ๊ฒ€์ฆ, ์ง€์†์ ์ธ ์„ฑ๋Šฅ ๋ชจ๋‹ˆํ„ฐ๋ง, ๊ทธ๋ฆฌ๊ณ  ๋‹คํ•™์ œ์  ๊ฑฐ๋ฒ„๋„Œ์Šค ์ฒด๊ณ„๋ฅผ ํ†ตํ•ด ๊ธฐ์กด ์›Œํฌํ”Œ๋กœ์šฐ์— ์–ผ๋งˆ๋‚˜ ํšจ๊ณผ์ ์œผ๋กœ ํ†ตํ•ฉ๋˜๋А๋ƒ์— ๋‹ฌ๋ ค ์žˆ๋‹ค๊ณ  ๊ฒฐ๋ก ์ง€์—ˆ๋‹ค.
Added: 2026-08-02 00:00View โ†—

15PRACTALL 2025: Artificial intelligence-application of allergy and immunology to patient care.

2026-07The Journal of allergy and clinical immunologyDOI 10.1016/j.jaci.2026.04.031

Artificial intelligence (AI), first defined in 1955 by John McCarthy, has transformed daily life across industries through applications such as chatbots, autonomous vehicles, and navigation systems. The 2022 release of ChatGPT marked a pivotal moment, highlighting AI's rapidly expanding potential. The health care industry is increasingly embracing AI-enabled tools across oncology, pathology, and radiology to augment disease screening and clinical workflows. Ambient listening technologies support clinical documentation, reduce administrative burden, and improve patient-physician interactions. Large language models combined with natural language processing are being evaluated for generating clinical summaries managing patient portal messaging and converting freehand notes to electronic health records, while also uncovering patterns in patient data to support more personalized treatments. Forward-thinking health systems are establishing informatics departments to optimize these models. Notably, with 1 in 6 adults sourcing health information from AI, rising to nearly one-quarter among individuals younger than 30 years, there is a growing need to ensure that these technologies provide accurate and reliable information to safeguard patient safety and support appropriate clinical use. PRACTALL, a collaboration between the American Academy of Allergy, Asthma & Immunology and the European Academy of Allergy & Clinical Immunology, aims to equip allergist-immunologists with essential AI insights highlighting tools for clinical practice, education, and research. By addressing potential pitfalls and biases, PRACTALL illustrates how AI can enhance efficiency, improve patient care, and alleviate administrative burdens in health care.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ PRACTALL 2025 ๋ฌธ์„œ๋Š” ์ธ๊ณต์ง€๋Šฅ(AI)์ด ์•Œ๋ ˆ๋ฅด๊ธฐยท๋ฉด์—ญํ•™ ์ž„์ƒ ์ง„๋ฃŒ์— ๋ฏธ์น˜๋Š” ์˜ํ–ฅ์„ ๊ฒ€ํ† ํ•˜๊ณ , ์•Œ๋ ˆ๋ฅด๊ธฐ-๋ฉด์—ญ ์ „๋ฌธ์˜๊ฐ€ AI ๋„๊ตฌ๋ฅผ ํšจ๊ณผ์ ์œผ๋กœ ํ™œ์šฉํ•  ์ˆ˜ ์žˆ๋„๋ก ํ•ต์‹ฌ ์ง€์‹์„ ์ œ๊ณตํ•˜๋Š” ๊ฒƒ์„ ๋ชฉ์ ์œผ๋กœ ํ•œ๋‹ค. ๋ฏธ๊ตญ ์•Œ๋ ˆ๋ฅด๊ธฐยท์ฒœ์‹ยท๋ฉด์—ญํ•™ํšŒ(AAAAI)์™€ ์œ ๋Ÿฝ ์•Œ๋ ˆ๋ฅด๊ธฐยท์ž„์ƒ๋ฉด์—ญํ•™ํšŒ(EAACI)์˜ ๊ณต๋™ ํ˜‘๋ ฅ์ธ PRACTALL์€ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ, ์ž์—ฐ์–ด ์ฒ˜๋ฆฌ, ์ฃผ๋ณ€ ์ฒญ์ทจ ๊ธฐ์ˆ  ๋“ฑ ๋‹ค์–‘ํ•œ AI ์ ์šฉ ์‚ฌ๋ก€๋ฅผ ์ž„์ƒ ๋ฌธ์„œํ™”, ํ™˜์ž ํฌํ„ธ ๋ฉ”์‹œ์ง€ ๊ด€๋ฆฌ, ๊ฐœ์ธ ๋งž์ถคํ˜• ์น˜๋ฃŒ ์ง€์› ์ธก๋ฉด์—์„œ ์ฒด๊ณ„์ ์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. AI๊ฐ€ ํ–‰์ • ๋ถ€๋‹ด ๊ฒฝ๊ฐ ๋ฐ ์ง„๋ฃŒ ํšจ์œจ ํ–ฅ์ƒ์— ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ๋Š” ๋ฐ˜๋ฉด, ์„ฑ์ธ 6๋ช… ์ค‘ 1๋ช…์ด AI๋ฅผ ๊ฑด๊ฐ• ์ •๋ณด ์ถœ์ฒ˜๋กœ ํ™œ์šฉํ•˜๋Š” ํ˜„์‹ค์„ ๊ณ ๋ คํ•  ๋•Œ ์ •๋ณด์˜ ์ •ํ™•์„ฑยท์‹ ๋ขฐ์„ฑ ํ™•๋ณด ๋ฐ ํŽธํ–ฅ ๋ฐฉ์ง€๋ฅผ ์œ„ํ•œ ๋น„ํŒ์  ์ ‘๊ทผ์ด ํ•„์ˆ˜์ ์ž„์„ ๊ฐ•์กฐํ•˜์˜€๋‹ค.
Added: 2026-08-02 00:00View โ†—

16Personalized Rule-Based Proofreading for Speech Recognition Errors in Radiology Reports: Development Using Artificial Intelligence Coding Assistance.

2026-07AJR. American journal of roentgenologyDOI 10.2214/ajr.26.34694

A radiologist without programming expertise used artificial intelligenceโ€“assisted coding with large language models to build a personalized rule-based proofreading system for correcting recurrent and persistent speech recognition report errors. Subjective reflections and most common corrections were summarized. The system provided a practical approach for correcting user-specific speech recognition errors unaddressed by existing solutions

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
ํ”„๋กœ๊ทธ๋ž˜๋ฐ ์ „๋ฌธ ์ง€์‹์ด ์—†๋Š” ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜๊ฐ€ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ ๊ธฐ๋ฐ˜์˜ ์ธ๊ณต์ง€๋Šฅ ์ฝ”๋”ฉ ๋ณด์กฐ ๋„๊ตฌ๋ฅผ ํ™œ์šฉํ•˜์—ฌ, ๋ฐฉ์‚ฌ์„  ํŒ๋…๋ฌธ์—์„œ ๋ฐ˜๋ณต์ ์œผ๋กœ ๋ฐœ์ƒํ•˜๋Š” ์Œ์„ฑ ์ธ์‹ ์˜ค๋ฅ˜๋ฅผ ๊ต์ •ํ•˜๊ธฐ ์œ„ํ•œ ๊ฐœ์ธํ™”๋œ ๊ทœ์น™ ๊ธฐ๋ฐ˜ ๊ต์ • ์‹œ์Šคํ…œ์„ ๊ฐœ๋ฐœํ•˜์˜€๋‹ค. ๊ฐœ๋ฐœ ๊ณผ์ •์—์„œ ์ฃผ๊ด€์  ๊ฒฝํ—˜๊ณผ ์ฃผ์š” ๊ต์ • ์‚ฌ๋ก€๋ฅผ ์ฒด๊ณ„์ ์œผ๋กœ ์ •๋ฆฌํ•˜์˜€์œผ๋ฉฐ, ๊ธฐ์กด ์†”๋ฃจ์…˜์œผ๋กœ๋Š” ํ•ด๊ฒฐ๋˜์ง€ ์•Š๋˜ ์‚ฌ์šฉ์ž ํŠน์ด์  ์Œ์„ฑ ์ธ์‹ ์˜ค๋ฅ˜๋ฅผ ํšจ๊ณผ์ ์œผ๋กœ ๊ต์ •ํ•˜๋Š” ์‹ค์šฉ์ ์ธ ์ ‘๊ทผ๋ฒ•์„ ์ œ์‹œํ•˜์˜€๋‹ค. ์ด ์—ฐ๊ตฌ๋Š” ๋น„์ „๋ฌธ๊ฐ€ ์˜์‚ฌ๋„ ์ธ๊ณต์ง€๋Šฅ ๋ณด์กฐ ์ฝ”๋”ฉ์„ ํ†ตํ•ด ์ž„์ƒ ํ˜„์žฅ์—์„œ ๋งž์ถคํ˜• ์ž๋™ ๊ต์ • ๋„๊ตฌ๋ฅผ ์ž์ฒด ๊ฐœ๋ฐœํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-08-02 00:00View โ†—

17Automated generation of impressions in abdominal radiology reports using an artificial intelligence-based tool: performance compared to manual impressions.

2026-07Abdominal radiology (New York)โญ Q1DOI 10.1007/s00261-026-05688-7
OBJECTIVE

Recent years have seen a rapid development of artificial intelligence (AI) tools to enhance radiologists' workflow, but most have focused on image acquisition and interpretation. Generating radiology reports is a critical element of practice but remains a source of inefficiency and cognitive load.

OBJECTIVE

To investigate the performance of an AI based tool integrated into dictation software to automate generation of report impressions and compare it with radiologist generated impressions.

METHODS

One hundred consecutive abdominal radiology reports were retrospectively selected in January 2024 from a single center. An AI-generated impression (GI) was created for each report and compared with the original radiologist impression (RI). Ten subspecialty abdominal radiologists evaluated the blinded, randomized pairs. Each impression was evaluated on a 5-point Likert scale for coherence, comprehensiveness, and factual consistency, and an overall preference for the impression was recorded. Cumulative link mixed models and binomial logistic regression were used where appropriate.

RESULTS

GI was preferred in 38% of reports (95% CI 32-45%), RI in 45% (95% CI 39-50%), with no preference in 17% of instances. Mixed-effects logistic regression demonstrated that GI was rated as equivalent or preferred to RI (odds ratio 1.34, 95% CI 1.004-1.788, pโ€‰=โ€‰0.02). Radiologists rated GI equivalent or superior to RI in 79% of cases for coherence, 66% for comprehensiveness, and 77% for factual consistency.

CONCLUSION

The AI generated impressions were clinically acceptable and rated equivalent or superior to radiologist-generated impressions in the majority of cases across key quality metrics.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ณต๋ถ€ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์˜ ์ธ์ƒ(impression) ์ž‘์„ฑ์„ ์ž๋™ํ™”ํ•˜๋Š” AI ๊ธฐ๋ฐ˜ ๋„๊ตฌ์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜๊ณ ์ž, ๋‹จ์ผ ๊ธฐ๊ด€์—์„œ ํ›„ํ–ฅ์ ์œผ๋กœ ์„ ๋ณ„๋œ 100๊ฑด์˜ ๋ณต๋ถ€ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ AI ์ƒ์„ฑ ์ธ์ƒ(GI)๊ณผ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜๊ฐ€ ์ง์ ‘ ์ž‘์„ฑํ•œ ์ธ์ƒ(RI)์„ 10๋ช…์˜ ์ „๋ฌธ์˜๊ฐ€ ๋งน๊ฒ€ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 5์  Likert ์ฒ™๋„๋ฅผ ์ด์šฉํ•œ ํ‰๊ฐ€์—์„œ GI๋Š” ์ผ๊ด€์„ฑ 79%, ํฌ๊ด„์„ฑ 66%, ์‚ฌ์‹ค์  ์ •ํ™•์„ฑ 77%์˜ ์‚ฌ๋ก€์—์„œ RI์™€ ๋™๋“ฑํ•˜๊ฑฐ๋‚˜ ์šฐ์ˆ˜ํ•œ ๊ฒƒ์œผ๋กœ ํ‰๊ฐ€๋˜์—ˆ์œผ๋ฉฐ, ํ˜ผํ•ฉํšจ๊ณผ ๋กœ์ง€์Šคํ‹ฑ ํšŒ๊ท€๋ถ„์„์—์„œ๋„ GI๊ฐ€ RI์— ๋™๋“ฑํ•˜๊ฑฐ๋‚˜ ์„ ํ˜ธ๋˜๋Š” ๋น„์œจ์ด ์œ ์˜ํ•˜๊ฒŒ ๋†’์•˜๋‹ค(OR 1.34, p = 0.02). ์ด ๊ฒฐ๊ณผ๋Š” AI ์ƒ์„ฑ ์ธ์ƒ์ด ์ž„์ƒ์ ์œผ๋กœ ์ˆ˜์šฉ ๊ฐ€๋Šฅํ•œ ์ˆ˜์ค€์ด๋ฉฐ, ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜์˜ ๋ณด๊ณ ์„œ ์ž‘์„ฑ ์›Œํฌํ”Œ๋กœ์šฐ ํšจ์œจํ™”์— ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ๋Š” ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-07-19 00:00View โ†—

18Examining explainable artificial intelligence in TNM staging with PET-CT: a user-centred observation study.

2026-07La Radiologia medicaDOI 10.1007/s11547-026-02226-9
OBJECTIVE

Artificial intelligence (AI) is increasingly proposed as a solution to improve efficiency in radiology and nuclear medicine, particularly in the context of workforce shortages. However, adoption of AI-based clinical decision support systemsย (AI-CDSS)ย remainsย slow, due to limited model transparency. Explainable AI (XAI) may improve clinician acceptance by supporting oversight and trust. This study evaluated the impact of different XAI explanation types on radiologists' willingness to adopt AI systems.

METHODS

Ten nuclear medicine radiologists from eight UK institutions performed lung cancer TNM staging using whole-body PET/CT scans supported by a simulated AI-CDSS. Three explanation approaches were assessed: input feature attribution, high-level concept explanations and global algorithmic transparency. Adoption likelihood and explanation usefulness were rated using Likert scales and analysed with nonparametric sign tests. Semi-structured interviews were additionally analysed through thematic evaluation supported by large language model-assisted coding with human verification.

RESULTS

All explanation approaches significantly increased radiologists' willingness to adopt the AI system comparedย toย a black-box model (pโ€‰<โ€‰0.05). Explanations were consistently considered useful in enabling participants to confirm or challenge AI staging recommendationsย (pโ€‰<โ€‰0.001). Qualitative findings highlighted the importance of clinical relevance, error detection and decision support value. A trade-off between explanation depth and usability was identified as a key factor influencing preferences.

CONCLUSION

Incorporatingย XAIย into nuclear medicineย CDSSย enhances radiologists' acceptance and provides clinically meaningful information for oversight of AI recommendations,ย in accordance withย the EU AI Act. These findings support the role of XAIย inย facilitatingย integration of AI tools into diagnostic workflows.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” PET/CT ๊ธฐ๋ฐ˜ ํ์•” TNM ๋ณ‘๊ธฐ ๊ฒฐ์ • ๊ณผ์ •์—์„œ ์„ค๋ช… ๊ฐ€๋Šฅํ•œ ์ธ๊ณต์ง€๋Šฅ(XAI)์˜ ์„ธ ๊ฐ€์ง€ ์„ค๋ช… ๋ฐฉ์‹โ€”์ž…๋ ฅ ํŠน์ง• ๊ท€์ธ, ๊ณ ์ˆ˜์ค€ ๊ฐœ๋… ์„ค๋ช…, ์ „์—ญ์  ์•Œ๊ณ ๋ฆฌ์ฆ˜ ํˆฌ๋ช…์„ฑโ€”์ด ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ์˜ AI ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ์‹œ์Šคํ…œ(AI-CDSS) ์ˆ˜์šฉ๋„์— ๋ฏธ์น˜๋Š” ์˜ํ–ฅ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์˜๊ตญ 8๊ฐœ ๊ธฐ๊ด€ ์†Œ์† ํ•ต์˜ํ•™ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ 10๋ช…์„ ๋Œ€์ƒ์œผ๋กœ ๋ชจ์˜ AI-CDSS๋ฅผ ํ™œ์šฉํ•œ ๊ด€์ฐฐ ์—ฐ๊ตฌ๋ฅผ ์ˆ˜ํ–‰ํ•˜์˜€์œผ๋ฉฐ, ๋ฆฌ์ปคํŠธ ์ฒ™๋„ ๋ฐ ๋น„๋ชจ์ˆ˜ ๋ถ€ํ˜ธ ๊ฒ€์ •๊ณผ ๋ฐ˜๊ตฌ์กฐํ™” ์ธํ„ฐ๋ทฐ๋ฅผ ํ†ตํ•ด ๊ฒฐ๊ณผ๋ฅผ ๋ถ„์„ํ•˜์˜€๋‹ค. ์„ธ ๊ฐ€์ง€ XAI ์„ค๋ช… ๋ฐฉ์‹ ๋ชจ๋‘ ๋ธ”๋ž™๋ฐ•์Šค ๋ชจ๋ธ ๋Œ€๋น„ AI ์‹œ์Šคํ…œ ๋„์ž… ์˜ํ–ฅ์„ ์œ ์˜ํ•˜๊ฒŒ ํ–ฅ์ƒ์‹œ์ผฐ๊ณ (p < 0.05), ์„ค๋ช…์˜ ์ž„์ƒ์  ์œ ์šฉ์„ฑ์ด AI ๊ถŒ๊ณ ์•ˆ์— ๋Œ€ํ•œ ํ™•์ธ ๋ฐ ์˜ค๋ฅ˜ ๊ฒ€์ถœ์„ ์ง€์›ํ•จ์œผ๋กœ์จ ์ง„๋‹จ ์›Œํฌํ”Œ๋กœ์šฐ ๋‚ด AI ํ†ตํ•ฉ์„ ์ด‰์ง„ํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-07-19 00:00View โ†—

19Patient and physician perspectives on large language model generated responses about brain aneurysm.

2026-07Child's nervous system : ChNS : official journal of the International Society for Pediatric Neurosurgery๐Ÿ”ท Q2DOI 10.1007/s00381-026-07392-9
OBJECTIVE

Large language models (LLMs) are becoming increasingly popular in medicine and neurosurgery. Because LLMs are not trained in specific subspecialties or diagnoses, a better understanding of the implications, effectiveness, and use of LLMs by users and clinicians is necessary. To better understand LLM's effectiveness in neurosurgery and aneurysms, we compared community and physician feedback on ChatGPT 4o and Gemini 1.5 Flash responses to frequently asked questions regarding brain aneurysms.

METHODS

External surveys were made available on the Brain Aneurysm Foundation page for patients and families to complete and internal surveys were distributed and completed by physicians in the department of neurosurgery at Boston Children's Hospital.

RESULTS

In the community survey assessing response usefulness and helpfulness, ChatGPT and Gemini provided different response quality despite similar AI sentiment. Clarity of procedure explanation (pโ€‰=โ€‰0.04), discussion of alternative procedures (pโ€‰=โ€‰0.02), and discussion of procedure risks (pโ€‰=โ€‰0.01) were different. The physician survey, assessing response accuracy, safety, and helpfulness, also found differences in multiple domains. Importantly, differences were found in consistency with current medical knowledge and practice guidelines (pโ€‰=โ€‰0.001), omittance of key points (pโ€‰<โ€‰0.001), and amount of clinically relevant detail included (pโ€‰<โ€‰0.001).

CONCLUSION

LLMs had variable performance across several key domains, consistent with previous research. Despite the apparent advantages of ChatGPT, physician feedback highlighted the continued need for information oversight. Interestingly, community participants consistently found LLM responses to be better than physician ones, while physicians found LLM responses to be similar or somewhat worse than the one they would have provided.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋‡Œ๋™๋งฅ๋ฅ˜์— ๊ด€ํ•œ ์ž์ฃผ ๋ฌป๋Š” ์งˆ๋ฌธ์— ๋Œ€ํ•ด ChatGPT 4o์™€ Gemini 1.5 Flash๊ฐ€ ์ƒ์„ฑํ•œ ์‘๋‹ต์˜ ์œ ์šฉ์„ฑ๊ณผ ์ •ํ™•์„ฑ์„ ํ™˜์ž ๋ฐ ์˜์‚ฌ ๊ด€์ ์—์„œ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ๋‡Œ๋™๋งฅ๋ฅ˜์žฌ๋‹จ ์›นํŽ˜์ด์ง€๋ฅผ ํ†ตํ•œ ํ™˜์žยท๊ฐ€์กฑ ๋Œ€์ƒ ์™ธ๋ถ€ ์„ค๋ฌธ๊ณผ ๋ณด์Šคํ„ด ์†Œ์•„๋ณ‘์› ์‹ ๊ฒฝ์™ธ๊ณผ ์˜์‚ฌ ๋Œ€์ƒ ๋‚ด๋ถ€ ์„ค๋ฌธ์„ ์‹œํ–‰ํ•˜์—ฌ ์‘๋‹ต์˜ ๋ช…ํ™•์„ฑ, ์•ˆ์ „์„ฑ, ํ˜„ํ–‰ ์ง„๋ฃŒ์ง€์นจ ์ผ์น˜๋„ ๋“ฑ ๋‹ค์–‘ํ•œ ์˜์—ญ์„ ๋ถ„์„ํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, ChatGPT๊ฐ€ ์ „๋ฐ˜์ ์œผ๋กœ ์šฐ์ˆ˜ํ•œ ๊ฒฝํ–ฅ์„ ๋ณด์˜€์œผ๋‚˜ ๋‘ ๋ชจ๋ธ ๋ชจ๋‘ ํ•ต์‹ฌ ๋‚ด์šฉ ๋ˆ„๋ฝ ๋ฐ ์ž„์ƒ์  ์„ธ๋ถ€ ์ •๋ณด ๋ถ€์กฑ ๋“ฑ์˜ ํ•œ๊ณ„๋ฅผ ๋‚˜ํƒ€๋ƒˆ์œผ๋ฉฐ, ํ™˜์ž๋“ค์€ LLM ์‘๋‹ต์„ ์˜์‚ฌ ์‘๋‹ต๋ณด๋‹ค ์šฐ์ˆ˜ํ•˜๊ฒŒ ํ‰๊ฐ€ํ•œ ๋ฐ˜๋ฉด ์˜์‚ฌ๋“ค์€ LLM ์‘๋‹ต์ด ๋™๋“ฑํ•˜๊ฑฐ๋‚˜ ๋‹ค์†Œ ์—ด๋“ฑํ•˜๋‹ค๊ณ  ํ‰๊ฐ€ํ•˜์—ฌ ์˜๋ฃŒ ์ •๋ณด ๊ฐ๋…์˜ ์ง€์†์ ์ธ ํ•„์š”์„ฑ์ด ๊ฐ•์กฐ๋˜์—ˆ๋‹ค.
Added: 2026-07-19 00:00View โ†—

20Patient understanding of AI-simplified oncology imaging reports requires further validation.

2026-07Cancer imaging : the official publication of the International Cancer Imaging Societyโญ Q1DOI 10.1186/s40644-026-01091-z

Ribeiro and colleagues offer timely evidence that large language models can make oncology imaging reports more accessible to radiologists and patient and public involvement representatives. Their findings also highlight an important distinction between a report that is easier to read and one that patients understand accurately. As the authors acknowledge, the patient-facing assessment included three representatives with previous experience in oncology imaging research and was restricted to reports that had received high factual-correctness scores from radiologists. This approach was suitable for comparing presentation preferences, although it provides limited evidence about comprehension in routine practice, particularly for lower-scoring or clinically ambiguous outputs. Further studies could examine unselected reports in patients with varied health and digital literacy, using outcomes such as comprehension, recognition of uncertainty, emotional response, and intended action. In future adequately powered studies, accounting for ratings clustered by reader and report may provide more precise estimates. Such evidence would help clarify the role of AI-generated explanations in patient-facing oncology care.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ๋…ผ๋ฌธ์€ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์ข…์–‘ํ•™ ์˜์ƒ ๋ณด๊ณ ์„œ๋ฅผ ํ™˜์ž์—๊ฒŒ ๋” ์‰ฝ๊ฒŒ ์ „๋‹ฌํ•  ์ˆ˜ ์žˆ๋‹ค๋Š” ์ตœ๊ทผ ์—ฐ๊ตฌ(Ribeiro ๋“ฑ)์— ๋Œ€ํ•œ ๋น„ํ‰์  ๋…ผํ‰์œผ๋กœ, ๊ฐ€๋…์„ฑ ํ–ฅ์ƒ์ด ๋ฐ˜๋“œ์‹œ ์ •ํ™•ํ•œ ํ™˜์ž ์ดํ•ด๋ฅผ ๋ณด์žฅํ•˜์ง€ ์•Š๋Š”๋‹ค๋Š” ์ ์„ ๊ฐ•์กฐํ•œ๋‹ค. ํ•ด๋‹น ์›์—ฐ๊ตฌ์˜ ํ™˜์ž ํ‰๊ฐ€๋Š” ์ข…์–‘ํ•™ ์˜์ƒ ์—ฐ๊ตฌ ๊ฒฝํ—˜์ด ์žˆ๋Š” 3๋ช…์˜ ๋Œ€ํ‘œ์ž์— ํ•œ์ •๋˜์—ˆ๊ณ , ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ๋กœ๋ถ€ํ„ฐ ๋†’์€ ์‚ฌ์‹ค ์ •ํ™•๋„ ์ ์ˆ˜๋ฅผ ๋ฐ›์€ ๋ณด๊ณ ์„œ๋งŒ์„ ๋Œ€์ƒ์œผ๋กœ ํ•˜์—ฌ ์‹ค์ œ ์ž„์ƒ ํ™˜๊ฒฝ์—์„œ์˜ ์ผ๋ฐ˜ํ™” ๊ฐ€๋Šฅ์„ฑ์ด ์ œํ•œ์ ์ด๋‹ค. ์ €์ž๋“ค์€ ํ–ฅํ›„ ์—ฐ๊ตฌ์—์„œ ๋‹ค์–‘ํ•œ ๊ฑด๊ฐ• ๋ฐ ๋””์ง€ํ„ธ ๋ฆฌํ„ฐ๋Ÿฌ์‹œ๋ฅผ ๊ฐ€์ง„ ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ ์„ ๋ณ„๋˜์ง€ ์•Š์€ ๋ณด๊ณ ์„œ๋ฅผ ํ™œ์šฉํ•˜์—ฌ ์ดํ•ด๋„, ๋ถˆํ™•์‹ค์„ฑ ์ธ์‹, ๊ฐ์ •์  ๋ฐ˜์‘ ๋ฐ ์˜๋„๋œ ํ–‰๋™ ๋“ฑ์„ ํ‰๊ฐ€ํ•˜๋Š” ๋ณด๋‹ค ์—„๋ฐ€ํ•œ ๊ทผ๊ฑฐ๋ฅผ ๋งˆ๋ จํ•  ๊ฒƒ์„ ์ด‰๊ตฌํ•œ๋‹ค.
Added: 2026-07-19 00:00View โ†—

21Integrating 3D Volumetric Segmentation and LLM-Based Classification for csPCa Detection on mpMRI: Multi-Institutional External Validation.

2026-07Academic radiologyโญ Q1DOI 10.1016/j.acra.2026.06.029

RATIONALE AND

OBJECTIVE

To develop and externally validate integrated models combining three-dimensional (3D) volumetric segmentation and large language model (LLM)-based slice-wise classification for improved detection of clinically significant Prostate Cancer (csPCa) on multiparametric Magnetic Resonance Imaging (mpMRI).

METHODS

This retrospective multi-institutional study included 5050 patients (3896 for model development and 1154 for external validation) who underwent mpMRI. A 3D V-Net was trained for voxel-wise csPCa segmentation, generating three patient-level metrics: positive volume, positive slice count, and positive slice rate. A four-billion-parameter MedGemma-IT LLM was fine-tuned for slice-level classification, producing positive slice count and rate metrics. The optimal V-Net and LLM metrics were integrated into two combined models: Combined Model 1 (logistic regression of V-Net and LLM positive slice rates) and Combined Model 2 (zero-adjusted Gamma regression of V-Net positive volume and LLM positive slice rate). Performance was evaluated in the external validation cohort using area under the receiver operating characteristic curve (AUROC), area under the precision-recall curve (AUPRC), decision curve analysis (DCA), net reclassification improvement (NRI), and integrated discrimination improvement (IDI).

RESULTS

In external validation, Combined Model 1 achieved an AUROC of 0.900 (95% CI: 0.882-0.918) and Combined Model 2 achieved 0.885 (95% CI: 0.865-0.905), significantly outperforming all individual V-Net and LLM metrics (ฮ”AUROC: 0.060-0.145; all P <ย 0.001). Both combined models demonstrated significant improvements in NRI and IDI for most comparisons (P <ย 0.05 for most) and provided superior standardized net benefit across clinically relevant risk thresholds (0.3-0.9) on DCA.

CONCLUSION

Integrating 3D volumetric segmentation with LLM-based slice-wise classification significantly improves csPCa detection accuracy and clinical utility compared to single-modality approaches, demonstrating promising generalizability in an independent, multi-institutional external validation cohort with substantial scanner and protocol heterogeneity.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋‹ค๊ธฐ๊ด€ mpMRI ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•˜์—ฌ 3์ฐจ์› ์ฒด์  ๋ถ„ํ•  ๋ชจ๋ธ(3D V-Net)๊ณผ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(MedGemma-IT) ๊ธฐ๋ฐ˜ ์Šฌ๋ผ์ด์Šค ๋‹จ์œ„ ๋ถ„๋ฅ˜๋ฅผ ํ†ตํ•ฉํ•œ ์ž„์ƒ์ ์œผ๋กœ ์œ ์˜ํ•œ ์ „๋ฆฝ์„ ์•”(csPCa) ๊ฒ€์ถœ ๋ชจ๋ธ์„ ๊ฐœ๋ฐœํ•˜๊ณ  ์™ธ๋ถ€ ๊ฒ€์ฆํ•˜์˜€๋‹ค. ์ด 5,050๋ช…์˜ ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ ๋‘ ๊ฐ€์ง€ ํ†ตํ•ฉ ๋ชจ๋ธ(๋กœ์ง€์Šคํ‹ฑ ํšŒ๊ท€ ๋ฐ ์˜-์กฐ์ • ๊ฐ๋งˆ ํšŒ๊ท€)์„ ๊ตฌ์ถ•ํ•˜์˜€์œผ๋ฉฐ, 1,154๋ช…์˜ ๋…๋ฆฝ ์™ธ๋ถ€ ๊ฒ€์ฆ ์ฝ”ํ˜ธํŠธ์—์„œ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์™ธ๋ถ€ ๊ฒ€์ฆ ๊ฒฐ๊ณผ, ํ†ตํ•ฉ ๋ชจ๋ธ 1๊ณผ 2๋Š” ๊ฐ๊ฐ AUROC 0.900 ๋ฐ 0.885๋ฅผ ๋‹ฌ์„ฑํ•˜์—ฌ ๋‹จ์ผ ๋ชจ๋‹ฌ๋ฆฌํ‹ฐ ๋ชจ๋ธ ๋Œ€๋น„ ์œ ์˜ํ•˜๊ฒŒ ์šฐ์ˆ˜ํ•œ ํŒ๋ณ„ ๋Šฅ๋ ฅ ๋ฐ ์ž„์ƒ์  ์ˆœ์ด์ต์„ ๋ณด์˜€์œผ๋ฉฐ, ๋‹ค์–‘ํ•œ ์Šค์บ๋„ˆ ๋ฐ ํ”„๋กœํ† ์ฝœ ํ™˜๊ฒฝ์—์„œ๋„ ๋†’์€ ์ผ๋ฐ˜ํ™” ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-07-12 00:00View โ†—

22PIPA: Prior-Driven Prompting with Diagnosis-Oriented Retrieval-Augmentation for 3D Radiology Report Generation.

2026-07IEEE transactions on medical imagingโญ Q1DOI 10.1109/tmi.2026.3710717

Automatic radiology report generation has gained increasing attention for its potential to assist in clinical reporting and reduce the workload of radiologists. Existing 3D radiology report generation methods employ multi-modal foundation model to encode volume-text inputs and produce diagnosis reports, while they ignore the characteristics of 3D volumes including much background regions and suffer from generating hallucinations, especially in medical domain that contains many uncommon professional terms. In this paper, we aim to efficiently adapt the pre-trained foundation model to specific 3D radiology report generation, and present a Prior-drIven Prompting with diagnosis-oriented retrieval-Augmentation (PIPA) framework. In PIPA, we design a Prior-drIven Prompting (PIP) strategy to exploit diagnostic knowledge from input volumes and a Diagnosis-oriented volume-report retrievalaugmentation Generation (DIG) module to explore beneficial knowledge from external database. Specifically, in PIP, to take full advantage of the patient's clinical information, e.g., age and symptoms, and the possible disease information, e.g., brain tumor, edema, we formulate them as the patient and disease priors to mine clinical relevant knowledge. Furthermore, we propose utilizing visual and textual embeddings as queries to retrieve similar external data by devising a diagnosis-oriented retrieval-augmentation scheme for leveraging more report resources as references for LLM to produce accuracy outcomes. With PIP and DIG, PIPA integrates clinical priors and external data to learn effective diagnostic representations for high-quality report generation. We evaluate the framework on both public and in-house 3D medical datasets with corresponding reports, demonstrating its strong performance in generating accurate diagnosis reports. Source codes have been published at https://github.com/CUHK-AIM-Group/PIPA/tree/ main.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 3D ๋ฐฉ์‚ฌ์„  ์˜์ƒ ํŒ๋… ๋ณด๊ณ ์„œ์˜ ์ž๋™ ์ƒ์„ฑ ์‹œ ๋ฐœ์ƒํ•˜๋Š” ํ™˜๊ฐ(hallucination) ๋ฌธ์ œ์™€ ๋ฐฐ๊ฒฝ ์˜์—ญ ๊ณผ๋‹ค๋กœ ์ธํ•œ ๋น„ํšจ์œจ์„ฑ์„ ํ•ด๊ฒฐํ•˜๊ธฐ ์œ„ํ•ด, ์‚ฌ์ „ ์ง€์‹ ๊ธฐ๋ฐ˜ ํ”„๋กฌํ”„ํŒ…๊ณผ ์ง„๋‹จ ์ง€ํ–ฅ ๊ฒ€์ƒ‰ ์ฆ๊ฐ•์„ ๊ฒฐํ•ฉํ•œ PIPA(Prior-drIven Prompting with diagnosis-oriented retrieval-Augmentation) ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ์ œ์•ˆํ•˜์˜€๋‹ค. PIPA๋Š” ํ™˜์ž์˜ ๋‚˜์ดยท์ฆ์ƒ ๋ฐ ์งˆํ™˜ ์ •๋ณด๋ฅผ ์ž„์ƒ ์‚ฌ์ „ ์ง€์‹(prior)์œผ๋กœ ํ™œ์šฉํ•˜๋Š” PIP ์ „๋žต๊ณผ, ์‹œ๊ฐยทํ…์ŠคํŠธ ์ž„๋ฒ ๋”ฉ์„ ์ฟผ๋ฆฌ๋กœ ์‚ฌ์šฉํ•˜์—ฌ ์™ธ๋ถ€ ๋ฐ์ดํ„ฐ๋ฒ ์ด์Šค์—์„œ ์œ ์‚ฌ ์ฆ๋ก€๋ฅผ ๊ฒ€์ƒ‰ยท์ฐธ์กฐํ•˜๋Š” DIG ๋ชจ๋“ˆ๋กœ ๊ตฌ์„ฑ๋œ๋‹ค. ๊ณต๊ฐœ ๋ฐ์ดํ„ฐ์…‹ ๋ฐ ์ž์ฒด 3D ์˜๋ฃŒ ๋ฐ์ดํ„ฐ์…‹์„ ๋Œ€์ƒ์œผ๋กœ ํ•œ ํ‰๊ฐ€์—์„œ PIPA๋Š” ๊ธฐ์กด ๋ฐฉ๋ฒ• ๋Œ€๋น„ ์šฐ์ˆ˜ํ•œ ์ง„๋‹จ ๋ณด๊ณ ์„œ ์ƒ์„ฑ ์„ฑ๋Šฅ์„ ๋ณด์—ฌ, ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜์˜ ์ž„์ƒ ์—…๋ฌด ๋ถ€๋‹ด ๊ฒฝ๊ฐ์— ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ์Œ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-07-12 00:00View โ†—

23Agentic Artificial Intelligence for the Automated Generation of Accurate Summary Podcasts of Radiology Research Papers.

2026-07Korean journal of radiologyโญ Q1DOI 10.3348/kjr.2026.0203
OBJECTIVE

To evaluate whether a custom agentic artificial intelligence (AI) pipeline can overcome the limitations of general-purpose large language model tools, when compared with a generic commercial tool (Google NotebookLM [NBLM]), for generating podcast-style summaries of radiology research articles.

METHODS

Twenty-two PDF-format original research articles published in the April 2025 issue of Radiology were processed using our Programmable, Phoneme-Aware PDF-to-Podcast Pipeline (P5) and NBLM to generate 44 audio episodes. P5 utilizes a multi-agent workflow for script generation, quality assurance, pronunciation enhancement, and audio synthesis. Four radiologists from a pool of 25 (7 generalists and 18 specialists) were randomly assigned to evaluate each blinded audio episode, yielding 176 total evaluations. The primary outcomes were the number of hallucinations (factual errors) per episode and the percentage of hallucination-free episodes. Secondary outcomes included the number of inappropriate statements, mispronunciations, and flow disruptions; the composite quality score (Quality Assessment of Educational Podcasts [QAEP]); the key results coverage score; and overall listener preference. Data were analyzed using generalized linear mixed models.

RESULTS

The P5 method produced significantly fewer hallucinations per episode compared with NBLM (mean, 0.32 vs. 0.93; P < 0.001) and a higher proportion of hallucination-free episodes (71.6% [63/88] vs. 56.8% [50/88]; P = 0.013), consistently across generalists and specialists. P5 demonstrated significantly fewer mispronunciations (mean, 0.11 vs. 1.62; P < 0.001) and flow disruptions (mean, 0.19 vs. 1.06; P < 0.001) per episode, as well as higher mean QAEP composite scores (4.62 vs. 4.28; P < 0.001), compared with NBLM, consistently across generalists and specialists. Raters preferred P5 over NBLM in 72.7% (64/88) of comparisons (P = 0.003).

CONCLUSION

Our custom agentic AI pipeline generated podcast-style summaries of radiology research articles with significantly higher quality and greater listener preference than the generic commercial tool.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„ ํ•™ ์—ฐ๊ตฌ ๋…ผ๋ฌธ์˜ ํŒŸ์บ์ŠคํŠธ ์š”์•ฝ ์ƒ์„ฑ์— ์žˆ์–ด ๋งž์ถคํ˜• ๋‹ค์ค‘ ์—์ด์ „ํŠธ AI ํŒŒ์ดํ”„๋ผ์ธ(P5)์ด ๋ฒ”์šฉ ์ƒ์—… ๋„๊ตฌ(Google NotebookLM)๋ฅผ ๋Šฅ๊ฐ€ํ•˜๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. *Radiology* 2025๋…„ 4์›”ํ˜ธ์— ๊ฒŒ์žฌ๋œ ์›์ € 22ํŽธ์„ ๋‘ ์‹œ์Šคํ…œ์œผ๋กœ ์ฒ˜๋ฆฌํ•˜์—ฌ ์ด 44๊ฐœ์˜ ์˜ค๋””์˜ค ์—ํ”ผ์†Œ๋“œ๋ฅผ ์ƒ์„ฑํ•˜๊ณ , 25๋ช…์˜ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜๊ฐ€ 176๊ฑด์˜ ํ‰๊ฐ€๋ฅผ ์ˆ˜ํ–‰ํ•˜์˜€๋‹ค. P5๋Š” NBLM ๋Œ€๋น„ ์—ํ”ผ์†Œ๋“œ๋‹น ํ™˜๊ฐ(์‚ฌ์‹ค ์˜ค๋ฅ˜) ๋ฐœ์ƒ ๊ฑด์ˆ˜๊ฐ€ ์œ ์˜ํ•˜๊ฒŒ ์ ์—ˆ์œผ๋ฉฐ(ํ‰๊ท  0.32 vs. 0.93, P < 0.001), ๋ฐœ์Œ ์˜ค๋ฅ˜ ๋ฐ ํ๋ฆ„ ๋‹จ์ ˆ๋„ ํ˜„์ €ํžˆ ๊ฐ์†Œํ•˜์˜€๊ณ , ์ „๋ฐ˜์ ์ธ ์ฒญ์ทจ ์„ ํ˜ธ๋„์—์„œ๋„ 72.7%์˜ ๋น„๊ต์—์„œ ์šฐ์œ„๋ฅผ ๋ณด์—ฌ ๋งž์ถคํ˜• ์—์ด์ „ํŠธ AI ํŒŒ์ดํ”„๋ผ์ธ์ด ๋ฐฉ์‚ฌ์„ ํ•™ ์—ฐ๊ตฌ ์š”์•ฝ ํŒŸ์บ์ŠคํŠธ ์ƒ์„ฑ์— ์žˆ์–ด ๋ฒ”์šฉ ์ƒ์—… ๋„๊ตฌ๋ณด๋‹ค ์œ ์˜ํ•˜๊ฒŒ ๋†’์€ ํ’ˆ์งˆ์„ ์ œ๊ณตํ•จ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-07-05 00:00View โ†—

24Certainty Language Use in Pediatric Radiology: A Single-Institution Analysis.

2026-07Journal of the American College of Radiology : JACRโญ Q1DOI 10.1016/j.jacr.2026.03.003
BACKGROUND

Radiologists often employ diagnostic certainty phrases (DCPs) to convey levels of confidence in imaging interpretations. Prior research in adult radiology demonstrated wide variability in DCP usage, potentially complicating communication with clinicians and patients. Little is known about these practices in pediatric radiology. We aimed to characterize DCP use among pediatric radiologists in a large academic institution.

METHODS

We retrospectively analyzed radiology reports from a freestanding pediatric hospital system between October 2023 and September 2024. From 309,751 pediatric imaging reports meeting inclusion criteria, 10,000 were randomly sampled to identify DCPs using a natural language processing pipeline (GPT-4o, OpenAI, San Francisco, California). After exclusions and manual review, 122 unique DCPs were identified. We then counted the occurrence of these 122 DCPs across the remaining 309,751 radiology report first impressions. Usage rates were evaluated across patient demographics, imaging modalities, clinical settings, and radiologist seniority. Statistical significance was assessed using two-sided one-sample t tests with significance assigned to values P โ‰ค .001.

RESULTS

Of the 309,751 analyzable reports, 20.9% contained at least one DCP in the first impression. The mean DCP frequency was 0.27 per report. MRI (0.54, P < .001), patients aged 1 to 5 years (0.34, P < .001), and emergency cases (0.31, P < .001) showed significantly higher usage. Midcareer and early-career radiologists employed more DCPs (0.30, P < .001) than senior radiologists (0.22, P < .001). The most common phrases included "likely," "could represent," and "cannot rule out."

CONCLUSION

DCPs are frequently used in pediatric radiology, with variability influenced by modality, patient age, clinical context, and radiologist seniority. Such variability may compound existing uncertainty of the clinical diagnostic process and possibly indicate an opportunity for standardized uncertainty reporting approaches.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋‹จ์ผ ์†Œ์•„ ์ „๋ฌธ ํ•™์ˆ  ๋ณ‘์›์—์„œ ๋ฐœํ–‰๋œ 309,751๊ฑด์˜ ์˜์ƒ ํŒ๋…๋ฌธ์„ ๋Œ€์ƒ์œผ๋กœ ์ž์—ฐ์–ด์ฒ˜๋ฆฌ(GPT-4o) ํŒŒ์ดํ”„๋ผ์ธ์„ ํ™œ์šฉํ•˜์—ฌ ์ง„๋‹จ์  ํ™•์‹ค์„ฑ ํ‘œํ˜„(DCP)์˜ ์‚ฌ์šฉ ์–‘์ƒ์„ ํ›„ํ–ฅ์ ์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ์ „์ฒด ํŒ๋…๋ฌธ์˜ 20.9%์—์„œ ์ตœ์†Œ 1๊ฐœ ์ด์ƒ์˜ DCP๊ฐ€ ์‚ฌ์šฉ๋˜์—ˆ์œผ๋ฉฐ, MRI ํŒ๋…, 1โ€“5์„ธ ํ™˜์ž๊ตฐ, ์‘๊ธ‰ ์ž„์ƒ ํ™˜๊ฒฝ, ๊ทธ๋ฆฌ๊ณ  ๊ฒฝ๋ ฅ ์ดˆยท์ค‘๋ฐ˜ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜์—์„œ DCP ์‚ฌ์šฉ ๋นˆ๋„๊ฐ€ ์œ ์˜ํ•˜๊ฒŒ ๋†’์•˜๊ณ , ๊ฐ€์žฅ ํ”ํžˆ ์‚ฌ์šฉ๋œ ํ‘œํ˜„์€ "likely," "could represent," "cannot rule out"์ด์—ˆ๋‹ค. ์ด๋Ÿฌํ•œ DCP ์‚ฌ์šฉ์˜ ๋†’์€ ๋ณ€๋™์„ฑ์€ ์ž„์ƒ์  ์ง„๋‹จ ๊ณผ์ •์˜ ๋ถˆํ™•์‹ค์„ฑ์„ ๊ฐ€์ค‘์‹œํ‚ฌ ์ˆ˜ ์žˆ์œผ๋ฉฐ, ์†Œ์•„ ์˜์ƒ์˜ํ•™ ๋ถ„์•ผ์—์„œ ๋ถˆํ™•์‹ค์„ฑ ํ‘œํ˜„์˜ ํ‘œ์ค€ํ™”๋œ ๋ณด๊ณ  ์ฒด๊ณ„ ๋งˆ๋ จ์ด ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-07-05 00:00View โ†—

25Integrating Generative AI into Nomograms for Breast Cancer Nodal Risk Predictions.

2026-07Annals of surgical oncologyโญ Q1DOI 10.1245/s10434-026-19482-8
BACKGROUND

Nomograms predicting the likelihood of sentinel lymph node (SLN) metastasis in early-stage breast cancer can aid surgical decision-making but are underused due to the burden of variable review and input. This study evaluated whether OpenAI's large language models (LLMs) can extract information required for nomogram use and reproduce the estimated rate of SLN metastasis using the Memorial Sloan Kettering and MD Anderson nomograms.

METHODS

The study analyzed the de-identified radiology and pathology notes for 20 patients. Three prompts were tested: (1) o1 prompted to generate SLN metastasis estimates without nomograms, (2) o1 prompted to use the nomogram from online calculators with and without serial corrections, and (3) GPT-4o prompted to use chain-of-thought reasoning from nomogram variables and the corresponding point values from a pictorial nomogram. Artificial intelligence (AI) estimates of SLN metastasis rates were compared with a physician-expert's manual use of nomograms.

RESULTS

OpenAI o1 captured all clinical variables in 65% of cases without serial correction (94-95% of individual variables) and 80% of cases with serial correction (96-97% of individual variables), with tumor size most frequently misidentified. Agreement between LLM-generated and physician-calculated risk prediction was low (exact matches in 0-10% of cases, near agreement in 20-45% of cases), indicating moderate rater reliability (intraclass correlation coefficient [ICC], 0.57-0.62). The optimized GPT-4o prompt demonstrated greater agreement (exact matches in 25% and near agreement in 90% of cases) and reliability (ICC, 0.95).

CONCLUSION

Large language models can reliably extract nomogram inputs from clinical notes but require task-specific prompt engineering for accurate SLN metastasis risk estimation. Currently, automating nomogram-based risk estimators with AI may not justify the significant resources required for optimization.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์กฐ๊ธฐ ์œ ๋ฐฉ์•” ํ™˜์ž์˜ ๊ฐ์‹œ ๋ฆผํ”„์ ˆ(SLN) ์ „์ด ์œ„ํ—˜์„ ์˜ˆ์ธกํ•˜๋Š” ๋…ธ๋ชจ๊ทธ๋žจ(MSK, MD Anderson)์— OpenAI์˜ ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ(LLM)์„ ํ†ตํ•ฉํ•˜์—ฌ ์ž„์ƒ ๋ณ€์ˆ˜ ์ถ”์ถœ ๋ฐ ์œ„ํ—˜๋„ ์‚ฐ์ถœ ์ž๋™ํ™”์˜ ๊ฐ€๋Šฅ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 20๋ช…์˜ ์ต๋ช…ํ™”๋œ ์˜์ƒยท๋ณ‘๋ฆฌ ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ o1 ๋ฐ GPT-4o ๋ชจ๋ธ์— ๋Œ€ํ•ด ๋‹ค์–‘ํ•œ ํ”„๋กฌํ”„ํŠธ ์ „๋žต์„ ์ ์šฉํ•˜์˜€์œผ๋ฉฐ, AI ์‚ฐ์ถœ ๊ฒฐ๊ณผ๋ฅผ ์ „๋ฌธ์˜์˜ ์ˆ˜๋™ ๋…ธ๋ชจ๊ทธ๋žจ ๊ณ„์‚ฐ๊ฐ’๊ณผ ๋น„๊ต ๋ถ„์„ํ•˜์˜€๋‹ค. LLM์€ ์ž„์ƒ ๋ณ€์ˆ˜ ์ถ”์ถœ์—์„œ ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์œผ๋‚˜, ํ‘œ์ค€ ์ ‘๊ทผ๋ฒ•์—์„œ์˜ ์œ„ํ—˜๋„ ์˜ˆ์ธก ์ผ์น˜๋„๋Š” ๋‚ฎ์•˜๊ณ (ICC 0.57โ€“0.62), ์ตœ์ ํ™”๋œ GPT-4o ํ”„๋กฌํ”„ํŠธ๋งŒ์ด ๋†’์€ ์‹ ๋ขฐ๋„(ICC 0.95)๋ฅผ ๋‹ฌ์„ฑํ•˜์—ฌ, ํ˜„์žฌ ์ˆ˜์ค€์˜ AI ๊ธฐ๋ฐ˜ ๋…ธ๋ชจ๊ทธ๋žจ ์ž๋™ํ™”๋Š” ์ตœ์ ํ™”์— ์š”๊ตฌ๋˜๋Š” ์ƒ๋‹นํ•œ ์ž์› ํˆฌ์ž๋ฅผ ์ •๋‹นํ™”ํ•˜๊ธฐ ์–ด๋ ต๋‹ค๋Š” ๊ฒฐ๋ก ์„ ์ œ์‹œํ•˜์˜€๋‹ค.
Added: 2026-07-05 00:00View โ†—

26Large Language Models for Cardiac MRI Diagnosis Based on Standardized Text Descriptions.

2026-07Journal of magnetic resonance imaging : JMRIโญ Q1DOI 10.1002/jmri.70327
BACKGROUND

MRI is important for cardiac disease evaluation, but accurate diagnosis remains challenging in less experienced centers. Although large language models (LLMs) have shown promise in medical imaging diagnosis, their application in cardiac MRI is limited. HYPOTHESIS: LLMs may be effective in achieving cardiac MRI diagnosis based on standardized descriptions. STUDY TYPE: Retrospective. POPULATION: A total of 203 hypertrophic cardiomyopathy, 186 dilated cardiomyopathy, 46 hypertensive heart disease, 198 ischemic cardiomyopathy, 38 constrictive pericarditis, 45 cardiac amyloidosis, 91 myocarditis, and 144 normal controls. FIELD STRENGTH/SEQUENCES: Balanced steady-state free-precession, short tau inversion recovery, and breath-hold inversion-recovery segmented gradient-echo sequences at 3.0โ€‰T. ASSESSMENT: Clinical and cardiac MRI information from each subject was converted into standardized descriptions and input into Generative Pre-trained Transformer-4.5 (GPT-4.5), GPT-4 Omni (GPT-4o), Deepseek-V3, and Deepseek-R1 LLMs. Cardiac MRI information included LV function, wall thickness and motion, and abnormalities on T2WI, perfusion and late gadolinium enhancement sequences. Each model was asked to generate an imaging diagnosis. In addition, a medical student (8โ€‰months experience) and three radiologists (junior, mid-level and senior: with 3, 6, and 10โ€‰years' experience, respectively) provided diagnoses based on cardiac MRI images and clinical information. STATISTIC TESTS: Frequency-weighted sensitivity and specificity were calculated. The diagnostic performances of the LLMs and human readers were compared using the McNemar test with Bonferroni correction. A p value <โ€‰0.05 was considered significant.

RESULTS

All LLMs showed excellent frequency-weighted specificity (0.973-0.983). The frequency-weighted sensitivities of all LLMs were not significantly different from that of the junior radiologist, were significantly higher than that of the medical student, and significantly inferior to those of the senior radiologist (GPT-4.5: 0.863, GPT-4o: 0.821, Deepseek-V3: 0.843, and Deepseek-R1: 0.851 vs. junior radiologist: 0.850, all adjusted pโ€‰=โ€‰1.000; vs. medical student: 0.731, all adjusted pโ€‰<โ€‰0.001; vs. senior radiologist: 0.942, all adjusted pโ€‰<โ€‰0.001). Additionally, the mid-level radiologist achieved a frequency-weighted sensitivity of 0.895, outperforming all LLMs except GPT-4.5. DATA

CONCLUSION

LLMs may generate accurate diagnoses from standardized cardiac MRI descriptions, potentially benefiting less experienced physicians. TECHNICAL EFFICACY: Stage 5. Reading cardiac MRI scans can be difficult, especially for doctors who have less clinical experience. New artificial intelligence tools called large language models (LLMs) may help support this process. In our study, we found that several LLMs were able to make diagnostic judgments at a level similar to junior radiologists. This suggests that LLMs could be used as supportive tools to provide an initial interpretation of cardiac MRI examinations. Such assistance may help improve diagnostic efficiency, reduce workload, and promote more consistent early evaluation, particularly in settings where experienced cardiac imaging specialists are not readily available.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ํ‘œ์ค€ํ™”๋œ ์‹ฌ์žฅ MRI ํ…์ŠคํŠธ ์„ค๋ช…์„ ๊ธฐ๋ฐ˜์œผ๋กœ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์‹ฌ์žฅ ์งˆํ™˜ ์ง„๋‹จ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ์ด 951๋ช…์˜ ๋‹ค์–‘ํ•œ ์‹ฌ์žฅ ์งˆํ™˜ ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ ์ž„์ƒ ๋ฐ ์‹ฌ์žฅ MRI ์ •๋ณด๋ฅผ ํ‘œ์ค€ํ™”๋œ ํ˜•์‹์œผ๋กœ ๋ณ€ํ™˜ํ•˜์—ฌ GPT-4.5, GPT-4o, Deepseek-V3, Deepseek-R1์— ์ž…๋ ฅํ•˜๊ณ , ์˜๋Œ€์ƒ ๋ฐ ๊ฒฝ๋ ฅ 3ยท6ยท10๋…„์ฐจ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜์˜ ์ง„๋‹จ ์„ฑ๋Šฅ๊ณผ ๋น„๊ตํ•˜์˜€๋‹ค. ๋ชจ๋“  LLM์€ ๋†’์€ ํŠน์ด๋„(0.973โ€“0.983)๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ๋ฏผ๊ฐ๋„๋Š” 3๋…„์ฐจ ์ „๊ณต์˜ ์ˆ˜์ค€๊ณผ ์œ ์‚ฌํ•˜๊ณ  ์˜๋Œ€์ƒ๋ณด๋‹ค ์œ ์˜ํ•˜๊ฒŒ ์šฐ์ˆ˜ํ•˜์˜€์œผ๋‚˜ 10๋…„์ฐจ ์‹œ๋‹ˆ์–ด ์ „๋ฌธ์˜์—๋Š” ๋ฏธ์น˜์ง€ ๋ชปํ•˜์—ฌ, LLM์ด ์‹ฌ์žฅ MRI ํŒ๋… ๊ฒฝํ—˜์ด ๋ถ€์กฑํ•œ ์˜๋ฃŒ ํ™˜๊ฒฝ์—์„œ ์ดˆ๊ธฐ ๋ณด์กฐ ์ง„๋‹จ ๋„๊ตฌ๋กœ ํ™œ์šฉ๋  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-07-05 00:00View โ†—

27Large Language Models for the Differentiation of Benign and Malignant Liver Nodules based on Multimodal Prompts in Liver US Cases.

2026-07Ultrasound in medicine & biologyโญ Q1DOI 10.1016/j.ultrasmedbio.2026.03.010
OBJECTIVE

Large language models (LLMs) that can process both images and text are increasingly being used in radiology. This study aimed to evaluate the performance of LLMs including GPT-4 Omni (GPT-4o), Claude-3.5-Sonnet (Claude), and Gemini 1.5 Pro (Gemini) in differentiating benign and malignant nodules in liver US cases and compare it with that of human readers.

METHODS

Four hundred liver US cases with pathologically confirmed liver nodules visible on B-mode US from January 2020 to November 2024 were randomly selected in this retrospective study. They were divided into a development set (n = 100) and a test set (n = 300). Five prompt groups for LLMs including US image [I-only], image description [D-only], image and description [I+D], image and liver US e-textbook [I+T], and image and medical history [I+H] were evaluated to identify the optimal input in development set. In test set, accuracy of LLMs in differentiating benign and malignant liver nodules was compared with that of human readers using McNemar's test.

RESULTS

In development set, the prompt group I+H for all LLMs exhibited the highest diagnostic accuracy in differentiating benign and malignant liver nodules, being considering as the optimal input (taking GPT-4o as an example, with I-only, 57.0% [as reference]; D-only, 62.0%, p = 0.55; I+D, 62.0%, p = 0.54; I+T, 62.0%, p = 0.36; I+H, 77.0%, p = 0.01). In test set, LLMs with I+H outperformed junior group and showed similar accuracy to senior group (Junior, 70.0% [as reference1]; Senior, 78.3% [as reference2]; GPT-4o, 83.3%, P1 < .001, P2 = .10; Claude, 77.0%, p1 = 0.04, p2 = 0.72; Gemini, 75.3%, p1 = 0.14, p2 = 0.36).

CONCLUSION

Large language models with US image and medical history inputs achieved accuracy comparable to senior radiologists and superior to junior radiologists in differentiating benign and malignant liver nodules.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” GPT-4o, Claude-3.5-Sonnet, Gemini 1.5 Pro ๋“ฑ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ๊ฐ„ ์ดˆ์ŒํŒŒ์—์„œ ์–‘์„ฑ ๋ฐ ์•…์„ฑ ๊ฒฐ์ ˆ ๊ฐ๋ณ„ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜๊ณ  ์ธ๊ฐ„ ํŒ๋…์ž์™€ ๋น„๊ตํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ๋ณ‘๋ฆฌํ•™์ ์œผ๋กœ ํ™•์ธ๋œ ๊ฐ„ ๊ฒฐ์ ˆ 400์˜ˆ(๊ฐœ๋ฐœ ์„ธํŠธ 100์˜ˆ, ํ…Œ์ŠคํŠธ ์„ธํŠธ 300์˜ˆ)๋ฅผ ๋Œ€์ƒ์œผ๋กœ ์ดˆ์ŒํŒŒ ์˜์ƒ ๋‹จ๋…, ์˜์ƒ ์„ค๋ช…๋ฌธ, ๊ต๊ณผ์„œ, ๋ณ‘๋ ฅ ๋“ฑ ๋‹ค์„ฏ ๊ฐ€์ง€ ํ”„๋กฌํ”„ํŠธ ์กฐํ•ฉ์„ ๋น„๊ตํ•œ ๊ฒฐ๊ณผ, ์ดˆ์ŒํŒŒ ์˜์ƒ๊ณผ ์ž„์ƒ ๋ณ‘๋ ฅ์„ ํ•จ๊ป˜ ์ž…๋ ฅ(I+H)ํ•˜๋Š” ๋ฐฉ์‹์ด ์ตœ์ ์˜ ์ง„๋‹จ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€๋‹ค. ํ…Œ์ŠคํŠธ ์„ธํŠธ์—์„œ I+H ์กฐํ•ฉ์„ ์‚ฌ์šฉํ•œ LLM๋“ค์€ ์ „๊ณต์˜ ์ˆ˜์ค€์˜ ์ดˆ๊ธ‰ ํŒ๋…์ž๋ณด๋‹ค ์œ ์˜ํ•˜๊ฒŒ ์šฐ์ˆ˜ํ•˜์˜€์œผ๋ฉฐ(GPT-4o 83.3% vs. ์ดˆ๊ธ‰ 70.0%, p<0.001), ์ „๋ฌธ์˜ ์ˆ˜์ค€์˜ ๊ณ ๊ธ‰ ํŒ๋…์ž์™€ ์œ ์‚ฌํ•œ ์ •ํ™•๋„(๊ณ ๊ธ‰ 78.3%, p=0.10)๋ฅผ ๋‹ฌ์„ฑํ•˜์˜€๋‹ค.
Added: 2026-07-05 00:00View โ†—

28Are LLM-generated plain language summaries truly understandable? A large-scale crowdsourced evaluation.

2026-07Journal of biomedical informaticsโญ Q1DOI 10.1016/j.jbi.2026.105038

Plain language summaries (PLSs) are essential for facilitating effective communication between clinicians and patients by making complex medical information easier for laypeople to understand and act upon. Large language models (LLMs) have recently shown promise in automating PLS generation, but their effectiveness in supporting health information comprehension remains unclear. Prior evaluations have generally relied on automated scores that do not measure understandability directly, or subjective ratings from convenience samples with limited generalizability. To address these gaps, we conducted a large-scale crowdsourced evaluation of LLM-generated PLSs using Amazon Mechanical Turk with 150 participants. We assessed PLS quality through subjective perceived ratings of simplicity, informativeness, coherence, and faithfulness; and task-based measures of reader performance, including multiple-choice accuracy as an operational proxy for comprehension and recall as a complementary measure of gist-level retention. Additionally, we examined the alignment between 10 automated evaluation metrics and human judgments. Our results show that participants often rated LLM-generated PLSs as similarly clear and coherent as human-authored summaries, but participants performed significantly better on the comprehension questions after reading human-written PLSs. This divergence between perceived quality and actual understanding suggests that fluent, trustworthy-sounding AI-generated summaries may engender confidence without reliably supporting comprehension. Furthermore, automated evaluation metrics fail to reflect human judgment, calling into question their suitability for evaluating PLSs. This is the first study to systematically evaluate LLM-generated PLSs based on both reader preferences and comprehension outcomes. Our findings highlight the need for evaluation frameworks that move beyond surface-level quality and for generation methods that explicitly optimize for layperson comprehension, as fluent AI-generated summaries may make readers feel confident without truly understanding the underlying health information.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์ƒ์„ฑํ•œ ์ผ๋ฐ˜์ธ์šฉ ์š”์•ฝ๋ฌธ(Plain Language Summary, PLS)์ด ์‹ค์ œ๋กœ ํ™˜์ž์˜ ๊ฑด๊ฐ• ์ •๋ณด ์ดํ•ด๋ฅผ ๋•๋Š”์ง€ ํ‰๊ฐ€ํ•˜๊ณ ์ž, Amazon Mechanical Turk๋ฅผ ํ†ตํ•ด 150๋ช…์˜ ์ฐธ์—ฌ์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ ํฌ๋ผ์šฐ๋“œ์†Œ์‹ฑ ๊ธฐ๋ฐ˜์˜ ๋Œ€๊ทœ๋ชจ ํ‰๊ฐ€๋ฅผ ์ˆ˜ํ–‰ํ•˜์˜€๋‹ค. ์ฃผ๊ด€์  ํ’ˆ์งˆ ํ‰๊ฐ€(๋‹จ์ˆœ์„ฑ, ์ •๋ณด์„ฑ, ์ผ๊ด€์„ฑ, ์ถฉ์‹ค๋„)์™€ ํ•จ๊ป˜ ๊ฐ๊ด€์‹ ๋ฌธํ•ญ์„ ํ™œ์šฉํ•œ ์ดํ•ด๋„ ๋ฐ ํ•ต์‹ฌ ๋‚ด์šฉ ํšŒ์ƒ ๊ณผ์ œ๋ฅผ ํ†ตํ•ด ๋…์ž ์ˆ˜ํ–‰ ๋Šฅ๋ ฅ์„ ์ธก์ •ํ•˜์˜€์œผ๋ฉฐ, 10๊ฐ€์ง€ ์ž๋™ํ™” ํ‰๊ฐ€ ์ง€ํ‘œ์™€ ์ธ๊ฐ„ ํ‰๊ฐ€ ๊ฐ„์˜ ์ผ์น˜๋„๋„ ๋ถ„์„ํ•˜์˜€๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ์ฐธ์—ฌ์ž๋“ค์€ LLM ์ƒ์„ฑ ์š”์•ฝ๋ฌธ์„ ์ธ๊ฐ„ ์ž‘์„ฑ ์š”์•ฝ๋ฌธ๊ณผ ์œ ์‚ฌํ•˜๊ฒŒ ๋ช…ํ™•ํ•˜๊ณ  ์ผ๊ด€์„ฑ ์žˆ๋‹ค๊ณ  ํ‰๊ฐ€ํ•˜์˜€์œผ๋‚˜, ์‹ค์ œ ์ดํ•ด๋„ ์ธก์ •์—์„œ๋Š” ์ธ๊ฐ„ ์ž‘์„ฑ ์š”์•ฝ๋ฌธ์„ ์ฝ์€ ๊ฒฝ์šฐ ์œ ์˜๋ฏธํ•˜๊ฒŒ ๋†’์€ ์„ฑ์ทจ๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ์ž๋™ํ™” ํ‰๊ฐ€ ์ง€ํ‘œ ์—ญ์‹œ ์ธ๊ฐ„์˜ ํŒ๋‹จ์„ ์ œ๋Œ€๋กœ ๋ฐ˜์˜ํ•˜์ง€ ๋ชปํ•˜๋Š” ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚˜, ์œ ์ฐฝํ•œ AI ์ƒ์„ฑ ์š”์•ฝ๋ฌธ์ด ์‹ค์งˆ์  ์ดํ•ด ์—†์ด ์‹ ๋ขฐ๊ฐ๋งŒ์„ ์œ ๋ฐœํ•  ์ˆ˜ ์žˆ๋‹ค๋Š” ์ ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-05-03 00:00View โ†—

29Automated identification of incidentalomas requiring follow-up: A multi-anatomy evaluation of LLM-based and supervised approaches.

2026-07Journal of biomedical informaticsโญ Q1DOI 10.1016/j.jbi.2026.105048
OBJECTIVE

To evaluate large language models (LLMs) against supervised baselines for fine-grained, lesion-level detection of incidentalomas requiring follow-up, addressing the limitations of current document-level classification systems.

METHODS

We utilized a dataset of 400 annotated radiology reports containing 1623 verified lesion findings. We compared two supervised transformer-based encoders (BioClinicalModernBERT, ModernBERT) against four generative LLM configurations (Llama 3.1-8B, Fine-tuned Llama 3.1-8b, GPT-4o, GPT-OSS-20B). We introduced a novel inference strategy using lesion-tagged inputs and anatomy-aware prompting to ground model reasoning. Performance was evaluated using class-specific F1-scores.

RESULTS

The anatomy-informed GPT-OSS-20B model achieved the highest performance, yielding an incidentaloma-positive macro-F1 of 0.79. This surpassed all supervised baselines (maximum macro-F1: 0.70) and closely matched the inter-annotator agreement of 0.76. Explicit anatomical grounding yielded statistically significant performance gains across GPT-based models (p<0.05), while a majority-vote ensemble of the top systems further improved the macro-F1 to 0.90. Error analysis revealed that anatomy-aware LLMs demonstrated superior contextual reasoning in distinguishing actionable findings from benign lesions.

CONCLUSION

Generative LLMs, when enhanced with structured lesion tagging and anatomical context, significantly outperform traditional supervised encoders and achieve performance comparable to human experts. This approach offers a reliable, interpretable pathway for automated incidental finding surveillance in radiology workflows.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์—์„œ ์ถ”์  ๊ด€์ฐฐ์ด ํ•„์š”ํ•œ ์šฐ์—ฐ ๋ฐœ๊ฒฌ ๋ณ‘๋ณ€(incidentaloma)์„ ๋ณ‘๋ณ€ ์ˆ˜์ค€์—์„œ ์ž๋™ ํƒ์ง€ํ•˜๊ธฐ ์œ„ํ•ด, ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)๊ณผ ์ง€๋„ ํ•™์Šต ๊ธฐ๋ฐ˜ ํŠธ๋žœ์Šคํฌ๋จธ ์ธ์ฝ”๋”์˜ ์„ฑ๋Šฅ์„ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 1,623๊ฐœ์˜ ๋ณ‘๋ณ€ ์†Œ๊ฒฌ์ด ์ฃผ์„๋œ 400๊ฑด์˜ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ๋ฅผ ํ™œ์šฉํ•˜์—ฌ, ๋ณ‘๋ณ€ ํƒœ๊น… ์ž…๋ ฅ ๋ฐ ํ•ด๋ถ€ํ•™์  ๋งฅ๋ฝ ํ”„๋กฌํ”„ํŒ… ์ „๋žต์„ ์ ์šฉํ•œ GPT-OSS-20B ๋ชจ๋ธ๊ณผ BioClinicalModernBERT ๋“ฑ ์ง€๋„ ํ•™์Šต ๋ชจ๋ธ์„ ๋น„๊ตํ•˜์˜€๋‹ค. ํ•ด๋ถ€ํ•™ ์ •๋ณด๊ฐ€ ๊ฐ•ํ™”๋œ GPT-OSS-20B ๋ชจ๋ธ์€ incidentaloma ์–‘์„ฑ macro-F1 0.79๋ฅผ ๋‹ฌ์„ฑํ•˜์—ฌ ์ง€๋„ ํ•™์Šต ์ตœ๊ณ  ์„ฑ๋Šฅ(0.70)์„ ์œ ์˜๋ฏธํ•˜๊ฒŒ ์ƒํšŒํ•˜์˜€์œผ๋ฉฐ, ์ฃผ์„์ž ๊ฐ„ ์ผ์น˜๋„(0.76)์— ๊ทผ์ ‘ํ•˜์˜€๊ณ , ์•™์ƒ๋ธ” ์ ์šฉ ์‹œ macro-F1 0.90๊นŒ์ง€ ํ–ฅ์ƒ๋˜์–ด ๋ฐฉ์‚ฌ์„  ํŒ๋… ์›Œํฌํ”Œ๋กœ์šฐ์—์„œ์˜ ์ž„์ƒ์  ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-05-03 00:00View โ†—

30Benchmarking large language models for quality control of chest radiographs and CT reports: a retrospective multimodal study.

2026-07European journal of radiologyโญ Q1DOI 10.1016/j.ejrad.2026.112840
OBJECTIVE

This study aims to establish a retrospective, single-centre, feasibility-oriented benchmark for medical imaging quality control(QC) and to evaluate the potential of multiple large language models for chest X-ray radiograph(CXR) technical QC and CT report consistency assessment, based on a relatively small, radiologist-annotated dataset derived from routine clinical practice.

METHODS

This retrospective, single-centre study included 161 CXRs and 219 structured CT reports from routine clinical practice. Twelve labels were used for CXR QC, including eleven radiologist-defined error categories and one error-free label, while nine labels were used for CT report evaluation, including eight inconsistency categories and one error-free label. All cases were annotated using a radiologist consensus reference standard. Multiple large language models(LLMs) and multimodal large language models(MLLMs) were evaluated using Micro-F1 and Macro-F1 for CXR QC and expert-based Micro-F1 for CT report QC.

RESULTS

For CXR QC, Gemini 2.0 Flash showed the strongest performance, achieving robust category-level generalization, while GPT-4o and Qwen2.5-VL-72B-Instruct demonstrated more balanced but weaker performance. In CT report QC, DeepSeek-R1 achieved the highest recall (62.23%) and the best overall performance. Across models, protocol-report mismatches and metric inconsistencies were the most common error types.

CONCLUSION

This study presents an initial, feasibility-oriented multimodal benchmark for medical imaging quality control, showing that LLM's performance is highly task- and modality-dependent. Given the limited sample size, single-centre design, and Chinese-language scope, the findings support future multicentral validation and workflow-integrated evaluation, rather than immediate clinical deployment.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ํ‰๋ถ€ X์„ (CXR) ๊ธฐ์ˆ ์  ํ’ˆ์งˆ๊ด€๋ฆฌ ๋ฐ CT ๋ณด๊ณ ์„œ ์ผ๊ด€์„ฑ ํ‰๊ฐ€์— ๋Œ€ํ•œ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด, ๋‹จ์ผ ๊ธฐ๊ด€์—์„œ ์ˆ˜์ง‘ํ•œ CXR 161๊ฑด ๋ฐ ๊ตฌ์กฐํ™” CT ๋ณด๊ณ ์„œ 219๊ฑด์„ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜ ํ•ฉ์˜ ๊ธฐ์ค€์œผ๋กœ ์ฃผ์„ ์ฒ˜๋ฆฌํ•œ ์†Œ๊ทœ๋ชจ ํšŒ๊ณ ์  ๋ฒค์น˜๋งˆํฌ๋ฅผ ๊ตฌ์ถ•ํ•˜์˜€๋‹ค. Micro-F1 ๋ฐ Macro-F1 ์ง€ํ‘œ๋ฅผ ํ™œ์šฉํ•˜์—ฌ ๋‹ค์ˆ˜์˜ LLM ๋ฐ ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ LLM์„ ํ‰๊ฐ€ํ•œ ๊ฒฐ๊ณผ, CXR ํ’ˆ์งˆ๊ด€๋ฆฌ์—์„œ๋Š” Gemini 2.0 Flash๊ฐ€ ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ๋ฒ”์ฃผ ์ˆ˜์ค€์˜ ์ผ๋ฐ˜ํ™” ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, CT ๋ณด๊ณ ์„œ ํ’ˆ์งˆ๊ด€๋ฆฌ์—์„œ๋Š” DeepSeek-R1์ด ์ตœ๊ณ  ์žฌํ˜„์œจ(62.23%) ๋ฐ ์ „๋ฐ˜์ ์œผ๋กœ ๊ฐ€์žฅ ๋†’์€ ์„ฑ๋Šฅ์„ ๋‹ฌ์„ฑํ•˜์˜€๋‹ค. ๋ณธ ์—ฐ๊ตฌ๋Š” LLM์˜ ์˜๋ฃŒ์˜์ƒ ํ’ˆ์งˆ๊ด€๋ฆฌ ์„ฑ๋Šฅ์ด ๊ณผ์ œ ์œ ํ˜• ๋ฐ ๋ฐ์ดํ„ฐ ์–‘์‹์— ๋”ฐ๋ผ ํฌ๊ฒŒ ๋‹ฌ๋ผ์ง์„ ์‹œ์‚ฌํ•˜๋ฉฐ, ์†Œ๊ทœ๋ชจ ๋‹จ์ผ ๊ธฐ๊ด€ยท์ค‘๊ตญ์–ด ๋ฐ์ดํ„ฐ์˜ ํ•œ๊ณ„๋ฅผ ๊ณ ๋ คํ•  ๋•Œ ์ฆ‰๊ฐ์ ์ธ ์ž„์ƒ ์ ์šฉ๋ณด๋‹ค๋Š” ๋‹ค๊ธฐ๊ด€ ๊ฒ€์ฆ ๋ฐ ์›Œํฌํ”Œ๋กœ์šฐ ํ†ตํ•ฉ ํ‰๊ฐ€๊ฐ€ ํ•„์š”ํ•จ์„ ๊ฐ•์กฐํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

31Real-world text-only inference of PI-RADS v2.1 from prostate MRI reports using large language models: a lesion-level, zone-aware study.

2026-07European journal of radiologyโญ Q1DOI 10.1016/j.ejrad.2026.112838
OBJECTIVE

To evaluate the feasibility and limitations of real-world, text-only inference of PI-RADS v2.1 categories from prostate MRI reports using large language models, with lesion-level and zone-aware analysis.

METHODS

This single-center retrospective study included 1,205 lesion-level entries from 1,118 patients derived from semi-structured prostate MRI reports after removal of all explicit PI-RADS elements. ChatGPT-4o was prompted to assign numeric PI-RADS categories based solely on report text. Agreement with radiologist-assigned reference categories was assessed using exact agreement, Cohen's ฮบ, and class-wise metrics. Analyses were performed overall, by zone (peripheral vs transition), and using collapsed risk strata (1-2/3/4-5). Discordant cases were reviewed to identify error mechanisms and severity. Human interobserver agreement, intra-model reproducibility, temporal stability, and a paired model-version sensitivity analysis comparing ChatGPT-4o with GPT-5.2 were also evaluated.

RESULTS

Overall exact agreement was 72.9% (ฮบย =ย 0.538; macro-F1ย =ย 61.2%), with a systematic tendency toward overcalling. Agreement was higher in the peripheral zone than in the transition zone (ฮบย =ย 0.476 vs 0.077, reference PI-RADS 3-5). PI-RADS 3 showed the lowest precision and recall, with frequent bidirectional misclassification. Collapsing categories improved agreement (ฮบย =ย 0.610). Incorrect diffusion-weighted imaging subscores were the most common error mechanism, with zone-specific differences. Clinically high-impact downgrades of PI-RADS 4-5 to 1-2 were rare (1.6%). Human interobserver agreement was excellent (ฮบย =ย 0.916-0.967). GPT-5.2 outperformed ChatGPT-4o in paired analyses but produced invalid outputs in a minority of cases.

CONCLUSION

Text-only large language models can infer radiologist-assigned PI-RADS v2.1 categories from real-world prostate MRI reports with moderate agreement, but performance is zone dependent and limited around PI-RADS 3, particularly in the transition zone. These models are best suited as supervised tools for quality control rather than autonomous decision-making.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์ „๋ฆฝ์„  MRI ๋ณด๊ณ ์„œ ํ…์ŠคํŠธ๋งŒ์„ ์ด์šฉํ•˜์—ฌ PI-RADS v2.1 ๋ฒ”์ฃผ๋ฅผ ์ž๋™์œผ๋กœ ์ถ”๋ก ํ•˜๋Š” ๊ฒƒ์˜ ๊ฐ€๋Šฅ์„ฑ๊ณผ ํ•œ๊ณ„๋ฅผ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 1,118๋ช… ํ™˜์ž์˜ 1,205๊ฐœ ๋ณ‘๋ณ€ ๋ฐ์ดํ„ฐ๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ ChatGPT-4o์— ๋ช…์‹œ์  PI-RADS ํ•ญ๋ชฉ์„ ์ œ๊ฑฐํ•œ ๋ฐ˜๊ตฌ์กฐํ™” ๋ณด๊ณ ์„œ๋ฅผ ์ž…๋ ฅํ•˜์—ฌ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ ํŒ๋… ๊ฒฐ๊ณผ์™€์˜ ์ผ์น˜๋„๋ฅผ ๋ณ‘๋ณ€ ์ˆ˜์ค€ ๋ฐ ๊ตฌ์—ญ๋ณ„๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. ์ „์ฒด ์ •ํ™• ์ผ์น˜์œจ์€ 72.9%(ฮบ=0.538)๋กœ ์ค‘๋“ฑ๋„ ์ˆ˜์ค€์ด์—ˆ์œผ๋ฉฐ, ๋ง์ดˆ ๊ตฌ์—ญ์—์„œ์˜ ์„ฑ๋Šฅ(ฮบ=0.476)์ด ์ดํ–‰ ๊ตฌ์—ญ(ฮบ=0.077)๋ณด๋‹ค ํ˜„์ €ํžˆ ์šฐ์ˆ˜ํ•˜์˜€๊ณ  PI-RADS 3 ๋ฒ”์ฃผ์—์„œ ๊ฐ€์žฅ ๋‚ฎ์€ ์ •ํ™•๋„๋ฅผ ๋ณด์—ฌ, ์ด ๋ชจ๋ธ์€ ์ž์œจ์  ํŒ๋‹จ๋ณด๋‹ค๋Š” ๊ฐ๋… ํ•˜ ํ’ˆ์งˆ ๊ด€๋ฆฌ ๋„๊ตฌ๋กœ ํ™œ์šฉํ•˜๋Š” ๊ฒƒ์ด ์ ํ•ฉํ•˜๋‹ค๊ณ  ๊ฒฐ๋ก ์ง€์—ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

32Ultrasound-based Detection and Malignancy Prediction of Breast Lesions Eligible for Biopsy: A Multi-center Clinical-scenario Study Using Nomograms, Large Language Models, and Radiologist Evaluation.

2026-07Academic radiologyโญ Q1DOI 10.1016/j.acra.2026.03.009

RATIONALE AND

OBJECTIVE

To develop and externally validate ultrasound nomograms combining BI-RADS features and quantitative morphometric characteristics, and to compare their performance with expert radiologists and large language models in biopsy recommendation and malignancy prediction for breast lesions.

METHODS

In this multi-center, multi-national study, 1747 women with breast lesions underwent ultrasound across three centers in Iran and Turkey. A total of 10 BIRADS and 26 morphological features were extracted from each lesion. Three nomograms based on BI-RADS, morphometric, and both feature sets were constructed. Three radiologists (one senior, two general) and two ChatGPTs including ChatGPT-o3 and o4-mini-high interpreted de-identified breast lesion images. Diagnostic performance for biopsy recommendation and malignancy prediction was assessed across all cohorts.

RESULTS

According to the pooled results, although the difference between the fused nomogram and the BI-RADS version was not statistically significant, the fused version consistently outperformed all models in biopsy recommendation and malignancy prediction (AUCs of 0.901 and 0.853, respectively) compared to BI-RADS nomogram (AUCs of 0.898 and 0.834), morphometric nomogram (AUCs of 0.825 and 0.708), radiologist1 (AUCs of 0.820 and 0.729), radiologist2 (AUCs of 0.605 and 0.719), radiologist3 (AUCs of 0.728 and 0.699), ChatGPT-o3 (AUCs of 0.729 and 0.689), and o4-mini-high (AUCs of 0.713 and 0.695).

CONCLUSION

The proposed BI-RADS-morphometric nomogram outperforms standalone nomogram models, LLMs, and radiologists in guiding biopsy decisions and predicting malignancy. The proposed novel fused nomogram has the potential to reduce unnecessary biopsies and enhance personalized decision-making in breast imaging.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
์ด ๋‹ค๊ธฐ๊ด€ ๋‹ค๊ตญ๊ฐ€ ์—ฐ๊ตฌ๋Š” ์ดˆ์ŒํŒŒ ๊ธฐ๋ฐ˜ BI-RADS ํŠน์ง•๊ณผ ์ •๋Ÿ‰์  ํ˜•ํƒœ๊ณ„์ธก ๋ณ€์ˆ˜๋ฅผ ๊ฒฐํ•ฉํ•œ ๋…ธ๋ชจ๊ทธ๋žจ์„ ๊ฐœ๋ฐœยท์™ธ๋ถ€ ๊ฒ€์ฆํ•˜๊ณ , ์œ ๋ฐฉ ๋ณ‘๋ณ€์˜ ์กฐ์ง๊ฒ€์‚ฌ ๊ถŒ๊ณ  ๋ฐ ์•…์„ฑ ์˜ˆ์ธก ์„ฑ๋Šฅ์„ ์ „๋ฌธ ์˜์ƒ์˜ํ•™๊ณผ ์˜์‚ฌ ๋ฐ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(ChatGPT-o3, o4-mini-high)๊ณผ ๋น„๊ตํ•˜์˜€๋‹ค. ์ด๋ž€ยทํ„ฐํ‚ค 3๊ฐœ ๊ธฐ๊ด€์—์„œ ์œ ๋ฐฉ ๋ณ‘๋ณ€์„ ๊ฐ€์ง„ ์—ฌ์„ฑ 1,747๋ช…์„ ๋Œ€์ƒ์œผ๋กœ 10๊ฐœ์˜ BI-RADS ๋ฐ 26๊ฐœ์˜ ํ˜•ํƒœํ•™์  ํŠน์ง•์„ ์ถ”์ถœํ•˜์—ฌ ์„ธ ๊ฐ€์ง€ ๋…ธ๋ชจ๊ทธ๋žจ์„ ๊ตฌ์ถ•ํ•˜๊ณ , 3๋ช…์˜ ์˜์ƒ์˜ํ•™๊ณผ ์˜์‚ฌ ๋ฐ ๋‘ ์ข…๋ฅ˜์˜ ChatGPT ๋ชจ๋ธ๊ณผ ์ง„๋‹จ ์„ฑ๋Šฅ์„ ๋น„๊ตํ•˜์˜€๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, BI-RADS์™€ ํ˜•ํƒœ๊ณ„์ธก ๋ณ€์ˆ˜๋ฅผ ์œตํ•ฉํ•œ ๋…ธ๋ชจ๊ทธ๋žจ์ด ์กฐ์ง๊ฒ€์‚ฌ ๊ถŒ๊ณ (AUC 0.901) ๋ฐ ์•…์„ฑ ์˜ˆ์ธก(AUC 0.853) ๋ชจ๋‘์—์„œ ๋‹จ๋… ๋…ธ๋ชจ๊ทธ๋žจ, ์˜์ƒ์˜ํ•™๊ณผ ์˜์‚ฌ, LLM ๋ชจ๋ธ์„ ๋Šฅ๊ฐ€ํ•˜์˜€์œผ๋ฉฐ, ๋ถˆํ•„์š”ํ•œ ์กฐ์ง๊ฒ€์‚ฌ ๊ฐ์†Œ์™€ ๊ฐœ์ธ ๋งž์ถคํ˜• ์œ ๋ฐฉ ์˜์ƒ ์˜์‚ฌ๊ฒฐ์ • ํ–ฅ์ƒ์— ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

33Radiologist-Large Language Model Collaboration in Dermatologic Ultrasound Reporting: Evaluating the Clinical Utility of ChatGPT.

2026-06Studies in health technology and informaticsDOI 10.3233/shti260814

Large language models (LLMs) are increasingly explored in medical imaging, but their reliability in independently interpreting images remains uncertain. This study evaluated the clinical utility of radiology reports generated under three reporting conditions using dermatologic ultrasound images: Condition 1 (radiologist's reporting), Condition 2 (LLM reporting based solely on the ultrasound image), and Condition 3 (LLM reporting using both the ultrasound image and the radiologist's report). A total of 202 dermatologic ultrasound images from a public dataset were analyzed. Reports were evaluated for diagnostic accuracy, appropriateness of next-step recommendations, and readability. Diagnostic accuracy was highest in Condition 3 (83.2%), compared with Condition 1 (55.4%) and Condition 2 (26.2%) (p<0.001). Next-step suggestion accuracy was also highest in Condition 3 (77.7%), followed by Condition 2 (59.9%) and Condition 1 (38.6%) (p<0.001). Report readability was also highest in Condition 3. Integrating LLM outputs with radiologist reports may improve clinical communication and decision support in dermatologic ultrasound.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ํ”ผ๋ถ€๊ณผ์  ์ดˆ์ŒํŒŒ ํŒ๋…์—์„œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์ž„์ƒ์  ์œ ์šฉ์„ฑ์„ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด 202๊ฐœ์˜ ์ดˆ์ŒํŒŒ ์˜์ƒ์„ ๋Œ€์ƒ์œผ๋กœ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ ๋‹จ๋… ํŒ๋…, LLM ๋‹จ๋… ํŒ๋…, ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ ๋ณด๊ณ ์„œ์™€ LLM์„ ๊ฒฐํ•ฉํ•œ ์„ธ ๊ฐ€์ง€ ์กฐ๊ฑด์„ ๋น„๊ตํ•˜์˜€๋‹ค. ์ง„๋‹จ ์ •ํ™•๋„๋Š” ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ ๋ณด๊ณ ์„œ์™€ LLM์„ ํ•จ๊ป˜ ํ™œ์šฉํ•œ ์กฐ๊ฑด 3์—์„œ 83.2%๋กœ ๊ฐ€์žฅ ๋†’์•˜์œผ๋ฉฐ, ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ ๋‹จ๋…(55.4%) ๋ฐ LLM ๋‹จ๋…(26.2%)๋ณด๋‹ค ์œ ์˜ํ•˜๊ฒŒ ์šฐ์ˆ˜ํ•˜์˜€๋‹ค(p<0.001). ์ด๋Ÿฌํ•œ ๊ฒฐ๊ณผ๋Š” LLM์ด ๋…๋ฆฝ์  ์˜์ƒ ํŒ๋…์—๋Š” ํ•œ๊ณ„๊ฐ€ ์žˆ์œผ๋‚˜, ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ์˜ ๋ณด๊ณ ์„œ์™€ ํ†ตํ•ฉ๋  ๊ฒฝ์šฐ ํ”ผ๋ถ€๊ณผ์  ์ดˆ์ŒํŒŒ์˜ ์ž„์ƒ์  ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ๋ฐ ์ง„๋ฃŒ ์ปค๋ฎค๋‹ˆ์ผ€์ด์…˜ ํ–ฅ์ƒ์— ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-07-05 00:00View โ†—

34JRadiEvo: A Japanese radiology report generation model enhanced by evolutionary optimization of model merging.

2026-06Artificial intelligence in medicineโญ Q1DOI 10.1016/j.artmed.2026.103482

Radiology report generation is an important application of artificial intelligence (AI), as the interpretation of medical images and the production of clinically relevant reports are time-consuming and cognitively demanding tasks, especially in high-volume settings such as chest X-ray screening. Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled substantial progress in automated medical report generation. However, most existing medical foundation models are trained primarily on English datasets, limiting their practicality in non-English-speaking regions such as Japan. Publicly available radiology datasets are overwhelmingly English, while constructing large-scale non-English datasets is costly and difficult because of translation effort and medical data privacy constraints. Existing adaptation methods, such as fine-tuning and continued pre-training, typically require large amounts of in-domain data. This makes them difficult to apply in low-resource medical language settings where large-scale annotated datasets are unavailable. Accordingly, this study asks: how can an accurate Japanese chest X-ray radiology report generator be developed without access to large-scale, curated Japanese medical image-report pairs? To address this challenge, we propose JRadiEvo, a Japanese chest X-ray report generation model built through evolutionary optimization of model merging. By combining pretrained models with complementary strengths in vision-language alignment, medical knowledge, and Japanese generation, JRadiEvo enables data-efficient adaptation without large-scale training. To the best of our knowledge, this is the first attempt to build a non-English medical vision-language model through evolutionary optimization of model merging. Despite using only 50 translated training samples from publicly available data, JRadiEvo outperforms CheXagent, a state-of-the-art model trained on approximately 8.5 million samples, in ROUGE-L and METEOR metrics. These results provide a proof of concept for extreme data-efficient adaptation in low-resource medical languages.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์ผ๋ณธ์–ด ์˜๋ฃŒ ์˜์ƒ-๋ณด๊ณ ์„œ ์Œ ๋ฐ์ดํ„ฐ ์—†์ด ์ผ๋ณธ์–ด ํ‰๋ถ€ X์„  ์˜์ƒ ํŒ๋… ๋ณด๊ณ ์„œ๋ฅผ ์ž๋™ ์ƒ์„ฑํ•˜๋Š” ๋ชจ๋ธ์„ ๊ฐœ๋ฐœํ•˜๋Š” ๊ฒƒ์„ ๋ชฉํ‘œ๋กœ ํ•œ๋‹ค. ์ด๋ฅผ ์œ„ํ•ด ์‹œ๊ฐ-์–ธ์–ด ์ •๋ ฌ, ์˜๋ฃŒ ์ง€์‹, ์ผ๋ณธ์–ด ์ƒ์„ฑ์— ๊ฐ๊ฐ ํŠนํ™”๋œ ์‚ฌ์ „ํ•™์Šต ๋ชจ๋ธ๋“ค์„ ์ง„ํ™”์  ์ตœ์ ํ™” ๊ธฐ๋ฐ˜ ๋ชจ๋ธ ๋ณ‘ํ•ฉ(evolutionary optimization of model merging) ๋ฐฉ์‹์œผ๋กœ ๊ฒฐํ•ฉํ•œ JRadiEvo๋ฅผ ์ œ์•ˆํ•˜์˜€๋‹ค. ๊ณต๊ฐœ ๋ฐ์ดํ„ฐ์—์„œ ๋ฒˆ์—ญ๋œ ๋‹จ 50๊ฐœ์˜ ํ›ˆ๋ จ ์ƒ˜ํ”Œ๋งŒ์„ ์‚ฌ์šฉํ•˜์˜€์Œ์—๋„ ๋ถˆ๊ตฌํ•˜๊ณ , JRadiEvo๋Š” ์•ฝ 850๋งŒ ๊ฐœ์˜ ์ƒ˜ํ”Œ๋กœ ํ›ˆ๋ จ๋œ ์ตœ์‹  ๋ชจ๋ธ์ธ CheXagent๋ฅผ ROUGE-L ๋ฐ METEOR ์ง€ํ‘œ์—์„œ ๋Šฅ๊ฐ€ํ•˜์—ฌ, ์ €์ž์› ์˜๋ฃŒ ์–ธ์–ด ํ™˜๊ฒฝ์—์„œ์˜ ๊ทน๋‹จ์  ๋ฐ์ดํ„ฐ ํšจ์œจ์  ์ ์‘ ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-06-28 00:00View โ†—

35Multimodal Bidirectional Direct Preference Optimization and Instruction Fine-Tuning for Medical Image Understanding and Generation.

2026-06IEEE journal of biomedical and health informaticsโญ Q1DOI 10.1109/jbhi.2026.3707092

Although multimodal large language models (MLLMs) are advancing rapidly in general vision-language tasks, their ability to capture the subtle nuances of medical images, especially in radiology, is limited. Current methods primarily apply supervised fine-tuning, which often leads to hallucinated results and renders them untrustworthy in clinical decision support. As a way of overcoming this drawback, we suggest a two step finetuning framework. The first stage uses a VQ-GAN-based visual tokenizer takes medical images and transforms them into discrete tokens, and then aligned with the language token format. Both image and text generation are considered autoregressive tasks, based on text or image inputs. This step conducts a visual-language instructional fine tuning, which allows the model to interpret and follow imaging specific instructions in a variety of imaging modalities. In the second stage, we present an improved Direct Preference Optimization (DPO) method. We deliberately distort images to induce halluci nations and generate dispreferred data, while genuine data serve as the preferred reference. This refined DPO strategy effectively mitigates hallucinations. Experimental results demonstrate that the proposed framework significantly improves the accuracy, faithfulness, and clinical relevance of generated radiology reports, while also enhancing the quality of medical image generation. These findings high light the potential of MLLMs for reliable multimodal clinical assistance, supporting precision diagnosis and advancing trustworthy AI applications in healthcare.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋‹ค์ค‘๋ชจ๋‹ฌ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(MLLM)์ด ๋ฐฉ์‚ฌ์„  ์˜์ƒ ๋“ฑ ์˜๋ฃŒ ์ด๋ฏธ์ง€ ํ•ด์„ ์‹œ ํ™˜๊ฐ(hallucination)์„ ์ƒ์„ฑํ•˜๋Š” ํ•œ๊ณ„๋ฅผ ๊ทน๋ณตํ•˜๊ณ  ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์›์˜ ์‹ ๋ขฐ์„ฑ์„ ๋†’์ด๊ณ ์ž, VQ-GAN ๊ธฐ๋ฐ˜ ์‹œ๊ฐ ํ† ํฌ๋‚˜์ด์ €๋ฅผ ํ™œ์šฉํ•œ ์‹œ๊ฐ-์–ธ์–ด ์ง€์‹œ ๋ฏธ์„ธ์กฐ์ •(1๋‹จ๊ณ„)๊ณผ ์˜๋„์ ์œผ๋กœ ์™œ๊ณก๋œ ์ด๋ฏธ์ง€๋ฅผ ๋น„์„ ํ˜ธ ๋ฐ์ดํ„ฐ๋กœ ํ™œ์šฉํ•˜๋Š” ๊ฐœ์„ ๋œ ์ง์ ‘ ์„ ํ˜ธ ์ตœ์ ํ™”(DPO, 2๋‹จ๊ณ„)๋ฅผ ๊ฒฐํ•ฉํ•œ ์ด์ค‘ ๋‹จ๊ณ„ ๋ฏธ์„ธ์กฐ์ • ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ์ œ์•ˆํ•˜์˜€๋‹ค. ์ œ์•ˆ๋œ ํ”„๋ ˆ์ž„์›Œํฌ๋Š” ํ…์ŠคํŠธ ๋ฐ ์ด๋ฏธ์ง€ ์ž…๋ ฅ ๋ชจ๋‘๋ฅผ ์ž๊ธฐํšŒ๊ท€ ๋ฐฉ์‹์œผ๋กœ ์ฒ˜๋ฆฌํ•˜๋ฉฐ, ์‹ค์ œ ๋ฐ์ดํ„ฐ๋ฅผ ์„ ํ˜ธ ์ฐธ์กฐ๋กœ ์„ค์ •ํ•จ์œผ๋กœ์จ ํ™˜๊ฐ์„ ํšจ๊ณผ์ ์œผ๋กœ ์–ต์ œํ•œ๋‹ค. ์‹คํ—˜ ๊ฒฐ๊ณผ, ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์˜ ์ •ํ™•์„ฑยท์ถฉ์‹ค๋„ยท์ž„์ƒ์  ๊ด€๋ จ์„ฑ์ด ์œ ์˜๋ฏธํ•˜๊ฒŒ ํ–ฅ์ƒ๋˜๊ณ  ์˜๋ฃŒ ์˜์ƒ ์ƒ์„ฑ ํ’ˆ์งˆ๋„ ๊ฐœ์„ ๋˜์–ด, ์ •๋ฐ€ ์ง„๋‹จ์„ ์ง€์›ํ•˜๋Š” ์‹ ๋ขฐ์„ฑ ๋†’์€ ์ž„์ƒ AI ๋„๊ตฌ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์ด ์ž…์ฆ๋˜์—ˆ๋‹ค.
Added: 2026-06-28 00:00View โ†—

36"Enhancing Patient Understanding of Radiology Reports Through LLM-Generated Summaries, Clickable Terms, and AI Videos".

2026-06Journal of the American College of Radiology : JACRโญ Q1DOI 10.1016/j.jacr.2026.06.001
OBJECTIVE

To evaluate how a custom web application integrating clinician-edited large language model (LLM)-generated summaries, clickable definitions, and artificial intelligence (AI)-generated videos affects radiology report comprehension, feature preferences, and overall sentiment toward AI-assisted report summaries.

METHODS

This prospective study recruited participants between May and July 2025 at a hospital-based outpatient imaging floor before their scheduled examinations at a tertiary university hospital. Following exam completion and report publication, patient-friendly AI summaries were generated and reviewed by a radiologist for accuracy. Participants were then shown a web application containing their own de-identified, AI-augmented reports featuring clinician-edited LLM-generated summaries with clickable terms and AI videos. Participants were surveyed on comprehension, feature usefulness, and attitudes toward LLM summaries.

RESULTS

Participants (n=101, 40 male/61 female, racially diverse) ranged from 20 to 82 (mean 58ยฑ15) years old. Overall comprehension improved significantly (median pre: 4.00, post: 5.00, p<0.001), with 47.52% (n=48) identifying LLM-summaries as most helpful. However, LLM-summaries required manual clinician edits (average per summary: 24.75 words removed; 0.13 words added, lexical similarity = 84.63%; semantic similarity = 98.25%). When asked if they were comfortable with LLM-summaries without clinician edits, most participants reported being only Somewhat comfortable (27.72%) or Very uncomfortable (25.74%).

CONCLUSION

This prospective study demonstrates that interactive, LLM-driven applications can significantly improve self-reported patient comprehension of complex radiology reports, emphasizing their potential to enhance patient-centered communication. However, patients had reservations about clinician-edited LLM-generated summaries, indicating that successful integration is contingent on professional oversight - an added workload that may limit scalable real-world implementation.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” LLM ๊ธฐ๋ฐ˜ ์š”์•ฝ๋ฌธ, ํด๋ฆญ ๊ฐ€๋Šฅํ•œ ์šฉ์–ด ์ •์˜, AI ์ƒ์„ฑ ์˜์ƒ์„ ํ†ตํ•ฉํ•œ ์›น ์• ํ”Œ๋ฆฌ์ผ€์ด์…˜์ด ํ™˜์ž์˜ ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ ์ดํ•ด๋„์— ๋ฏธ์น˜๋Š” ์˜ํ–ฅ์„ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด, 101๋ช…์˜ ์™ธ๋ž˜ ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ ์ „ํ–ฅ์ ์œผ๋กœ ์ˆ˜ํ–‰๋˜์—ˆ๋‹ค. ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜๊ฐ€ ๊ฒ€ํ† ยทํŽธ์ง‘ํ•œ AI ์š”์•ฝ๋ณธ์„ ํฌํ•จํ•œ ์›น ์• ํ”Œ๋ฆฌ์ผ€์ด์…˜ ์‚ฌ์šฉ ์ „ํ›„ ์ดํ•ด๋„๋ฅผ ๋น„๊ตํ•œ ๊ฒฐ๊ณผ, ์ดํ•ด๋„๊ฐ€ ์œ ์˜ํ•˜๊ฒŒ ํ–ฅ์ƒ๋˜์—ˆ์œผ๋ฉฐ(์ค‘์•™๊ฐ’ 4.00 โ†’ 5.00, p<0.001), ์ฐธ๊ฐ€์ž์˜ 47.52%๊ฐ€ LLM ์š”์•ฝ๋ฌธ์„ ๊ฐ€์žฅ ์œ ์šฉํ•œ ๊ธฐ๋Šฅ์œผ๋กœ ์„ ํƒํ•˜์˜€๋‹ค. ๊ทธ๋Ÿฌ๋‚˜ LLM ์ƒ์„ฑ ์š”์•ฝ๋ฌธ์€ ํ‰๊ท  24.75๊ฐœ ๋‹จ์–ด์˜ ์ˆ˜๋™ ์ˆ˜์ •์ด ํ•„์š”ํ•˜์˜€๊ณ , ์ž„์ƒ์˜ ๊ฒ€ํ†  ์—†์ด AI ์š”์•ฝ๋งŒ์„ ์‹ ๋ขฐํ•˜๋Š” ๊ฒƒ์— ๋Œ€ํ•ด ๋‹ค์ˆ˜์˜ ํ™˜์ž๊ฐ€ ๋ถˆํŽธํ•จ์„ ํ‘œ๋ช…ํ•˜์—ฌ, ์‹ค์ œ ์ž„์ƒ ํ˜„์žฅ์—์„œ์˜ ํ™•์žฅ์  ๋„์ž…์—๋Š” ์ „๋ฌธ๊ฐ€ ๊ฐ๋… ์ฒด๊ณ„ ๊ตฌ์ถ•์ด ์„ ๊ฒฐ ๊ณผ์ œ์ž„์ด ์‹œ์‚ฌ๋˜์—ˆ๋‹ค.
Added: 2026-06-14 00:00View โ†—

37Extraction of distant recurrence sites for breast cancer patients from free-text clinical notes using large language models.

2026-06Journal of biomedical informaticsโญ Q1DOI 10.1016/j.jbi.2026.105032
OBJECTIVE

Accurate documentation of distant recurrence sites in breast cancer is essential for evaluating treatment effectiveness and outcomes research. However, such information is embedded in unstructured clinical notes, making manual abstraction labor-intensive. Large language models (LLMs) offer a scalable solution for extracting complex information from heterogeneous clinical narratives; however, generic LLMs often lack the specialized clinical reasoning needed for accurate interpretation of oncologic documentation. This study aims to develop an efficient LLM-based framework to automatically extract distant recurrence sites from free-text documentation. MATERIALS &

METHODS

We used clinical notes, pathology and radiology reports from recurrent breast cancer patients at Mayo Clinic (nย =ย 766) for model development and evaluated generalizability on internal hold-out samples (nย =ย 112) and an external Stanford Medicine cohort (nย =ย 110). For cross-disease domain adaptation, we further validated on prostate cancer patients (nย =ย 49). Our proposed framework employs BioLinkBERT, a pretrained language model (PLM) backbone, with weak supervision and an epoch-wise entropy optimization to address limited labeled data and class imbalance across recurrence sites. The fine-tuned model was compared against state-of-the-art models, including Llama2-7B, Llama-3-8B and MedAlpaca, using precision, recall, and F1-score.

RESULTS

The fine-tuned model outperformed generic and domain-specific LLM baselines, with notable gains in identifying multi-site distant recurrence. In-domain validation showed consistent F1-score improvement (average 0.78), particularly for rare recurrence sites. The model also demonstrated strong performance on the external Stanford cohort and on prostate cancer, achieving F1-score of 0.83 and 0.93, respectively.

CONCLUSION

This study presents an efficient, weakly supervised LLM framework that accurately extracts metastatic recurrence sites, reducing reliance on manual chart review. The results demonstrate that relatively small LLMs, optimized with domain-aware weak supervision, can outperform larger models for complex oncologic information extraction. The model is released as a platform-independent Docker image to support seamless cancer registry integration.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์œ ๋ฐฉ์•” ํ™˜์ž์˜ ๋น„์ •ํ˜• ์ž„์ƒ ๊ธฐ๋ก์—์„œ ์›๊ฒฉ ์ „์ด ๋ถ€์œ„๋ฅผ ์ž๋™์œผ๋กœ ์ถ”์ถœํ•˜๊ธฐ ์œ„ํ•ด BioLinkBERT ๊ธฐ๋ฐ˜์˜ ์•ฝ์ง€๋„ ํ•™์Šต(weakly supervised) LLM ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ๊ฐœ๋ฐœํ•˜์˜€์Šต๋‹ˆ๋‹ค. ํ•ด๋‹น ๋ชจ๋ธ์€ ๋‚ด๋ถ€ ๋ฐ ์™ธ๋ถ€ ๊ฒ€์ฆ ๋ฐ์ดํ„ฐ์…‹์—์„œ ๊ธฐ์กด LLM ๋Œ€๋น„ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ํŠนํžˆ ๋‹ค๋ฐœ์„ฑ ๋ฐ ํฌ๊ท€ ์ „์ด ๋ถ€์œ„ ์‹๋ณ„์—์„œ ๋†’์€ ์ •ํ™•๋„๋ฅผ ์ž…์ฆํ–ˆ์Šต๋‹ˆ๋‹ค. ์ด๋Š” ์ˆ˜๋™ ์ฐจํŠธ ๋ฆฌ๋ทฐ์˜ ๋ถ€๋‹ด์„ ์ค„์ด๊ณ  ์•” ๋“ฑ๋ก ์ฒด๊ณ„์˜ ํšจ์œจ์„ฑ์„ ๋†’์ด๋Š” ๋ฐ ๊ธฐ์—ฌํ•  ๊ฒƒ์œผ๋กœ ๊ธฐ๋Œ€๋ฉ๋‹ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

38Large Language Models in Clinical Decision Support: A Comparative Analysis of Chat-GPT and Breast Radiologists on ACR Appropriateness Criteria.

2026-06Academic radiologyโญ Q1DOI 10.1016/j.acra.2026.03.015

RATIONALE AND

OBJECTIVE

This study evaluates the performance of ChatGPT, a large language model (LLM), in selecting appropriate imaging modalities for breast imaging scenarios using the American College of Radiology (ACR) Appropriateness Criteria (AC). We aim to compare the agreement of ChatGPT with the ACR AC to that of breast radiologists at a single institution in selecting appropriate imaging modalities. METHODS/MATERIALS: The study utilized ten randomly selected clinical variants from the ACR AC breast imaging category. Outputs were obtained from ChatGPT-3.5, ChatGPT-4, and ChatGPT-4o using the versions available in July 2024. The ChatGPT versions and four breast radiologists rated the appropriateness of 81 imaging decisions on a scale from 1 to 9. For each imaging option within a clinical scenario, the ratings provided by the radiologists and four independent samplings of ChatGPT's responses were aggregated. Agreement between ratings from ChatGPT, radiologists, and the ACR AC was analyzed using generalized estimating equations (GEEs) and Bland-Altman plots to assess consistency and bias.

RESULTS

Radiologists had the lowest overall mean bias (0.2438) relative to the ACR (p = 0.489). All versions of ChatGPT had larger mean biases that were significant (GPT-4o: 2.463; GPT-4: 1.7623, GPT-3.5: 2.4691, all p<0.001). All had a slope bias (p < 0.001), but radiologists had the smallest slope bias. In summary, radiologists were closer to the ACR AC and were oftentimes as variable or even less variable as a group than the same ChatGPT version at the same time.

CONCLUSION

ChatGPT shows promise as an AI tool for imaging decision-making, but current versions lack the accuracy, consistency, and reproducibility demonstrated by experienced radiologists. The study underscores the importance for human oversight in clinical applications and the need for further development to improve ChatGPT's and other LLMs' reliability and alignment with established guidelines.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ChatGPT(GPT-3.5, GPT-4, GPT-4o)๊ฐ€ ๋ฏธ๊ตญ๋ฐฉ์‚ฌ์„ ํ•™ํšŒ(ACR) ์ ์ •์„ฑ ๊ธฐ์ค€(Appropriateness Criteria)์— ๋”ฐ๋ฅธ ์œ ๋ฐฉ ์˜์ƒ ๊ฒ€์‚ฌ ์„ ํƒ์— ์žˆ์–ด ์œ ๋ฐฉ ์˜์ƒ ์ „๋ฌธ์˜์™€ ๋น„๊ตํ•˜์—ฌ ์–ด๋А ์ˆ˜์ค€์˜ ์„ฑ๋Šฅ์„ ๋ณด์ด๋Š”์ง€ ํ‰๊ฐ€ํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ACR ์œ ๋ฐฉ ์˜์ƒ ์นดํ…Œ๊ณ ๋ฆฌ์—์„œ ๋ฌด์ž‘์œ„๋กœ ์„ ์ •๋œ 10๊ฐœ์˜ ์ž„์ƒ ๋ณ€ํ˜•์— ๋Œ€ํ•ด 81๊ฐœ์˜ ์˜์ƒ ๊ฒฐ์ • ํ•ญ๋ชฉ์„ 1~9์  ์ฒ™๋„๋กœ ํ‰๊ฐ€ํ•˜๊ณ , ์ผ๋ฐ˜ํ™” ์ถ”์ • ๋ฐฉ์ •์‹(GEE) ๋ฐ Bland-Altman ๋ถ„์„์œผ๋กœ ACR ๊ธฐ์ค€๊ณผ์˜ ์ผ์น˜๋„๋ฅผ ๋น„๊ตํ•˜์˜€๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ์ „๋ฌธ ์˜์ƒ์˜ํ•™๊ณผ ์˜์‚ฌ๋“ค์€ ACR ๊ธฐ์ค€ ๋Œ€๋น„ ํŽธํ–ฅ์ด ์œ ์˜ํ•˜์ง€ ์•Š์€ ๋ฐ˜๋ฉด(ํ‰๊ท  ํŽธํ–ฅ 0.24, p=0.489), ๋ชจ๋“  ChatGPT ๋ฒ„์ „์€ ํ†ต๊ณ„์ ์œผ๋กœ ์œ ์˜ํ•œ ๊ณผ๋Œ€ ํŽธํ–ฅ์„ ๋ณด์—ฌ(GPT-4o: 2.46, p<0.001), ํ˜„์žฌ์˜ LLM์€ ์ž„์ƒ ์ ์šฉ ์‹œ ์ •ํ™•์„ฑยท์ผ๊ด€์„ฑยท์žฌํ˜„์„ฑ ์ธก๋ฉด์—์„œ ์ „๋ฌธ์˜ ์ˆ˜์ค€์— ๋ฏธ์น˜์ง€ ๋ชปํ•˜๋ฉฐ ์ธ๊ฐ„์˜ ๊ฐ๋…์ด ํ•„์ˆ˜์ ์ž„์„ ๊ฐ•์กฐํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

39Read like a radiologist: Efficient vision-language model for 3D medical imaging interpretation.

2026-06Medical image analysisโญ Q1DOI 10.1016/j.media.2026.104077

Recent medical vision-language models (VLMs) have shown promise in 2D medical image interpretation. However extending them to 3D medical imaging has been challenging due to computational complexities and data scarcity. Although a few recent VLMs specified for 3D medical imaging have emerged, all are limited to learning volumetric representation of a 3D medical image as a set of sub-volumetric features. Such process introduces overly correlated representations along the z-axis that neglect slice-specific clinical details, particularly for 3D medical images where adjacent slices have low redundancy. To address this limitation, we introduce MS-VLM that mimic radiologists' workflow in 3D medical image interpretation. Specifically, radiologists analyze 3D medical images by examining individual slices sequentially and synthesizing information across slices and views. Likewise, MS-VLM leverages self-supervised 2D transformer encoders to learn a volumetric representation that capture inter-slice dependencies from a sequence of slice-specific features. Unbound by sub-volumetric patchification, MS-VLM is capable of obtaining useful volumetric representations from 3D medical images with any slice length and from multiple images acquired from different planes and phases. We evaluate MS-VLM on publicly available chest CT dataset CT-RATE and in-house rectal MRI dataset. In both scenarios, MS-VLM surpasses existing methods in radiology report generation, producing more coherent and clinically relevant reports. These findings highlight the potential of MS-VLM to advance 3D medical image interpretation and improve the robustness of medical VLMs.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ธฐ์กด 3D ์˜๋ฃŒ ์˜์ƒ ํŠนํ™” ์‹œ๊ฐ-์–ธ์–ด ๋ชจ๋ธ(VLM)์ด z์ถ• ๋ฐฉํ–ฅ์œผ๋กœ ๊ณผ๋„ํ•˜๊ฒŒ ์ƒ๊ด€๋œ ํ‘œํ˜„์„ ํ•™์Šตํ•˜์—ฌ ์Šฌ๋ผ์ด์Šค๋ณ„ ์ž„์ƒ ์„ธ๋ถ€ ์ •๋ณด๋ฅผ ๊ฐ„๊ณผํ•œ๋‹ค๋Š” ํ•œ๊ณ„๋ฅผ ๊ทน๋ณตํ•˜๊ณ ์ž, ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜์˜ ํŒ๋… ๋ฐฉ์‹์„ ๋ชจ๋ฐฉํ•œ MS-VLM์„ ์ œ์•ˆํ•˜์˜€๋‹ค. MS-VLM์€ ์ž๊ธฐ์ง€๋„ํ•™์Šต ๊ธฐ๋ฐ˜ 2D ํŠธ๋žœ์Šคํฌ๋จธ ์ธ์ฝ”๋”๋ฅผ ํ™œ์šฉํ•˜์—ฌ ์Šฌ๋ผ์ด์Šค๋ณ„ ํŠน์ง•์„ ์ˆœ์ฐจ์ ์œผ๋กœ ์ฒ˜๋ฆฌํ•˜๊ณ  ์Šฌ๋ผ์ด์Šค ๊ฐ„ ์˜์กด์„ฑ์„ ํฌ์ฐฉํ•จ์œผ๋กœ์จ, ๋‹ค์–‘ํ•œ ์Šฌ๋ผ์ด์Šค ๊ธธ์ด ๋ฐ ๋‹ค์ค‘ ์ดฌ์˜ ํ‰๋ฉดยท์œ„์ƒ ์˜์ƒ์—๋„ ์œ ์—ฐํ•˜๊ฒŒ ์ ์šฉ ๊ฐ€๋Šฅํ•œ ์ฒด์  ํ‘œํ˜„์„ ํ•™์Šตํ•œ๋‹ค. ๊ณต๊ฐœ ํ‰๋ถ€ CT ๋ฐ์ดํ„ฐ์…‹(CT-RATE) ๋ฐ ์ง์žฅ MRI ๋ฐ์ดํ„ฐ์…‹์„ ๋Œ€์ƒ์œผ๋กœ ํ•œ ํ‰๊ฐ€์—์„œ MS-VLM์€ ๊ธฐ์กด ๋ฐฉ๋ฒ•๋“ค์„ ๋Šฅ๊ฐ€ํ•˜๋Š” ์ผ๊ด€์„ฑ ์žˆ๊ณ  ์ž„์ƒ์ ์œผ๋กœ ์œ ์˜๋ฏธํ•œ ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ๋ฅผ ์ƒ์„ฑํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

40How green are large language models for radiology report labelling? Comparing human, rule-based and hybrid workflows.

2026-05Insights into imagingโญ Q1DOI 10.1186/s13244-026-02289-2
OBJECTIVE

To address limited quantitative data on sustainable use of large language models (LLMs) in radiology, we quantified the resource footprint of LLMs for labelling CT pulmonary embolism reports and assessed how a hybrid rule-based-LLM workflow changes time, cost and carbon emissions compared with manual labelling.

METHODS

In this single-centre retrospective study, 2923 structured CT reports were labelled using four workflows: a rule-based extractor (RBE), an LLM-only pipeline using 18 open-weight and four proprietary models, a hybrid RBE-LLM pipeline that routed RBE failures to an LLM, and full manual labelling by radiologists. Ground truth was based on radiologist adjudication. For each LLM, we measured per-report latency, estimated CO2 emissions and cost. Radiologists recorded the labelling time per report.

RESULTS

Manual labelling required 32.8โ€‰h for 2923 reports (40.4โ€‰s/report; โ‚ฌ0.42/report) with 95.0% accuracy (95% CI: 93.7-96.2). LLM-only pipelines were less accurate (85.1%; 95% CI: 84.9-85.5) but reduced labelling time to 12.4โ€‰h and cost to โ‚ฌ2.60 (both pโ€‰<โ€‰0.001). Hybrid RBE-LLM workflows yielded the highest accuracy (98.5%) and lowest resource use: across 22 models, switching from LLM-only to hybrid reduced time (6.7 to 0.97โ€‰h), cost (โ‚ฌ1.19 to โ‚ฌ0.17), and CO2 (0.82 to 0.12โ€‰kg; all pโ€‰<โ€‰0.001).

CONCLUSION

LLM-only labelling reduced labour time and direct costs compared with manual annotation. A hybrid RBE-LLM pipeline that forwards rule-based failures to an LLM concentrated compute where needed and markedly decreased time, cost and emissions, supporting targeted deployment of LLMs for sustainable data-annotation workflows in radiology. CRITICAL RELEVANCE: By quantifying time, cost and carbon emissions of manual, rule-based, LLM and hybrid report labelling, this study identifies sustainable workflows for deploying LLMs in routine radiology reporting. KEY POINTS: Manual expert labelling of CT pulmonary embolism reports is time-intensive and costly. Mid-sized LLM configurations provide favourable trade-offs between performance and resource use. Hybrid rule-based-LLM workflows sustain accuracy while reducing resource demands.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ํ›„ํ–ฅ์  ๋‹จ์ผ๊ธฐ๊ด€ ์—ฐ๊ตฌ๋Š” CT ํ์ƒ‰์ „์ฆ ๋ณด๊ณ ์„œ 2,923๊ฑด์„ ๋Œ€์ƒ์œผ๋กœ ์ˆ˜๋™ ๋ ˆ์ด๋ธ”๋ง, ๊ทœ์น™ ๊ธฐ๋ฐ˜ ์ถ”์ถœ๊ธฐ(RBE), LLM ๋‹จ๋…, ๊ทธ๋ฆฌ๊ณ  ํ•˜์ด๋ธŒ๋ฆฌ๋“œ RBE-LLM ์›Œํฌํ”Œ๋กœ์šฐ์˜ ์‹œ๊ฐ„ยท๋น„์šฉยทํƒ„์†Œ ๋ฐฐ์ถœ๋Ÿ‰์„ ์ •๋Ÿ‰์ ์œผ๋กœ ๋น„๊ตํ•˜์˜€๋‹ค. LLM ๋‹จ๋… ํŒŒ์ดํ”„๋ผ์ธ์€ ์ˆ˜๋™ ๋ ˆ์ด๋ธ”๋ง ๋Œ€๋น„ ์ž‘์—… ์‹œ๊ฐ„๊ณผ ๋น„์šฉ์„ ์ค„์˜€์œผ๋‚˜ ์ •ํ™•๋„๊ฐ€ 85.1%๋กœ ๋‚ฎ์•˜๋˜ ๋ฐ˜๋ฉด, RBE ์‹คํŒจ ์‚ฌ๋ก€๋งŒ LLM์œผ๋กœ ์ฒ˜๋ฆฌํ•˜๋Š” ํ•˜์ด๋ธŒ๋ฆฌ๋“œ ๋ฐฉ์‹์€ ์ •ํ™•๋„ 98.5%๋กœ ๊ฐ€์žฅ ๋†’์œผ๋ฉด์„œ ์‹œ๊ฐ„(6.7โ†’0.97์‹œ๊ฐ„), ๋น„์šฉ(โ‚ฌ1.19โ†’โ‚ฌ0.17), COโ‚‚ ๋ฐฐ์ถœ๋Ÿ‰(0.82โ†’0.12 kg)์„ ๋ชจ๋‘ ์œ ์˜ํ•˜๊ฒŒ ๊ฐ์†Œ์‹œ์ผฐ๋‹ค. ์ด ์—ฐ๊ตฌ๋Š” ๊ทœ์น™ ๊ธฐ๋ฐ˜๊ณผ LLM์„ ๊ฒฐํ•ฉํ•œ ํ•˜์ด๋ธŒ๋ฆฌ๋“œ ์›Œํฌํ”Œ๋กœ์šฐ๊ฐ€ ๋ฐฉ์‚ฌ์„ ๊ณผ ๋ฐ์ดํ„ฐ ์ฃผ์„ ์ž‘์—…์—์„œ ๋†’์€ ์ •ํ™•๋„๋ฅผ ์œ ์ง€ํ•˜๋ฉด์„œ๋„ ์ž์› ์†Œ๋น„๋ฅผ ์ตœ์†Œํ™”ํ•˜๋Š” ์ง€์† ๊ฐ€๋Šฅํ•œ ์ „๋žต์ž„์„ ์ œ์‹œํ•œ๋‹ค.
Added: 2026-05-31 00:01View โ†—

41Who labels best? Radiologists, rules, or large language models for CT reports on pulmonary embolism.

2026-05European radiology experimentalโญ Q1DOI 10.1186/s41747-026-00738-7
OBJECTIVE

To compare open-weight and proprietary large language models (LLMs), a rule-based extractor (RBE) and radiologists for labelling pulmonary embolism CT reports, and to test whether a hybrid RBE-LLM workflow improves labelling performance.

METHODS

This single-centre retrospective study included structured CT reports from October 2021 to March 2025. Three labelling pipelines were evaluated: an RBE; a model-agnostic LLM extractor (18 open-weight, four GPT-4 variants); and a hybrid pipeline routing only RBE failures to an LLM. Ground truth was defined at the report-text level by deterministic schema matching for initially RBE-valid fields and blinded adjudication of RBE-invalid fields by two attending radiologists. Eight radiologists provided a human baseline. Outcomes included F1 scores, accuracy, LLM-based salvage of RBE failures, and labelling time.

RESULTS

In total, 2,923 reports from 2,923 patients (mean age 66โ€‰ยฑโ€‰17 years; 1,465 women) were included. Falcon3-10b and GPT-4.1-mini achieved similar item-level performance (F1 0.98 [95% CI, 0.97-0.98] for both; pโ€‰=โ€‰0.70) and both exceeded the RBE (F1 0.81 [95% CI, 0.80-0.82]; pโ€‰<โ€‰0.001). Salvage of RBE failures was comparable between open-weight and proprietary models (88.1% vs 91.9%; pโ€‰=โ€‰0.12). The hybrid RBE-LLM workflow achieved 99.8% accuracy and F1 0.99 (0.98-0.99), exceeding both the RBE and pooled radiologists (F1 0.92 [95% CI, 0.90-0.93]; all pโ€‰<โ€‰0.001).

CONCLUSION

Schema-constrained open-weight and proprietary LLMs exceeded rule-based extraction and, at the upper end of performance, matched a pooled radiologist label-transfer baseline. A rules-first, targeted LLM workflow enabled near-perfect extraction from finalised structured pulmonary embolism CT reports. RELEVANCE STATEMENT: A rules-first LLM workflow can automate high-fidelity extraction of structured CT findings from finalised radiology reports, enabling scalable, auditable, and more consistent cohort curation for clinical research, registries, and quality improvement. KEY POINTS: A hybrid rules-first workflow combining a rule-based extractor (RBE) with targeted large language model (LLM) salvage achieved the highest overall performance for labelling of pulmonary embolism CT reports (F1, 0.99; accuracy, 99.8%). The top standalone open-weight and proprietary LLMs (Falcon3-10b and GPT-4.1-mini) both exceeded the RBE and, at the upper end of performance, matched a pooled radiologist label-transfer baseline. The hybrid workflow reduced cohort-curation time from 32.2โ€‰h for radiologists to 1.0โ€‰h while reducing LLM calls by 85.6%, because the LLM was only triggered for rule-failed fields.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ๋‹จ์ผ๊ธฐ๊ด€ ํ›„ํ–ฅ์  ์—ฐ๊ตฌ๋Š” 2,923๊ฑด์˜ ํ์ƒ‰์ „์ฆ CT ๊ตฌ์กฐํ™” ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ ๊ทœ์น™ ๊ธฐ๋ฐ˜ ์ถ”์ถœ๊ธฐ(RBE), ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM, ์˜คํ”ˆ ์†Œ์Šค 18์ข… ๋ฐ GPT-4 ๊ณ„์—ด 4์ข…), ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ 8๋ช…์˜ ๋ ˆ์ด๋ธ”๋ง ์„ฑ๋Šฅ์„ ๋น„๊ตํ•˜๊ณ , RBE ์‹คํŒจ ํ•ญ๋ชฉ์„ LLM์œผ๋กœ ๋ณด์™„ํ•˜๋Š” ํ•˜์ด๋ธŒ๋ฆฌ๋“œ ํŒŒ์ดํ”„๋ผ์ธ์˜ ํšจ์šฉ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ๋‹จ๋… LLM(Falcon3-10b, GPT-4.1-mini)์€ RBE(F1 0.81)๋ฅผ ์œ ์˜ํ•˜๊ฒŒ ์ƒํšŒํ•˜์˜€์œผ๋ฉฐ(F1 0.98, p<0.001), ์˜คํ”ˆ ์†Œ์Šค ๋ชจ๋ธ๊ณผ ์ƒ์šฉ ๋ชจ๋ธ ๊ฐ„ ์„ฑ๋Šฅ ์ฐจ์ด๋Š” ํ†ต๊ณ„์ ์œผ๋กœ ์œ ์˜ํ•˜์ง€ ์•Š์•˜๋‹ค. ๊ทœ์น™ ์šฐ์„  ์ฒ˜๋ฆฌ ํ›„ ์‹คํŒจ ํ•ญ๋ชฉ์—๋งŒ LLM์„ ์ ์šฉํ•˜๋Š” ํ•˜์ด๋ธŒ๋ฆฌ๋“œ ์›Œํฌํ”Œ๋กœ์šฐ๋Š” F1 0.99, ์ •ํ™•๋„ 99.8%๋กœ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ ์ง‘๋‹จ(F1 0.92)์„ ์ดˆ๊ณผํ•˜์˜€์œผ๋ฉฐ, ์ฝ”ํ˜ธํŠธ ๊ตฌ์ถ• ์†Œ์š” ์‹œ๊ฐ„์„ 32.2์‹œ๊ฐ„์—์„œ 1.0์‹œ๊ฐ„์œผ๋กœ ๋‹จ์ถ•ํ•˜๊ณ  LLM ํ˜ธ์ถœ์„ 85.6% ๊ฐ์†Œ์‹œ์ผœ ์ž„์ƒ ์—ฐ๊ตฌ ๋ฐ ๋ ˆ์ง€์ŠคํŠธ๋ฆฌ๋ฅผ ์œ„ํ•œ ํ™•์žฅ ๊ฐ€๋Šฅํ•˜๊ณ  ๊ฐ์‚ฌ ๊ฐ€๋Šฅํ•œ ์ž๋™ํ™” ์ถ”์ถœ ์ „๋žต์œผ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-05-31 00:01View โ†—

42Glass-box agentic-style workflow for multiclass cine cardiac magnetic resonance imaging classification with a large language model.

2026-05Diagnostic and interventional radiology (Ankara, Turkey)๐Ÿ”ท Q2DOI 10.4274/dir.2026.264016
OBJECTIVE

To develop and evaluate a glass-box, agentic-style radiology pipeline that separates perception from reasoning for auditable multiclass diagnosis on cine cardiac magnetic resonance imaging (MRI), and to quantify accuracy, robustness across decoding temperatures, and fidelity/safety of generated narrative explanations.

METHODS

Using the labeled Automated Cardiac Diagnosis Challenge training cohort (n = 100; five diagnostic classes), cine bSSFP images were segmented at end-diastole (ED) and end-systole (ES) with a pretrained nnU-Net, and 17 clinically interpretable biomarkers were extracted. A large language model (LLM) (GPT-OSS-120B) queried prompts under three different prompt strategies (V1-V3) with majority-vote self-consistency after a stratified split into prompt development (n = 20) and independent evaluation (n = 80). Temperatures (T = 0.1, 1.0, and 2.0) were tested for stability. A decoupled narrative module generated radiologist-style reports. Narratives underwent radiologist audit for numeric fidelity and clinical safety. Machine learning algorithms [Random Forest, Support Vector Machine (SVM), Logistic Regression, Decision Tree] were trained on the same biomarker set for benchmarking.

RESULTS

Automated segmentation showed high agreement with reference masks [Dice at ED: right ventricle [RV] cavity 0.984 ยฑ 0.004, left ventricle (LV)] myocardium 0.965 ยฑ 0.009, LV cavity 0.989 ยฑ 0.003; ES: RV cavity 0.979 ยฑ 0.013, LV myocardium 0.975 ยฑ 0.009, LV cavity 0.985 ยฑ 0.005). The hierarchical veto-logic strategy (V3) achieved an accuracy of 0.925 (95% confidence interval: 0.863-0.975) and a macro-F1 of 0.924, remaining stable across temperatures, outperforming V2 (accuracy 0.787-0.800) and V1 (0.562-0.600). Reproducibility was highest for V3 at T = 0.1 (Fleiss' kappa: 0.969) with a low failure rate (0.83%). Narrative generation produced 97.5% valid reports with 100% numeric fidelity and audited safety โ‰ฅ 97.5%. Performance was comparable to supervised models (Random Forest accuracy 0.938; SVM/Logistic Regression accuracy 0.925).

CONCLUSION

In this single-dataset internal evaluation, a glass-box workflow combining automated segmentation-derived biomarkers with an LLM enables robust multiclass cardiac MRI diagnosis while producing numerically faithful, safety-audited narratives, supporting auditability and governance for radiology artificial intelligence (AI). External multicenter validation is needed to confirm generalizability. CLINICAL

CONCLUSION

A glass-box, biomarker-driven agentic-style workflow enables auditable cine cardiac MRI classification with numerically grounded explanations, addressing interpretability and stability barriers that limit translation of radiology AI into routine practice.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์‹ฌ์žฅ MRI(cine CMR)์˜ ๋‹ค์ค‘ ๋ถ„๋ฅ˜ ์ง„๋‹จ์„ ์œ„ํ•ด ์ž๋™ ๋ถ„ํ• (nnU-Net)๋กœ ์ถ”์ถœํ•œ 17๊ฐœ์˜ ์ž„์ƒ ๋ฐ”์ด์˜ค๋งˆ์ปค์™€ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(GPT-OSS-120B)์„ ๊ฒฐํ•ฉํ•œ '๊ธ€๋ž˜์Šค๋ฐ•์Šค(glass-box)' ์—์ด์ „ํ‹ฑ ํŒŒ์ดํ”„๋ผ์ธ์„ ๊ฐœ๋ฐœํ•˜๊ณ  ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ACDC ์ฝ”ํ˜ธํŠธ(n=100, 5๊ฐœ ์ง„๋‹จ ํด๋ž˜์Šค)๋ฅผ ๋Œ€์ƒ์œผ๋กœ ๊ณ„์ธต์  ๊ฑฐ๋ถ€ ๋…ผ๋ฆฌ(veto-logic) ํ”„๋กฌํ”„ํŠธ ์ „๋žต(V3)๊ณผ ๋‹ค์ˆ˜๊ฒฐ ์ž๊ธฐ์ผ๊ด€์„ฑ ๋ฐฉ์‹์„ ์ ์šฉํ•œ ๊ฒฐ๊ณผ, ๋…๋ฆฝ ํ‰๊ฐ€ ์„ธํŠธ์—์„œ ์ •ํ™•๋„ 0.925, macro-F1 0.924๋ฅผ ๋‹ฌ์„ฑํ•˜์˜€์œผ๋ฉฐ ๋‹ค์–‘ํ•œ ๋””์ฝ”๋”ฉ ์˜จ๋„์—์„œ๋„ ๋†’์€ ์žฌํ˜„์„ฑ(Fleiss' ฮบ=0.969)์„ ์œ ์ง€ํ•˜์˜€๋‹ค. ๋˜ํ•œ ์ƒ์„ฑ๋œ ์„œ์ˆ ํ˜• ๋ณด๊ณ ์„œ๋Š” ์ˆ˜์น˜ ์ถฉ์‹ค๋„ 100%, ์ž„์ƒ ์•ˆ์ „์„ฑ โ‰ฅ97.5%๋กœ ๊ฐ์‚ฌ ๊ฒ€์ฆ์„ ํ†ต๊ณผํ•˜์—ฌ, ๋ฐฉ์‚ฌ์„  AI์˜ ํ•ด์„ ๊ฐ€๋Šฅ์„ฑ๊ณผ ๊ฑฐ๋ฒ„๋„Œ์Šค๋ฅผ ๋™์‹œ์— ์ถฉ์กฑํ•˜๋Š” ๊ฐ์‚ฌ ๊ฐ€๋Šฅํ•œ ์ง„๋‹จ ์›Œํฌํ”Œ๋กœ์šฐ์˜ ์ž„์ƒ ์ ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ์ œ์‹œํ•˜์˜€๋‹ค.
Added: 2026-05-17 00:00View โ†—

43Foundation models for radiology: fundamentals, applications, opportunities, challenges, risks, and prospects.

2026-05Diagnostic and interventional radiology (Ankara, Turkey)๐Ÿ”ท Q2DOI 10.4274/dir.2025.253445

Foundation models (FMs) represent a significant evolution in artificial intelligence (AI), impacting diverse fields. Within radiology, this evolution offers greater adaptability, multimodal integration, and improved generalizability compared with traditional narrow AI. Utilizing large-scale pre-training and efficient fine-tuning, FMs can support diverse applications, including image interpretation, report generation, integrative diagnostics combining imaging with clinical/laboratory data, and synthetic data creation, holding significant promise for advancements in precision medicine. However, clinical translation of FMs faces several substantial challenges. Key concerns include the inherent opacity of model decision-making processes, environmental and social sustainability issues, risks to data privacy, complex ethical considerations, such as bias and fairness, and navigating the uncertainty of regulatory frameworks. Moreover, rigorous validation is essential to address inherent stochasticity and the risk of hallucination. This international collaborative effort provides a comprehensive overview of the fundamentals, applications, opportunities, challenges, and prospects of FMs, aiming to guide their responsible and effective adoption in radiology and healthcare.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„ ํ•™ ๋ถ„์•ผ์—์„œ ํŒŒ์šด๋ฐ์ด์…˜ ๋ชจ๋ธ(FM)์˜ ๊ธฐ์ดˆ ์›๋ฆฌ, ์ ์šฉ ๊ฐ€๋Šฅ์„ฑ ๋ฐ ์ž„์ƒ์  ์ „๋ง์„ ๊ตญ์ œ ๊ณต๋™์—ฐ๊ตฌ๋ฅผ ํ†ตํ•ด ์ข…ํ•ฉ์ ์œผ๋กœ ๊ณ ์ฐฐํ•˜์˜€๋‹ค. FM์€ ๋Œ€๊ทœ๋ชจ ์‚ฌ์ „ํ•™์Šต๊ณผ ํšจ์œจ์ ์ธ ๋ฏธ์„ธ์กฐ์ •์„ ๊ธฐ๋ฐ˜์œผ๋กœ ์˜์ƒ ํŒ๋…, ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ์ž๋™ ์ƒ์„ฑ, ์ž„์ƒยท๊ฒ€์‚ฌ ๋ฐ์ดํ„ฐ์™€์˜ ํ†ตํ•ฉ ์ง„๋‹จ, ํ•ฉ์„ฑ ๋ฐ์ดํ„ฐ ์ƒ์„ฑ ๋“ฑ ๋‹ค์–‘ํ•œ ๋ฐฉ์‚ฌ์„ ๊ณผ ์—…๋ฌด์— ํ™œ์šฉ ๊ฐ€๋Šฅํ•˜๋ฉฐ, ๊ธฐ์กด ํ˜‘์˜ AI ๋Œ€๋น„ ๋†’์€ ์ ์‘์„ฑ๊ณผ ๋‹ค์ค‘๋ชจ๋‹ฌ ํ†ตํ•ฉ ๋Šฅ๋ ฅ์„ ์ œ๊ณตํ•œ๋‹ค. ๊ทธ๋Ÿฌ๋‚˜ ๋ชจ๋ธ ์˜์‚ฌ๊ฒฐ์ • ๊ณผ์ •์˜ ๋ถˆํˆฌ๋ช…์„ฑ, ๋ฐ์ดํ„ฐ ํ”„๋ผ์ด๋ฒ„์‹œ ์œ„ํ—˜, ํŽธํ–ฅ ๋ฐ ๊ณต์ •์„ฑ ๋“ฑ์˜ ์œค๋ฆฌ์  ๋ฌธ์ œ, ๊ทœ์ œ ๋ถˆํ™•์‹ค์„ฑ, ๊ทธ๋ฆฌ๊ณ  ํ™˜๊ฐ(hallucination) ์œ„ํ—˜์— ๋Œ€ํ•œ ์—„๊ฒฉํ•œ ๊ฒ€์ฆ์˜ ํ•„์š”์„ฑ ๋“ฑ ์ž„์ƒ ์ ์šฉ์„ ์œ„ํ•ด ํ•ด๊ฒฐํ•ด์•ผ ํ•  ์ค‘์š”ํ•œ ๊ณผ์ œ๋“ค์ด ์กด์žฌํ•œ๋‹ค.
Added: 2026-05-10 00:00View โ†—

44GenAI-Supported Virtual Patients in Health Care Education: Systematic Review.

2026-05Journal of medical Internet researchโญ Q1DOI 10.2196/82756
BACKGROUND

Generative artificial intelligence (GenAI) is enhancing virtual patient simulations in health care education by enabling dynamic, adaptive interactions, reshaping how clinical skills are taught. A synthesis of the current evidence is needed to guide implementation and future research, given the pace of technological advancement.

OBJECTIVE

This systematic review aims to synthesize empirical research on the design, implementation, and educational impact of GenAI-supported virtual patients in health care education.

METHODS

A systematic search was conducted across 5 databases (CINAHL, Medline, Embase, Scopus, and Web of Science) from their inception to March 19, 2026. Reference lists of included studies and relevant systematic reviews were also screened. Peer-reviewed studies in English that evaluated GenAI-supported virtual patients using quantitative or mixed methods were included. Two reviewers independently screened studies and extracted data. Study quality and risk of bias were assessed critically using JBI (Joanna Briggs Institute) checklists, with disagreements resolved by consensus.

RESULTS

A total of 15 studies met the inclusion criteria (total participants N=645), spanning health care disciplines, including nursing, medicine, pharmacy, radiography, and medical first-responder training. The virtual patients varied in design; input modalities included text (9 studies), voice (5 studies), or hybrid (1 study); output was text (9 studies), speech (5 studies), or both (1 study); 6 studies used 3D-embodied avatars, while 9 used nonembodied interfaces. A total of 13 studies used OpenAI GPT models (eg, ChatGPT), 1 used a fine-tuned model from a different provider, and 1 evaluated multiple model families (Claude, GPT, and open-source). Further, 6 studies used controlled experimental designs, including 3 randomized controlled trials (RCTs); the remainder were cross-sectional or prepost evaluations. Primary outcomes included user perceptions (14 studies), communication skills (4 studies), clinical reasoning (3 studies), and performance (7 studies). In controlled comparisons, GenAI-supported virtual patients consistently improved outcomes relative to control conditions: for example, enhanced clinical decision-making (RCT, n=21), ophthalmology history-taking skills (RCT, n=26), and medical history-taking performance (crossover RCT, n=20). The evidence base is characterized by brief intervention durations, a predominant reliance on single-session interactions, and a general lack of underpinning educational theory. No meta-analysis was performed due to the limited number of studies and significant heterogeneity in designs, interventions, and outcome measures.

CONCLUSION

The evidence supports the feasibility and acceptability of GenAI-supported virtual patients, with positive learner perceptions and promising outcomes for skills development. However, critical limitations remain in emotional-behavioral complexity, simulation adaptability, and research design rigor (eg, limited use of control groups and validated instruments). The review offers educators, instructional designers, and policymakers actionable insights for integrating dynamic, artificial intelligence-driven simulations while identifying crucial gaps-such as the need for theoretical grounding, longitudinal studies, and standardized design protocols-that must be addressed for safe and effective implementation.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
์ด ์ฒด๊ณ„์  ๋ฌธํ—Œ๊ณ ์ฐฐ์€ ์ƒ์„ฑํ˜• ์ธ๊ณต์ง€๋Šฅ(GenAI) ๊ธฐ๋ฐ˜ ๊ฐ€์ƒ ํ™˜์ž ์‹œ๋ฎฌ๋ ˆ์ด์…˜์ด ์˜๋ฃŒ ๊ต์œก์—์„œ ๊ฐ–๋Š” ์„ค๊ณ„ยท๊ตฌํ˜„ยท๊ต์œก์  ํšจ๊ณผ๋ฅผ ์ข…ํ•ฉํ•˜๊ธฐ ์œ„ํ•ด, 5๊ฐœ ์ฃผ์š” ๋ฐ์ดํ„ฐ๋ฒ ์ด์Šค์—์„œ ์ด 15ํŽธ์˜ ์—ฐ๊ตฌ(์ฐธ์—ฌ์ž N=645)๋ฅผ ์„ ์ •ํ•˜์—ฌ ๋ถ„์„ํ•˜์˜€๋‹ค. ๊ฐ„ํ˜ธ, ์˜ํ•™, ์•ฝํ•™, ๋ฐฉ์‚ฌ์„ ํ•™ ๋“ฑ ๋‹ค์–‘ํ•œ ๋ณด๊ฑด์˜๋ฃŒ ๋ถ„์•ผ์— ๊ฑธ์ณ ์ฃผ๋กœ OpenAI GPT ๊ณ„์—ด ๋ชจ๋ธ์ด ํ™œ์šฉ๋˜์—ˆ์œผ๋ฉฐ, ๋ฌด์ž‘์œ„ ๋Œ€์กฐ์‹œํ—˜์„ ํฌํ•จํ•œ ํ†ต์ œ ๋น„๊ต ์—ฐ๊ตฌ์—์„œ GenAI ๊ฐ€์ƒ ํ™˜์ž๋Š” ์ž„์ƒ์  ์˜์‚ฌ๊ฒฐ์ •, ๋ณ‘๋ ฅ์ฒญ์ทจ ๊ธฐ์ˆ  ๋ฐ ์ˆ˜ํ–‰ ๋Šฅ๋ ฅ ํ–ฅ์ƒ์— ์žˆ์–ด ๋Œ€์กฐ๊ตฐ ๋Œ€๋น„ ์ผ๊ด€๋˜๊ฒŒ ์šฐ์ˆ˜ํ•œ ๊ฒฐ๊ณผ๋ฅผ ๋ณด์˜€๋‹ค. ๋‹ค๋งŒ ํ˜„ ๊ทผ๊ฑฐ๋Š” ๋‹จํšŒ์„ฑ ๊ฐœ์ž…์— ํŽธ์ค‘๋˜์–ด ์žˆ๊ณ  ๊ต์œก ์ด๋ก ์  ๊ทผ๊ฑฐ๊ฐ€ ๋ถ€์กฑํ•˜๋ฉฐ ์—ฐ๊ตฌ ์„ค๊ณ„์˜ ์—„๋ฐ€์„ฑ์ด ๋ฏธํกํ•˜์—ฌ, ์•ˆ์ „ํ•˜๊ณ  ํšจ๊ณผ์ ์ธ ์ž„์ƒ ๊ต์œก ํ†ตํ•ฉ์„ ์œ„ํ•ด์„œ๋Š” ์ด๋ก ์  ํ† ๋Œ€ ํ™•๋ฆฝ, ์žฅ๊ธฐ ์ถ”์  ์—ฐ๊ตฌ, ํ‘œ์ค€ํ™”๋œ ์„ค๊ณ„ ํ”„๋กœํ† ์ฝœ ๊ฐœ๋ฐœ์ด ์‹œ๊ธ‰ํžˆ ์š”๊ตฌ๋œ๋‹ค.
Added: 2026-05-10 00:00View โ†—

45BI-RADS-compliant structured mammography reporting using locally deployed large language models under privacy constraints.

2026-05European radiologyโญ Q1DOI 10.1007/s00330-025-12147-2
OBJECTIVE

To develop a privacy-preserving method for structuring free-text mammography reports using a locally fine-tuned, open-source large language model (LLM).

METHODS

In this multicenter study, 7161 unstructured mammography reports were collected from three institutions. The open-source Llama-3 model was fine-tuned via supervised learning using pseudo-labels from the commercial Qwen-Max model with low-rank adaptation. All labels were pseudo-labels generated by the commercial Qwen-Max model rather than human annotations. Structured outputs followed a BI-RADS-oriented nested JSON schema. Performance was evaluated across 23 features using Precision, Recall, and F1-score. Structural integrity was assessed using the JSON format accuracy (JFA) and field integrity accuracy (FIA) metrics. Statistical comparisons were performed using the paired Wilcoxon signed-rank test and Cohen's d effect size.

RESULTS

A total of 7161 reports were retrospectively obtained from three institutions and analyzed. The fine-tuned model achieved strong performance at epoch 10 (Precision 0.942, Recall 0.929, F1-score 0.932), with JFA and FIA reaching 0.964 and 1.000, respectively, showing significant gains over the base model (pโ€‰<โ€‰0.05, Cohen's dโ€‰>โ€‰0.8). While slightly below Qwen-Max overall, the model exhibited moderate yet statistically significant differences (pโ€‰<โ€‰0.05; 0.5โ€‰<โ€‰Cohen's dโ€‰<โ€‰0.8), particularly in the "special signs" category (F1โ€‰=โ€‰0.737 vs 0.947).

CONCLUSION

This method effectively converts mammography reports into structured data using a locally fine-tuned, open-source LLM. Although there is a slight performance trade-off, it improves privacy and can be deployed locally. Its accuracy, clinical relevance, and compliance make it a practical solution for medical institutions. KEY POINTS: Question Free-text mammography reports lack standardization, making structured extraction difficult, while existing solutions often compromise privacy, adaptability, or require costly commercial tools. Findings Our fine-tuned LLaMA-3 model achieved high extraction accuracy (F1-score: 0.932) and complete structural integrity (FIA: 1.000) within a fully local deployment pipeline. Clinical relevance This method enables standardized and privacy-preserving mammography reporting without disruption to the clinical workflow, supporting safer AI integration, enhanced data quality, and compliance with regulations such as the General Data Protection Regulation (GDPR).

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ฐœ์ธ์ •๋ณด ๋ณดํ˜ธ ๊ทœ์ •์„ ์ค€์ˆ˜ํ•˜๋ฉด์„œ ์ž์œ  ํ˜•์‹์˜ ์œ ๋ฐฉ์ดฌ์˜์ˆ  ํŒ๋…๋ฌธ์„ BI-RADS ๊ธฐ๋ฐ˜์˜ ๊ตฌ์กฐํ™”๋œ ๋ฐ์ดํ„ฐ๋กœ ์ž๋™ ๋ณ€ํ™˜ํ•˜๊ธฐ ์œ„ํ•ด, ์˜คํ”ˆ์†Œ์Šค Llama-3 ๋ชจ๋ธ์„ ์ƒ์—…์šฉ Qwen-Max ๋ชจ๋ธ์˜ ์˜์‚ฌ ๋ ˆ์ด๋ธ”(pseudo-label)๋กœ ๋กœ์ปฌ ํ™˜๊ฒฝ์—์„œ ๋ฏธ์„ธ ์กฐ์ •ํ•˜๋Š” ๋ฐฉ๋ฒ•์„ ๊ฐœ๋ฐœํ•˜์˜€๋‹ค. 3๊ฐœ ๊ธฐ๊ด€์—์„œ ์ˆ˜์ง‘ํ•œ 7,161๊ฑด์˜ ๋น„๊ตฌ์กฐํ™” ํŒ๋…๋ฌธ์„ ํ™œ์šฉํ•˜์—ฌ ์ €์ˆœ์œ„ ์ ์‘(LoRA) ๊ธฐ๋ฒ•์œผ๋กœ ํ•™์Šตํ•œ ๊ฒฐ๊ณผ, ๋ฏธ์„ธ ์กฐ์ • ๋ชจ๋ธ์€ F1-score 0.932, JSON ํ˜•์‹ ์ •ํ™•๋„(JFA) 0.964, ํ•„๋“œ ๋ฌด๊ฒฐ์„ฑ ์ •ํ™•๋„(FIA) 1.000์„ ๋‹ฌ์„ฑํ•˜๋ฉฐ ๊ธฐ๋ณธ ๋ชจ๋ธ ๋Œ€๋น„ ํ†ต๊ณ„์ ์œผ๋กœ ์œ ์˜ํ•œ ์„ฑ๋Šฅ ํ–ฅ์ƒ์„ ๋ณด์˜€๋‹ค(p < 0.05, Cohen's d > 0.8). ์ƒ์—…์šฉ ๋ชจ๋ธ ๋Œ€๋น„ "ํŠน์ˆ˜ ์†Œ๊ฒฌ" ํ•ญ๋ชฉ์—์„œ ๋‹ค์†Œ ๋‚ฎ์€ ์„ฑ๋Šฅ(F1 0.737 vs 0.947)์„ ๋ณด์˜€์œผ๋‚˜, ์™„์ „ํ•œ ๋กœ์ปฌ ๋ฐฐํฌ๋ฅผ ํ†ตํ•ด GDPR ๋“ฑ ๊ฐœ์ธ์ •๋ณด ๋ณดํ˜ธ ๊ทœ์ •์„ ์ค€์ˆ˜ํ•˜๋ฉด์„œ๋„ ์ž„์ƒ ์›Œํฌํ”Œ๋กœ์šฐ์— ๋Œ€ํ•œ ๋ถ€๋‹ด ์—†์ด ํ‘œ์ค€ํ™”๋œ ์œ ๋ฐฉ์ดฌ์˜์ˆ  ๋ณด๊ณ ๊ฐ€ ๊ฐ€๋Šฅํ•จ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-05-03 00:00View โ†—

46Comparing performance of seven fine-tuned open-source large language models in summarizing and predicting outcome-relevant information from mechanical thrombectomy reports in patients with acute ischemic stroke.

2026-05European radiologyโญ Q1DOI 10.1007/s00330-025-12122-x
OBJECTIVE

This study evaluates seven open-source Large Language Models (LLMs) in summarizing radiology reports of acute ischemic stroke patients treated with mechanical thrombectomy and predicting angiography-based outcome measures relevant to post-thrombectomy reperfusion.

METHODS

2000 mechanical thrombectomy reports (findings and summarizing impression section as gold standard) were split into training set (Nโ€‰=โ€‰1900) for model fine-tuning and test set (Nโ€‰=โ€‰100). A two-step evaluation was performed: (1) Quantitative analyses of seven LLMs with metrics ROUGE-1, -2, -L, METEOR, BERTScore (F1) and BLEU comparing LLM-generated summaries against gold-standard impressions. (2) Qualitative manual evaluation of the four best-performing models by two radiologists, assessing correctness and completeness across key parameters: outcome-relevant scores, vessel information, occlusion side, number of passes, relevant additional information, hallucinations, and grammar quality. Statistical significance was assessed via a two-tailed, four-sample ฯ‡ยฒ test, followed by post hoc pairwise ฯ‡ยฒ comparisons.

RESULTS

BioMistral-7b scored highest across most quantitative metrics (ROUGE-1: 0.47, ROUGE-2: 0.30, ROUGE-L: 0.43, METEOR: 0.46, BERTScore (F1): 0.82). Manual evaluation revealed gemma-2-9b most frequently documented pass counts (56 out of 100 cases (56%); pโ€‰<โ€‰0.02 vs. Llama-3.1-8b/mistral-7b-instruct), while mistral-7b-instruct described them most often correctly (29 out of 38 mentioned passes (76.32%); pโ€‰<โ€‰0.02 vs. BioMistral-7b and pโ€‰<โ€‰0.01 vs. gemma-2-9b). All four manually evaluated LLMs performed moderately well in predicting "Thrombolysis-In-Cerebral-Ischemia (TICI)" Score (correctness rate ranging from 66 to 71%; pโ€‰=โ€‰0.89).

CONCLUSION

All four manually evaluated LLMs effectively summarized thrombectomy reports and demonstrated moderate accuracy predicting TICI scores. Their integration into radiology workflows could enhance efficiency, warranting further clinical validation. KEY POINTS: Question Specifically fine-tuned Large Language Models (LLMs) can improve radiology workflow by automatically summarizing thrombectomy reports and inferring angiographic classifications from textual descriptions. Findings Fine-tuned LLMs achieve similar performance in summarizing thrombectomy reports, with each model performing best in specific categories and showing moderate accuracy in correct "Thrombolysis-In-Cerebral-Ischemia (TICI)" Score prediction (66-71%). Clinical relevance Integrating fine-tuned LLMs into radiology workflows may accelerate decision-making and improve patient outcomes by automatically summarizing reports and assessing recanalization success, while future work should enhance contextual understanding, address ambiguous inputs, and limit hallucinations.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ธ‰์„ฑ ํ—ˆํ˜ˆ์„ฑ ๋‡Œ์กธ์ค‘ ํ™˜์ž์˜ ๊ธฐ๊ณ„์  ํ˜ˆ์ „์ œ๊ฑฐ์ˆ  ๋ณด๊ณ ์„œ ์š”์•ฝ ๋ฐ ํ˜ˆ๊ด€์žฌ๊ฐœํ†ต ๊ฒฐ๊ณผ ์˜ˆ์ธก์— ์žˆ์–ด ๋ฏธ์„ธ์กฐ์ •๋œ ์˜คํ”ˆ์†Œ์Šค ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM) 7์ข…์˜ ์„ฑ๋Šฅ์„ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 1,900๊ฑด์˜ ๋ณด๊ณ ์„œ๋กœ ๋ชจ๋ธ์„ ๋ฏธ์„ธ์กฐ์ •ํ•˜๊ณ  100๊ฑด์˜ ํ…Œ์ŠคํŠธ ์„ธํŠธ์— ๋Œ€ํ•ด ROUGE, METEOR, BERTScore ๋“ฑ ์ •๋Ÿ‰์  ์ง€ํ‘œ์™€ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜ 2์ธ์˜ ์ •์„ฑ์  ํ‰๊ฐ€๋ฅผ ๋ณ‘ํ–‰ํ•˜์˜€๋‹ค. BioMistral-7b๊ฐ€ ์ •๋Ÿ‰์  ์ง€ํ‘œ์—์„œ ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ์ •์„ฑ์  ํ‰๊ฐ€์—์„œ ์ƒ์œ„ 4๊ฐœ ๋ชจ๋ธ ๋ชจ๋‘ ํ˜ˆ์ „์ œ๊ฑฐ์ˆ  ๋ณด๊ณ ์„œ๋ฅผ ํšจ๊ณผ์ ์œผ๋กœ ์š”์•ฝํ•˜๊ณ  TICI ์ ์ˆ˜๋ฅผ ์ค‘๋“ฑ๋„ ์ •ํ™•๋„(66~71%)๋กœ ์˜ˆ์ธกํ•จ์œผ๋กœ์จ, ๋ฐฉ์‚ฌ์„ ๊ณผ ์ž„์ƒ ์›Œํฌํ”Œ๋กœ์šฐ ํ†ตํ•ฉ ์‹œ ์ง„๋ฃŒ ํšจ์œจ ํ–ฅ์ƒ ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-05-03 00:00View โ†—

47Patient and lesion characteristics associated with follow-up completion for pancreatic cystic lesions detected on MRI.

2026-05Abdominal radiology (New York)โญ Q1DOI 10.1007/s00261-025-05230-1
OBJECTIVE

To evaluate the association of patient characteristics, community-level social determinants of health, and cyst risk categories with completion of follow-up recommendations for incidental Pancreatic Cystic Lesions (PCLs).

METHODS

We retrospectively identified consecutive patients (2013-2023) whose MRI radiology reports described PCLs. A fine-tuned LLaMA-3.1 8B Instruct large language model was used to extract PCL features. Lesions were classified using the 2017 ACR white paper: Category 1 (low risk), Category 2 (worrisome features), or Category 3 (high-risk stigmata). We recorded demographics and follow-up imaging or endoscopic ultrasound dates. Community-level factors were characterized by the 2020 CDC Social Vulnerability Index (SVI), stratified into quartiles. The primary outcome, "inappropriate follow-up," combined late and no follow-up. Multivariable binomial regression was applied to evaluate associations with inappropriate follow-up.

RESULTS

In 7,745 patients (mean age 66.3ย years; 4,796 women), 92.9% (7,198/7,745) of cysts were Category 1, 6.4% (498/7,745) were Category 2, and 0.6% (49/7,745) were Category 3. Only 36.3% of patients completed appropriate follow-up, 12.1% were late, and 51.6% were lost to follow-up. Inappropriate follow-up was high in every cyst category: 64.2% in Category 1, 59.4% in Category 2 and 49.0% in Category 3. In multivariable analysis, non-English primary language (RR 1.08; 95% CI, 1.02-1.14) and residing in more vulnerable communities of the 3rd quartiles of the socioeconomic Social Vulnerability Index subcategory (RR 1.07; 95% CI, 1.02-1.12) were associated with inappropriate follow-up. Higher age-adjusted Charlson Comorbidity Index (CCIโ€‰โ‰ฅโ€‰4) (RR .84; 95% CI, .79-.88), CCI 2-3 (RR .84; 95% CI, .79-.88), and higher-risk cysts in patients under 65ย years of age (RR .76; 95% CI, .65-.89) were associated with completed follow-up.

CONCLUSION

Follow-up completion for incidental PCLs was low. Factors most consistently associated with follow-up completion were language barriers, residence in socioeconomically vulnerable communities, age-adjusted CCI and higher-risk features among those under 65ย years.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” MRI์—์„œ ์šฐ์—ฐํžˆ ๋ฐœ๊ฒฌ๋œ ์ทŒ์žฅ ๋‚ญ์„ฑ ๋ณ‘๋ณ€(PCL) ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ ์ถ”์  ๊ด€์ฐฐ ์™„๋ฃŒ ์—ฌ๋ถ€์™€ ๊ด€๋ จ๋œ ํ™˜์ž ํŠน์„ฑ, ๋ณ‘๋ณ€ ์œ„ํ—˜ ๋ถ„๋ฅ˜, ์ง€์—ญ์‚ฌํšŒ ์ˆ˜์ค€์˜ ์‚ฌํšŒ์  ๊ฑด๊ฐ• ๊ฒฐ์ • ์š”์ธ์„ ํ›„ํ–ฅ์ ์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. 2013๋…„๋ถ€ํ„ฐ 2023๋…„๊นŒ์ง€ 7,745๋ช…์˜ ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ LLaMA ๊ธฐ๋ฐ˜ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ์„ ํ™œ์šฉํ•ด ๋ณ‘๋ณ€ ํŠน์„ฑ์„ ์ถ”์ถœํ•˜๊ณ , ACR 2017 ๊ธฐ์ค€์— ๋”ฐ๋ผ ๋ณ‘๋ณ€์„ ๋ถ„๋ฅ˜ํ•œ ๋’ค ๋‹ค๋ณ€๋Ÿ‰ ์ดํ•ญ ํšŒ๊ท€๋ถ„์„์„ ์‹œํ–‰ํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ ์ ์ ˆํ•œ ์ถ”์  ๊ด€์ฐฐ ์™„๋ฃŒ์œจ์€ 36.3%์— ๋ถˆ๊ณผํ•˜์˜€์œผ๋ฉฐ, ์˜์–ด๊ฐ€ ๋ชจ๊ตญ์–ด๊ฐ€ ์•„๋‹Œ ๊ฒฝ์šฐ ๋ฐ ์‚ฌํšŒ๊ฒฝ์ œ์ ์œผ๋กœ ์ทจ์•ฝํ•œ ์ง€์—ญ ๊ฑฐ์ฃผ๊ฐ€ ์ถ”์  ์‹คํŒจ์™€ ์œ ์˜ํ•˜๊ฒŒ ์—ฐ๊ด€๋œ ๋ฐ˜๋ฉด, ๋†’์€ ๋™๋ฐ˜ ์งˆํ™˜ ์ง€์ˆ˜(CCI โ‰ฅ 2)์™€ 65์„ธ ๋ฏธ๋งŒ์˜ ๊ณ ์œ„ํ—˜ ๋ณ‘๋ณ€์€ ์ถ”์  ๊ด€์ฐฐ ์™„๋ฃŒ์™€ ๊ด€๋ จ์ด ์žˆ์—ˆ๋‹ค.
Added: 2026-05-03 00:00View โ†—

48Predicting molecular types of adult-type diffuse gliomas based on MRI reports with large language models.

2026-05European radiologyโญ Q1DOI 10.1007/s00330-025-12211-x
OBJECTIVE

To evaluate the performance of large language models (LLMs) in predicting molecular types of adult-type diffuse gliomas according to the 2021 WHO classification using MRI radiology reports.

METHODS

This retrospective study included 2169 patients diagnosed with adult-type diffuse gliomas (294 oligodendrogliomas, 295 IDH-mutant astrocytomas, and 1580 IDH-wildtype glioblastomas) between July 2005 and March 2024 from four hospitals in Asia and Europe. Seven proprietary and open-source LLMs were assessed: GPT-4o-mini, GPT-4.1-mini, Llama 3.1 8B, Llama 3.1 70B, Qwen2.5 7B, Deepseek-r1 8B, and Mistal 7B. The performance of LLMs in classifying molecular types was compared based on the provision of relevant knowledge of glioma imaging findings (knowledge-based vs. naรฏve prompt). The impact of radiologists' subspecialization in neuro-oncology, report quality, and reporting language on LLMs' performance was also evaluated.

RESULTS

LLMs achieved significantly higher (naรฏve vs. knowledge-based; GPT-4o-mini, 77.0% vs. 79.1%, pโ€‰<โ€‰0.001; Qwen2.5 7B, 75.9% vs. 79.5%, pโ€‰<โ€‰0.001; Deepseek-r1 8B, 66.0% vs. 73.2%, pโ€‰<โ€‰0.001) or comparable accuracy (GPT-4.1-mini, 78.7% vs. 78.6%; Llama 3.1 70B, 78.0% vs. 78.1%; Mistral 7B, 58.4% vs. 57.4%) using knowledge-based prompt compared to naรฏve prompt, except for Llama 3.1 8B (65.4% vs. 44.6%, pโ€‰<โ€‰0.001). Differences in accuracy were more pronounced in smaller-sized LLMs. Additionally, the accuracy was significantly higher with reports by neuro-oncology specialists and high-quality reports in all LLMs (pโ€‰<โ€‰0.001).

CONCLUSION

LLMs may provide preoperative information on the tumor types of adult-type diffuse gliomas from MRI reports by providing relevant knowledge in the prompt. Informative and descriptive reports could further enhance LLMs' performance. KEY POINTS: Question Our study aimed to evaluate large language models' (LLMs) ability to efficiently predict molecular types of adult-type diffuse gliomas according to the 2021 WHO classification. Findings Larger models generally showed better accuracy and were less sensitive to domain-specific knowledge. Their performance improved when using high-quality, longer reports or reports by neuro-oncology specialists. Clinical relevance These findings highlight the potential role of LLMs in predicting glioma molecular types, underscoring the importance of informative and descriptive reports in enhancing their performance.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 2021๋…„ WHO ๋ถ„๋ฅ˜์— ๋”ฐ๋ฅธ ์„ฑ์ธํ˜• ๋ฏธ๋งŒ์„ฑ ์‹ ๊ฒฝ๊ต์ข…์˜ ๋ถ„์ž ์œ ํ˜•์„ MRI ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์˜ˆ์ธกํ•  ์ˆ˜ ์žˆ๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์•„์‹œ์•„ ๋ฐ ์œ ๋Ÿฝ 4๊ฐœ ๋ณ‘์›์—์„œ ์ˆ˜์ง‘๋œ 2,169๋ช…์˜ ํ™˜์ž ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•˜์—ฌ GPT-4o-mini, Llama, Qwen ๋“ฑ 7์ข…์˜ LLM ์„ฑ๋Šฅ์„ ์‹ ๊ฒฝ์ข…์–‘ํ•™ ๊ด€๋ จ ์ง€์‹ ์ œ๊ณต ์—ฌ๋ถ€(์ง€์‹ ๊ธฐ๋ฐ˜ vs. ๋‹จ์ˆœ ํ”„๋กฌํ”„ํŠธ)์— ๋”ฐ๋ผ ๋น„๊ต ๋ถ„์„ํ•˜์˜€๋‹ค. ์ง€์‹ ๊ธฐ๋ฐ˜ ํ”„๋กฌํ”„ํŠธ ์ ์šฉ ์‹œ ๋Œ€๋ถ€๋ถ„์˜ ๋ชจ๋ธ์—์„œ ๋ถ„๋ฅ˜ ์ •ํ™•๋„๊ฐ€ ์œ ์˜ํ•˜๊ฒŒ ํ–ฅ์ƒ๋˜์—ˆ์œผ๋ฉฐ(์ตœ๋Œ€ 79.5%), ์‹ ๊ฒฝ์ข…์–‘ ์ „๋ฌธ ์˜์ƒ์˜ํ•™๊ณผ ์˜์‚ฌ๊ฐ€ ์ž‘์„ฑํ•œ ๊ณ ํ’ˆ์งˆ ๋ณด๊ณ ์„œ๋ฅผ ํ™œ์šฉํ•  ๊ฒฝ์šฐ ๋ชจ๋“  ๋ชจ๋ธ์—์„œ ์ •ํ™•๋„๊ฐ€ ์ถ”๊ฐ€๋กœ ํ–ฅ์ƒ๋˜์–ด, ์ƒ์„ธํ•˜๊ณ  ์ •๋ณด๊ฐ€ ํ’๋ถ€ํ•œ ํŒ๋…๋ฌธ์ด LLM์˜ ์ˆ˜์ˆ  ์ „ ์‹ ๊ฒฝ๊ต์ข… ๋ถ„์ž ์œ ํ˜• ์˜ˆ์ธก ์„ฑ๋Šฅ์„ ํฌ๊ฒŒ ๋†’์ผ ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-05-03 00:00View โ†—

49ONCO-RADS-guided Large Language Models for Extraction and Classification of Incidental Findings on Whole-Body Imaging Reports.

2026-05Radiology. Imaging cancerโญ Q1DOI 10.1148/rycan.250484

Purpose To evaluate large language model (LLM)-based strategy performance for extraction and classification of incidental findings from whole-body (WB) imaging reports, particularly strategies incorporating Oncologically Relevant Findings Reporting and Data System (ONCO-RADS). Materials and Methods In this retrospective bicenter study, authors included all WB MRI reports from January 2016 to December 2023 at a referral center (internal dataset). Two observers extracted all incidental findings, and patient records were used to confirm final diagnoses. First, authors evaluated ONCO-RADS performance and the reproducibility of its incidental finding classifications by six radiologists. Then, authors evaluated the accuracy of three LLM-based strategies: (a) a fine-tuned DeBERTa/medical named entity recognition (NER) model; (b) zero-shot LLMs (ChatGPT-o1 [OpenAI], Gemini-2.5-Pro [Google]); and (c) reference-guided prompting of these LLMs using ONCO-RADS. Authors then expanded these strategies to an external dataset of 605 reports with multiple imaging techniques (405 WB MRI; 100 fluorodeoxyglucose PET/CT; and 100 chest-abdomen-pelvis CT acquisitions) from January 2022 to January 2025. Results The internal dataset included 823 patients (mean age, 63.7 years ยฑ 11.7 [SD]; 457 male patients) with 1488 WB MRI reports. The average interobserver reproducibility of ONCO-RADS incidental finding classifications was excellent (Cohen ฮบ, 0.87). The per-report accuracies of ONCO-RADS-guided LLMs (95.6% [151 of 158] and 86.7% [137 of 158] for ChatGPT-o1 and Gemini-2.5-Pro, respectively) were higher than those of the medical NER (69.0% [109 of 158]) and zero-shot LLMs (57.0% [90 of 158] and 70.9% [112 of 158] for ChatGPT-o1 and Gemini-2.5-Pro, respectively) (P < .001). In the external test set (mean age, 60.6 years ยฑ 12.9; 330 male patients), the per-report accuracies of ONCO-RADS-guided ChatGPT-o1 (83.5% [505 of 605]) and Gemini-2.5-Pro (82.0% [496 of 605]) were higher than those of the models without ONCO-RADS prompting (63.1% [382 of 605] and 61.2% [370 of 605], respectively) and the medical NER (55.7% [337 of 605]) (P < .001). Conclusion Reference-guided prompting of the LLMs ChatGPT-o1 and Gemini-2.5-Pro with ONCO-RADS improved their performance in extracting and classifying incidental findings on WB imaging reports compared with zero-shot prompting and medical NER. Keywords: Large Language Models, Incidental Findings, Whole-Body MRI Supplemental material is available for this article. ยฉ RSNA, 2026.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์ „์‹ (whole-body) ์˜์ƒ ํŒ๋…๋ฌธ์—์„œ ์šฐ๋ฐœ์  ์†Œ๊ฒฌ(incidental finding)์„ ์ถ”์ถœยท๋ถ„๋ฅ˜ํ•˜๋Š” ๋ฐ ์žˆ์–ด ONCO-RADS ๊ธฐ๋ฐ˜ ์ฐธ์กฐ ์œ ๋„ ํ”„๋กฌํ”„ํŒ…(reference-guided prompting)์„ ์ ์šฉํ•œ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ๋‚ด๋ถ€ ๋ฐ์ดํ„ฐ์…‹(WB MRI ๋ณด๊ณ ์„œ 1,488๊ฑด)๊ณผ ์™ธ๋ถ€ ๋ฐ์ดํ„ฐ์…‹(605๊ฑด, WB MRIยทFDG PET/CTยทํ‰๋ณต๋ถ€ CT ํฌํ•จ)์„ ๋Œ€์ƒ์œผ๋กœ ๋ฏธ์„ธ์กฐ์ • NER ๋ชจ๋ธ, ์ œ๋กœ์ƒท LLM(ChatGPT-o1, Gemini-2.5-Pro), ๊ทธ๋ฆฌ๊ณ  ONCO-RADS ์ฐธ์กฐ ์œ ๋„ LLM์˜ ์„ฑ๋Šฅ์„ ๋น„๊ต ๋ถ„์„ํ•˜์˜€๋‹ค. ONCO-RADS ์œ ๋„ ChatGPT-o1์€ ๋‚ด๋ถ€ ๋ฐ์ดํ„ฐ์…‹์—์„œ 95.6%, ์™ธ๋ถ€ ๋ฐ์ดํ„ฐ์…‹์—์„œ 83.5%์˜ ๋ณด๊ณ ์„œ๋ณ„ ์ •ํ™•๋„๋ฅผ ๋‹ฌ์„ฑํ•˜์—ฌ ์ œ๋กœ์ƒท ํ”„๋กฌํ”„ํŒ… ๋ฐ ์˜๋ฃŒ NER ๋ชจ๋ธ์— ๋น„ํ•ด ์œ ์˜ํ•˜๊ฒŒ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ(P < .001), ํ‘œ์ค€ํ™”๋œ ๋ณด๊ณ  ์ฒด๊ณ„๋ฅผ LLM ํ”„๋กฌํ”„ํŒ…์— ํ†ตํ•ฉํ•˜๋Š” ๊ฒƒ์ด ์ž„์ƒ ์˜์ƒ ํŒ๋…๋ฌธ์˜ ์ž๋™ํ™” ๋ถ„์„์— ํšจ๊ณผ์ ์ž„์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

50Objective quality assessment of neuroradiology reports using large language models.

2026-05Clinical radiology๐Ÿ”ท Q2DOI 10.1016/j.crad.2026.107291
OBJECTIVE

Radiology reports are essential for clinical decision-making and must meet standards of clarity, completeness, and diagnostic accuracy. However, quality assessment is often subjective, time-consuming, and dependent on expert reviewers. Large language models (LLMs) offer a promising alternative for automating this process. This study evaluates whether LLMs can objectively and consistently assess the formal and diagnostic quality of neuroradiology reports.

METHODS

We analysed 277 neuroradiology reports, originally authored by 10 radiologists and subsequently evaluated by a different radiologist to assess formal and diagnostic quality. Reports were annotated using six quality criteria: diagnostic discrepancy, report structure, clarity, completeness, grading/staging system use, and recommendation of additional tests. Three locally deployed LLMs (LLaMA 3.2, DeepSeek R1:7B, and Gemma 3:4B) were evaluated for their ability to classify reports according to these criteria.

RESULTS

LLMs showed strong performance in recognising positive categories, with DeepSeek achieving the highest accuracy in report structure (70.40%) and test recommendations (68.59%). However, performance declined significantly in detecting issues such as report completeness or diagnostic discrepancies, reflected in low F1-scores for those categories.

CONCLUSION

While current models have limitations in identifying subtle errors or complex clinical nuances, LLMs may assist in automating aspects of radiology report quality control. With further refinement, they could help improve consistency and support clinical decision-making.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์‹ ๊ฒฝ๋ฐฉ์‚ฌ์„ ๊ณผ ๋ณด๊ณ ์„œ์˜ ํ˜•์‹์ ยท์ง„๋‹จ์  ํ’ˆ์งˆ์„ ๊ฐ๊ด€์ ์ด๊ณ  ์ผ๊ด€๋˜๊ฒŒ ํ‰๊ฐ€ํ•  ์ˆ˜ ์žˆ๋Š”์ง€ ๊ฒ€ํ† ํ•˜์˜€๋‹ค. 10๋ช…์˜ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ๊ฐ€ ์ž‘์„ฑํ•œ 277๊ฑด์˜ ์‹ ๊ฒฝ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ, LLaMA 3.2, DeepSeek R1:7B, Gemma 3:4B ์„ธ ๊ฐ€์ง€ LLM์„ ํ™œ์šฉํ•˜์—ฌ ์ง„๋‹จ ๋ถˆ์ผ์น˜, ๋ณด๊ณ ์„œ ๊ตฌ์กฐ, ๋ช…ํ™•์„ฑ, ์™„๊ฒฐ์„ฑ, ๋“ฑ๊ธ‰/๋ณ‘๊ธฐ ์‹œ์Šคํ…œ ์‚ฌ์šฉ, ์ถ”๊ฐ€ ๊ฒ€์‚ฌ ๊ถŒ๊ณ  ๋“ฑ 6๊ฐ€์ง€ ํ’ˆ์งˆ ๊ธฐ์ค€์œผ๋กœ ํ‰๊ฐ€๋ฅผ ์ˆ˜ํ–‰ํ•˜์˜€๋‹ค. LLM์€ ๋ณด๊ณ ์„œ ๊ตฌ์กฐ ๋ฐ ์ถ”๊ฐ€ ๊ฒ€์‚ฌ ๊ถŒ๊ณ ์™€ ๊ฐ™์€ ์–‘์„ฑ ๋ฒ”์ฃผ ์ธ์‹์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋‚˜, ๋ณด๊ณ ์„œ ์™„๊ฒฐ์„ฑ ๋ถ€์กฑ์ด๋‚˜ ์ง„๋‹จ ๋ถˆ์ผ์น˜ ๊ฐ™์€ ๋ฏธ๋ฌ˜ํ•œ ์˜ค๋ฅ˜ ํƒ์ง€์—์„œ๋Š” ์„ฑ๋Šฅ์ด ํ˜„์ €ํžˆ ์ €ํ•˜๋˜์–ด, ํ˜„์žฌ ์ˆ˜์ค€์—์„œ๋Š” ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ํ’ˆ์งˆ ๊ด€๋ฆฌ์˜ ๋ถ€๋ถ„์  ์ž๋™ํ™” ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

51Automated identification of radiotherapy treatment sites from unstructured physician notes.

2026-04Journal of applied clinical medical physicsโญ Q1DOI 10.1002/acm2.70558
OBJECTIVE

Ambiguous or incomplete documentation is a recurrent bottleneck in radiation oncology workflows, leading to inefficiencies in communication and potential treatment delays. Large language models (LLMs) pose a solution to addressing these ambiguities without added burden to clinical staff. We aim to assess the effectiveness of Meta's open-source Llama 3.3 model in using physician consultation notes to isolate and classify anatomical treatment sites and create helpful extractive summaries for each patient.

METHODS

Semi-structured interviews with five radiation therapists revealed that CT simulation orders lack the necessary details to acquire the appropriate image. A retrospective cohort of 100 patient notes was used for iterative prompt engineering. The final model was evaluated on an independent test cohort of 52 patient notes. The LLM's accuracy in identifying the treatment site was benchmarked against two human observers (a medical physicist and a physician) as well as the final delivered treatment plan (ground truth). The helpfulness and accuracy of the AI-generated summaries were also rated by both observers on a 5-point Likert scale.

RESULTS

Llama 3.3 achieved a weighted accuracy of 94.2% [95%CI: 89.4%-98.1%] when compared to sites isolated by either observer. When compared to the sites isolated from the retrospectively delivered plans, the model reached a weighted accuracy of 92.3% [95% CI: 87.5%-97.1%]. The model classified the anatomical sites with a weighted accuracy of 96.2% [95%CI: 87.0% -98.9%]. The AI-generated summaries were highly rated by both observers (Observer 1: 4.96 [95%CI: 4.87-5.00] and Observer 2: 4.58 [95% CI: 4.38-4.73]).

CONCLUSION

This pilot study provides foundational evidence that LLMs can classify data with high accuracy, achieve benchmarks comparable to human experts when isolating anatomical treatment sites, and produce clinically helpful summaries. Our results suggest that LLMs can be effectively integrated to streamline complex radiotherapy workflows in the clinic.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„  ์ข…์–‘ํ•™ ์›Œํฌํ”Œ๋กœ์šฐ์—์„œ ๋ฐ˜๋ณต์ ์œผ๋กœ ๋ฐœ์ƒํ•˜๋Š” ๋ฌธ์„œ ๋ถˆ์™„์ „์„ฑ ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๊ธฐ ์œ„ํ•ด, Meta์˜ ์˜คํ”ˆ์†Œ์Šค ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์ธ Llama 3.3์„ ํ™œ์šฉํ•˜์—ฌ ์˜์‚ฌ ํ˜‘์ง„ ๋…ธํŠธ๋กœ๋ถ€ํ„ฐ ํ•ด๋ถ€ํ•™์  ์น˜๋ฃŒ ๋ถ€์œ„๋ฅผ ์ž๋™์œผ๋กœ ์‹๋ณ„ยท๋ถ„๋ฅ˜ํ•˜๊ณ  ํ™˜์ž๋ณ„ ์š”์•ฝ๋ฌธ์„ ์ƒ์„ฑํ•˜๋Š” ์‹œ์Šคํ…œ์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 100๋ช…์˜ ํ›„ํ–ฅ์  ํ™˜์ž ๋…ธํŠธ๋กœ ํ”„๋กฌํ”„ํŠธ ์—”์ง€๋‹ˆ์–ด๋ง์„ ์ˆ˜ํ–‰ํ•œ ํ›„ 52๋ช…์˜ ๋…๋ฆฝ ๊ฒ€์ฆ ์ฝ”ํ˜ธํŠธ์— ์ ์šฉํ•œ ๊ฒฐ๊ณผ, ์น˜๋ฃŒ ๋ถ€์œ„ ์‹๋ณ„ ์ •ํ™•๋„๋Š” ์ธ๊ฐ„ ๊ด€์ฐฐ์ž ๋Œ€๋น„ 94.2%, ์‹ค์ œ ์น˜๋ฃŒ ๊ณ„ํš ๋Œ€๋น„ 92.3%์˜ ๊ฐ€์ค‘ ์ •ํ™•๋„๋ฅผ ๋‹ฌ์„ฑํ•˜์˜€๋‹ค. AI ์ƒ์„ฑ ์š”์•ฝ๋ฌธ์— ๋Œ€ํ•œ ์ž„์ƒ์  ์œ ์šฉ์„ฑ ํ‰๊ฐ€์—์„œ๋„ ๋‘ ๊ด€์ฐฐ์ž ๋ชจ๋‘ 5์  ๋งŒ์  ๊ธฐ์ค€ 4.58~4.96์˜ ๋†’์€ ์ ์ˆ˜๋ฅผ ๋ถ€์—ฌํ•˜์—ฌ, LLM์ด ์ „๋ฌธ๊ฐ€ ์ˆ˜์ค€์˜ ์ •ํ™•๋„๋กœ ๋ฐฉ์‚ฌ์„  ์น˜๋ฃŒ ์›Œํฌํ”Œ๋กœ์šฐ๋ฅผ ํšจ๊ณผ์ ์œผ๋กœ ์ง€์›ํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

52Comparison of AI-generated radiology impressions: a multi-stakeholder evaluation.

2026-04NPJ digital medicineโญ Q1DOI 10.1038/s41746-026-02586-6

A retrospective, blinded evaluation of 200 oncologic computed tomography reports compared original radiologist-authored impressions, impressions generated by a custom domain-specific AI model fine-tuned on institutional data, and impressions generated by a general-purpose large language model. Ten clinicians, including original radiologists (nโ€‰=โ€‰4), independent radiologists (nโ€‰=โ€‰3), and oncologists (nโ€‰=โ€‰3), rated impressions for completeness, correctness, conciseness, clarity, clinical utility, and patient harm. Original and independent radiologists assigned lower preference to generic model impressions (Cohen's h 1.04-1.22 and 0.66-0.69, pโ€‰<โ€‰0.001). Original radiologists slightly preferred their own impressions to the custom model (hโ€‰=โ€‰0.18, pโ€‰=โ€‰0.0716), while independent radiologists showed no preference (hโ€‰=โ€‰-0.03, pโ€‰=โ€‰0.78). Oncologists demonstrated no significant preference among impression types (hโ€‰=โ€‰0.04-0.12, all pโ€‰>โ€‰0.20). Custom model impressions achieved near parity with human impressions; original radiologists rated their own impressions slightly more complete (rโ€‰=โ€‰0.22, pโ€‰=โ€‰0.0016). Generic model impressions were longer (75.1โ€‰ยฑโ€‰20.4 words), slightly more complete (rโ€‰=โ€‰0.18-0.39, pโ€‰<โ€‰0.001-0.01), but significantly less concise (rโ€‰=โ€‰0.85-0.87, pโ€‰<โ€‰0.001). Patient harm ratings were uniformly low (likelihood 1.01-1.14; extent 1.05-1.21). Inter-rater reliability ranged from -0.09 to 0.67 (ฮฑโ€‰=โ€‰0.67 conciseness; ฮฑโ€‰=โ€‰-0.09-0.03 clinical utility/correctness).

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์ข…์–‘ํ•™์  ๋ณต๋ถ€ ์ „์‚ฐํ™”๋‹จ์ธต์ดฌ์˜(CT) ํŒ๋…๋ฌธ 200๊ฑด์„ ๋Œ€์ƒ์œผ๋กœ, ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜๊ฐ€ ์ง์ ‘ ์ž‘์„ฑํ•œ ์ธ์ƒ(impression)๊ณผ ๊ธฐ๊ด€ ๋ฐ์ดํ„ฐ๋กœ ๋ฏธ์„ธ์กฐ์ •๋œ ๋„๋ฉ”์ธ ํŠนํ™” AI ๋ชจ๋ธ ๋ฐ ๋ฒ”์šฉ ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ์ด ์ƒ์„ฑํ•œ ์ธ์ƒ์„ ๋‹ค์ง๊ตฐ ํ‰๊ฐ€์ž(์˜์ƒ์˜ํ•™๊ณผ ์˜์‚ฌ 7๋ช…, ์ข…์–‘๋‚ด๊ณผ ์˜์‚ฌ 3๋ช…)๊ฐ€ ๋งน๊ฒ€ ๋ฐฉ์‹์œผ๋กœ ๋น„๊ตยทํ‰๊ฐ€ํ•˜์˜€๋‹ค. ๋„๋ฉ”์ธ ํŠนํ™” ๋ชจ๋ธ์˜ ์ธ์ƒ์€ ์ „๋ฌธ์˜ ์ž‘์„ฑ ์ธ์ƒ๊ณผ ๊ฑฐ์˜ ๋™๋“ฑํ•œ ์ˆ˜์ค€์„ ๋‹ฌ์„ฑํ•œ ๋ฐ˜๋ฉด, ๋ฒ”์šฉ ๋ชจ๋ธ์˜ ์ธ์ƒ์€ ์™„๊ฒฐ์„ฑ์€ ๋‹ค์†Œ ๋†’์•˜์œผ๋‚˜ ๊ฐ„๊ฒฐ์„ฑ์ด ํ˜„์ €ํžˆ ๋‚ฎ์•„ ์˜์ƒ์˜ํ•™๊ณผ ์˜์‚ฌ๋“ค๋กœ๋ถ€ํ„ฐ ์œ ์˜ํ•˜๊ฒŒ ๋‚ฎ์€ ์„ ํ˜ธ๋„๋ฅผ ๋ฐ›์•˜๋‹ค(Cohen's h 0.66โ€“1.22, p < 0.001). ์ข…์–‘๋‚ด๊ณผ ์˜์‚ฌ๋“ค์€ ์ธ์ƒ ์œ ํ˜• ๊ฐ„ ์œ ์˜๋ฏธํ•œ ์„ ํ˜ธ ์ฐจ์ด๋ฅผ ๋ณด์ด์ง€ ์•Š์•˜์œผ๋ฉฐ, ๋ชจ๋“  ๊ตฐ์—์„œ ํ™˜์ž ์œ„ํ•ด ๊ฐ€๋Šฅ์„ฑ ํ‰์ ์€ ์ผ๊ด€๋˜๊ฒŒ ๋‚ฎ์•„ ์ž„์ƒ์  ์•ˆ์ „์„ฑ์€ ์–‘ํ˜ธํ•œ ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ๋‹ค.
Added: 2026-04-21 16:11View โ†—

53Context-Aware Sentence Classification of Radiology Reports Using Synthetic Data: Development and Validation Study.

2026-04Journal of medical Internet researchโญ Q1DOI 10.2196/86365
BACKGROUND

Automated structuring of radiology reports is essential for data utilization and the development of medical artificial intelligence models. However, manual annotation by experts is labor-intensive, and processing real clinical data through commercial large language models (LLMs) presents significant privacy risks. These challenges are particularly pronounced for non-English languages like Japanese, where specialized medical corpora are scarce. While synthetic data generation offers a potential privacy-preserving alternative, its effectiveness in capturing complex clinical nuances-such as negation and contextual dependencies-to train robust classification models without any real-world training data has not been fully established.

OBJECTIVE

This study aimed to develop a context-aware sentence classification model for Japanese radiology reports using an entirely synthetic training pipeline, thereby eliminating reliance on real-world clinical data during the development phase. Furthermore, we sought to evaluate the generalizability of this approach by validating the model's performance on diverse, multi-institutional, real-world reports.

METHODS

Japanese radiology reports (n=3104) were generated using GPT-4.1 and automatically annotated at the sentence level into 4 categories (background, positive finding, negative finding, and continuation) using GPT-4.1-mini. The synthetic data were partitioned into training (n=2670), validation (n=334), and test (n=100) sets. We fine-tuned several models, including lightweight local LLMs (Qwen3 and Llama 3.2 series) using low-rank adaptation and Japanese text classification models (Bidirectional Encoder Representations from Transformers [BERT]-base Japanese v3, Japanese Medical Robustly Optimized BERT Pretraining Approach [JMedRoBERTa]-base, and ModernBERT-Ja-130M). External validation was performed using 280 real-world reports (3477 sentences) from 7 institutions in the Japan Medical Image Database, with ground-truth labels established by board-certified radiologists. Evaluation metrics included accuracy, macro-averaged F1 (macro F1) score, and positive predictive value for positive findings (PPV_1).

RESULTS

All models achieved high performance on the synthetic test set (accuracy: 0.938-0.951; macro F1-score: 0.924-0.940). Overall performance declined on the external validation dataset (accuracy: 0.783-0.813; macro F1-score: 0.761-0.790), reflecting distributional differences between synthetic and real-world reports; however, PPV_1 remained stable and high across datasets (eg, 0.957 on the synthetic test set vs 0.952 on the external validation dataset for Qwen3 [4B]). Parsing errors occurred in LLM-based approaches (19-260 sentences, 0.55%-7.48% in the external dataset).

CONCLUSION

This study demonstrates the feasibility of developing context-aware sentence classification models for Japanese radiology reports using a training pipeline based entirely on synthetic data. The stability of PPV_1 indicates that the models successfully captured the essential clinical terminology and linguistic patterns required to identify positive findings in real-world reports, despite the observed performance degradation during external validation. This approach substantially reduces manual annotation requirements and privacy risks, providing a scalable foundation for constructing structured radiology datasets to support the development of clinically relevant medical artificial intelligence models.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ฐœ์ธ์ •๋ณด ๋ณดํ˜ธ ๋ฌธ์ œ์™€ ์ˆ˜์ž‘์—… ์ฃผ์„์˜ ๋ถ€๋‹ด์„ ์ค„์ด๊ธฐ ์œ„ํ•ด, ์‹ค์ œ ์ž„์ƒ ๋ฐ์ดํ„ฐ ์—†์ด GPT-4.1๋กœ ์ƒ์„ฑํ•œ ํ•ฉ์„ฑ ์ผ๋ณธ์–ด ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ(3,104๊ฑด)๋งŒ์„ ํ™œ์šฉํ•˜์—ฌ ๋ฌธ์žฅ ์ˆ˜์ค€์˜ ๋งฅ๋ฝ ์ธ์‹ ๋ถ„๋ฅ˜ ๋ชจ๋ธ์„ ๊ฐœ๋ฐœํ•˜์˜€๋‹ค. ๊ฒฝ๋Ÿ‰ LLM(Qwen3, Llama 3.2 ๊ณ„์—ด) ๋ฐ BERT ๊ธฐ๋ฐ˜ ์ผ๋ณธ์–ด ์˜๋ฃŒ ๋ชจ๋ธ์„ ๋ฏธ์„ธ์กฐ์ •ํ•œ ํ›„, 7๊ฐœ ๊ธฐ๊ด€์˜ ์‹ค์ œ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ 280๊ฑด(3,477๋ฌธ์žฅ)์œผ๋กœ ์™ธ๋ถ€ ๊ฒ€์ฆ์„ ์ˆ˜ํ–‰ํ•˜์˜€๋‹ค. ํ•ฉ์„ฑ ํ…Œ์ŠคํŠธ์…‹์—์„œ๋Š” ๋†’์€ ์„ฑ๋Šฅ(์ •ํ™•๋„ 0.938โ€“0.951, macro F1 0.924โ€“0.940)์ด ํ™•์ธ๋˜์—ˆ์œผ๋ฉฐ, ์™ธ๋ถ€ ๊ฒ€์ฆ์—์„œ ์ „๋ฐ˜์  ์„ฑ๋Šฅ์€ ๋‹ค์†Œ ์ €ํ•˜๋˜์—ˆ์œผ๋‚˜ ์–‘์„ฑ ์†Œ๊ฒฌ์— ๋Œ€ํ•œ ์–‘์„ฑ ์˜ˆ์ธก๋„(PPV)๋Š” 0.952๋กœ ์•ˆ์ •์ ์œผ๋กœ ์œ ์ง€๋˜์–ด, ํ•ฉ์„ฑ ๋ฐ์ดํ„ฐ ๊ธฐ๋ฐ˜ ํŒŒ์ดํ”„๋ผ์ธ์ด ์ž„์ƒ์ ์œผ๋กœ ํ™œ์šฉ ๊ฐ€๋Šฅํ•œ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ๊ตฌ์กฐํ™” ๋ชจ๋ธ ๊ฐœ๋ฐœ์— ์‹ค์šฉ์ ์ธ ๋Œ€์•ˆ์ž„์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

54Improving patient understanding of oncology imaging: radiologist and patient evaluation of summarised versus full-length AI-simplified reports from a tertiary cancer centre.

2026-04Cancer imaging : the official publication of the International Cancer Imaging Societyโญ Q1DOI 10.1186/s40644-026-01031-x
Added: 2026-04-21 16:11View โ†—

55Initial Insights Into an Institutional Secure Large Language Model for Magnetic Resonance Imaging Examination Requests: Retrospective Study.

2026-04Journal of medical Internet researchโญ Q1DOI 10.2196/82579
BACKGROUND

Incomplete clinical details on magnetic resonance imaging (MRI) examination requests (MERs) can lead to suboptimal protocol selection. An institutional secure large language model (sLLM) with access to manually retrieved salient data from the electronic medical record (EMR) may improve request completeness and protocol accuracy across multiple MRI subspecialties.

OBJECTIVE

The objective of this study was to compare clinician MERs with sLLM-augmented MERs for information quality and to evaluate the protocoling accuracy of the sLLM versus board-certified radiologists across body, musculoskeletal, and neuroradiology MRI.

METHODS

This retrospective study included 608 random outpatient MRI examinations performed between September 2023 and July 2024 (body 206, musculoskeletal 203, neuroradiology 199). The cohort comprised 528 patients (mean 51.2 years, SD 19.2; range 4-93; n=279, 52.8% women, n=249, 47.2% men). MERs without EMR access were excluded. A privately hosted Anthropic Claude 3.5 model (temperature 0) augmented each MER with manually retrieved salient EMR data and, via rule-based parsing, mapped the extracted elements onto predefined institutional criteria to recommend region or coverage and contrast use. Two experienced radiologists established a consensus reference standard. Two board-certified general radiologists (Rad 3 and Rad 4) and the sLLM were compared with this standard. Clinical information quality was graded using the Reason-for-Exam Imaging Reporting and Data System (RI-RADS). Interrater reliability was quantified with Gwet AC1. Paired accuracies were compared with the McNemar test to determine whether there was a statistically significant difference.

RESULTS

Interreader agreement for RI-RADS was almost perfect for sLLM-augmented MERs (AC1 0.97, 95% CI 0.94-0.99) and moderate for clinician MERs (AC1 0.43, 95% CI 0.34-0.52). Limited or deficient clinical information (RI-RADS C/D) fell to 0% to 0.7% (0/608 to 4/608) with sLLM augmentation vs 4.1% to 20.4% (25/608 to 124/608) for clinician MERs. Overall protocol accuracy was 93.1% (566/608; 95% CI 89.6-96.6) for the sLLM, 91.4% (556/608; 95% CI 87.6-95.3) for Rad 3, and 92.1% (560/608; 95% CI 88.4-95.8) for Rad 4 (sLLM vs Rad 3 P=.23 vs Rad 4 P=.40). Region or coverage accuracy was similar (sLLM: 579/608, 95.2%; Rad 3: 585/608, 96.2%; Rad 4: 573/608, 94.2%; P=.46 and P=.36). Contrast decisions were more accurate using the sLLM at 94.4% (574/608; 95% CI 91.3-97.5) vs Rad 3 at 92.1% (560/608; 95% CI 88.4-95.8; P=.027) and were not significantly different to Rad 4 at 92.9% (565/608; 95% CI 89.4-96.4; P=.16). Subspecialty analyses showed similar patterns, with the sLLM outperforming Rad 4 for musculoskeletal MRI contrast decisions (96.6% vs 91.1%; P=.006) and matching readers elsewhere. Manual review indicated that sLLM improvements arose from EMR details not listed on the MER (infection/inflammation, tumor history, prior surgery). No clinically significant hallucinations were identified in a manual review of discordant cases.

CONCLUSION

Across body, musculoskeletal, and neuroradiology MRI, sLLM-augmented examination requests improved clinical context and enhanced contrast selection while demonstrating accuracy comparable to general radiologists for region or coverage. Integrating sLLMs into routine vetting workflows may reduce manual workload in protocol selection for more efficient, standardized protocoling.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ํ›„ํ–ฅ์  ์—ฐ๊ตฌ๋Š” ์ „์ž์˜๋ฌด๊ธฐ๋ก(EMR) ๋ฐ์ดํ„ฐ๋กœ ์ฆ๊ฐ•๋œ ๊ธฐ๊ด€ ๋‚ด ๋ณด์•ˆ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(sLLM, Claude 3.5)์ด MRI ๊ฒ€์‚ฌ ์˜๋ขฐ์„œ์˜ ์ž„์ƒ ์ •๋ณด ์งˆ๊ณผ ํ”„๋กœํ† ์ฝœ ์„ ํƒ ์ •ํ™•๋„๋ฅผ ํ–ฅ์ƒ์‹œํ‚ฌ ์ˆ˜ ์žˆ๋Š”์ง€ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด, ์ฒด๋ถ€ยท๊ทผ๊ณจ๊ฒฉยท์‹ ๊ฒฝ๋ฐฉ์‚ฌ์„  ๋ถ„์•ผ ์™ธ๋ž˜ MRI 608๊ฑด์„ ๋Œ€์ƒ์œผ๋กœ sLLM๊ณผ ์ „๋ฌธ์˜ 2๋ช…์˜ ํŒ๋…์„ ๋น„๊ตํ•˜์˜€๋‹ค. sLLM ์ฆ๊ฐ• ํ›„ ๋ถˆ์ถฉ๋ถ„ํ•œ ์ž„์ƒ ์ •๋ณด(RI-RADS C/D) ๋น„์œจ์ด ์ž„์ƒ์˜ ์˜๋ขฐ์„œ์˜ 4.1~20.4%์—์„œ 0~0.7%๋กœ ๋Œ€ํญ ๊ฐ์†Œํ•˜์˜€์œผ๋ฉฐ, ์ „์ฒด ํ”„๋กœํ† ์ฝœ ์ •ํ™•๋„๋Š” sLLM 93.1%, ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜ 91.4~92.1%๋กœ ํ†ต๊ณ„์ ์œผ๋กœ ์œ ์˜ํ•œ ์ฐจ์ด๊ฐ€ ์—†์—ˆ๋‹ค. ํŠนํžˆ ์กฐ์˜์ œ ์‚ฌ์šฉ ๊ฒฐ์ •์—์„œ sLLM(94.4%)์ด ์ „๋ฌธ์˜ 1์ธ(92.1%, P=.027)๋ณด๋‹ค ์œ ์˜ํ•˜๊ฒŒ ์šฐ์ˆ˜ํ•˜์—ฌ, sLLM์„ MRI ํ”„๋กœํ† ์ฝœ ๊ฒ€ํ†  ์›Œํฌํ”Œ๋กœ์šฐ์— ํ†ตํ•ฉํ•˜๋ฉด ์ˆ˜์ž‘์—… ๋ถ€๋‹ด์„ ์ค„์ด๊ณ  ํ‘œ์ค€ํ™”๋œ ํ”„๋กœํ† ์ฝœ ์„ ํƒ์— ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

56Multidimensional evaluation of large language models in radiology report readability.

2026-04NPJ digital medicineโญ Q1DOI 10.1038/s41746-026-02589-3

This study systematically investigated the influence of demographic characteristics on the readability of patient-centric radiology reports and compared the performance of different large language models (LLMs) in generating patient-centered reports. Adopting a sequential two-stage design, the research first conducted a retrospective evaluation involving 320 radiology reports followed by a clinical setting validation with 800 patients. Results suggested that all three LLMs significantly improved the readability of radiology reports (Pโ€‰<โ€‰0.05), with DeepSeek-R1 showing potentially superior performance within this specific cohort. Demographic analysis revealed significant interactive effects: higher education and older age (within consistent educational levels) were associated with better comprehension. Clinical setting validation further indicated that reading simplified reports suggesting the potential to significantly improved patients' subjective and objective comprehension while significantly alleviating medical anxiety (Pโ€‰<โ€‰0.05). However, limitations persist, including inconsistent model outputs, missing anatomical details, and comprehension variances driven by demographic factors. Consequently, LLMs should be integrated as auxiliary communication tools for radiologists rather than standalone solutions, necessitating personalized interventions tailored to specific demographic profiles.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ํ™˜์ž ์ค‘์‹ฌ ๋ฐฉ์‚ฌ์„  ํŒ๋…๋ฌธ์˜ ๊ฐ€๋…์„ฑ ํ–ฅ์ƒ์— ๋ฏธ์น˜๋Š” ์˜ํ–ฅ์„ ์ฒด๊ณ„์ ์œผ๋กœ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด, 320๊ฑด์˜ ํ›„ํ–ฅ์  ํŒ๋…๋ฌธ ํ‰๊ฐ€์™€ 800๋ช… ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ ํ•œ ์ž„์ƒ ๊ฒ€์ฆ์˜ 2๋‹จ๊ณ„ ์ˆœ์ฐจ ์„ค๊ณ„๋ฅผ ์ฑ„ํƒํ•˜์˜€๋‹ค. ์„ธ ๊ฐ€์ง€ LLM ๋ชจ๋‘ ๋ฐฉ์‚ฌ์„  ํŒ๋…๋ฌธ์˜ ๊ฐ€๋…์„ฑ์„ ์œ ์˜ํ•˜๊ฒŒ ๊ฐœ์„ ํ•˜์˜€์œผ๋ฉฐ(P < 0.05), ํŠนํžˆ DeepSeek-R1์ด ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€๊ณ , ๊ณ ํ•™๋ ฅ ๋ฐ ๋™์ผ ๊ต์œก ์ˆ˜์ค€ ๋‚ด ๊ณ ๋ น ํ™˜์ž์—์„œ ๋” ๋†’์€ ์ดํ•ด๋„๊ฐ€ ๊ด€์ฐฐ๋˜๋Š” ์ธ๊ตฌํ†ต๊ณ„ํ•™์  ์ƒํ˜ธ์ž‘์šฉ ํšจ๊ณผ๊ฐ€ ํ™•์ธ๋˜์—ˆ๋‹ค. ๋‹จ์ˆœํ™”๋œ ํŒ๋…๋ฌธ์€ ํ™˜์ž์˜ ์ฃผ๊ด€์ ยท๊ฐ๊ด€์  ์ดํ•ด๋„๋ฅผ ์œ ์˜ํ•˜๊ฒŒ ํ–ฅ์ƒ์‹œํ‚ค๊ณ  ์˜๋ฃŒ ๋ถˆ์•ˆ์„ ๊ฒฝ๊ฐ์‹œ์ผฐ์œผ๋‚˜, ๋ชจ๋ธ ์ถœ๋ ฅ์˜ ๋น„์ผ๊ด€์„ฑ, ํ•ด๋ถ€ํ•™์  ์ •๋ณด ๋ˆ„๋ฝ, ์ธ๊ตฌํ†ต๊ณ„ํ•™์  ์š”์ธ์— ๋”ฐ๋ฅธ ์ดํ•ด๋„ ์ฐจ์ด ๋“ฑ์˜ ํ•œ๊ณ„๊ฐ€ ์กด์žฌํ•˜์—ฌ LLM์€ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ์˜ ๋…๋ฆฝ์  ๋Œ€์ฒด ์ˆ˜๋‹จ์ด ์•„๋‹Œ ๋ณด์กฐ์  ํ™˜์ž ์†Œํ†ต ๋„๊ตฌ๋กœ ํ™œ์šฉ๋˜์–ด์•ผ ํ•˜๋ฉฐ ๊ฐœ์ธ ๋งž์ถคํ˜• ์ ‘๊ทผ์ด ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

57Towards Automated FIGO Staging in Radiology: The Role of LLMs in Cervical and Endometrial Cancer.

2026-04Academic radiologyโญ Q1DOI 10.1016/j.acra.2026.01.024

RATIONALE AND

OBJECTIVE

Staging gynecological malignancies is a complex process, and radiologists should be familiar with the evolution of FIGO staging criteria. Large Language Models (LLMs) offer potential to support radiologists by automating classification tasks from free-text MRI reports.

METHODS

We conducted a retrospective study using two curated datasets of pelvic MRI reports from patients with cervical (n = 261, FIGO 2018) and endometrial cancer (n = 555, FIGO 2023). A general-purpose LLM (Cohere Command-A) was evaluated under three prompting strategies (zero-shot, guided, and chain-of-thought [CoT]), using exact stage accuracy, an ordinal FIGO distance metric, and the rate of severe errors. The Cohere Command-A model was chosen for its long-context reasoning, instruction-following capabilities, reproducible fixed version, and secure handling of sensitive clinical data. While alternative LLMs (eg, GPT-4o, Gemini, Llama-3, DeepSeek) could offer complementary insights, access, resources, and compliance constraints limited broader comparisons.

RESULTS

For cervical cancer, CoT prompting achieved the highest accuracy (80.5%) and the lowest FIGO distance, with 23 severe misclassifications (โ‰ฅ2-stage deviation), outperforming guided and zero-shot prompting. For endometrial cancer, all strategies performed appropriately, with CoT again yielding the best results (accuracy, 90.6%) and the lowest number of severe misclassifications (37 cases), compared with guided and zero-shot prompting. In a small subset of cases with no agreement between any prompting strategy and the reference label, manual review showed that only a minority presented potentially suboptimal annotations, suggesting that CoT-based predictions may also help flag doubtful reports.

CONCLUSION

The LLMs used demonstrated strong performance in automatically assigning FIGO stages for cervical and endometrial cancers from MRI reports. Their integration could reduce workload and improve consistency in staging. Further validation is needed before clinical implementation.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์ž๊ถ๊ฒฝ๋ถ€์•” ๋ฐ ์ž๊ถ๋‚ด๋ง‰์•” ํ™˜์ž์˜ ๊ณจ๋ฐ˜ MRI ๋ณด๊ณ ์„œ์—์„œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ํ™œ์šฉํ•˜์—ฌ FIGO ๋ณ‘๊ธฐ ๋ถ„๋ฅ˜๋ฅผ ์ž๋™ํ™”ํ•  ์ˆ˜ ์žˆ๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ž๊ถ๊ฒฝ๋ถ€์•” 261๋ก€(FIGO 2018)์™€ ์ž๊ถ๋‚ด๋ง‰์•” 555๋ก€(FIGO 2023)์˜ MRI ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ Cohere Command-A ๋ชจ๋ธ์— ์ œ๋กœ์ƒท, ๊ฐ€์ด๋“œ, ์‚ฌ๊ณ  ์—ฐ์‡„(Chain-of-Thought, CoT) ์„ธ ๊ฐ€์ง€ ํ”„๋กฌํ”„ํŒ… ์ „๋žต์„ ์ ์šฉํ•˜์—ฌ ๋ณ‘๊ธฐ ์ •ํ™•๋„์™€ ์ค‘์ฆ ์˜ค๋ถ„๋ฅ˜์œจ์„ ๋น„๊ตํ•˜์˜€๋‹ค. CoT ํ”„๋กฌํ”„ํŒ…์ด ์ž๊ถ๊ฒฝ๋ถ€์•”์—์„œ 80.5%, ์ž๊ถ๋‚ด๋ง‰์•”์—์„œ 90.6%์˜ ์ตœ๊ณ  ์ •ํ™•๋„๋ฅผ ๋‹ฌ์„ฑํ•˜๋ฉฐ ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€๊ณ , ์ด๋Š” LLM์ด ๋ฐฉ์‚ฌ์„ ๊ณผ ๋ณด๊ณ ์„œ ๊ธฐ๋ฐ˜ FIGO ๋ณ‘๊ธฐ ์ž๋™ํ™”์— ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์ด ๋†’์Œ์„ ์‹œ์‚ฌํ•˜๋‚˜ ์ž„์ƒ ์ ์šฉ ์ „ ์ถ”๊ฐ€ ๊ฒ€์ฆ์ด ํ•„์š”ํ•˜๋‹ค.
Added: 2026-04-21 16:11View โ†—

58A RRA Perspective on AI and Machine Learning Applications in Radiology: From Experimental to Clinically Viable Solutions.

2026-03Academic radiologyโญ Q1DOI 10.1016/j.acra.2025.11.021

This article is the first in a seven-part Radiology Research Alliance (RRA) review series on emerging technologies in radiology. It examines the role of artificial intelligence and machine learning applications across three domains: diagnostic interpretation, workflow optimization, and report generation. Advances in deep learning, multimodal large language models, and natural language processing have delivered improvements in accuracy, efficiency, and reporting quality. Yet important challenges remain, including variable performance, limited generalizability, and barriers to workflow integration. Current evidence shows that artificial intelligence is most effective when used to augment human expertise, with radiologist-AI collaboration producing the strongest outcomes. This review highlights the transition of artificial intelligence from experimental innovation to a clinically viable technology poised to enhance radiology practice when thoughtfully implemented with appropriate oversight.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
์ด ๋…ผ๋ฌธ์€ ๋ฐฉ์‚ฌ์„ ํ•™์—์„œ ์ธ๊ณต์ง€๋Šฅ(AI) ๋ฐ ๋จธ์‹ ๋Ÿฌ๋‹ ์‘์šฉ ๋ถ„์•ผ๋ฅผ ์ง„๋‹จ ํ•ด์„, ์›Œํฌํ”Œ๋กœ์šฐ ์ตœ์ ํ™”, ๋ณด๊ณ ์„œ ์ƒ์„ฑ์˜ ์„ธ ๊ฐ€์ง€ ์˜์—ญ์— ๊ฑธ์ณ ์ฒด๊ณ„์ ์œผ๋กœ ๊ฒ€ํ† ํ•œ Radiology Research Alliance(RRA) ๋ฆฌ๋ทฐ ์‹œ๋ฆฌ์ฆˆ์˜ ์ฒซ ๋ฒˆ์งธ ๋…ผ๋ฌธ์ด๋‹ค. ๋”ฅ๋Ÿฌ๋‹, ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ, ์ž์—ฐ์–ด ์ฒ˜๋ฆฌ ๊ธฐ์ˆ ์˜ ๋ฐœ์ „์ด ์ง„๋‹จ ์ •ํ™•๋„, ์—…๋ฌด ํšจ์œจ์„ฑ, ๋ณด๊ณ  ํ’ˆ์งˆ ํ–ฅ์ƒ์— ๊ธฐ์—ฌํ•˜์˜€์œผ๋‚˜, ์„ฑ๋Šฅ์˜ ๊ฐ€๋ณ€์„ฑ, ์ผ๋ฐ˜ํ™” ํ•œ๊ณ„, ์›Œํฌํ”Œ๋กœ์šฐ ํ†ตํ•ฉ์˜ ์žฅ๋ฒฝ ๋“ฑ ํ•ด๊ฒฐํ•ด์•ผ ํ•  ๊ณผ์ œ๊ฐ€ ์—ฌ์ „ํžˆ ์กด์žฌํ•œ๋‹ค. ํ˜„์žฌ๊นŒ์ง€์˜ ๊ทผ๊ฑฐ์— ๋”ฐ๋ฅด๋ฉด AI๋Š” ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜์˜ ์ „๋ฌธ์„ฑ์„ ๋ณด์™„ํ•˜๋Š” ํ˜•ํƒœ๋กœ ํ™œ์šฉ๋  ๋•Œ ๊ฐ€์žฅ ํšจ๊ณผ์ ์ด๋ฉฐ, ์ ์ ˆํ•œ ๊ฐ๋… ์ฒด๊ณ„์™€ ํ•จ๊ป˜ ์‹ ์ค‘ํ•˜๊ฒŒ ๋„์ž…๋  ๊ฒฝ์šฐ ์‹คํ—˜์  ํ˜์‹ ์—์„œ ์ž„์ƒ์ ์œผ๋กœ ์‹ค์šฉ ๊ฐ€๋Šฅํ•œ ๊ธฐ์ˆ ๋กœ์˜ ์ „ํ™˜์ด ๊ฐ€๋Šฅํ•จ์„ ๊ฐ•์กฐํ•˜๊ณ  ์žˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

59Adaptive, Privacy-Preserving Small Language Models for Multi-Task Clinical Assistance.

2026-03Journal of imaging informatics in medicineโญ Q1DOI 10.1007/s10278-026-01912-4

The purpose of this study is to evaluate whether a single, fine-tuned SLM can match or exceed the performance of LLMs across diverse clinical tasks, enabling hospitals to build tailored, privacy-preserving, efficient, and deployable language models that do not require managing multiple task-specific systems. We used SLMs of varying sizes and applied low-rank adaptation (LoRA) for fine-tuning across three clinical tasks: (1) medical report labeling, (2) DICOM series description harmonization, and (3) impression generation from findings. These tasks were constructed using two datasets: the public Open-i Indiana University Chest X-ray Dataset and an in-house brain MRI DICOM metadata dataset. We compared single-task SLMs, a multi-task SLM (representing our proposed configuration), and GPT-4o using zero-shot and few-shot prompting. We found OPT-350ย m to be the optimal SLM. In medical report labeling, the multi-task SLM achieved an F1 score of 0.894 compared to additional prompt-engineered GPT-4o's 0.728. In DICOM series description harmonization, the multi-task achieved an accuracy of 0.975 compared to additional prompt-engineered GPT-4o's 0.878. In impression generation from findings, the multi-task SLM achieved an average Likert scale score of 4.39โ€‰ยฑโ€‰1.00, compared to GPT-4o's 3.65โ€‰ยฑโ€‰1.00 (pโ€‰=โ€‰0.0008). This study demonstrates that a single fine-tuned SLM can serve as a general-purpose clinical assistant, offering performance on par with or better than larger models. With lower resource requirements, greater customizability, privacy protection, and strong task generalization, fine-tuning one SLM to support multiple clinical tasks meets the practical demands of clinical AI deployment in both high-resource and resource-limited healthcare settings.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋‹จ์ผ ์†Œํ˜• ์–ธ์–ด ๋ชจ๋ธ(SLM)์„ ์ €์ˆœ์œ„ ์ ์‘(LoRA) ๊ธฐ๋ฒ•์œผ๋กœ ๋ฏธ์„ธ ์กฐ์ •ํ•˜์—ฌ ์˜๋ฃŒ ๋ณด๊ณ ์„œ ๋ ˆ์ด๋ธ”๋ง, DICOM ์‹œ๋ฆฌ์ฆˆ ์„ค๋ช… ํ‘œ์ค€ํ™”, ์†Œ๊ฒฌ์œผ๋กœ๋ถ€ํ„ฐ ์ธ์ƒ ์ƒ์„ฑ ๋“ฑ ์„ธ ๊ฐ€์ง€ ์ž„์ƒ ๊ณผ์ œ๋ฅผ ๋™์‹œ์— ์ˆ˜ํ–‰ํ•  ์ˆ˜ ์žˆ๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ํ‰๋ถ€ X์„  ๋ฐ ๋‡Œ MRI ๋ฐ์ดํ„ฐ์…‹์„ ํ™œ์šฉํ•˜์—ฌ ๋ฉ€ํ‹ฐํƒœ์Šคํฌ SLM(OPT-350m)๊ณผ GPT-4o์˜ ์„ฑ๋Šฅ์„ ๋น„๊ตํ•œ ๊ฒฐ๊ณผ, ๋ฉ€ํ‹ฐํƒœ์Šคํฌ SLM์€ ์˜๋ฃŒ ๋ณด๊ณ ์„œ ๋ ˆ์ด๋ธ”๋ง์—์„œ F1 ์ ์ˆ˜ 0.894(GPT-4o: 0.728), DICOM ํ‘œ์ค€ํ™”์—์„œ ์ •ํ™•๋„ 0.975(GPT-4o: 0.878), ์ธ์ƒ ์ƒ์„ฑ์—์„œ Likert ํ‰๊ท  4.39(GPT-4o: 3.65, p=0.0008)๋ฅผ ๋‹ฌ์„ฑํ•˜๋ฉฐ GPT-4o๋ฅผ ์ „๋ฐ˜์ ์œผ๋กœ ๋Šฅ๊ฐ€ํ•˜์˜€๋‹ค. ์ด ์—ฐ๊ตฌ๋Š” ๋‹จ์ผ ๋ฏธ์„ธ ์กฐ์ • SLM์ด ๋Œ€ํ˜• ๋ชจ๋ธ๊ณผ ๋™๋“ฑํ•˜๊ฑฐ๋‚˜ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ฐœํœ˜ํ•˜๋ฉด์„œ๋„ ๋‚ฎ์€ ์ž์› ์š”๊ตฌ๋Ÿ‰, ๋†’์€ ๊ฐœ์ธ์ •๋ณด ๋ณดํ˜ธ ์ˆ˜์ค€, ๊ฐ•๋ ฅํ•œ ๊ณผ์ œ ์ผ๋ฐ˜ํ™” ๋Šฅ๋ ฅ์„ ๊ฐ–์ถ”์–ด ์˜๋ฃŒ ํ˜„์žฅ์—์„œ ์‹ค์šฉ์ ์ธ ์ž„์ƒ AI ์†”๋ฃจ์…˜์œผ๋กœ ํ™œ์šฉ๋  ์ˆ˜ ์žˆ์Œ์„ ์ž…์ฆํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

60Automated Prediction of Radiological Protocols Using Retrieval Augmented Generation.

2026-03Journal of imaging informatics in medicineโญ Q1DOI 10.1007/s10278-025-01822-x

Radiological protocol selection is a critical but time-consuming step in clinical workflow, requiring radiologists to match patient indications with an appropriate MRI or CT protocol. Manual selection can be prone to delays or potential errors, and automated approaches must contend with substantial class imbalance, site-specific variation, and evolving nomenclature. We investigated whether a large language model (LLM) can support reliable protocol selection at scale and whether retrievalaugmented generation (RAG) offers operational advantages over direct fine-tuning. Using patient reports collected across three Mayo Clinic sites (Arizona, Florida, and Rochester) spanning six radiological divisions, we trained site-specific Llama 3.2 3B models for use with and without retrieval augmentation. Division-scoped Facebook AI Similarity Search (FAISS) indexes constructed from procedure and diagnosis text were used to supply contextual evidence in the RAG framework. Both fine-tuned non-RAG and RAG-augmented models achieved strong baseline performance across sites. Paired bootstrap analyses revealed that RAG improved macro F1 at two of three sites (Arizona:: ฮ” =0.0306, p=0.0074; Florida: ฮ” =0.0245, p=0.0217) while maintaining equivalent weighted F1. However, at Rochester, RAG showed no macro F1 improvement and significantly degraded weighted F1 ( ฮ” =-0.0180, p=1.0000), indicating site-specific heterogeneity in RAG effectiveness. RAG introduced an interpretable abstention mechanism with low baseline rates (1-2.5protocol classification without sacrificing common protocol accuracy at most sites, though site-specific tuning may be necessary. Retrieval indexes can be refreshed far more easily than retraining LLMs, enabling continual adaptation to evolving clinical workflows. Future prospective deployment should evaluate real-time accuracy, investigate site-specific performance drivers, and assess abstention as a safety mechanism in clinical decision support.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” MRI ๋ฐ CT ํ”„๋กœํ† ์ฝœ ์„ ํƒ ์ž๋™ํ™”๋ฅผ ์œ„ํ•ด ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์— ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG) ๊ธฐ๋ฒ•์„ ์ ์šฉํ•˜๋Š” ๋ฐฉ๋ฒ•๋ก ์„ ํƒ๊ตฌํ•˜์˜€์œผ๋ฉฐ, Mayo Clinic 3๊ฐœ ๊ธฐ๊ด€์—์„œ ์ˆ˜์ง‘๋œ ํ™˜์ž ๋ณด๊ณ ์„œ๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ Llama 3.2 3B ๋ชจ๋ธ์„ ๊ธฐ๊ด€๋ณ„๋กœ ํŒŒ์ธํŠœ๋‹ํ•˜์—ฌ RAG ์ ์šฉ ์—ฌ๋ถ€์— ๋”ฐ๋ฅธ ์„ฑ๋Šฅ์„ ๋น„๊ตํ•˜์˜€๋‹ค. RAG ๋ชจ๋ธ์€ Arizona ๋ฐ Florida ๊ธฐ๊ด€์—์„œ ๋น„RAG ๋ชจ๋ธ ๋Œ€๋น„ macro F1 ์ง€ํ‘œ๋ฅผ ์œ ์˜๋ฏธํ•˜๊ฒŒ ํ–ฅ์ƒ์‹œ์ผฐ์œผ๋‚˜(๊ฐ๊ฐ ฮ”=0.0306, p=0.0074; ฮ”=0.0245, p=0.0217), Rochester ๊ธฐ๊ด€์—์„œ๋Š” ์˜คํžˆ๋ ค weighted F1์ด ์œ ์˜ํ•˜๊ฒŒ ์ €ํ•˜๋˜์–ด ๊ธฐ๊ด€๋ณ„ ์ด์งˆ์„ฑ์ด ํ™•์ธ๋˜์—ˆ๋‹ค. RAG ๊ธฐ๋ฐ˜ ์ ‘๊ทผ๋ฒ•์€ LLM ์žฌํ•™์Šต ์—†์ด ๊ฒ€์ƒ‰ ์ธ๋ฑ์Šค ๊ฐฑ์‹ ๋งŒ์œผ๋กœ ์ง„ํ™”ํ•˜๋Š” ์ž„์ƒ ํ”„๋กœํ† ์ฝœ์— ์ ์‘ํ•  ์ˆ˜ ์žˆ๋Š” ์žฅ์ ์„ ์ง€๋‹ˆ๋ฉฐ, ์‘๋‹ต ๋ณด๋ฅ˜(abstention) ๊ธฐ์ „์„ ํ†ตํ•ด ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์›์—์„œ์˜ ์•ˆ์ „์„ฑ ํ™•๋ณด ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

61Etiologic classification of suspected MINOCA using cardiovascular magnetic resonance reports: a comparison of a large language model and human readers.

2026-03The international journal of cardiovascular imagingDOI 10.1007/s10554-026-03689-7

The study explored the feasibility of using a large language model (LLM) for etiologic classification in patients with suspected myocardial infarction with non-obstructive coronary arteries (MINOCA) based on cardiovascular magnetic resonance (CMR) reports. We included 156 patients with MINOCA from a prospective (nโ€‰=โ€‰50) and a retrospective (nโ€‰=โ€‰106) pooled cohort. A large language model and three human readers with different experience levels independently classified CMR reports into eight predefined diagnostic categories, using the final expert CMR diagnosis as the reference standard. Performance and agreement were assessed using standard multiclass classification metrics. In the pooled cohort, the LLM achieved an exact-match diagnostic accuracy of 67.3% (95% CI 59.6-74.2%), lower than the expert reader (80.1%, 95% CI 73.2-85.6) but comparable to the intermediate and junior readers (both 68.6%, 95% CI 60.9-75.4%). Diagnosis-specific analysis showed consistently high specificity for the LLM (mean 94.9%), with sensitivities up to 86.0% and F1-scores up to 85.1% for common etiologies. Agreement between the LLM and the reference standard was substantial (ICC 0.84, 95% CI: 0.79-0.89) and comparable to experienced readers, whereas agreement with the junior reader was markedly lower, indicating greater diagnostic variability despite similar accuracy. A large language model shows promise as a supportive tool for etiologic classification in suspected MINOCA from CMR reports, with performance comparable to less experienced readers and high agreement with the final expert CMR diagnosis. Integration into structured reporting workflows may enhance diagnostic consistency in routine clinical practice.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋น„ํ์‡„์„ฑ ๊ด€์ƒ๋™๋งฅ ์‹ฌ๊ทผ๊ฒฝ์ƒ‰(MINOCA) ์˜์‹ฌ ํ™˜์ž์˜ ์‹ฌ์žฅ ์ž๊ธฐ๊ณต๋ช…์˜์ƒ(CMR) ๋ณด๊ณ ์„œ๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ํ™œ์šฉํ•œ ๋ณ‘์ธ ๋ถ„๋ฅ˜์˜ ๊ฐ€๋Šฅ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ „ํ–ฅ์  ๋ฐ ํ›„ํ–ฅ์  ์ฝ”ํ˜ธํŠธ๋ฅผ ํฌํ•จํ•œ 156๋ช…์˜ ํ™˜์ž๋ฅผ ๋Œ€์ƒ์œผ๋กœ LLM๊ณผ ๊ฒฝํ—˜ ์ˆ˜์ค€์ด ๋‹ค๋ฅธ ์„ธ ๋ช…์˜ ์ธ๊ฐ„ ํŒ๋…์ž๊ฐ€ CMR ๋ณด๊ณ ์„œ๋ฅผ 8๊ฐœ์˜ ์‚ฌ์ „ ์ •์˜๋œ ์ง„๋‹จ ๋ฒ”์ฃผ๋กœ ๋…๋ฆฝ ๋ถ„๋ฅ˜ํ•˜์˜€์œผ๋ฉฐ, ์ „๋ฌธ๊ฐ€ CMR ์ง„๋‹จ์„ ๊ธฐ์ค€์œผ๋กœ ์„ฑ๋Šฅ์„ ๋น„๊ตํ•˜์˜€๋‹ค. LLM์˜ ์ •ํ™•๋„๋Š” 67.3%๋กœ ์ „๋ฌธ๊ฐ€ ํŒ๋…์ž(80.1%)๋ณด๋‹ค ๋‚ฎ์•˜์œผ๋‚˜ ์ค‘๊ธ‰ ๋ฐ ์ดˆ๊ธ‰ ํŒ๋…์ž(68.6%)์™€ ์œ ์‚ฌํ•˜์˜€๊ณ , ์ „๋ฌธ๊ฐ€ ์ง„๋‹จ๊ณผ์˜ ์ผ์น˜๋„๋Š” ๋†’์•„(ICC 0.84) ๊ตฌ์กฐํ™”๋œ CMR ๋ณด๊ณ  ์›Œํฌํ”Œ๋กœ์šฐ์— ๋ณด์กฐ ๋„๊ตฌ๋กœ ํ†ตํ•ฉ ์‹œ ์ง„๋‹จ ์ผ๊ด€์„ฑ ํ–ฅ์ƒ์— ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

62LLM-assisted systematic review of large language models in clinical medicine.

2026-03Nature medicineโญ Q1DOI 10.1038/s41591-026-04229-5

Clinical evaluations of large language models (LLMs) have rapidly expanded since 2022, yet their evidence base remains opaque. The overwhelming volume of studies creates challenges for manual curation and review. However, LLMs themselves offer the scalability and capability to evaluate the ever-growing evidence base. This LLM-assisted review identified 4,609 peer-reviewed studies in clinical medicine between January 2022 and September 2025, equating to roughly 3.2 papers per day. Only 1,048 studies used real-world patient data and of these only 19 were prospective randomized trials; most addressed simulated scenarios (nโ€‰=โ€‰1,857) or exam-style tasks (nโ€‰=โ€‰1,704). ChatGPT and related OpenAI models constitute 65.7% of evaluated models, with Gemini/Bard a distant second constituting 13.1% of evaluated models. Patient-facing communication and education comprised 17% of tasks, followed by knowledge retrieval, and education and assessment simulation. Across 1,046 head-to-head comparisons, LLMs outperformed humans in 33% of comparisons, with a strong dependency on task realism and level of training. At least 25% of studies had sample sizes less than 30. Despite the growth of LLMs in medicine, rigorous, patient-centered evidence remains scarce, underscoring the need for larger prospective trials before clinical adoption.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 2022๋…„ 1์›”๋ถ€ํ„ฐ 2025๋…„ 9์›”๊นŒ์ง€ ์ž„์ƒ์˜ํ•™ ๋ถ„์•ผ์—์„œ ๋ฐœํ‘œ๋œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM) ๊ด€๋ จ ๋™๋ฃŒ ์‹ฌ์‚ฌ ๋…ผ๋ฌธ 4,609ํŽธ์„ LLM ๋ณด์กฐ ์ฒด๊ณ„์  ๋ฌธํ—Œ๊ณ ์ฐฐ ๋ฐฉ๋ฒ•์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ์‹ค์ œ ํ™˜์ž ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•œ ์—ฐ๊ตฌ๋Š” 1,048ํŽธ์— ๋ถˆ๊ณผํ•˜์˜€์œผ๋ฉฐ ๊ทธ ์ค‘ ๋ฌด์ž‘์œ„ ๋Œ€์กฐ ์‹œํ—˜์€ 19ํŽธ์— ๊ทธ์ณค๊ณ , ๋Œ€๋ถ€๋ถ„์˜ ์—ฐ๊ตฌ๋Š” ์‹œ๋ฎฌ๋ ˆ์ด์…˜ ์‹œ๋‚˜๋ฆฌ์˜ค๋‚˜ ์‹œํ—˜ ํ˜•์‹์˜ ๊ณผ์ œ์— ์ง‘์ค‘๋˜์–ด ์žˆ์—ˆ์œผ๋ฉฐ ํ‰๊ฐ€ ๋ชจ๋ธ์˜ 65.7%๊ฐ€ ChatGPT ๋ฐ OpenAI ๊ณ„์—ด์ด์—ˆ๋‹ค. ์ธ๊ฐ„ ๋Œ€๋น„ LLM ์„ฑ๋Šฅ ๋น„๊ต์—์„œ LLM์ด ์šฐ์„ธํ•œ ๊ฒฝ์šฐ๋Š” 33%์— ๋ถˆ๊ณผํ•˜์˜€๊ณ  ์—ฐ๊ตฌ์˜ ์ตœ์†Œ 25%๋Š” ํ‘œ๋ณธ ํฌ๊ธฐ๊ฐ€ 30 ๋ฏธ๋งŒ์ด์–ด์„œ, ์ž„์ƒ ๋„์ž…์„ ์œ„ํ•ด์„œ๋Š” ์—„๊ฒฉํ•˜๊ณ  ํ™˜์ž ์ค‘์‹ฌ์ ์ธ ์ „ํ–ฅ์  ๋Œ€๊ทœ๋ชจ ์—ฐ๊ตฌ๊ฐ€ ์‹œ๊ธ‰ํžˆ ํ•„์š”ํ•จ์„ ๊ฐ•์กฐํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

63Large language models for structured cardiovascular data extraction: a foundation for scalable research and clinical applications.

2026-03European heart journal. Digital healthโญ Q1DOI 10.1093/ehjdh/ztaf127
OBJECTIVE

Automated extraction of information from cardiac reports would benefit both clinical reporting and research. Large language models (LLMs) hold promise for such automation, but their clinical performance and practical implementation across various computational environments remain unclear. This study aims to evaluate the feasibility and performance of LLM-based classification of echocardiogram and invasive coronary angiography reports, using real-world clinical data across local, high-performance computing and cloud-based platforms. METHODS AND

RESULTS

The angiography and echocardiography reports of 1000 patients, admitted with acute coronary syndrome, were labelled for multiple key diagnostic elements, including left ventricular function (LVF), culprit vessel, and acute occlusions. Report classification models were developed using LLMs via (i) prompt-based and (ii) fine-tuning approaches. Performance was assessed across different model types and compute infrastructures, with attention to class imbalance, ambiguous label annotations, and implementation costs. Large language models demonstrated strong performance in extracting structured diagnostic information from cardiac reports. Cloud-based models (such as GPT-4o) achieved the highest accuracy (0.87 for culprit vessel and 1.0 for LVF) and generalizability, but also smaller models run on a local high-performance cluster achieved reasonable accuracy, especially for less complex tasks (0.634 for culprit vessel and 0.984 for LVF). Classification was feasible with minimal pre-processing, enabling potential integration into electronic health record systems or research pipelines. Class imbalance, reflective of real-world prevalence, had a greater impact on fine-tuning approaches.

CONCLUSION

Large language models can reliably classify structured cardiology reports across diverse computed infrastructures. Their accuracy and adaptability support their use in clinical and research settings, particularly for scalable report structuring and dataset generation.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ธ‰์„ฑ ๊ด€์ƒ๋™๋งฅ ์ฆํ›„๊ตฐ์œผ๋กœ ์ž…์›ํ•œ 1,000๋ช… ํ™˜์ž์˜ ์‹ฌ์ดˆ์ŒํŒŒ ๋ฐ ์นจ์Šต์  ๊ด€์ƒ๋™๋งฅ ์กฐ์˜์ˆ  ๋ณด๊ณ ์„œ์—์„œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ํ™œ์šฉํ•ด ์ขŒ์‹ฌ์‹ค ๊ธฐ๋Šฅ, ์ฑ…์ž„ํ˜ˆ๊ด€, ๊ธ‰์„ฑ ํ์ƒ‰ ๋“ฑ ํ•ต์‹ฌ ์ง„๋‹จ ์š”์†Œ๋ฅผ ์ž๋™ ์ถ”์ถœํ•˜๋Š” ํƒ€๋‹น์„ฑ๊ณผ ์„ฑ๋Šฅ์„ ๋‹ค์–‘ํ•œ ์ปดํ“จํŒ… ํ™˜๊ฒฝ(๋กœ์ปฌ, ๊ณ ์„ฑ๋Šฅ ํด๋Ÿฌ์Šคํ„ฐ, ํด๋ผ์šฐ๋“œ)์—์„œ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ํ”„๋กฌํ”„ํŠธ ๊ธฐ๋ฐ˜ ๋ฐ ํŒŒ์ธํŠœ๋‹ ๋ฐฉ์‹์œผ๋กœ ๋ถ„๋ฅ˜ ๋ชจ๋ธ์„ ๊ฐœ๋ฐœํ•œ ๊ฒฐ๊ณผ, ํด๋ผ์šฐ๋“œ ๊ธฐ๋ฐ˜ GPT-4o๊ฐ€ ์ขŒ์‹ฌ์‹ค ๊ธฐ๋Šฅ ์ •ํ™•๋„ 1.0, ์ฑ…์ž„ํ˜ˆ๊ด€ 0.87๋กœ ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ๋กœ์ปฌ ๊ณ ์„ฑ๋Šฅ ํด๋Ÿฌ์Šคํ„ฐ์˜ ์†Œํ˜• ๋ชจ๋ธ๋„ ๋‹จ์ˆœํ•œ ๊ณผ์ œ์—์„œ๋Š” ์ขŒ์‹ฌ์‹ค ๊ธฐ๋Šฅ ์ •ํ™•๋„ 0.984๋กœ ์ค€์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋‹ฌ์„ฑํ•˜์˜€๋‹ค. LLM์€ ์ตœ์†Œํ•œ์˜ ์ „์ฒ˜๋ฆฌ๋งŒ์œผ๋กœ ์‹ฌ์žฅ ๋ณด๊ณ ์„œ์˜ ๊ตฌ์กฐํ™”๋œ ๋ถ„๋ฅ˜๊ฐ€ ๊ฐ€๋Šฅํ•˜์—ฌ ์ „์ž์˜๋ฌด๊ธฐ๋ก ์‹œ์Šคํ…œ ํ†ตํ•ฉ ๋ฐ ๋Œ€๊ทœ๋ชจ ์—ฐ๊ตฌ ๋ฐ์ดํ„ฐ์…‹ ๊ตฌ์ถ•์— ์ž„์ƒ์ ์œผ๋กœ ์œ ์šฉํ•˜๊ฒŒ ํ™œ์šฉ๋  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

64The Rise of Deepfake Medical Imaging: Radiologists' Diagnostic Accuracy in Detecting ChatGPT-generated Radiographs.

2026-03Radiologyโญ Q1DOI 10.1148/radiol.252094

Background Large language models (LLMs) can generate realistic synthetic medical images (deepfakes), which raise concerns about potential misuse. Purpose To assess the ability of radiologists and multimodal LLMs to distinguish ChatGPT-generated synthetic radiographs from authentic clinical images. Materials and Methods This retrospective diagnostic accuracy study conducted between April and August 2025 included 17 practicing radiologists from six countries with varying experience levels. In phase 1, the radiologists, blinded to the purpose of the study, assessed image quality and provided diagnoses for 154 radiographs from multiple anatomic regions (77 synthetic images generated using ChatGPT [GPT-4o; OpenAI] and 77 authentic images). In phase 2, after being informed of the study's purpose, the radiologists determined whether randomly presented radiographs were GPT-4o-generated or authentic. The same classification task was performed by four LLMs: GPT-4o, GPT-5 (OpenAI), Gemini 2.5 Pro (Google), and Llama 4 Maverick (Meta). In phase 3, an additional set of 110 chest radiographs (55 synthetic images generated using RoentGen and 55 authentic images) was analyzed to evaluate the performance of readers and LLMs in distinguishing synthetic versus authentic images. The McNemar test and t test were used for comparisons. Results Forty-one percent (seven of 17) of purpose-blinded radiologists spontaneously identified artificial intelligence-generated radiographs as being present in the dataset. After being informed that some radiographs were synthetic, there was no evidence of a difference in overall accuracy among all 17 radiologists in distinguishing synthetic images in the GPT-4o dataset (75% [95% CI: 68, 81]) versus in the RoentGen dataset (70% [95% CI: 62, 78]; P = .07). No tested LLM detected all synthetic radiographs in either dataset; however, GPT-4o-generated radiographs were more accurately differentiated from authentic ones by GPT-4o (accuracy, 85%) and GPT-5 (accuracy, 83%) compared with Llama 4 Maverick (accuracy, 59%) and Gemini 2.5 Pro (accuracy, 56%) (all P < .001). Common features of synthetic radiographs included bilateral symmetry, uniform grain or noise patterns, subtly unnatural soft-tissue textures, and overly smooth bone surfaces. Conclusion Synthetic radiographs (deepfakes) generated using an LLM were not easily distinguishable from authentic radiographs by either radiologists or LLMs. Training physicians and LLMs to recognize synthetic images is essential to mitigate risks. To support training, a curated deepfake dataset is available: https://noneedanick.github.io/DeepFakeXRay/. ยฉ RSNA, 2026 Supplemental material is available for this article. See also the editorial by Bhayana and Krishna in this issue.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
์ด ์—ฐ๊ตฌ๋Š” ChatGPT(GPT-4o)๋กœ ์ƒ์„ฑ๋œ ํ•ฉ์„ฑ ๋ฐฉ์‚ฌ์„  ์‚ฌ์ง„(๋”ฅํŽ˜์ดํฌ)์„ ์‹ค์ œ ์ž„์ƒ ์˜์ƒ๊ณผ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜ ๋ฐ ๋‹ค์ค‘๋ชจ๋‹ฌ ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ(LLM)์ด ์–ผ๋งˆ๋‚˜ ์ •ํ™•ํ•˜๊ฒŒ ๊ตฌ๋ณ„ํ•  ์ˆ˜ ์žˆ๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 6๊ฐœ๊ตญ 17๋ช…์˜ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜๋ฅผ ๋Œ€์ƒ์œผ๋กœ GPT-4o ์ƒ์„ฑ ํ•ฉ์„ฑ ์˜์ƒ 77์žฅ๊ณผ ์‹ค์ œ ์˜์ƒ 77์žฅ์„ ํฌํ•จํ•œ 154์žฅ์˜ ๋ฐฉ์‚ฌ์„  ์‚ฌ์ง„์„ ๋‹จ๊ณ„์ ์œผ๋กœ ํ‰๊ฐ€ํ•˜์˜€์œผ๋ฉฐ, GPT-4o, GPT-5, Gemini 2.5 Pro, Llama 4 Maverick ๋“ฑ 4์ข…์˜ LLM๋„ ๋™์ผํ•œ ๋ถ„๋ฅ˜ ๊ณผ์ œ๋ฅผ ์ˆ˜ํ–‰ํ•˜์˜€๋‹ค. ์—ฐ๊ตฌ ๋ชฉ์ ์„ ์ธ์ง€ํ•œ ํ›„์—๋„ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜์˜ ํ•ฉ์„ฑ ์˜์ƒ ํŒ๋ณ„ ์ •ํ™•๋„๋Š” ์•ฝ 75%์— ๊ทธ์ณค๊ณ , LLM ์ค‘์—์„œ๋Š” GPT-4o(85%)์™€ GPT-5(83%)๊ฐ€ Gemini 2.5 Pro(56%) ๋ฐ Llama 4 Maverick(59%)๋ณด๋‹ค ์œ ์˜ํ•˜๊ฒŒ ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์œผ๋‚˜, ํ•ฉ์„ฑ ์˜์ƒ์„ ์™„์ „ํžˆ ํƒ์ง€ํ•œ ๋ชจ๋ธ์€ ์—†์–ด ๋”ฅํŽ˜์ดํฌ ์˜๋ฃŒ ์˜์ƒ ์‹๋ณ„์„ ์œ„ํ•œ ์ „๋ฌธ ํ›ˆ๋ จ์˜ ํ•„์š”์„ฑ์ด ๊ฐ•์กฐ๋œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

65A comparative evaluation of large language models for simplifying prostate cancer pathology reports: ChatGPT and Gemini.

2026-02International journal of surgery (London, England)DOI 10.1097/js9.0000000000004454
OBJECTIVE

To evaluate the application value of three ChatGPT versions and Gemini in pathology report simplification tasks for prostate cancer.

METHODS

This retrospective study assessed GPT-3.5, GPT-4.0, GPT-4o, and Gemini on pathology reports from 228 prostate cancer patients across two institutions. Data were split into internal (center 1, n =ย 171) and external (center 2, n =ย 57) cohorts. Using specific prompts, models generated simplified texts. The evaluation of outputs included three main dimensions: (1) human scoring by patients, clinicians, and pathologists; (2) readability scores; and (3) BERT-based semantic similarity scores. Statistical comparisons employed paired t -tests or Wilcoxon signed-rank tests. Statistical consistency between raters was assessed using squared weighted kappa, intraclass correlation coefficient(3,1), and percent agreement, with 95% confidence intervals calculated for all metrics.

RESULTS

GPT-4o (Few-Shot) achieved the highest accuracy and comprehensiveness scores from pathologists, while Gemini demonstrated the best understandability. Patient and clinician understandability ratings were consistently high across models. Mean Reading Grade Level scores varied between internal and external datasets, with GPT-4o Few-Shot performing best overall. BERT-based semantic similarity scores demonstrated distinct trends across models, reflecting differences in text simplification strategies.

CONCLUSION

LLMs adopt distinct trade-off strategies between simplifying pathology reports and preserving their structure and logic, influenced by prompt design and textual style. Their application shows potential to enhance patient comprehension and clinical communication. Future work should focus on domain-specific fine-tuning to ensure safe and reliable clinical integration.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ํ›„ํ–ฅ์  ์—ฐ๊ตฌ๋Š” ์ „๋ฆฝ์„ ์•” ๋ณ‘๋ฆฌ ๋ณด๊ณ ์„œ ๊ฐ„์†Œํ™” ๊ณผ์ œ์—์„œ GPT-3.5, GPT-4.0, GPT-4o, Gemini ๋“ฑ 4๊ฐ€์ง€ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ 228๋ช…์˜ ํ™˜์ž ๋ฐ์ดํ„ฐ๋ฅผ ๋Œ€์ƒ์œผ๋กœ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ํ™˜์žยท์ž„์ƒ์˜ยท๋ณ‘๋ฆฌ์˜์‚ฌ์˜ ์ธ๊ฐ„ ํ‰๊ฐ€, ๊ฐ€๋…์„ฑ ์ ์ˆ˜, BERT ๊ธฐ๋ฐ˜ ์˜๋ฏธ์  ์œ ์‚ฌ๋„ ์ ์ˆ˜์˜ ์„ธ ๊ฐ€์ง€ ์ฐจ์›์—์„œ ๋ชจ๋ธ ์ถœ๋ ฅ๋ฌผ์„ ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, GPT-4o(Few-Shot ํ”„๋กฌํ”„ํŠธ)๋Š” ์ •ํ™•์„ฑ ๋ฐ ํฌ๊ด„์„ฑ์—์„œ ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ํ‰๊ฐ€๋ฅผ ๋ฐ›์•˜์œผ๋ฉฐ, Gemini๋Š” ์ดํ•ด ์šฉ์ด์„ฑ ์ธก๋ฉด์—์„œ ์ตœ๊ณ  ์ ์ˆ˜๋ฅผ ๊ธฐ๋กํ•˜์˜€๋‹ค. LLM์€ ํ”„๋กฌํ”„ํŠธ ์„ค๊ณ„ ๋ฐ ํ…์ŠคํŠธ ์Šคํƒ€์ผ์— ๋”ฐ๋ผ ๋ณด๊ณ ์„œ ๊ฐ„์†Œํ™”์™€ ๊ตฌ์กฐยท๋…ผ๋ฆฌ ๋ณด์กด ์‚ฌ์ด์—์„œ ์„œ๋กœ ๋‹ค๋ฅธ ์ ˆ์ถฉ ์ „๋žต์„ ์ทจํ•˜๋ฉฐ, ํ–ฅํ›„ ์ž„์ƒ ์ ์šฉ์˜ ์•ˆ์ „์„ฑ๊ณผ ์‹ ๋ขฐ์„ฑ ํ™•๋ณด๋ฅผ ์œ„ํ•ด ๋„๋ฉ”์ธ ํŠนํ™” ๋ฏธ์„ธ ์กฐ์ •์ด ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

66A fine-tuned large language model chatbot for multi-scenario radiology cancer care: randomized controlled trial on interaction optimization, emotional support, and provider burnout reduction.

2026-02Journal of translational medicineโญ Q1DOI 10.1186/s12967-026-07738-6

IMPORTANCE: Cancer patients are more prone to depression and anxiety symptoms compared to those with chronic diseases. Amidst surging clinical demands and constrained medical resources, the traditional radiology workflows, plagued by inefficient communication, exacerbates both patients' psychological distress and healthcare providers' burnout.

OBJECTIVE

To develop and validate a fine-tuned DeepSeek R1-based Radiology Examination Chatbot (REC) to optimize clinical interaction between cancer patients and radiology healthcare providers (RHPs), thereby effectively providing emotional support for cancer patients and reducing burnout among RHPs. DESIGN, SETTING, AND

METHODS

Audio recordings of multi-scenarios (appointment triage (AT), pre-examination preparation (PP), radiology clinic services (RCS)) were collected from the radiology departments of three tertiary hospitals (nโ€‰=โ€‰36,511ย min). This study conducts two independent randomized controlled sub-trials for distinct patient groups: Sub-trial 1 evaluates AT/PP participants (1,424 patients, 1:1 randomized to RHPโ€‰+โ€‰REC or RHP), while Sub-trial 2 assesses RCS participants (638 patients, 1:1 randomized to the same groups). Due to differing patient populations, the sub-trials were designed and implemented separately. INTERVENTION: The REC was fine-tuned using domain-specific dialogue data (80% for training) and scenario-specific prompts, with GPT-o1 as a comparative benchmark. Sub-trials randomized patients to RHPโ€‰+โ€‰REC or RHP groups. MAIN OUTCOME AND MEASURES: The primary outcome included dialogue quality (empathy, frustration, emotional regulation, factuality, integrity, and satisfaction), while the secondary outcomes comprised burnout (exhaustion, depersonalization, and personal achievement) and image quality (CT/MRI), all assessed via Likert scales and statistical tests.

RESULTS

RHPโ€‰+โ€‰REC group demonstrated superior dialogue quality in AT (factuality: 4.12โ€‰ยฑโ€‰0.86 vs. 3.39โ€‰ยฑโ€‰1.21, Pโ€‰<โ€‰0.001) and PP (satisfaction: 3.73โ€‰ยฑโ€‰0.11 vs. 3.19โ€‰ยฑโ€‰0.18, Pโ€‰<โ€‰0.001), with reduced burnout (exhaustion: 1.85โ€‰ยฑโ€‰0.91 vs. 2.40โ€‰ยฑโ€‰1.22, Pโ€‰<โ€‰0.01). CT image quality improved significantly (4.35โ€‰ยฑโ€‰0.51 vs. 4.00โ€‰ยฑโ€‰0.52, Pโ€‰<โ€‰0.01), and similar results were achieved in MRI examinations (4.12โ€‰ยฑโ€‰0.51 vs. 3.79โ€‰ยฑโ€‰0.58, Pโ€‰=โ€‰0.02). However, REC underperformed in empathy and emotional regulation during emotionally complex RCS (3.88โ€‰ยฑโ€‰0.67 vs. 4.42โ€‰ยฑโ€‰0.53, Pโ€‰=โ€‰0.002; 3.87โ€‰ยฑโ€‰0.19 vs. 4.12โ€‰ยฑโ€‰0.27, Pโ€‰=โ€‰0.004). Ablation studies confirmed the necessity of fine-tuning and scenario-specific prompts for performance. CONCLUSION AND RELEVANCE: Overall, the DeepSeek R1-based REC synergistically enhanced multi-scenarios clinical interaction, provided emotional support for cancer patients, and reduced RHPs' burnout, offering a scalable solution to optimize radiology workflows. TRIAL REGISTRATION: Chinese Clinical Trial Registry Registration number: (ChiCTR2500102740|| http://www.chictr.org.cn/ ), Registration Date: 2025-05-26. KEY POINTS: QUESTION: Can a DeepSeek R1 fine-tuned chatbot (REC) optimize clinical interactions, provide emotional support to patients, and reduce burnout among radiology healthcare providers (RHPs) across multiple scenarios in radiology cancer care (appointment triage (AT), pre-examination preparation (PP), radiology clinic services (RCS))?

RESULTS

REC significantly improved dialogue quality (factuality, integrity, satisfaction) and imaging quality (CT/MRI scans) in both AT and PP scenarios, while simultaneously reducing occupational burnout among RHPs. In emotionally complex scenario like the RCS, although REC outperformed humans in factuality and integrity, its empathy capabilities and emotional regulation lagged behind humans. Ablation experiments confirmed that domain-specific fine-tuning and scenario-adapted prompt templates were crucial to RECโ€™s performance gains. MEANING: REC functions as an efficient workflow aid in radiology departmentsโ€”enhancing efficiency for structured tasks (scheduling, procedural guidance), alleviating clinical burden, and delivering standardized patient support for cancer cases. Its limitations in emotionally sensitive contexts suggest future integration should combine human providersโ€™ emotional intelligence or incorporate multimodal technologies (e.g., facial expression/tone analysis) for model optimization. This study establishes a benchmark for contextual validation of medical LLMs, advancing safe AI implementation in cancer care.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์•” ํ™˜์ž์˜ ์‹ฌ๋ฆฌ์  ๋ถ€๋‹ด ์ฆ๊ฐ€์™€ ๋ฐฉ์‚ฌ์„  ์˜๋ฃŒ์ง„์˜ ์†Œ์ง„ ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๊ธฐ ์œ„ํ•ด, DeepSeek R1 ๊ธฐ๋ฐ˜์˜ ๋ฐฉ์‚ฌ์„ ๊ณผ ๊ฒ€์‚ฌ ์ฑ—๋ด‡(REC)์„ ๋„๋ฉ”์ธ ํŠนํ™” ๋Œ€ํ™” ๋ฐ์ดํ„ฐ๋กœ ํŒŒ์ธํŠœ๋‹ํ•˜์—ฌ ์˜ˆ์•ฝ ํŠธ๋ฆฌ์•„์ง€, ๊ฒ€์‚ฌ ์ „ ์ค€๋น„, ๋ฐฉ์‚ฌ์„  ํด๋ฆฌ๋‹‰ ์„œ๋น„์Šค ๋“ฑ ๋‹ค์ค‘ ์‹œ๋‚˜๋ฆฌ์˜ค์—์„œ์˜ ์ž„์ƒ ์ƒํ˜ธ์ž‘์šฉ ์ตœ์ ํ™” ํšจ๊ณผ๋ฅผ 3๊ฐœ 3์ฐจ ๋ณ‘์›์—์„œ ๋ฌด์ž‘์œ„ ๋Œ€์กฐ์‹œํ—˜(์ด 2,062๋ช…)์œผ๋กœ ๊ฒ€์ฆํ•˜์˜€๋‹ค. RHP+REC ๋ณ‘ํ•ฉ๊ตฐ์€ ์˜ˆ์•ฝ ํŠธ๋ฆฌ์•„์ง€ ๋ฐ ๊ฒ€์‚ฌ ์ „ ์ค€๋น„ ๋‹จ๊ณ„์—์„œ ๋Œ€ํ™”์˜ ์‚ฌ์‹ค์„ฑยท๋งŒ์กฑ๋„ ๋ฐ CT/MRI ์˜์ƒ ํ’ˆ์งˆ์„ ์œ ์˜ํ•˜๊ฒŒ ํ–ฅ์ƒ์‹œ์ผฐ์œผ๋ฉฐ(๋ชจ๋‘ P<0.05), ์˜๋ฃŒ์ง„์˜ ์ง์—…์  ์†Œ์ง„ ์ง€ํ‘œ(exhaustion) ๋˜ํ•œ ์œ ์˜ํ•˜๊ฒŒ ๊ฐ์†Œ์‹œ์ผฐ๋‹ค. ๋‹ค๋งŒ ๊ฐ์ •์ ์œผ๋กœ ๋ณต์žกํ•œ ๋ฐฉ์‚ฌ์„  ํด๋ฆฌ๋‹‰ ์„œ๋น„์Šค ์ƒํ™ฉ์—์„œ๋Š” REC์˜ ๊ณต๊ฐ ๋Šฅ๋ ฅ๊ณผ ๊ฐ์ • ์กฐ์ ˆ ์—ญ๋Ÿ‰์ด ์ธ๊ฐ„ ์˜๋ฃŒ์ง„์— ๋น„ํ•ด ์—ด๋“ฑํ•˜์—ฌ, ๊ฐ์ •์ ์œผ๋กœ ๋ฏผ๊ฐํ•œ ๋งฅ๋ฝ์—์„œ๋Š” ์ธ๊ฐ„์˜ ๊ฐ์„ฑ ์ง€๋Šฅ ๋˜๋Š” ๋‹ค์ค‘๋ชจ๋‹ฌ ๊ธฐ์ˆ ๊ณผ์˜ ํ†ตํ•ฉ์ด ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

67Can AI write reports like a radiologist? A blinded evaluation of large language model-generated lumbar spine MRI reports.

2026-02European radiology experimentalโญ Q1DOI 10.1186/s41747-026-00682-6
BACKGROUND

To compare the quality and clinical usefulness of large language model (LLM)-generated lumbar spine magnetic resonance imaging (MRI) reports with radiologist-written ones and assess whether medical professionals can distinguish between them.

METHODS

This retrospective observational single-center study was approved by the local ethics committee. A total of 125 lumbar spine MRI reports (104 human-written, 21 LLM-generated using ChatGPT-4o) were anonymized, randomized, and blindly evaluated by five medical professionals (one board-certified radiologist, two radiology residents, one general practitioner, one orthopedic surgeon), all with basic familiarity with LLM. Each report was scored on a five-point Likert scale for clinical relevance, clarity, completeness, diagnostic accuracy, and intelligibility, whereas general practitioner andย orthopedic surgeon evaluated intelligibility only. Evaluators also classified each report as AI-generated or human-written. Accuracy was defined as the proportion of correctly classified reports in distinguishing LLM-generated from radiologist-written texts. Mann-Whitney U or Student's t-tests were used.

RESULTS

Radiologists' reports consistently received higher median scores across all domains (pโ€‰<โ€‰0.001). No differences were found in the description of the imaging technique (pโ€‰>โ€‰0.175). No clinically false statements were identified in the LLM-generated reports. Identification accuracy varied widely among evaluators: Board-certified radiologist achieved 88.0% accuracy (sensitivity 66.7%, specificity 92.3%), Resident 1 65.6% (14.3%, 76.0%), Resident 2 94.4% (66.7%, 100%), orthopedic surgeon 78.4% (90.5%, 76.0%) and general practitioner 65.6% (81.0%, 62.5%).

CONCLUSION

Radiologist-written lumbar spine MRI reports outperform LLM-generated reports in quality and structure. However, some AI-generated reports were indistinguishable from human ones, particularly for non-specialized readers. LLMs may support radiologists in structured reporting and improve workflow efficiency, while maintaining diagnostic reliability. RELEVANCE STATEMENT: Large language models can draft lumbar spine MRI reports, but currently lack the quality and consistency of radiologist reports. With radiologist supervision, large language models may improve reporting efficiency while preserving diagnostic reliability and supporting clinical decision-making. KEY POINTS: LLM-generated reports are clinically coherent and stylistically comparable to those written by expert radiologists. Radiologist-written reports scored significantly higher for clinical relevance, findings, and structure. LLM-generated reports were sometimes misclassified as human-written by clinicians.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM, ChatGPT-4o)์ด ์ƒ์„ฑํ•œ ์š”์ถ” MRI ํŒ๋…๋ฌธ์˜ ์ž„์ƒ์  ์œ ์šฉ์„ฑ๊ณผ ํ’ˆ์งˆ์„ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜๊ฐ€ ์ž‘์„ฑํ•œ ํŒ๋…๋ฌธ๊ณผ ๋น„๊ตํ•˜๊ณ , ์˜๋ฃŒ ์ „๋ฌธ๊ฐ€๋“ค์ด ์–‘์ž๋ฅผ ๊ตฌ๋ณ„ํ•  ์ˆ˜ ์žˆ๋Š”์ง€ ํ‰๊ฐ€ํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ์ด 125๊ฑด์˜ ํŒ๋…๋ฌธ(์ธ๊ฐ„ ์ž‘์„ฑ 104๊ฑด, LLM ์ƒ์„ฑ 21๊ฑด)์„ ์ต๋ช…ํ™”ํ•˜์—ฌ 5๋ช…์˜ ํ‰๊ฐ€์ž(๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜ 1๋ช…, ์ „๊ณต์˜ 2๋ช…, ์ผ๋ฐ˜์˜ 1๋ช…, ์ •ํ˜•์™ธ๊ณผ ์ „๋ฌธ์˜ 1๋ช…)๊ฐ€ 5์  ๋ฆฌ์ปคํŠธ ์ฒ™๋„๋กœ ๋งน๊ฒ€ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜๊ฐ€ ์ž‘์„ฑํ•œ ํŒ๋…๋ฌธ์€ ์ž„์ƒ์  ๊ด€๋ จ์„ฑ, ๋ช…ํ™•์„ฑ, ์™„๊ฒฐ์„ฑ, ์ง„๋‹จ ์ •ํ™•๋„ ๋“ฑ ๋ชจ๋“  ์˜์—ญ์—์„œ ์œ ์˜ํ•˜๊ฒŒ ๋†’์€ ์ ์ˆ˜๋ฅผ ๋ฐ›์•˜์œผ๋‚˜(p < 0.001), ์ผ๋ถ€ LLM ์ƒ์„ฑ ํŒ๋…๋ฌธ์€ ํŠนํžˆ ๋น„์ „๋ฌธ ํ‰๊ฐ€์ž์— ์˜ํ•ด ์ธ๊ฐ„ ์ž‘์„ฑ๋ฌผ๋กœ ์˜ค๋ถ„๋ฅ˜๋˜์—ˆ์œผ๋ฉฐ, LLM์€ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜์˜ ๊ฐ๋… ํ•˜์— ๊ตฌ์กฐํ™”๋œ ํŒ๋… ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์ด ์žˆ๋Š” ๊ฒƒ์œผ๋กœ ๊ฒฐ๋ก ์ง€์–ด์กŒ๋‹ค.
Added: 2026-04-21 16:11View โ†—

68Evaluation of GPT-5 for Esophageal Cancer Staging Using Fluorodeoxyglucose Positron Emission Tomography Maximum-Intensity Projection Images: Comparative Pilot Study.

2026-02JMIR cancer๐Ÿ”ท Q2DOI 10.2196/86630
BACKGROUND

Accurate esophageal cancer staging relies on 18F fluorodeoxyglucose positron emission tomography (18F FDG-PET), but its interpretation is complex and time-intensive. This diagnostic burden is exacerbated by significant workforce shortages in both radiology and surgery, thus necessitating automated support systems. The emergence of advanced large language models (LLMs) has raised expectations for their potential to fulfill this role in complex medical tasks.

OBJECTIVE

We evaluated the diagnostic accuracy of LLMs for staging esophageal cancer using 18F FDG-PET images, with a focus on their ability to assess lymph nodes (LNs; clinical N [cN]) and distant metastases (clinical M [cM]) for automated radiology reporting.

METHODS

This retrospective study included 120 consecutive adult patients who were diagnosed with esophageal squamous cell carcinoma and underwent 18F FDG-PET/computed tomography at Tohoku University Hospital between January 2019 and December 2021. Patients with prior treatment, nonsquamous cell carcinoma histology, or blood glucose levels โ‰ฅ200 mg/dL were excluded. Frontal maximum-intensity projection positron emission tomography images were extracted, standardized, and analyzed along with information regarding the tumor location. Six LLMs (GPT-5, GPT-4.5, GPT-4.1, OpenAI-o3, -o1, and GPT-4 Turbo) and 4 blinded human evaluators (a nuclear medicine specialist, a gastrointestinal surgeon, and 2 radiology residents) assessed the presence of thoracic and abdominal LN metastases on a region-level basis and determined cN and cM staging on a patient-level basis. The model analyses were performed using the application programming interface in a zero-shot setting. Radiology reports served as the reference standard. Diagnostic agreement and accuracy were evaluated using Cohen ฮบ and the Cochran Q test. Additionally, to account for the class imbalance in the dataset, the Matthews Correlation Coefficient was calculated as a robust metric for binary classification performance. Post hoc McNemar tests were performed with Bonferroni correction; statistical significance for pairwise comparisons was set at P<.0083 (adjusted from P<.05) using JMP Pro (version 18.0; SAS Institute Inc).

RESULTS

The average accuracy was 41/120 (34%) to 94/120 (78%) for LLMs and 72/120 (60%) to 102/120 (85%) for physicians, with significantly higher accuracy for physicians (P<.05) in the thoracic LN, abdominal LN, and cN stages. Interrater reliability was slight to fair for LLMs (ฮบ: -0.07 to 0.25) and fair to substantial for physicians (ฮบ: 0.27 to 0.74). Matthews Correlation Coefficient scores were consistently higher for physicians (0.28 to 0.75) than for LLMs (-0.07 to 0.32). Among the LLMs, GPT-5 demonstrated the highest overall accuracy, with newer LLMs showing improved diagnostic accuracy when compared with previous models in identifying abdominal LN metastases and cM staging, though they showed weaker consistency for cN staging. For example, in thoracic LN detection, GPT-5 achieved 76/120 (63%) accuracy, whereas other LLMs achieved 72/120 (60%) or lower accuracy.

CONCLUSION

Although current LLMs have not yet reached physician-level accuracy in comprehensive staging, recent models show promise in assisting with specific diagnostic tasks.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์‹๋„ ํŽธํ‰์„ธํฌ์•” ํ™˜์ž 120๋ช…์˜ 18F-FDG-PET ์ตœ๋Œ€๊ฐ•๋„ํˆฌ์˜ ์˜์ƒ์„ ์ด์šฉํ•˜์—ฌ GPT-5๋ฅผ ํฌํ•จํ•œ 6์ข…์˜ ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ๋ฆผํ”„์ ˆ ์ „์ด ๋ฐ ์›๊ฒฉ์ „์ด ๋ณ‘๊ธฐ ํŒ๋… ์ •ํ™•๋„๋ฅผ 4๋ช…์˜ ์ž„์ƒ์˜(ํ•ต์˜ํ•™ ์ „๋ฌธ์˜, ์œ„์žฅ๊ด€ ์™ธ๊ณผ์˜, ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๊ณต์˜ 2๋ช…)์™€ ํ›„ํ–ฅ์ ์œผ๋กœ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. LLM์˜ ํ‰๊ท  ์ •ํ™•๋„๋Š” 34~78%, ์ž„์ƒ์˜๋Š” 60~85%๋กœ ์ž„์ƒ์˜๊ฐ€ ํ‰๋ถ€ ๋ฆผํ”„์ ˆ, ๋ณต๋ถ€ ๋ฆผํ”„์ ˆ ๋ฐ cN ๋ณ‘๊ธฐ ํŒ๋…์—์„œ ์œ ์˜ํ•˜๊ฒŒ ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์œผ๋ฉฐ(P<.05), ํ‰๊ฐ€์ž ๊ฐ„ ์‹ ๋ขฐ๋„ ์—ญ์‹œ ์ž„์ƒ์˜(ฮบ: 0.27~0.74)๊ฐ€ LLM(ฮบ: โˆ’0.07~0.25)๋ณด๋‹ค ์šฐ์ˆ˜ํ•˜์˜€๋‹ค. ํ˜„์žฌ์˜ LLM์€ ์•„์ง ์ž„์ƒ์˜ ์ˆ˜์ค€์˜ ์ข…ํ•ฉ์  ๋ณ‘๊ธฐ ๊ฒฐ์ • ๋Šฅ๋ ฅ์—๋Š” ๋ฏธ์น˜์ง€ ๋ชปํ•˜๋‚˜, GPT-5๋ฅผ ๋น„๋กฏํ•œ ์ตœ์‹  ๋ชจ๋ธ์€ ๋ณต๋ถ€ ๋ฆผํ”„์ ˆ ์ „์ด ๋ฐ cM ๋ณ‘๊ธฐ ํŒ๋… ๋“ฑ ํŠน์ • ์ง„๋‹จ ๊ณผ์ œ์—์„œ ๊ฐœ์„ ๋œ ์„ฑ๋Šฅ์„ ๋ณด์—ฌ ํ–ฅํ›„ ๋ฐฉ์‚ฌ์„  ํŒ๋… ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

69Integrating Fine-Tuning and Retrieval-Augmented Generation for Healthcare AI Systems: A Scoping Review.

2026-02Bioengineering (Basel, Switzerland)๐Ÿ”ท Q2DOI 10.3390/bioengineering13020225

(1)

BACKGROUND

Large language models (LLMs) show promise in healthcare but are constrained by hallucinations, static knowledge, and limited domain specificity. Fine-tuning (FT) and retrieval-augmented generation (RAG) offer complementary solutions, with FT embedding domain reasoning and RAG enabling dynamic, up-to-date knowledge access. Hybrid FT + RAG frameworks have been proposed to improve factual accuracy and clinical reliability. This scoping review synthesizes current evidence on such hybrids in healthcare AI. (2)

METHODS

The search across PubMed, IEEE Xplore, Google Scholar, and Embase identified studies implementing explicit FT + RAG hybrids in healthcare or biomedical tasks. Eligible studies reported empirical evaluations of LLM performance or behavior. Data were extracted on base models, FT strategies, RAG architectures, applications, and performance outcomes. (3)

RESULTS

Seven studies met inclusion criteria. FT + RAG systems consistently outperformed FT-only or RAG-only approaches across QA, clinical summarization, report generation, and decision support tasks. Parameter-efficient FT methods (e.g., LoRA) were common, while RAG implementations varied (dense, hybrid, hierarchical, multimodal, federated). Reported benefits included improved accuracy, reduced hallucination, and greater clinician preference and feasibility in protected settings. (4)

CONCLUSION

FT + RAG frameworks represent a promising direction for clinically grounded healthcare AI, combining domain-specific reasoning with transparent, up-to-date retrieval. Future work should prioritize standardized evaluation, workflow integration, and governance to enable safe deployment.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์Šค์ฝ”ํ•‘ ๋ฆฌ๋ทฐ๋Š” ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ํ™˜๊ฐ(hallucination), ์ •์  ์ง€์‹, ๋„๋ฉ”์ธ ํŠน์ด์„ฑ ํ•œ๊ณ„๋ฅผ ๊ทน๋ณตํ•˜๊ธฐ ์œ„ํ•ด ๋ฏธ์„ธ ์กฐ์ •(FT)๊ณผ ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG)์„ ๊ฒฐํ•ฉํ•œ ํ•˜์ด๋ธŒ๋ฆฌ๋“œ ํ”„๋ ˆ์ž„์›Œํฌ์˜ ์˜๋ฃŒ AI ์ ์šฉ ํ˜„ํ™ฉ์„ PubMed, IEEE Xplore, Google Scholar, Embase๋ฅผ ํ†ตํ•ด ์ฒด๊ณ„์ ์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. ์ตœ์ข… ์„ ์ •๋œ 7๊ฐœ ์—ฐ๊ตฌ์—์„œ FT+RAG ํ•˜์ด๋ธŒ๋ฆฌ๋“œ ์‹œ์Šคํ…œ์€ ์งˆ์˜์‘๋‹ต, ์ž„์ƒ ์š”์•ฝ, ๋ณด๊ณ ์„œ ์ƒ์„ฑ, ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ๋“ฑ ์ „ ์˜์—ญ์—์„œ FT ๋‹จ๋… ๋˜๋Š” RAG ๋‹จ๋… ๋ฐฉ์‹๋ณด๋‹ค ์ผ๊ด€๋˜๊ฒŒ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, LoRA ๋“ฑ ํŒŒ๋ผ๋ฏธํ„ฐ ํšจ์œจ์  ๋ฏธ์„ธ ์กฐ์ • ๊ธฐ๋ฒ•๊ณผ ๋ฐ€์ง‘ํ˜•ยท๊ณ„์ธตํ˜•ยท๋‹ค๋ชจ๋‹ฌยท์—ฐํ•ฉ RAG ๊ตฌ์กฐ๊ฐ€ ์ฃผ๋กœ ํ™œ์šฉ๋˜์—ˆ๋‹ค. ์ €์ž๋“ค์€ FT+RAG ํ”„๋ ˆ์ž„์›Œํฌ๊ฐ€ ๋„๋ฉ”์ธ ํŠนํ™” ์ถ”๋ก ๊ณผ ์ตœ์‹  ์ง€์‹ ๊ฒ€์ƒ‰์„ ๊ฒฐํ•ฉํ•œ ์œ ๋งํ•œ ์ž„์ƒ AI ๋ฐฉํ–ฅ์ž„์„ ๊ฒฐ๋ก ์ง€์œผ๋ฉฐ, ์•ˆ์ „ํ•œ ์ž„์ƒ ๋ฐฐํฌ๋ฅผ ์œ„ํ•ด ํ‘œ์ค€ํ™”๋œ ํ‰๊ฐ€ ์ฒด๊ณ„, ์›Œํฌํ”Œ๋กœ์šฐ ํ†ตํ•ฉ, ๊ฑฐ๋ฒ„๋„Œ์Šค ์ˆ˜๋ฆฝ์ด ํ–ฅํ›„ ๊ณผ์ œ๋กœ ํ•„์š”ํ•จ์„ ๊ฐ•์กฐํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

70Retrieval-augmented generation-enhanced large language models for comprehensive CAD-RADSย 2.0 categorization from structured coronary CTA reports.

2026-02Radiologie (Heidelberg, Germany)DOI 10.1007/s00117-026-01580-z
OBJECTIVE

To evaluate the performance of large language models (LLMs), including retrieval-augmented generation (RAG)-based approaches, in extracting components and management recommendations from structured coronary computed tomography angiography (CCTA) reports according to the Coronary Artery Disease Reporting and Data System (CAD-RADSย 2.0).

METHODS

Aย total of 320 fully structured CCTA reports were analyzed using LLM. Closed-source standard ChatGPTโ€‘5, NotebookLM (RAG-based model), and aย RAG-adapted ChatGPTโ€‘5 model (ChatGPT-5-RAG) were used. Each model extracted the CAD-RADS category, plaque burden, presence of high-risk plaque (HRP), other modifiers, full score, and management recommendations in accordance with the CAD-RADS 2.0ย guidelines. We compared LLM outputs with reference standards determined by two expert cardiovascular radiologists.

RESULTS

ChatGPT-5-RAG showed the highest accuracy for CAD-RADS classification (0.959, 95% CI: 0.932-0.976), plaque burden (0.912, 95% CI: 0.876-0.939), HRP detection (0.988, 95% CI: 0.968-0.995), other modifiers (0.950, 95% CI: 0.920-0.969), and full score (0.828, 95% CI: 0.783-0.866). Closed-source ChatGPTโ€‘5 showed the weakest performance across all components. Significant statistical differences were found among the three models (pโ€ฏ<โ€‰0.001). Management recommendations were qualitatively rated on aย three-point Likert scale; although agreement between models was low, ChatGPT-5-RAG and NotebookLM performed almost perfectly (median 3ย points).

CONCLUSION

This study demonstrates that RAG-enhanced LLMs significantly improve accuracy and reliability in extracting CAD-RADSย 2.0 components and generating clinical management recommendations. The findings highlight the potential of RAG-based LLMs as innovative, explainable tools for automated and standardized CCTA reporting in clinical radiology workflows. ZUSAMMENFASSUNG: ZIEL: Ziel der vorliegenden Arbeit war es, die Leistungsfรคhigkeit von Large-Language-Modellen (LLM) bei der Herausfilterung von Komponenten und Therapieempfehlungen aus strukturierten Befunden von Computertomographie-Angiographien der Koronarien (CCTA) gemรครŸ Coronary Artery Disease Reporting and Data System (CAD-RADSย 2.0) zu untersuchen, einschlieรŸlich Ansรคtzen auf der Grundlage von durch Abfrage von Informationen verbesserter Generierung von Texten (โ€žretrieval-augmented generationโ€œ, RAG). MATERIAL UND METHODE: Es wurden 320 vollstรคndig strukturierte CCTA-Befunde mittels LLM ausgewertet. Dafรผr wurden Closed-Source-Standard-ChatGPTโ€‘5, NotebookLM (RAG-basiertes Modell) und ein RAG-adaptiertes ChatGPT-5-Modell (ChatGPT-5-RAG) eingesetzt. Jedes Modell extrahierte die CAD-RADS-Kategorie, Plaquelast, das Vorliegen von Hochrisikoplaques (HRP), sonstige Modifikatoren, den vollstรคndigen Score und die Therapieempfehlungen in รœbereinstimmung mit den CAD-RADSโ€‘2.0โ€‘Leitlinien. Die LLM-Ergebnisse wurden mit Referenzstandards verglichen, die durch 2ย Fachรคrzte fรผr den Bereich kardiovaskulรคre Radiologie festgelegt worden waren. ERGEBNISSE: ChatGPT-5-RAG wies die grรถรŸte Genauigkeit auf bei der CAD-RADS-Klassifizierung (0,959; 95%-Konfidenzintervall, 95%-KI: 0,932โ€“0,976), Plaquelast (0,912; 95%-KI: 0,876โ€“0,939), HRP-Erkennung (0,988; 95%-KI: 0,968โ€“0,995), sonstigen Modifikatoren (0,950; 95%-KI: 0,920โ€“0,969) und vollstรคndigem Score (0,828; 95%-KI: 0,783โ€“0,866). Fรผr Closed-Source-ChatGPTโ€‘5 zeigte sich die schwรคchste Leistung hinsichtlich aller Komponenten. Zwischen den 3ย Modellen wurden signifikante statistische Unterschiede festgestellt (pโ€ฏ<โ€‰0,001). Die Therapieempfehlungen wurden qualitativ auf einer 3โ€‘Punkt-Likert-Skala bewertet; auch wenn die รœbereinstimmung zwischen den Modellen gering war, war die Leistung von ChatGPT-5-RAG und NotebookLM fast perfekt (im Mittel 3ย Punkte). SCHLUSSFOLGERUNG: In der vorliegenden Studie wurde gezeigt, dass mittels RAG verbesserte LLM die Genauigkeit und Verlรคsslichkeit bei der Herausfilterung von CAD-RADSโ€‘2.0โ€‘Komponenten und der Generierung klinischer Therapieempfehlungen signifikant verbessern. Diese Ergebnisse unterstreichen das Potenzial von RAG-basierten LLM als innovative, erklรคrbare Instrumente fรผr die automatisierte und standardisierte CCTA-Befunderstellung in den Arbeitsablรคufen der klinischen Radiologie.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ด€์ƒ๋™๋งฅ CT ํ˜ˆ๊ด€์กฐ์˜(CCTA) ๋ณด๊ณ ์„œ์—์„œ CAD-RADS 2.0 ๊ตฌ์„ฑ ์š”์†Œ ๋ฐ ์ž„์ƒ ๊ด€๋ฆฌ ๊ถŒ๊ณ ์‚ฌํ•ญ์„ ์ž๋™ ์ถ”์ถœํ•˜๋Š” ๋ฐ ์žˆ์–ด ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM), ํŠนํžˆ ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG) ๊ธฐ๋ฐ˜ ์ ‘๊ทผ๋ฒ•์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ด 320๊ฑด์˜ ๊ตฌ์กฐํ™”๋œ CCTA ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ ํ‘œ์ค€ ChatGPT-5, NotebookLM(RAG ๊ธฐ๋ฐ˜), ๊ทธ๋ฆฌ๊ณ  RAG ์ ์šฉ ChatGPT-5(ChatGPT-5-RAG) ์„ธ ๋ชจ๋ธ์„ ๋น„๊ตํ•˜์˜€์œผ๋ฉฐ, ๋‘ ๋ช…์˜ ์‹ฌํ˜ˆ๊ด€ ์˜์ƒ์˜ํ•™ ์ „๋ฌธ์˜๊ฐ€ ์„ค์ •ํ•œ ๊ธฐ์ค€๊ฐ’๊ณผ์˜ ์ผ์น˜๋„๋ฅผ ๋ถ„์„ํ•˜์˜€๋‹ค. ChatGPT-5-RAG๊ฐ€ CAD-RADS ๋ถ„๋ฅ˜(์ •ํ™•๋„ 0.959), ์ฃฝ์ƒ๋ฐ˜ ๋ถ€๋‹ด, ๊ณ ์œ„ํ—˜ ์ฃฝ์ƒ๋ฐ˜ ๊ฒ€์ถœ(์ •ํ™•๋„ 0.988) ๋“ฑ ๋ชจ๋“  ํ•ญ๋ชฉ์—์„œ ๊ฐ€์žฅ ๋†’์€ ์„ฑ๋Šฅ์„ ๋ณด์—ฌ, RAG ๊ธฐ๋ฐ˜ LLM์ด ํ‘œ์ค€ํ™”๋œ CCTA ๋ณด๊ณ ์„œ ์ž๋™ํ™”๋ฅผ ์œ„ํ•œ ์ •ํ™•ํ•˜๊ณ  ์„ค๋ช… ๊ฐ€๋Šฅํ•œ ์ž„์ƒ ๋„๊ตฌ๋กœ์„œ์˜ ๋†’์€ ์ž ์žฌ๋ ฅ์„ ์ง€๋‹˜์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

71Case-matched retrieval improves textual alignment of LLM-generated radiology impressions.

2026PloS oneโญ Q1DOI 10.1371/journal.pone.0354688
BACKGROUND

Radiology impressions guide clinical care. Large Language Models (LLMs)-drafted impressions can drift into generic, off-style text. Retrieval-augmented generation (RAG) enables context-aware few-shot prompting during inference.

METHODS

This retrospective IRB-approved study included 11,998 CT pulmonary angiography (CTPA) reports. We built a retrieval bank from 11,399 reports and reserved 599 reports for testing. GPT-4o and LLaMA 3.1-70B generated impressions from the "findings" section using three setups: zero-shot, fixed random few-shot, and dynamic retrieval-selected few-shot (top-k semantic matches; kโ€‰=โ€‰3/5/10). We ran temperatures 0, 0.7, 1. We scored outputs against the original impressions with ROUGE and BERTScore F1, report mean scores with 95% confidence intervals, and tested for statistical significance using Wilcoxon signed-rank test.

RESULTS

Dynamic retrieval-based few-shot prompting outperformed zero-shot and fixed few-shot prompting across all configurations (all pโ€‰<โ€‰0.05). The highest scores were observed at temperature 0 and kโ€‰=โ€‰10. ROUGE-1 F1 increased to 0.44-0.47 for GPT-4o and 0.37-0.50 for LLaMA, versus 0.35-0.37 and 0.25-0.37, respectively, in zero-shot prompting. Lower temperature and larger k were associated with higher similarity scores.

CONCLUSION

Dynamic, case-matched retrieval improved alignment of LLM-generated CTPA impressions with reference impressions on automated text-similarity metrics. Scores remained moderate, and radiologists' verification is still required before clinical deployment.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ(LLM)์ด ์ƒ์„ฑํ•œ ๋ฐฉ์‚ฌ์„  ํŒ๋… ์†Œ๊ฒฌ์„œ์˜ ๋ฌธ์ฒด ์ผ๊ด€์„ฑ์„ ํ–ฅ์ƒ์‹œํ‚ค๊ธฐ ์œ„ํ•ด ๋™์  ์‚ฌ๋ก€ ๊ธฐ๋ฐ˜ ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG) ๊ธฐ๋ฒ•์˜ ํšจ์šฉ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 11,998๊ฑด์˜ CT ํ๋™๋งฅ์กฐ์˜์ˆ (CTPA) ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ GPT-4o ๋ฐ LLaMA 3.1-70B ๋ชจ๋ธ์— ์ œ๋กœ์ƒท, ๊ณ ์ • ํ“จ์ƒท, ๋™์  ๊ฒ€์ƒ‰ ๊ธฐ๋ฐ˜ ํ“จ์ƒท(์ƒ์œ„ k=3/5/10 ์˜๋ฏธ๋ก ์  ์œ ์‚ฌ ์‚ฌ๋ก€) ๋ฐฉ์‹์„ ์ ์šฉํ•˜์—ฌ ROUGE ๋ฐ BERTScore F1์œผ๋กœ ์›๋ณธ ์†Œ๊ฒฌ๊ณผ์˜ ์œ ์‚ฌ๋„๋ฅผ ๋น„๊ตํ•˜์˜€๋‹ค. ๋™์  ๊ฒ€์ƒ‰ ๊ธฐ๋ฐ˜ ํ“จ์ƒท ๋ฐฉ์‹์ด ๋ชจ๋“  ์„ค์ •์—์„œ ๋‚˜๋จธ์ง€ ๋ฐฉ์‹๋ณด๋‹ค ์œ ์˜ํ•˜๊ฒŒ ๋†’์€ ์œ ์‚ฌ๋„ ์ ์ˆ˜๋ฅผ ๋‚˜ํƒ€๋ƒˆ์œผ๋‚˜(p < 0.05), ์ ์ˆ˜๋Š” ์—ฌ์ „ํžˆ ์ค‘๊ฐ„ ์ˆ˜์ค€์— ๋จธ๋ฌผ๋Ÿฌ ์ž„์ƒ ์ ์šฉ ์ „ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜์˜ ์ตœ์ข… ๊ฒ€ํ† ๊ฐ€ ํ•„์ˆ˜์ ์ž„์„ ํ™•์ธํ•˜์˜€๋‹ค.
Added: 2026-08-02 00:00View โ†—

72Risk-Aware Prompt-Guided Framework for Brain MRI Diagnostic Report Generation with Radiologist Feedback Integration

2026International Journal of Drug Delivery TechnologyDOI 10.25258/ijddt.16.47s.150
Added: 2026-06-14 00:00View โ†—

73Accuracy and reproducibility of large language model measurements of liver metastases: comparison with radiologist measurements

2026Japanese Journal of Radiologyโญ Q1DOI 10.1007/s11604-025-01884-5
Added: 2026-04-21 16:11View โ†—

74Application of Artificial Intelligence in Medical Education: A Systematic and Narrative Review of Pedagogical Potential and Ethical Implications.

2026Advances in medical education and practice๐Ÿ”ท Q2DOI 10.2147/amep.s567190

Artificial intelligence (AI) is rapidly transforming medical education through large language models (LLMs), virtual reality (VR), intelligent tutoring systems, and decision-support platforms. These tools enable adaptive instruction, immersive simulation, and real-time feedback, showing strong potential to improve outcomes across health professions training. To explore both opportunities and risks, we conducted a systematic review of PubMed, EMBASE, Web of Science, and Scopus for English-language studies published between January 2015 and May 2025, following the PRISMA framework. Nineteen studies met eligibility criteria. AI modalities identified included LLMs such as ChatGPT, VR-based simulation systems, automated tutoring platforms, and clinical decision-support tools, spanning specialties including radiology, surgery, and psychiatry. Across contexts, AI enhanced examination performance, procedural competence, self-directed learning, engagement, and motivation relative to traditional methods. Students and faculty expressed strong interest and optimism but reported limited formal AI training, favoring interactive practice over didactic lectures. Despite these benefits, concerns consistently emerged regarding algorithmic bias, inaccuracy, data security, and the necessity of human oversight in educational and clinical settings. Ethical issues such as job displacement, the erosion of humanistic care, and the impact on the patient-physician relationship were also highlighted. Limited formal AI training, uneven institutional readiness, and gaps in faculty expertise were common challenges across regions.To harness its transformative potential responsibly, investment is required in faculty development, structured curricula addressing both technical and ethical competencies, and governance frameworks that ensure equitable, transparent, and accountable use. Properly integrated, AI can not only personalize learning and expand access but also support a more inclusive and ethically grounded vision for the future of medical education.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
์ด ์—ฐ๊ตฌ๋Š” ์˜ํ•™๊ต์œก์—์„œ ์ธ๊ณต์ง€๋Šฅ(AI) ํ™œ์šฉ์˜ ๊ต์œก์  ์ž ์žฌ๋ ฅ๊ณผ ์œค๋ฆฌ์  ํ•จ์˜๋ฅผ ํƒ์ƒ‰ํ•˜๊ธฐ ์œ„ํ•ด 2015๋…„ 1์›”๋ถ€ํ„ฐ 2025๋…„ 5์›”๊นŒ์ง€ ๋ฐœํ‘œ๋œ ๋ฌธํ—Œ์„ PRISMA ์ง€์นจ์— ๋”ฐ๋ผ ์ฒด๊ณ„์ ์œผ๋กœ ๊ฒ€ํ† ํ•˜์˜€์œผ๋ฉฐ, ์ตœ์ข… 19๊ฐœ ์—ฐ๊ตฌ๋ฅผ ๋ถ„์„ ๋Œ€์ƒ์œผ๋กœ ์„ ์ •ํ•˜์˜€๋‹ค. ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ(LLM), ๊ฐ€์ƒํ˜„์‹ค ๊ธฐ๋ฐ˜ ์‹œ๋ฎฌ๋ ˆ์ด์…˜, ์ง€๋Šฅํ˜• ํŠœํ„ฐ๋ง ์‹œ์Šคํ…œ ๋“ฑ ๋‹ค์–‘ํ•œ AI ๋„๊ตฌ๋“ค์ด ์‹œํ—˜ ์„ฑ์  ํ–ฅ์ƒ, ์ˆ ๊ธฐ ์—ญ๋Ÿ‰ ๊ฐ•ํ™”, ์ž๊ธฐ์ฃผ๋„ ํ•™์Šต ๋ฐ ํ•™์Šต ๋™๊ธฐ ์ฆ์ง„์— ์žˆ์–ด ๊ธฐ์กด ๊ต์œก ๋ฐฉ์‹ ๋Œ€๋น„ ์œ ์˜๋ฏธํ•œ ํšจ๊ณผ๋ฅผ ๋ณด์˜€๋‹ค. ๊ทธ๋Ÿฌ๋‚˜ ์•Œ๊ณ ๋ฆฌ์ฆ˜ ํŽธํ–ฅ, ๋ถ€์ •ํ™•์„ฑ, ๋ฐ์ดํ„ฐ ๋ณด์•ˆ, ์ธ๊ฐ„์  ๋Œ๋ด„์˜ ์•ฝํ™”, ํ™˜์ž-์˜์‚ฌ ๊ด€๊ณ„์— ๋Œ€ํ•œ ์˜ํ–ฅ ๋“ฑ ์œค๋ฆฌ์  ์šฐ๋ ค๊ฐ€ ์ง€์†์ ์œผ๋กœ ์ œ๊ธฐ๋˜์—ˆ์œผ๋ฉฐ, AI๋ฅผ ์˜ํ•™๊ต์œก์— ์ฑ…์ž„๊ฐ ์žˆ๊ฒŒ ํ†ตํ•ฉํ•˜๊ธฐ ์œ„ํ•ด์„œ๋Š” ๊ต์ˆ˜ ์—ญ๋Ÿ‰ ๊ฐœ๋ฐœ, ๊ธฐ์ˆ ์ ยท์œค๋ฆฌ์  ์—ญ๋Ÿ‰์„ ์•„์šฐ๋ฅด๋Š” ์ฒด๊ณ„์  ๊ต์œก๊ณผ์ •, ๊ทธ๋ฆฌ๊ณ  ํ˜•ํ‰์„ฑ๊ณผ ํˆฌ๋ช…์„ฑ์„ ๋ณด์žฅํ•˜๋Š” ๊ฑฐ๋ฒ„๋„Œ์Šค ์ฒด๊ณ„ ๊ตฌ์ถ•์ด ํ•„์š”ํ•˜๋‹ค๊ณ  ๊ฐ•์กฐํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

75Large language modelโ€“assisted radiology reporting in a single-radiologist implementation: a retrospective cohort study interpreted through a UTAUT lens

2026Abdominal Radiologyโญ Q1DOI 10.1007/s00261-026-05524-y
Added: 2026-04-21 16:11View โ†—

76Semiautomated breast ultrasound report generation using multimodal large language models and deep learning.

2026Frontiers in medicineโญ Q1DOI 10.3389/fmed.2026.1679203
BACKGROUND

Breast ultrasound (US) imaging is essential for early breast cancer detection, yet generating diagnostic reports is labor-intensive, particularly when incorporating multimodal elastography.

METHODS

This study presents a novel framework that combines multimodal large language models and deep learning to generate semiautomated breast US reports. This framework bridges the gap between manual and fully automated workflows by integrating radiologist annotations with advanced image classification and structured report compilation. A total of 2,119 elastography images and 60 annotated patient cases were retrospectively collected from two US machines.

RESULTS

The system demonstrated robust performance in elastography classification, achieving areas under the receiver operating characteristic curve of 0.92, 0.91, and 0.88 for shear-wave, strain, and Doppler images, respectively. In the evaluated dataset, the report generation module correctly identified all suspicious masses across both US machines, achieving 100% sensitivity in lesion detection, with an average report generation time of 31 s per patient using the GE Healthcare machine and 36 s using the Supersonic Image machine.

CONCLUSION

The proposed framework enables accurate, efficient, and device-adaptable breast US report generation by combining multimodal DL and prompt-based LLM inference. It significantly reduces radiologist workload and demonstrates potential for scalable deployment in real-world clinical workflows.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋‹ค์ค‘๋ชจ๋‹ฌ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)๊ณผ ๋”ฅ๋Ÿฌ๋‹์„ ๊ฒฐํ•ฉํ•˜์—ฌ ์œ ๋ฐฉ ์ดˆ์ŒํŒŒ ์˜์ƒ ํŒ๋…๋ฌธ์„ ๋ฐ˜์ž๋™์œผ๋กœ ์ƒ์„ฑํ•˜๋Š” ์ƒˆ๋กœ์šด ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ์ œ์•ˆํ•˜์˜€์œผ๋ฉฐ, ํƒ„์„ฑ์ดˆ์ŒํŒŒ(elastography) ์ด๋ฏธ์ง€ ๋ถ„๋ฅ˜์™€ ๊ตฌ์กฐํ™”๋œ ๋ณด๊ณ ์„œ ์ž‘์„ฑ์„ ํ†ตํ•ฉํ•˜์˜€๋‹ค. 2,119๊ฐœ์˜ ํƒ„์„ฑ์ดˆ์ŒํŒŒ ์ด๋ฏธ์ง€์™€ 60๊ฑด์˜ ์ฃผ์„ ์ฒ˜๋ฆฌ๋œ ํ™˜์ž ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•œ ๊ฒ€์ฆ ๊ฒฐ๊ณผ, ์ „๋‹จํŒŒยท๋ณ€ํ˜•๋ฅ ยท๋„ํ”Œ๋Ÿฌ ์˜์ƒ ๋ถ„๋ฅ˜์—์„œ ๊ฐ๊ฐ AUC 0.92, 0.91, 0.88์˜ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ์˜์‹ฌ ๋ณ‘๋ณ€ ํƒ์ง€ ๋ฏผ๊ฐ๋„ 100%๋ฅผ ๋‹ฌ์„ฑํ•˜์˜€๋‹ค. ํ™˜์ž๋‹น ํ‰๊ท  31โ€“36์ดˆ์˜ ์‹ ์†ํ•œ ๋ณด๊ณ ์„œ ์ƒ์„ฑ์ด ๊ฐ€๋Šฅํ•˜์—ฌ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ์˜ ์—…๋ฌด ๋ถ€๋‹ด์„ ํ˜„์ €ํžˆ ์ค„์ด๊ณ  ์‹ค์ œ ์ž„์ƒ ํ™˜๊ฒฝ์—์„œ์˜ ํ™•์žฅ ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

77Coarse-to-Fine Personalized LLM Impressions for Streamlined Radiology Reports

2025-01-01SSRN Electronic JournalDOI 10.2139/ssrn.5374739

Preprint on personalized radiology impression generation using open-source LLMs with parameter-efficient fine-tuning and human preference learning to adapt to individual radiologist style.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ฐœ๋ณ„ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜์˜ ์ž‘์„ฑ ์Šคํƒ€์ผ์— ๋งž์ถคํ™”๋œ ์˜์ƒ ํŒ๋… ์†Œ๊ฒฌ๋ฌธ(impression) ์ž๋™ ์ƒ์„ฑ ์‹œ์Šคํ…œ์„ ๊ฐœ๋ฐœํ•˜๋Š” ๊ฒƒ์„ ๋ชฉ์ ์œผ๋กœ ํ•œ๋‹ค. ์˜คํ”ˆ์†Œ์Šค ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์— ํŒŒ๋ผ๋ฏธํ„ฐ ํšจ์œจ์  ๋ฏธ์„ธ์กฐ์ •(parameter-efficient fine-tuning) ๋ฐ ์ธ๊ฐ„ ์„ ํ˜ธ๋„ ํ•™์Šต(human preference learning)์„ ์ ์šฉํ•˜๋Š” ๊ฑฐ์นœ-์ •๋ฐ€(coarse-to-fine) ์ ‘๊ทผ๋ฒ•์„ ์ฑ„ํƒํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, ๊ฐœ๋ณ„ ํŒ๋…์˜์˜ ๊ณ ์œ ํ•œ ํ‘œํ˜„ ๋ฐฉ์‹๊ณผ ์ž„์ƒ์  ๊ฐ•์กฐ์ ์„ ๋ฐ˜์˜ํ•œ ๊ฐœ์ธํ™”๋œ ์†Œ๊ฒฌ๋ฌธ ์ƒ์„ฑ์ด ๊ฐ€๋Šฅํ•จ์„ ํ™•์ธํ•˜์—ฌ, ์˜์ƒ ํŒ๋… ์—…๋ฌด์˜ ํšจ์œจํ™” ๊ฐ€๋Šฅ์„ฑ์„ ์ œ์‹œํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

78Radiology Report Generation via Multi-objective Preference Optimization

2025-01-01Proceedings of the AAAI Conference on Artificial IntelligenceDOI 10.1609/aaai.v39i8.32936

Preference optimization approach for radiology report generation that models differing clinician preferences, such as fluency versus clinical accuracy, within a personalized report generation framework.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ์ž๋™ ์ƒ์„ฑ ์‹œ ์ž„์ƒ์˜๋งˆ๋‹ค ์ƒ์ดํ•œ ์„ ํ˜ธ๋„(์œ ์ฐฝ์„ฑ ๋Œ€ ์ž„์ƒ์  ์ •ํ™•์„ฑ ๋“ฑ)๋ฅผ ๋ฐ˜์˜ํ•˜๊ธฐ ์œ„ํ•ด ๋‹ค๋ชฉ์  ์„ ํ˜ธ๋„ ์ตœ์ ํ™”(multi-objective preference optimization) ๊ธฐ๋ฒ•์„ ์ œ์•ˆํ•˜์˜€๋‹ค. ๊ฐœ์ธํ™”๋œ ๋ณด๊ณ ์„œ ์ƒ์„ฑ ํ”„๋ ˆ์ž„์›Œํฌ ๋‚ด์—์„œ ๋ณต์ˆ˜์˜ ์ž„์ƒ ๋ชฉํ‘œ๋ฅผ ๋™์‹œ์— ๋ชจ๋ธ๋งํ•จ์œผ๋กœ์จ, ๋‹จ์ผ ์ตœ์ ํ™” ๊ธฐ์ค€์˜ ํ•œ๊ณ„๋ฅผ ๊ทน๋ณตํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ์ด๋ฅผ ํ†ตํ•ด ์ž„์ƒ์˜ ๊ฐœ์ธ์˜ ์ž‘์„ฑ ์Šคํƒ€์ผ๊ณผ ์ง„๋‹จ ์ค‘์  ์‚ฌํ•ญ์„ ๋ฐ˜์˜ํ•œ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ์ƒ์„ฑ์ด ๊ฐ€๋Šฅํ•จ์„ ๋ณด์—ฌ์ฃผ์—ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

79BP-LLM: Belief Propagation for Binary Feedback in Large Language Model Alignment

2025UnknownDOI 10.36227/techrxiv.176521139.94618209/v1
Added: 2026-04-21 16:11View โ†—

80Comments on โ€œStructured Report Generation for Breast Cancer Imaging Based on Large Language Modeling: A Comparative Analysis of GPT-4 and DeepSeekโ€

2025Academic Radiologyโญ Q1DOI 10.1016/j.acra.2025.12.017
Added: 2026-04-21 16:11View โ†—

81Comparison of a Specialized Large Language Model with GPT-4o for CT and MRI Radiology Report Summarization | Tweetorial

2025Radiologyโญ Q1DOI 10.1148/radiol.243774.tweetorial
Added: 2026-04-21 16:11View โ†—

82Engaging Preference Optimization Alignment in Large Language Model for Continual Radiology Report Generation: A Hybrid Approach

2025Cognitive Computationโญ Q1DOI 10.1007/s12559-025-10404-6
Added: 2026-04-21 16:11View โ†—

83HUMAN-IN-THE-LOOP AI FOR RADIOLOGY REPORT GENERATION: A METHODOLOGY FRAMEWORK FOR CONTINUOUS LEARNING FROM EXPERT FEEDBACK

2025Proceedings of the 22nd International Conference on Applied Computing 2025 and 24th International Conference on WWW/Internet 2025DOI 10.33965/ac_icwi_2025_202508l006
Added: 2026-04-21 16:11View โ†—

84Improving CXR Report Labeling Through LLM Fine-Tuning and Human Feedback

2025UnknownDOI 10.20944/preprints202504.1668.v1
Added: 2026-04-21 16:11View โ†—

85Large Language Models in radiology: A technical and clinical perspective

2025European Journal of Radiology Artificial IntelligenceDOI 10.1016/j.ejrai.2025.100021

<h2>Abstract</h2> Large Language Models (LLMs) are transformer-based deep learning models trained on vast text corpora, enabling advanced natural language understanding and generation. This technical note presents a comprehensive overview of LLMs in the field of radiology, following a style akin to high-impact journals. We discuss applications of LLMs in radiology, including image interpretation, automated report generation, and workflow efficiency improvements. We provide in-depth technical insight into transformer architectures and specific models such as GPT-4, PaLM, and Med-PaLM, including key mathematical formulations and algorithms that underpin LLM functionality. We examine evaluation metrics for radiology LLM tasks (accuracy, language metrics, etc.), along with critical considerations of bias, ethical issues, and regulatory guidelines. Finally, we address practical considerations for radiologists and AI researchers, offering perspectives on integrating LLMs into clinical practice and research. The aim is to combine technical rigor with clear explanations to serve a mixed audience of clinicians and AI scientists. Figures and tables are included as placeholders for clarity, and references are provided in a numbered format for further reading.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ๋…ผ๋ฌธ์€ ํŠธ๋žœ์Šคํฌ๋จธ ๊ธฐ๋ฐ˜ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ๋ฐฉ์‚ฌ์„ ํ•™ ๋ถ„์•ผ ์ ์šฉ ํ˜„ํ™ฉ์„ ๊ธฐ์ˆ ์ ยท์ž„์ƒ์  ๊ด€์ ์—์„œ ์ข…ํ•ฉ์ ์œผ๋กœ ๊ณ ์ฐฐํ•˜๋Š” ๊ฒƒ์„ ๋ชฉ์ ์œผ๋กœ ํ•œ๋‹ค. GPT-4, PaLM, Med-PaLM ๋“ฑ ์ฃผ์š” ๋ชจ๋ธ์˜ ์•„ํ‚คํ…์ฒ˜์™€ ํ•ต์‹ฌ ์•Œ๊ณ ๋ฆฌ์ฆ˜์„ ๋ถ„์„ํ•˜๊ณ , ์˜์ƒ ํŒ๋…, ์ž๋™ ๋ณด๊ณ ์„œ ์ƒ์„ฑ, ์›Œํฌํ”Œ๋กœ์šฐ ํšจ์œจํ™” ๋“ฑ ๋ฐฉ์‚ฌ์„ ํ•™ ์ž„์ƒ ์ ์šฉ ์‚ฌ๋ก€๋ฅผ ์ฒด๊ณ„์ ์œผ๋กœ ๊ฒ€ํ† ํ•˜์˜€๋‹ค. ์•„์šธ๋Ÿฌ ํ‰๊ฐ€ ์ง€ํ‘œ, ํŽธํ–ฅ, ์œค๋ฆฌ์  ๋ฌธ์ œ, ๊ทœ์ œ ์ง€์นจ ๋“ฑ ์‹ค์ œ ์ž„์ƒ ๋„์ž… ์‹œ ๊ณ ๋ คํ•ด์•ผ ํ•  ํ•ต์‹ฌ ์‚ฌํ•ญ์„ ์ œ์‹œํ•จ์œผ๋กœ์จ, ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ์™€ AI ์—ฐ๊ตฌ์ž ๋ชจ๋‘์—๊ฒŒ LLM์˜ ์ž„์ƒ ํ†ตํ•ฉ์„ ์œ„ํ•œ ์‹ค์งˆ์  ์ง€์นจ์„ ์ œ๊ณตํ•˜๊ณ  ์žˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

86Large language models in methodological quality evaluation of radiomics research based on METRICS: ChatGPT vs NotebookLM vs radiologist

2025European Journal of Radiologyโญ Q1DOI 10.1016/j.ejrad.2025.111960
Added: 2026-04-21 16:11View โ†—

87Large language models-powered clinical decision support: enhancing or replacing human expertise?

2025Intelligent Medicineโญ Q1DOI 10.1016/j.imed.2025.01.001

This editorial presents an optimistic yet cautious perspective on the development, deployment, and regulation of large language models (LLMs) in the field of medicine. It is essential to strike a balance between embracing the benefits of artificial intelligence-driven solutions and preserving the human touch that is vital for providing compassionate care. The exponential growth of medical data has paved the way for the integration of LLMs into healthcare, offering unprecedented opportunities to enhance clinical decision-making and alleviate physicians' workloads. Recently, LLMs have exhibited remarkable potential across various clinical scenarios, including streamlining diagnostic processes, optimizing radiology reports, and providing personalized treatment recommendations. However, the implementation of LLMs in healthcare is not without its challenges. Issues such as the scarcity of high-quality annotated data, privacy concerns, and the risk of generating misleading or overconfident information are significant hurdles that must be addressed. Moreover, while LLMs can replace certain basic tasks traditionally performed by humans, it is crucial to recognize that senior clinicians play an irreplaceable role in complex decision-making and providing emotional support to patients. By harnessing the power of LLMs to augment human capabilities while maintaining essential human elements within healthcare, we might shape a future where artificial intelligence and human intelligence coexist harmoniously. Prioritizing ethical development and deployment for artificial intelligence, empowering healthcare professionals, and safeguarding patient privacy will be key to realizing the full potential of LLMs in revolutionizing healthcare delivery. Through ongoing research, collaboration, and adaptation, responsible integration of LLMs holds promise for elevating both quality and accessibility globally, ultimately creating a more efficient, personalized, and patient-centric healthcare system.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—๋””ํ† ๋ฆฌ์–ผ์€ ์˜๋ฃŒ ๋ถ„์•ผ์—์„œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ๊ฐœ๋ฐœ, ๋„์ž… ๋ฐ ๊ทœ์ œ์— ๊ด€ํ•ด ๋‚™๊ด€์ ์ด๋ฉด์„œ๋„ ์‹ ์ค‘ํ•œ ๊ด€์ ์„ ์ œ์‹œํ•˜๋ฉฐ, AI ๊ธฐ๋ฐ˜ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์›์ด ์ง„๋‹จ ๊ฐ„์†Œํ™”, ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ์ตœ์ ํ™”, ๊ฐœ์ธ ๋งž์ถคํ˜• ์น˜๋ฃŒ ๊ถŒ๊ณ  ๋“ฑ ๋‹ค์–‘ํ•œ ์ž„์ƒ ์‹œ๋‚˜๋ฆฌ์˜ค์—์„œ ์˜๋ฏธ ์žˆ๋Š” ๊ฐ€๋Šฅ์„ฑ์„ ๋ณด์ด๊ณ  ์žˆ์Œ์„ ๊ณ ์ฐฐํ•˜์˜€๋‹ค. ๋‹ค๋งŒ ๊ณ ํ’ˆ์งˆ ์ฃผ์„ ๋ฐ์ดํ„ฐ ๋ถ€์กฑ, ํ™˜์ž ํ”„๋ผ์ด๋ฒ„์‹œ ์นจํ•ด ์šฐ๋ ค, ์˜ค์ •๋ณด ์ƒ์„ฑ ์œ„ํ—˜ ๋“ฑ์˜ ๊ณผ์ œ๊ฐ€ ์กด์žฌํ•˜๋ฉฐ, LLM์ด ์ผ๋ถ€ ๊ธฐ๋ณธ์  ์—…๋ฌด๋ฅผ ๋Œ€์ฒดํ•  ์ˆ˜ ์žˆ์–ด๋„ ๋ณต์žกํ•œ ์˜์‚ฌ๊ฒฐ์ • ๋ฐ ํ™˜์ž์— ๋Œ€ํ•œ ์ •์„œ์  ์ง€์ง€์—์„œ ์ˆ™๋ จ๋œ ์ž„์ƒ์˜์˜ ์—ญํ• ์€ ๋Œ€์ฒด ๋ถˆ๊ฐ€ํ•จ์„ ๊ฐ•์กฐํ•˜์˜€๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ, LLM์€ ์ธ๊ฐ„์˜ ์—ญ๋Ÿ‰์„ ๋ณด์™„ํ•˜๋Š” ๋ฐฉํ–ฅ์œผ๋กœ ์œค๋ฆฌ์ ์ด๊ณ  ์ฑ…์ž„๊ฐ ์žˆ๊ฒŒ ํ†ตํ•ฉ๋  ๋•Œ ์˜๋ฃŒ ์„œ๋น„์Šค์˜ ์งˆ๊ณผ ์ ‘๊ทผ์„ฑ์„ ์ „ ์„ธ๊ณ„์ ์œผ๋กœ ํ–ฅ์ƒ์‹œํ‚ค๋Š” ๋ฐ ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ๋‹ค๊ณ  ์ œ์–ธํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

88The radiology report: a clinician and a radiologist perspective in a resource-limited setting.

2025Journal Africain d Imagerie Mรฉdicale (J Afr Imag Mรฉd) Journal Officiel de la Sociรฉtรฉ de Radiologie dโ€™Afrique Noire Francophone (SRANF)DOI 10.55715/jaim.v17i1.750
Added: 2026-04-21 16:11View โ†—

89Adapted large language models can outperform medical experts in clinical text summarization

2024Nature Medicineโญ Q1DOI 10.1038/s41591-024-02855-5
Added: 2026-04-21 16:11View โ†—

90Evaluation and mitigation of the limitations of large language models in clinical decision-making

2024Nature Medicineโญ Q1DOI 10.1038/s41591-024-03097-1

Clinical decision-making is one of the most impactful parts of a physician's responsibilities and stands to benefit greatly from artificial intelligence solutions and large language models (LLMs) in particular. However, while LLMs have achieved excellent performance on medical licensing exams, these tests fail to assess many skills necessary for deployment in a realistic clinical decision-making environment, including gathering information, adhering to guidelines, and integrating into clinical workflows. Here we have created a curated dataset based on the Medical Information Mart for Intensive Care database spanning 2,400 real patient cases and four common abdominal pathologies as well as a framework to simulate a realistic clinical setting. We show that current state-of-the-art LLMs do not accurately diagnose patients across all pathologies (performing significantly worse than physicians), follow neither diagnostic nor treatment guidelines, and cannot interpret laboratory results, thus posing a serious risk to the health of patients. Furthermore, we move beyond diagnostic accuracy and demonstrate that they cannot be easily integrated into existing workflows because they often fail to follow instructions and are sensitive to both the quantity and order of information. Overall, our analysis reveals that LLMs are currently not ready for autonomous clinical decision-making while providing a dataset and framework to guide future studies.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ์‹ค์ œ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ํ™˜๊ฒฝ์—์„œ ํ™œ์šฉ ๊ฐ€๋Šฅํ•œ์ง€๋ฅผ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด, ์ค‘ํ™˜์ž ์˜๋ฃŒ ์ •๋ณด ๋ฐ์ดํ„ฐ๋ฒ ์ด์Šค(MIMIC)๋ฅผ ๊ธฐ๋ฐ˜์œผ๋กœ 4๊ฐ€์ง€ ๋ณต๋ถ€ ๋ณ‘๋ฆฌ๋ฅผ ํฌํ•จํ•œ 2,400๊ฑด์˜ ์‹ค์ œ ํ™˜์ž ์‚ฌ๋ก€ ๋ฐ์ดํ„ฐ์…‹๊ณผ ํ˜„์‹ค์ ์ธ ์ž„์ƒ ํ™˜๊ฒฝ ์‹œ๋ฎฌ๋ ˆ์ด์…˜ ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ๊ตฌ์ถ•ํ•˜์˜€๋‹ค. ํ‰๊ฐ€ ๊ฒฐ๊ณผ, ์ตœ์‹  LLM๋“ค์€ ์ „๋ฐ˜์ ์ธ ์ง„๋‹จ ์ •ํ™•๋„์—์„œ ์˜์‚ฌ๋ณด๋‹ค ์œ ์˜ํ•˜๊ฒŒ ๋‚ฎ์€ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ์ง„๋‹จ ๋ฐ ์น˜๋ฃŒ ๊ฐ€์ด๋“œ๋ผ์ธ์„ ์ค€์ˆ˜ํ•˜์ง€ ๋ชปํ•˜๊ณ  ๊ฒ€์‚ฌ์‹ค ๊ฒฐ๊ณผ ํ•ด์„์—๋„ ํ•œ๊ณ„๋ฅผ ๋“œ๋Ÿฌ๋ƒˆ๋‹ค. ๋˜ํ•œ ์ง€์‹œ ์ดํ–‰ ์‹คํŒจ ๋ฐ ์ •๋ณด์˜ ์–‘ยท์ˆœ์„œ์— ๋Œ€ํ•œ ๋ฏผ๊ฐ์„ฑ์œผ๋กœ ์ธํ•ด ๊ธฐ์กด ์ž„์ƒ ์›Œํฌํ”Œ๋กœ์šฐ์—์˜ ํ†ตํ•ฉ์ด ์–ด๋ ต๊ณ , ํ˜„ ์‹œ์ ์—์„œ LLM์˜ ์ž์œจ์  ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ํ™œ์šฉ์€ ํ™˜์ž ์•ˆ์ „์— ์‹ฌ๊ฐํ•œ ์œ„ํ—˜์„ ์ดˆ๋ž˜ํ•  ์ˆ˜ ์žˆ์Œ์„ ๊ฒฐ๋ก ์ง€์—ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

91GPT-Driven Radiology Report Generation with Fine-Tuned Llama 3

2024Bioengineering๐Ÿ”ท Q2DOI 10.3390/bioengineering11101043

The integration of deep learning into radiology has the potential to enhance diagnostic processes, yet its acceptance in clinical practice remains limited due to various challenges. This study aimed to develop and evaluate a fine-tuned large language model (LLM), based on Llama 3-8B, to automate the generation of accurate and concise conclusions in magnetic resonance imaging (MRI) and computed tomography (CT) radiology reports, thereby assisting radiologists and improving reporting efficiency. A dataset comprising 15,000 radiology reports was collected from the University of Medicine and Pharmacy of Craiova's Imaging Center, covering a diverse range of MRI and CT examinations made by four experienced radiologists. The Llama 3-8B model was fine-tuned using transfer-learning techniques, incorporating parameter quantization to 4-bit precision and low-rank adaptation (LoRA) with a rank of 16 to optimize computational efficiency on consumer-grade GPUs. The model was trained over five epochs using an NVIDIA RTX 3090 GPU, with intermediary checkpoints saved for monitoring. Performance was evaluated quantitatively using Bidirectional Encoder Representations from Transformers Score (BERTScore), Recall-Oriented Understudy for Gisting Evaluation (ROUGE), Bilingual Evaluation Understudy (BLEU), and Metric for Evaluation of Translation with Explicit Ordering (METEOR) metrics on a held-out test set. Additionally, a qualitative assessment was conducted, involving 13 independent radiologists who participated in a Turing-like test and provided ratings for the AI-generated conclusions. The fine-tuned model demonstrated strong quantitative performance, achieving a BERTScore F1 of 0.8054, a ROUGE-1 F1 of 0.4998, a ROUGE-L F1 of 0.4628, and a METEOR score of 0.4282. In the human evaluation, the artificial intelligence (AI)-generated conclusions were preferred over human-written ones in approximately 21.8% of cases, indicating that the model's outputs were competitive with those of experienced radiologists. The average rating of the AI-generated conclusions was 3.65 out of 5, reflecting a generally favorable assessment. Notably, the model maintained its consistency across various types of reports and demonstrated the ability to generalize to unseen data. The fine-tuned Llama 3-8B model effectively generates accurate and coherent conclusions for MRI and CT radiology reports. By automating the conclusion-writing process, this approach can assist radiologists in reducing their workload and enhancing report consistency, potentially addressing some barriers to the adoption of deep learning in clinical practice. The positive evaluations from independent radiologists underscore the model's potential utility. While the model demonstrated strong performance, limitations such as dataset bias, limited sample diversity, a lack of clinical judgment, and the need for large computational resources require further refinement and real-world validation. Future work should explore the integration of such models into clinical workflows, address ethical and legal considerations, and extend this approach to generate complete radiology reports.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” Llama 3-8B ๋ชจ๋ธ์„ ๋ฏธ์„ธ ์กฐ์ •(fine-tuning)ํ•˜์—ฌ MRI ๋ฐ CT ์˜์ƒ์˜ ํŒ๋…๋ฌธ ๊ฒฐ๋ก ์„ ์ž๋™์œผ๋กœ ์ƒ์„ฑํ•˜๋Š” ์‹œ์Šคํ…œ์„ ๊ฐœ๋ฐœํ•˜๊ณ  ๊ทธ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. 15,000๊ฑด์˜ ๋ฐ์ดํ„ฐ๋ฅผ ํ•™์Šตํ•œ ๋ชจ๋ธ์€ ์ •๋Ÿ‰์  ์ง€ํ‘œ์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ๋…๋ฆฝ์ ์ธ ์˜์ƒ์˜ํ•™ ์ „๋ฌธ์˜ ํ‰๊ฐ€์—์„œ๋„ ์ธ๊ฐ„์ด ์ž‘์„ฑํ•œ ํŒ๋…๋ฌธ๊ณผ ๋Œ€๋“ฑํ•œ ์ˆ˜์ค€์˜ ์ •ํ™•๋„์™€ ์ผ๊ด€์„ฑ์„ ์ž…์ฆํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ด ๋ชจ๋ธ์€ ํŒ๋… ์—…๋ฌด์˜ ํšจ์œจ์„ฑ์„ ๋†’์ด๊ณ  ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ์˜ ํ‘œ์ค€ํ™”๋ฅผ ์ง€์›ํ•˜๋Š” ์ž„์ƒ ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์ œ์‹œํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

92Large Language Model (LLM) Monthly Report (2024 Apr)

2024UnknownDOI 10.55277/researchhub.0ps6xenm
Added: 2026-04-21 16:11View โ†—

93Large language models (LLMs) in radiology exams for medical students: Performance and consequences

2024RรถFo - Fortschritte auf dem Gebiet der Rรถntgenstrahlen und der bildgebenden VerfahrenDOI 10.1055/a-2437-2067

The evolving field of medical education is being shaped by technological advancements, including the integration of Large Language Models (LLMs) like ChatGPT. These models could be invaluable resources for medical students, by simplifying complex concepts and enhancing interactive learning by providing personalized support. LLMs have shown impressive performance in professional examinations, even without specific domain training, making them particularly relevant in the medical field. This study aims to assess the performance of LLMs in radiology examinations for medical students, thereby shedding light on their current capabilities and implications.This study was conducted using 151 multiple-choice questions, which were used for radiology exams for medical students. The questions were categorized by type and topic and were then processed using OpenAI's GPT-3.5 and GPT- 4 via their API, or manually put into Perplexity AI with GPT-3.5 and Bing. LLM performance was evaluated overall, by question type and by topic.GPT-3.5 achieved a 67.6% overall accuracy on all 151 questions, while GPT-4 outperformed it significantly with an 88.1% overall accuracy (p<0.001). GPT-4 demonstrated superior performance in both lower-order and higher-order questions compared to GPT-3.5, Perplexity AI, and medical students, with GPT-4 particularly excelling in higher-order questions. All GPT models would have successfully passed the radiology exam for medical students at our university.In conclusion, our study highlights the potential of LLMs as accessible knowledge resources for medical students. GPT-4 performed well on lower-order as well as higher-order questions, making ChatGPT-4 a potentially very useful tool for reviewing radiology exam questions. Radiologists should be aware of ChatGPT's limitations, including its tendency to confidently provide incorrect responses. ยท ChatGPT demonstrated remarkable performance, achieving a passing grade on a radiology examination for medical students that did not include image questions.. ยท GPT-4 exhibits significantly improved performance compared to its predecessors GPT-3.5 and Perplexity AI with 88% of questions answered correctly.. ยท Radiologists as well as medical students should be aware of ChatGPT's limitations, including its tendency to confidently provide incorrect responses.. ยท Gotta J, Le Hong QA, Koch V et al. Large language models (LLMs) in radiology exams for medical students: Performance and consequences. Rofo 2025; 197: 1057-1067.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜๋Œ€์ƒ ๋Œ€์ƒ ์˜์ƒ์˜ํ•™๊ณผ ์‹œํ—˜ ๋ฌธ์ œ 151๊ฐœ๋ฅผ ํ™œ์šฉํ•˜์—ฌ GPT-3.5์™€ GPT-4 ๋“ฑ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, GPT-4๋Š” 88.1%์˜ ์ •ํ™•๋„๋กœ GPT-3.5 ๋ฐ Perplexity AI๋ณด๋‹ค ์šฐ์ˆ˜ํ•œ ์„ฑ์ ์„ ๊ฑฐ๋‘๋ฉฐ ๋ชจ๋“  ๋ชจ๋ธ์ด ํ•ฉ๊ฒฉ ๊ธฐ์ค€์„ ์ƒํšŒํ•˜์˜€์Šต๋‹ˆ๋‹ค. LLM์€ ์˜์ƒ์˜ํ•™๊ณผ ํ•™์Šต ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ ๋†’์€ ์ž ์žฌ๋ ฅ์„ ๋ณด์˜€์œผ๋‚˜, ์ž˜๋ชป๋œ ์ •๋ณด๋ฅผ ํ™•์‹ ์— ์ฐจ์„œ ๋‹ต๋ณ€ํ•  ์ˆ˜ ์žˆ๋Š” ํ•œ๊ณ„์ ์— ๋Œ€ํ•œ ์ฃผ์˜๊ฐ€ ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

94Lung Cancer Staging Using Chest CT and FDG PET/CT Free-Text Reports: Comparison Among Three ChatGPT Large Language Models and Six Human Readers of Varying Experience

2024American Journal of Roentgenologyโญ Q1DOI 10.2214/ajr.24.31696

<b>BACKGROUND.</b> Although radiology reports are commonly used for lung cancer staging, this task can be challenging given radiologists' variable reporting styles as well as reports' potentially ambiguous and/or incomplete staging-related information. <b>OBJECTIVE.</b> The purpose of this study was to compare the performance of ChatGPT large language models (LLMs) and human readers of varying experience in lung cancer staging using chest CT and FDG PET/CT free-text reports. <b>METHODS.</b> This retrospective study included 700 patients (mean age, 73.8 ยฑ 29.5 [SD] years; 509 men, 191 women) from four institutions in Korea who underwent chest CT or FDG PET/CT for non-small cell lung cancer initial staging from January 2020 to December 2023. Examinations' reports used a free-text format, written exclusively in English or in mixed English and Korean. Two thoracic radiologists in consensus determined the overall stage group (IA, IB, IIA, IIB, IIIA, IIIB, IIIC, IVA, or IVB) for each report using the 8th-edition <i>AJCC Cancer Staging Manual</i> to establish the reference standard. Three ChatGPT models (GPT-4o, GPT-4, GPT-3.5) determined an overall stage group for each report using a script-based application programming interface, zero-shot learning, and a prompt incorporating a staging system summary. The code for this web application was made publicly available through a GitHub repository (https://github.com/elmidion/GPT_Information_Extractor). Six human readers (two fellowship-trained radiologists with less experience than the radiologists who determined the reference standard, two fellows, and two residents) also independently determined overall stage groups. GPT-4o's overall accuracy for determining the correct stage among the nine groups was compared with that of the other LLMs and human readers using McNemar tests. <b>RESULTS.</b> GPT-4o had an overall staging accuracy of 74.1%, significantly better than the accuracy of GPT-4 (70.1%, <i>p</i> = .02), GPT-3.5 (57.4%, <i>p</i> < .001), and resident 2 (65.7%, <i>p</i> < .001); significantly worse than the accuracy of fellowship-trained radiologist 1 (82.3%, <i>p</i> < .001) and fellowship-trained radiologist 2 (85.4%, <i>p</i> < .001); and not significantly different from the accuracy of fellow 1 (77.7%, <i>p</i> = .09), fellow 2 (75.6%, <i>p</i> = .53), and resident 1 (72.3%, <i>p</i> = .42). <b>CONCLUSION.</b> The best-performing model, GPT-4o, showed no significant difference in staging accuracy versus fellows but showed significantly worse performance versus fellowship-trained radiologists. The findings do not support use of LLMs for lung cancer staging in place of expert health care professionals. <b>CLINICAL IMPACT.</b> The findings indicate the importance of domain expertise for performing complex specialized tasks such as cancer staging.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ํ›„ํ–ฅ์  ์—ฐ๊ตฌ๋Š” ๋น„์†Œ์„ธํฌํ์•” ์ดˆ๊ธฐ ๋ณ‘๊ธฐ ๊ฒฐ์ • ์‹œ ChatGPT ๊ธฐ๋ฐ˜ ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ(LLM) ์„ธ ์ข…๋ฅ˜(GPT-4o, GPT-4, GPT-3.5)์™€ ๊ฒฝํ—˜ ์ˆ˜์ค€์ด ๋‹ค๋ฅธ ์ธ๊ฐ„ ํŒ๋…์ž 6๋ช…์˜ ์„ฑ๋Šฅ์„ ํ‰๋ถ€ CT ๋ฐ FDG PET/CT ์ž์œ ํ˜•์‹ ๋ณด๊ณ ์„œ๋ฅผ ์ด์šฉํ•˜์—ฌ ๋น„๊ตํ•˜์˜€๋‹ค. ํ•œ๊ตญ 4๊ฐœ ๊ธฐ๊ด€์—์„œ ์ˆ˜์ง‘๋œ 700๋ช…์˜ ํ™˜์ž ๋ณด๊ณ ์„œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ AJCC 8ํŒ ๊ธฐ์ค€์— ๋”ฐ๋ผ 9๊ฐœ ๋ณ‘๊ธฐ ๊ทธ๋ฃน์œผ๋กœ ๋ถ„๋ฅ˜ํ•˜๋Š” ์ •ํ™•๋„๋ฅผ McNemar ๊ฒ€์ •์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. ์ตœ๊ณ  ์„ฑ๋Šฅ ๋ชจ๋ธ์ธ GPT-4o์˜ ๋ณ‘๊ธฐ ๊ฒฐ์ • ์ •ํ™•๋„๋Š” 74.1%๋กœ ์ „์ž„์˜(fellow) ์ˆ˜์ค€๊ณผ ์œ ์˜ํ•œ ์ฐจ์ด๊ฐ€ ์—†์—ˆ์œผ๋‚˜, ํ‰๋ถ€์˜์ƒ ์ „๋ฌธ ์ˆ˜๋ จ ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜(82.3โ€“85.4%)์— ๋น„ํ•ด ์œ ์˜ํ•˜๊ฒŒ ๋‚ฎ์•„, ํ์•” ๋ณ‘๊ธฐ ๊ฒฐ์ •๊ณผ ๊ฐ™์€ ๋ณต์žกํ•œ ์ „๋ฌธ ์ž„์ƒ ๊ณผ์ œ์—์„œ LLM์ด ์ „๋ฌธ๊ฐ€๋ฅผ ๋Œ€์ฒดํ•˜๊ธฐ์—๋Š” ์•„์ง ํ•œ๊ณ„๊ฐ€ ์žˆ์Œ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

95PMC-LLaMA: toward building open-source language models for medicine

2024Journal of the American Medical Informatics Associationโญ Q1DOI 10.1093/jamia/ocae045

In this article, we systematically investigate the process of building up an open-source medical-specific LLM, PMC-LLaMA.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜ํ•™ ๋ถ„์•ผ์— ํŠนํ™”๋œ ์˜คํ”ˆ์†Œ์Šค ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ธ PMC-LLaMA์˜ ๊ตฌ์ถ• ๊ณผ์ •์„ ์ฒด๊ณ„์ ์œผ๋กœ ํƒ๊ตฌํ•˜๋Š” ๊ฒƒ์„ ๋ชฉํ‘œ๋กœ ํ•œ๋‹ค. PubMed Central(PMC)์˜ ๋ฐฉ๋Œ€ํ•œ ์˜ํ•™ ๋ฌธํ—Œ ๋ฐ ์˜ํ•™ ๊ต๊ณผ์„œ๋ฅผ ํ™œ์šฉํ•˜์—ฌ LLaMA ๊ธฐ๋ฐ˜ ๋ชจ๋ธ์„ ์ง€์†์  ์‚ฌ์ „ํ•™์Šต ๋ฐ ์ง€์‹œ ์กฐ์ •(instruction tuning) ๋ฐฉ์‹์œผ๋กœ ํ›ˆ๋ จํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, PMC-LLaMA๋Š” ์˜ํ•™ ์งˆ์˜์‘๋‹ต ๋ฒค์น˜๋งˆํฌ์—์„œ ํ–ฅ์ƒ๋œ ์„ฑ๋Šฅ์„ ๋ณด์—ฌ, ํ์‡„ํ˜• ์ƒ์šฉ ๋ชจ๋ธ์— ๋Œ€ํ•œ ์˜คํ”ˆ์†Œ์Šค ์˜ํ•™ LLM์˜ ์‹คํ˜„ ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

96Personalized Impression Generation for PET Reports Using Large Language Models

2024Journal of Imaging Informatics in Medicineโญ Q1DOI 10.1007/s10278-024-00985-3

Large language models (LLMs) have shown promise in accelerating radiology reporting by summarizing clinical findings into impressions. However, automatic impression generation for whole-body PET reports presents unique challenges and has received little attention. Our study aimed to evaluate whether LLMs can create clinically useful impressions for PET reporting. To this end, we fine-tuned twelve open-source language models on a corpus of 37,370 retrospective PET reports collected from our institution. All models were trained using the teacher-forcing algorithm, with the report findings and patient information as input and the original clinical impressions as reference. An extra input token encoded the reading physician's identity, allowing models to learn physician-specific reporting styles. To compare the performances of different models, we computed various automatic evaluation metrics and benchmarked them against physician preferences, ultimately selecting PEGASUS as the top LLM. To evaluate its clinical utility, three nuclear medicine physicians assessed the PEGASUS-generated impressions and original clinical impressions across 6 quality dimensions (3-point scales) and an overall utility score (5-point scale). Each physician reviewed 12 of their own reports and 12 reports from other physicians. When physicians assessed LLM impressions generated in their own style, 89% were considered clinically acceptable, with a mean utility score of 4.08/5. On average, physicians rated these personalized impressions as comparable in overall utility to the impressions dictated by other physicians (4.03, Pโ€‰=โ€‰0.41). In summary, our study demonstrated that personalized impressions generated by PEGASUS were clinically useful in most cases, highlighting its potential to expedite PET reporting by automatically drafting impressions.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 37,370๊ฑด์˜ PET ํŒ๋… ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•˜์—ฌ ํŒ๋…์˜๋ณ„ ๊ณ ์œ  ์Šคํƒ€์ผ์„ ํ•™์Šตํ•œ ๋งž์ถคํ˜• ์ธ์ƒ(impression) ์ƒ์„ฑ ๋ชจ๋ธ์ธ PEGASUS๋ฅผ ๊ฐœ๋ฐœํ•˜๊ณ  ๊ทธ ์ž„์ƒ์  ์œ ์šฉ์„ฑ์„ ํ‰๊ฐ€ํ–ˆ์Šต๋‹ˆ๋‹ค. ํ‰๊ฐ€ ๊ฒฐ๊ณผ, ๋ชจ๋ธ์ด ์ƒ์„ฑํ•œ ๋งž์ถคํ˜• ์ธ์ƒ์€ 89%์˜ ์ž„์ƒ์  ์ˆ˜์šฉ๋„๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ํ•ต์˜ํ•™๊ณผ ์ „๋ฌธ์˜๊ฐ€ ์ง์ ‘ ์ž‘์„ฑํ•œ ํŒ๋…๋ฌธ๊ณผ ๋Œ€๋“ฑํ•œ ์ˆ˜์ค€์˜ ์œ ์šฉ์„ฑ์„ ๋‚˜ํƒ€๋ƒˆ์Šต๋‹ˆ๋‹ค. ์ด๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์ด PET ํŒ๋… ๊ณผ์ •์—์„œ ํšจ์œจ์ ์ด๊ณ  ๊ฐœ์ธํ™”๋œ ์ดˆ์•ˆ์„ ์ œ๊ณตํ•จ์œผ๋กœ์จ ์—…๋ฌด ๋ถ€๋‹ด์„ ๊ฒฝ๊ฐํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

97Preface for Special Issue on Explainable/Reliable Artificial Intelligence, and Generative Artificial Intelligence with Large Language Model for Radiologist

2024Journal of the Korean Society of RadiologyDOI 10.3348/jksr.2024.0123
Added: 2026-04-21 16:11View โ†—

98Radiologist and Radiology Practice Wellbeing: A Report of the 2023 ARRS Wellness Summit

2024Academic Radiologyโญ Q1DOI 10.1016/j.acra.2023.08.025
Added: 2026-04-21 16:11View โ†—

99Survey on the radiology report at Chris Hani Baragwanath Academic Hospital: Clinician and radiologist perspectives

2024South African Journal of RadiologyDOI 10.4102/sajr.v28i1.2954
Added: 2026-04-21 16:11View โ†—

100Toward Clinical-Grade Evaluation of Large Language Models

2024International Journal of Radiation Oncology*Biology*PhysicsDOI 10.1016/j.ijrobp.2023.11.012
Added: 2026-04-21 16:11View โ†—

101Style-Aware Radiology Report Generation with RadGraph and Few-Shot Prompting

2023-12-01Findings of EMNLP 2023DOI 10.18653/v1/2023.findings-emnlp.977

Study on reproducing radiologist-specific report style using RadGraph-based structured content extraction and few-shot in-context prompting without model fine-tuning.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ชจ๋ธ ํŒŒ์ธํŠœ๋‹ ์—†์ด ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ ๊ฐœ์ธ๋ณ„ ๋ณด๊ณ ์„œ ์ž‘์„ฑ ์Šคํƒ€์ผ์„ ์žฌํ˜„ํ•˜๋Š” ๊ฒƒ์„ ๋ชฉํ‘œ๋กœ ํ•˜์˜€๋‹ค. RadGraph ๊ธฐ๋ฐ˜์˜ ๊ตฌ์กฐํ™”๋œ ์ฝ˜ํ…์ธ  ์ถ”์ถœ๊ณผ ํ“จ์ƒท ์ธ์ปจํ…์ŠคํŠธ ํ”„๋กฌํ”„ํŒ… ๊ธฐ๋ฒ•์„ ๊ฒฐํ•ฉํ•˜์—ฌ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์˜ ์Šคํƒ€์ผ ์ธ์‹ ์ƒ์„ฑ ๋ฐฉ๋ฒ•๋ก ์„ ์ œ์•ˆํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, ๋ณ„๋„์˜ ๋ชจ๋ธ ํ•™์Šต ์—†์ด๋„ ํŠน์ • ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์‚ฌ์˜ ์„œ์ˆ  ๋ฐฉ์‹๊ณผ ํ‘œํ˜„ ์Šคํƒ€์ผ์„ ํšจ๊ณผ์ ์œผ๋กœ ๋ชจ๋ฐฉํ•  ์ˆ˜ ์žˆ์Œ์„ ํ™•์ธํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

102A Context-based Chatbot Surpasses Radiologists and Generic ChatGPT in Following the ACR Appropriateness Guidelines

2023Radiologyโญ Q1DOI 10.1148/radiol.230970

Background Radiological imaging guidelines are crucial for accurate diagnosis and optimal patient care as they result in standardized decisions and thus reduce inappropriate imaging studies. Purpose In the present study, we investigated the potential to support clinical decision-making using an interactive chatbot designed to provide personalized imaging recommendations from American College of Radiology (ACR) appropriateness criteria documents using semantic similarity processing. Methods We utilized 209 ACR appropriateness criteria documents as specialized knowledge base and employed LlamaIndex, a framework that allows to connect large language models with external data, and the ChatGPT 3.5-Turbo to create an appropriateness criteria contexted chatbot (accGPT). Fifty clinical case files were used to compare the accGPT's performance against general radiologists at varying experience levels and to generic ChatGPT 3.5 and 4.0. Results All chatbots reached at least human performance level. For the 50 case files, the accGPT performed best in providing correct recommendations that were "usually appropriate" according to the ACR criteria and also did provide the highest proportion of consistently correct answers in comparison with generic chatbots and radiologists. Further, the chatbots provided substantial time and cost savings, with an average decision time of 5 minutes and a cost of 0.19 โ‚ฌ for all cases, compared to 50 minutes and 29.99 โ‚ฌ for radiologists (both p < 0.01). Conclusion ChatGPT-based algorithms have the potential to substantially improve the decision-making for clinical imaging studies in accordance with ACR guidelines. Specifically, a context-based algorithm performed superior to its generic counterpart, demonstrating the value of tailoring AI solutions to specific healthcare applications.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฏธ๊ตญ์˜์ƒ์˜ํ•™ํšŒ(ACR) ์ ์ •์„ฑ ๊ธฐ์ค€์„ ํ•™์Šต์‹œํ‚จ ๋ฌธ๋งฅ ๊ธฐ๋ฐ˜ ์ฑ—๋ด‡(accGPT)์˜ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. 50๊ฐœ์˜ ์ž„์ƒ ์‚ฌ๋ก€๋ฅผ ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, accGPT๋Š” ์ผ๋ฐ˜ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜ ๋ฐ ๋ฒ”์šฉ ChatGPT ๋ชจ๋ธ๋ณด๋‹ค ACR ๊ฐ€์ด๋“œ๋ผ์ธ์— ๋ถ€ํ•ฉํ•˜๋Š” ์ •ํ™•ํ•œ ๊ถŒ๊ณ ์•ˆ์„ ๋” ์ผ๊ด€๋˜๊ฒŒ ์ œ์‹œํ•˜์˜€์œผ๋ฉฐ, ์˜์‚ฌ๊ฒฐ์ •์— ์†Œ์š”๋˜๋Š” ์‹œ๊ฐ„๊ณผ ๋น„์šฉ์„ ์œ ์˜๋ฏธํ•˜๊ฒŒ ์ ˆ๊ฐํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ด๋Š” ํŠนํ™”๋œ ์˜๋ฃŒ ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•œ ๋ฌธ๋งฅ ๊ธฐ๋ฐ˜ AI ๋ชจ๋ธ์ด ์˜์ƒ ๊ฒ€์‚ฌ ์ ์ •์„ฑ ํŒ๋‹จ์˜ ํšจ์œจ์„ฑ๊ณผ ์ •ํ™•๋„๋ฅผ ๋†’์ด๋Š” ๋ฐ ํšจ๊ณผ์ ์ธ ๋„๊ตฌ๊ฐ€ ๋  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

103A SWOT analysis of ChatGPT: Implications for educational practice and research

2023Innovations in Education and Teaching Internationalโญ Q1DOI 10.1080/14703297.2023.2195846

ChatGPT is an AI tool that has sparked debates about its potential implications for education. We used the SWOT analysis framework to outline ChatGPTโ€™s strengths and weaknesses and to discuss its opportunities for and threats to education. The strengths include using a sophisticated natural language model to generate plausible answers, self-improving capability, and providing personalised and real-time responses. As such, ChatGPT can increase access to information, facilitate personalised and complex learning, and decrease teaching workload, thereby making key processes and tasks more efficient. The weaknesses are a lack of deep understanding, difficulty in evaluating the quality of responses, a risk of bias and discrimination, and a lack of higher-order thinking skills. Threats to education include a lack of understanding of the context, threatening academic integrity, perpetuating discrimination in education, democratising plagiarism, and declining high-order cognitive skills. We provide agenda for educational practice and research in times of ChatGPT.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” SWOT ๋ถ„์„ ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ํ™œ์šฉํ•˜์—ฌ AI ๋„๊ตฌ์ธ ChatGPT๊ฐ€ ๊ต์œก ๋ถ„์•ผ์— ๋ฏธ์น˜๋Š” ์ž ์žฌ์  ์˜ํ–ฅ์„ ์ฒด๊ณ„์ ์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. ChatGPT์˜ ๊ฐ•์ ์œผ๋กœ๋Š” ์ •๊ตํ•œ ์ž์—ฐ์–ด ์ƒ์„ฑ, ๊ฐœ์ธํ™”๋œ ์‹ค์‹œ๊ฐ„ ์‘๋‹ต, ๊ต์ˆ˜ ๋ถ€๋‹ด ๊ฒฝ๊ฐ ๋“ฑ์ด ํ™•์ธ๋˜์—ˆ์œผ๋ฉฐ, ์•ฝ์ ์œผ๋กœ๋Š” ์‹ฌ์ธต์  ์ดํ•ด ๋ถ€์กฑ, ์‘๋‹ต ํ’ˆ์งˆ ํ‰๊ฐ€์˜ ์–ด๋ ค์›€, ํŽธํ–ฅ ๋ฐ ์ฐจ๋ณ„์˜ ์œ„ํ—˜์„ฑ์ด ์ œ์‹œ๋˜์—ˆ๋‹ค. ๋˜ํ•œ ํ•™๋ฌธ์  ์ง„์‹ค์„ฑ ํ›ผ์†, ํ‘œ์ ˆ์˜ ๋Œ€์ค‘ํ™”, ๊ณ ์ฐจ์›์  ์ธ์ง€ ๋Šฅ๋ ฅ์˜ ์ €ํ•˜ ๋“ฑ ๊ต์œก์— ๋Œ€ํ•œ ์œ„ํ˜‘ ์š”์ธ๋„ ๊ทœ๋ช…๋˜์–ด, ์ €์ž๋“ค์€ ChatGPT ์‹œ๋Œ€์˜ ๊ต์œก ์‹ค์ฒœ ๋ฐ ์—ฐ๊ตฌ๋ฅผ ์œ„ํ•œ ํ–ฅํ›„ ๊ณผ์ œ๋ฅผ ์ œ์–ธํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

104Adapted Large Language Models Can Outperform Medical Experts in Clinical Text Summarization

2023arXiv (Cornell University)DOI 10.48550/arxiv.2309.07430

Analyzing vast textual data and summarizing key information from electronic health records imposes a substantial burden on how clinicians allocate their time. Although large language models (LLMs) have shown promise in natural language processing (NLP), their effectiveness on a diverse range of clinical summarization tasks remains unproven. In this study, we apply adaptation methods to eight LLMs, spanning four distinct clinical summarization tasks: radiology reports, patient questions, progress notes, and doctor-patient dialogue. Quantitative assessments with syntactic, semantic, and conceptual NLP metrics reveal trade-offs between models and adaptation methods. A clinical reader study with ten physicians evaluates summary completeness, correctness, and conciseness; in a majority of cases, summaries from our best adapted LLMs are either equivalent (45%) or superior (36%) compared to summaries from medical experts. The ensuing safety analysis highlights challenges faced by both LLMs and medical experts, as we connect errors to potential medical harm and categorize types of fabricated information. Our research provides evidence of LLMs outperforming medical experts in clinical text summarization across multiple tasks. This suggests that integrating LLMs into clinical workflows could alleviate documentation burden, allowing clinicians to focus more on patient care.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ, ํ™˜์ž ์งˆ๋ฌธ, ๊ฒฝ๊ณผ ๊ธฐ๋ก, ์˜์‚ฌ-ํ™˜์ž ๋Œ€ํ™” ๋“ฑ ๋„ค ๊ฐ€์ง€ ์ž„์ƒ ์š”์•ฝ ๊ณผ์ œ์—์„œ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์ ์‘ ๊ธฐ๋ฒ•์„ ์ ์šฉํ•˜์—ฌ ๊ทธ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. 8๊ฐœ์˜ LLM์— ์ ์‘ ๋ฐฉ๋ฒ•์„ ์ ์šฉํ•˜๊ณ  ๊ตฌ๋ฌธ์ ยท์˜๋ฏธ์ ยท๊ฐœ๋…์  NLP ์ง€ํ‘œ๋กœ ์ •๋Ÿ‰ ํ‰๊ฐ€ํ•œ ํ›„, 10๋ช…์˜ ์˜์‚ฌ๊ฐ€ ์ฐธ์—ฌํ•œ ์ž„์ƒ ํŒ๋…์ž ์—ฐ๊ตฌ๋ฅผ ํ†ตํ•ด ์š”์•ฝ์˜ ์™„์ „์„ฑ, ์ •ํ™•์„ฑ, ๊ฐ„๊ฒฐ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, ์ตœ์ ํ™”๋œ LLM์ด ์ƒ์„ฑํ•œ ์š”์•ฝ์€ ๋Œ€๋‹ค์ˆ˜์˜ ๊ฒฝ์šฐ ์˜๋ฃŒ ์ „๋ฌธ๊ฐ€์˜ ์š”์•ฝ๊ณผ ๋™๋“ฑ(45%)ํ•˜๊ฑฐ๋‚˜ ์šฐ์ˆ˜(36%)ํ•œ ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์œผ๋ฉฐ, ์ด๋Š” LLM์„ ์ž„์ƒ ์›Œํฌํ”Œ๋กœ์šฐ์— ํ†ตํ•ฉํ•จ์œผ๋กœ์จ ์˜๋ฃŒ์ง„์˜ ๋ฌธ์„œ ์ž‘์—… ๋ถ€๋‹ด์„ ๊ฒฝ๊ฐํ•˜๊ณ  ํ™˜์ž ์น˜๋ฃŒ์— ๋” ์ง‘์ค‘ํ•  ์ˆ˜ ์žˆ๋Š” ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

105ChatGPT for Education and Research: Opportunities, Threats, and Strategies

2023Applied SciencesDOI 10.3390/app13095783

In recent years, the rise of advanced artificial intelligence technologies has had a profound impact on many fields, including education and research. One such technology is ChatGPT, a powerful large language model developed by OpenAI. This technology offers exciting opportunities for students and educators, including personalized feedback, increased accessibility, interactive conversations, lesson preparation, evaluation, and new ways to teach complex concepts. However, ChatGPT poses different threats to the traditional education and research system, including the possibility of cheating on online exams, human-like text generation, diminished critical thinking skills, and difficulties in evaluating information generated by ChatGPT. This study explores the potential opportunities and threats that ChatGPT poses to overall education from the perspective of students and educators. Furthermore, for programming learning, we explore how ChatGPT helps students improve their programming skills. To demonstrate this, we conducted different coding-related experiments with ChatGPT, including code generation from problem descriptions, pseudocode generation of algorithms from texts, and code correction. The generated codes are validated with an online judge system to evaluate their accuracy. In addition, we conducted several surveys with students and teachers to find out how ChatGPT supports programming learning and teaching. Finally, we present the survey results and analysis.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์ธ ChatGPT๊ฐ€ ๊ต์œก ๋ฐ ์—ฐ๊ตฌ ๋ถ„์•ผ์— ๋ฏธ์น˜๋Š” ๊ธฐํšŒ์™€ ์œ„ํ˜‘์„ ํ•™์ƒ ๋ฐ ๊ต์œก์ž์˜ ๊ด€์ ์—์„œ ํƒ์ƒ‰ํ•˜์˜€๋‹ค. ์—ฐ๊ตฌํŒ€์€ ์ฝ”๋“œ ์ƒ์„ฑ, ์•Œ๊ณ ๋ฆฌ์ฆ˜ ์˜์‚ฌ์ฝ”๋“œ ๋ณ€ํ™˜, ์ฝ”๋“œ ์˜ค๋ฅ˜ ์ˆ˜์ • ๋“ฑ ๋‹ค์–‘ํ•œ ํ”„๋กœ๊ทธ๋ž˜๋ฐ ๊ด€๋ จ ์‹คํ—˜์„ ์ˆ˜ํ–‰ํ•˜๊ณ  ์˜จ๋ผ์ธ ์ฑ„์  ์‹œ์Šคํ…œ์œผ๋กœ ์ •ํ™•๋„๋ฅผ ๊ฒ€์ฆํ•˜์˜€์œผ๋ฉฐ, ํ•™์ƒ๊ณผ ๊ต์‚ฌ๋ฅผ ๋Œ€์ƒ์œผ๋กœ ์„ค๋ฌธ์กฐ์‚ฌ๋ฅผ ์‹ค์‹œํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ ChatGPT๋Š” ๊ฐœ์ธ ๋งž์ถคํ˜• ํ”ผ๋“œ๋ฐฑ ์ œ๊ณต, ํ”„๋กœ๊ทธ๋ž˜๋ฐ ํ•™์Šต ์ง€์› ๋“ฑ ๊ต์œก์  ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์ด ๋†’์€ ๋ฐ˜๋ฉด, ์˜จ๋ผ์ธ ์‹œํ—˜ ๋ถ€์ •ํ–‰์œ„ ์กฐ์žฅ ๋ฐ ๋น„ํŒ์  ์‚ฌ๊ณ  ๋Šฅ๋ ฅ ์ €ํ•˜์™€ ๊ฐ™์€ ์œ„ํ—˜์„ฑ๋„ ๋‚ดํฌํ•˜๊ณ  ์žˆ์Œ์ด ํ™•์ธ๋˜์—ˆ๋‹ค.
Added: 2026-04-21 16:11View โ†—

106ChatRadio-Valuer: A Chat Large Language Model for Generalizable Radiology Report Generation Based on Multi-institution and Multi-system Data

2023arXiv (Cornell University)DOI 10.48550/arxiv.2310.05242

Radiology report generation, as a key step in medical image analysis, is critical to the quantitative analysis of clinically informed decision-making levels. However, complex and diverse radiology reports with cross-source heterogeneity pose a huge generalizability challenge to the current methods under massive data volume, mainly because the style and normativity of radiology reports are obviously distinctive among institutions, body regions inspected and radiologists. Recently, the advent of large language models (LLM) offers great potential for recognizing signs of health conditions. To resolve the above problem, we collaborate with the Second Xiangya Hospital in China and propose ChatRadio-Valuer based on the LLM, a tailored model for automatic radiology report generation that learns generalizable representations and provides a basis pattern for model adaptation in sophisticated analysts' cases. Specifically, ChatRadio-Valuer is trained based on the radiology reports from a single institution by means of supervised fine-tuning, and then adapted to disease diagnosis tasks for human multi-system evaluation (i.e., chest, abdomen, muscle-skeleton, head, and maxillofacial $\&amp;$ neck) from six different institutions in clinical-level events. The clinical dataset utilized in this study encompasses a remarkable total of \textbf{332,673} observations. From the comprehensive results on engineering indicators, clinical efficacy and deployment cost metrics, it can be shown that ChatRadio-Valuer consistently outperforms state-of-the-art models, especially ChatGPT (GPT-3.5-Turbo) and GPT-4 et al., in terms of the diseases diagnosis from radiology reports. ChatRadio-Valuer provides an effective avenue to boost model generalization performance and alleviate the annotation workload of experts to enable the promotion of clinical AI applications in radiology reports.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ธฐ๊ด€ ๊ฐ„ ์Šคํƒ€์ผ ๋ฐ ํ‘œ์ค€ ์ฐจ์ด๋กœ ์ธํ•œ ์ด์งˆ์„ฑ ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๊ณ ์ž, ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM) ๊ธฐ๋ฐ˜์˜ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ์ž๋™ ์ƒ์„ฑ ์‹œ์Šคํ…œ์ธ ChatRadio-Valuer๋ฅผ ์ œ์•ˆํ•˜์˜€๋‹ค. ์ค‘๊ตญ ์ œ2์ƒน์•ผ๋ณ‘์›๊ณผ์˜ ํ˜‘๋ ฅ ํ•˜์— ๋‹จ์ผ ๊ธฐ๊ด€ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ๋กœ ์ง€๋„ ๋ฏธ์„ธ ์กฐ์ •(supervised fine-tuning)์„ ์ˆ˜ํ–‰ํ•œ ํ›„, ํ‰๋ถ€ยท๋ณต๋ถ€ยท๊ทผ๊ณจ๊ฒฉ๊ณ„ยท๋‘๋ถ€ยท์•…์•ˆ๋ฉด ๋ฐ ๊ฒฝ๋ถ€๋ฅผ ํฌํ•จํ•œ 5๊ฐœ ์‹ ์ฒด ๊ณ„ํ†ต์— ๊ฑธ์ณ 6๊ฐœ ๊ธฐ๊ด€์˜ ์ด 332,673๊ฑด ์ž„์ƒ ๋ฐ์ดํ„ฐ์— ์ ์‘์‹œ์ผฐ๋‹ค. ๊ณตํ•™์  ์ง€ํ‘œ, ์ž„์ƒ ํšจ์šฉ์„ฑ, ๋ฐฐํฌ ๋น„์šฉ ์ธก๋ฉด์˜ ์ข…ํ•ฉ ํ‰๊ฐ€์—์„œ ChatRadio-Valuer๋Š” ChatGPT(GPT-3.5-Turbo) ๋ฐ GPT-4๋ฅผ ํฌํ•จํ•œ ์ตœ์‹  ๋ชจ๋ธ๋“ค์„ ์ผ๊ด€๋˜๊ฒŒ ๋Šฅ๊ฐ€ํ•˜์—ฌ, ๋‹ค๊ธฐ๊ด€ ํ™˜๊ฒฝ์—์„œ์˜ ๋ชจ๋ธ ์ผ๋ฐ˜ํ™” ์„ฑ๋Šฅ ํ–ฅ์ƒ ๋ฐ ์ „๋ฌธ๊ฐ€ ์ฃผ์„ ๋ถ€๋‹ด ๊ฒฝ๊ฐ์— ํšจ๊ณผ์ ์ž„์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

107Clinical Text Summarization: Adapting Large Language Models Can Outperform Human Experts

2023Research SquareDOI 10.21203/rs.3.rs-3483777/v1
Added: 2026-04-21 16:11View โ†—

108ESR paper on structured reporting in radiologyโ€”update 2023

2023Insights into Imagingโญ Q1DOI 10.1186/s13244-023-01560-0

Structured reporting in radiology continues to hold substantial potential to improve the quality of service provided to patients and referring physicians. Despite many physicians' preference for structured reports and various efforts by radiological societies and some vendors, structured reporting has still not been widely adopted in clinical routine.While in many countries national radiological societies have launched initiatives to further promote structured reporting, cross-institutional applications of report templates and incentives for usage of structured reporting are lacking. Various legislative measures have been taken in the USA and the European Union to promote interoperable data formats such as Fast Healthcare Interoperability Resources (FHIR) in the context of the EU Health Data Space (EHDS) which will certainly be relevant for the future of structured reporting. Lastly, recent advances in artificial intelligence and large language models may provide innovative and efficient approaches to integrate structured reporting more seamlessly into the radiologists' workflow.The ESR will remain committed to advancing structured reporting as a key component towards more value-based radiology. Practical solutions for structured reporting need to be provided by vendors. Policy makers should incentivize the usage of structured radiological reporting, especially in cross-institutional setting.Critical relevance statement Over the past years, the benefits of structured reporting in radiology have been widely discussed and agreed upon; however, implementation in clinical routine is lacking due-policy makers should incentivize the usage of structured radiological reporting, especially in cross-institutional setting.Key points1. Various national societies have established initiatives for structured reporting in radiology.2. Almost no monetary or structural incentives exist that favor structured reporting.3. A consensus on technical standards for structured reporting is still missing.4. The application of large language models may help structuring radiological reports.5. Policy makers should incentivize the usage of structured radiological reporting.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ๋…ผ๋ฌธ์€ ๋ฐฉ์‚ฌ์„ ๊ณผ ์˜์—ญ์—์„œ ๊ตฌ์กฐํ™” ๋ณด๊ณ ์„œ(structured reporting)์˜ ํ˜„ํ™ฉ๊ณผ ํ–ฅํ›„ ๋ฐœ์ „ ๋ฐฉํ–ฅ์„ ๊ฒ€ํ† ํ•˜๊ธฐ ์œ„ํ•ด ๊ฐ๊ตญ ํ•™ํšŒ์˜ ์ถ”์ง„ ํ˜„ํ™ฉ, ๊ด€๋ จ ๋ฒ•์ œ, ๊ธฐ์ˆ  ํ‘œ์ค€ํ™” ๋…ธ๋ ฅ ๋ฐ ์ธ๊ณต์ง€๋Šฅ ๊ธฐ์ˆ ์˜ ์ ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ์ข…ํ•ฉ์ ์œผ๋กœ ๋ถ„์„ํ•˜์˜€๋‹ค. ๊ตฌ์กฐํ™” ๋ณด๊ณ ์„œ์˜ ์ž„์ƒ์  ์œ ์ต์„ฑ์— ๋Œ€ํ•œ ๊ด‘๋ฒ”์œ„ํ•œ ํ•ฉ์˜์—๋„ ๋ถˆ๊ตฌํ•˜๊ณ , ๊ธฐ์ˆ  ํ‘œ์ค€ ๋ถ€์žฌ, ๊ธฐ๊ด€ ๊ฐ„ ์ƒํ˜ธ์šด์šฉ์„ฑ ๋ฏธํก, ๊ฒฝ์ œ์ ยท๊ตฌ์กฐ์  ์ธ์„ผํ‹ฐ๋ธŒ ๋ถ€์กฑ์œผ๋กœ ์ธํ•ด ์ž„์ƒ ํ˜„์žฅ ๋„์ž…์€ ์—ฌ์ „ํžˆ ์ €์กฐํ•œ ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ๋‹ค. ํ–ฅํ›„ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM) ๊ธฐ๋ฐ˜ ์ธ๊ณต์ง€๋Šฅ์˜ ํ™œ์šฉ์ด ๋ฐฉ์‚ฌ์„ ์‚ฌ์˜ ์›Œํฌํ”Œ๋กœ์šฐ์— ๊ตฌ์กฐํ™” ๋ณด๊ณ ์„œ๋ฅผ ์ž์—ฐ์Šค๋Ÿฝ๊ฒŒ ํ†ตํ•ฉํ•˜๋Š” ํ˜์‹ ์  ์ˆ˜๋‹จ์ด ๋  ์ˆ˜ ์žˆ์œผ๋ฉฐ, ํŠนํžˆ ๊ธฐ๊ด€ ๊ฐ„ ์ง„๋ฃŒ ํ™˜๊ฒฝ์—์„œ์˜ ํ™œ์„ฑํ™”๋ฅผ ์œ„ํ•ด ์ •์ฑ… ์ž…์•ˆ์ž์˜ ์ ๊ทน์ ์ธ ์ œ๋„์  ์ธ์„ผํ‹ฐ๋ธŒ ๋งˆ๋ จ์ด ํ•„์š”ํ•˜๋‹ค.
Added: 2026-04-21 16:11View โ†—

109Exploring the Boundaries of GPT-4 in Radiology

2023arXiv (Cornell University)DOI 10.48550/arxiv.2310.14573

The recent success of general-domain large language models (LLMs) has significantly changed the natural language processing paradigm towards a unified foundation model across domains and applications. In this paper, we focus on assessing the performance of GPT-4, the most capable LLM so far, on the text-based applications for radiology reports, comparing against state-of-the-art (SOTA) radiology-specific models. Exploring various prompting strategies, we evaluated GPT-4 on a diverse range of common radiology tasks and we found GPT-4 either outperforms or is on par with current SOTA radiology models. With zero-shot prompting, GPT-4 already obtains substantial gains ($\approx$ 10% absolute improvement) over radiology models in temporal sentence similarity classification (accuracy) and natural language inference ($F_1$). For tasks that require learning dataset-specific style or schema (e.g. findings summarisation), GPT-4 improves with example-based prompting and matches supervised SOTA. Our extensive error analysis with a board-certified radiologist shows GPT-4 has a sufficient level of radiology knowledge with only occasional errors in complex context that require nuanced domain knowledge. For findings summarisation, GPT-4 outputs are found to be overall comparable with existing manually-written impressions.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฒ”์šฉ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ์ธ GPT-4์˜ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ๊ด€๋ จ ํ…์ŠคํŠธ ๊ธฐ๋ฐ˜ ๊ณผ์ œ์—์„œ์˜ ์„ฑ๋Šฅ์„ ์ตœ์‹  ๋ฐฉ์‚ฌ์„ ๊ณผ ํŠนํ™” ๋ชจ๋ธ๋“ค๊ณผ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ์ œ๋กœ์ƒท ๋ฐ ์˜ˆ์‹œ ๊ธฐ๋ฐ˜ ํ”„๋กฌํ”„ํŒ… ๋“ฑ ๋‹ค์–‘ํ•œ ์ „๋žต์„ ์ ์šฉํ•˜์—ฌ ์‹œ๊ฐ„์  ๋ฌธ์žฅ ์œ ์‚ฌ๋„ ๋ถ„๋ฅ˜, ์ž์—ฐ์–ด ์ถ”๋ก , ํŒ๋… ์†Œ๊ฒฌ ์š”์•ฝ ๋“ฑ ์—ฌ๋Ÿฌ ๊ณผ์ œ๋ฅผ ๋Œ€์ƒ์œผ๋กœ ์‹คํ—˜์„ ์ˆ˜ํ–‰ํ•˜์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, GPT-4๋Š” ์ œ๋กœ์ƒท ํ”„๋กฌํ”„ํŒ…๋งŒ์œผ๋กœ๋„ ์‹œ๊ฐ„์  ๋ถ„๋ฅ˜ ๋ฐ ์ž์—ฐ์–ด ์ถ”๋ก  ๊ณผ์ œ์—์„œ ๊ธฐ์กด ๋ชจ๋ธ ๋Œ€๋น„ ์•ฝ 10%์˜ ์ ˆ๋Œ€์  ์„ฑ๋Šฅ ํ–ฅ์ƒ์„ ๋ณด์˜€์œผ๋ฉฐ, ํŒ๋… ์†Œ๊ฒฌ ์š”์•ฝ ๊ณผ์ œ์—์„œ๋„ ์˜ˆ์‹œ ๊ธฐ๋ฐ˜ ํ”„๋กฌํ”„ํŒ…์„ ํ†ตํ•ด ์ง€๋„ํ•™์Šต ๊ธฐ๋ฐ˜ ์ตœ์‹  ๋ชจ๋ธ๊ณผ ๋™๋“ฑํ•œ ์ˆ˜์ค€์— ๋„๋‹ฌํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

110Foundation models for generalist medical artificial intelligence

2023Natureโญ Q1DOI 10.1038/s41586-023-05881-4

The exceptionally rapid development of highly flexible, reusable artificial intelligence (AI) models is likely to usher in newfound capabilities in medicine. We propose a new paradigm for medical AI, which we refer to as generalist medical AI (GMAI). GMAI models will be capable of carrying out a diverse set of tasks using very little or no task-specific labelled data. Built through self-supervision on large, diverse datasets, GMAI will flexibly interpret different combinations of medical modalities, including data from imaging, electronic health records, laboratory results, genomics, graphs or medical text. Models will in turn produce expressive outputs such as free-text explanations, spoken recommendations or image annotations that demonstrate advanced medical reasoning abilities. Here we identify a set of high-impact potential applications for GMAI and lay out specific technical capabilities and training datasets necessary to enable them. We expect that GMAI-enabled applications will challenge current strategies for regulating and validating AI devices for medicine and will shift practices associated with the collection of large medical datasets.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
์ด ๋…ผ๋ฌธ์€ ๊ณผ์—…๋ณ„ ๋ ˆ์ด๋ธ” ๋ฐ์ดํ„ฐ ์—†์ด๋„ ๋‹ค์–‘ํ•œ ์˜๋ฃŒ ๊ณผ์—…์„ ์ˆ˜ํ–‰ํ•  ์ˆ˜ ์žˆ๋Š” ๋ฒ”์šฉ ์˜๋ฃŒ ์ธ๊ณต์ง€๋Šฅ(GMAI, Generalist Medical AI)์ด๋ผ๋Š” ์ƒˆ๋กœ์šด ํŒจ๋Ÿฌ๋‹ค์ž„์„ ์ œ์•ˆํ•˜๋ฉฐ, ๋Œ€๊ทœ๋ชจ ๋‹ค์–‘ํ•œ ๋ฐ์ดํ„ฐ์…‹์„ ํ™œ์šฉํ•œ ์ž๊ธฐ์ง€๋„ํ•™์Šต(self-supervision)์„ ํ•ต์‹ฌ ๊ตฌ์ถ• ๋ฐฉ๋ฒ•์œผ๋กœ ์ œ์‹œํ•œ๋‹ค. GMAI ๋ชจ๋ธ์€ ์˜๋ฃŒ ์˜์ƒ, ์ „์ž๊ฑด๊ฐ•๊ธฐ๋ก, ๊ฒ€์‚ฌ ๊ฒฐ๊ณผ, ์œ ์ „์ฒด ๋ฐ์ดํ„ฐ, ์˜๋ฃŒ ํ…์ŠคํŠธ ๋“ฑ ๋ณต์ˆ˜์˜ ์˜๋ฃŒ ๋ชจ๋‹ฌ๋ฆฌํ‹ฐ๋ฅผ ํ†ตํ•ฉ ํ•ด์„ํ•˜๊ณ , ์ž์œ  ํ…์ŠคํŠธ ์„ค๋ช…ยท์Œ์„ฑ ๊ถŒ๊ณ ยท์ด๋ฏธ์ง€ ์ฃผ์„ ๋“ฑ ๊ณ ์ฐจ์›์  ์˜๋ฃŒ ์ถ”๋ก  ๊ฒฐ๊ณผ๋ฌผ์„ ์ถœ๋ ฅํ•  ์ˆ˜ ์žˆ๋‹ค. ์ €์ž๋“ค์€ GMAI์˜ ๊ณ ์˜ํ–ฅ ์ž ์žฌ ์‘์šฉ ๋ถ„์•ผ์™€ ์ด๋ฅผ ๊ตฌํ˜„ํ•˜๊ธฐ ์œ„ํ•œ ๊ธฐ์ˆ ์  ์š”๊ฑด ๋ฐ ํ›ˆ๋ จ ๋ฐ์ดํ„ฐ์…‹์„ ๊ตฌ์ฒด์ ์œผ๋กœ ์ œ์‹œํ•˜๋ฉฐ, ์ด ๊ธฐ์ˆ ์ด ํ˜„ํ–‰ ์˜๋ฃŒ AI ๊ทœ์ œยท๊ฒ€์ฆ ์ฒด๊ณ„์™€ ๋Œ€๊ทœ๋ชจ ์˜๋ฃŒ ๋ฐ์ดํ„ฐ ์ˆ˜์ง‘ ๊ด€ํ–‰์— ๊ทผ๋ณธ์ ์ธ ๋ณ€ํ™”๋ฅผ ์ด‰๊ตฌํ•  ๊ฒƒ์œผ๋กœ ์ „๋งํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

111From human writing to artificial intelligence generated text: examining the prospects and potential threats of ChatGPT in academic writing

2023Biology of Sportโญ Q1DOI 10.5114/biolsport.2023.125623

Natural language processing (NLP) has been studied in computing for decades. Recent technological advancements have led to the development of sophisticated artificial intelligence (AI) models, such as Chat Generative Pre-trained Transformer (ChatGPT). These models can perform a range of language tasks and generate human-like responses, which offers exciting prospects for academic efficiency. This manuscript aims at (i) exploring the potential benefits and threats of ChatGPT and other NLP technologies in academic writing and research publications; (ii) highlights the ethical considerations involved in using these tools, and (iii) consider the impact they may have on the authenticity and credibility of academic work. This study involved a literature review of relevant scholarly articles published in peer-reviewed journals indexed in Scopus as quartile 1. The search used keywords such as "ChatGPT," "AI-generated text," "academic writing," and "natural language processing." The analysis was carried out using a quasi-qualitative approach, which involved reading and critically evaluating the sources and identifying relevant data to support the research questions. The study found that ChatGPT and other NLP technologies have the potential to enhance academic writing and research efficiency. However, their use also raises concerns about the impact on the authenticity and credibility of academic work. The study highlights the need for comprehensive discussions on the potential use, threats, and limitations of these tools, emphasizing the importance of ethical and academic principles, with human intelligence and critical thinking at the forefront of the research process. This study highlights the need for comprehensive debates and ethical considerations involved in their use. The study also recommends that academics exercise caution when using these tools and ensure transparency in their use, emphasizing the importance of human intelligence and critical thinking in academic work.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ChatGPT๋ฅผ ๋น„๋กฏํ•œ ์ž์—ฐ์–ด ์ฒ˜๋ฆฌ(NLP) ๊ธฐ์ˆ ์ด ํ•™์ˆ  ์ €์ˆ  ๋ฐ ์—ฐ๊ตฌ ์ถœํŒ์— ๋ฏธ์น˜๋Š” ์ž ์žฌ์  ์ด์ ๊ณผ ์œ„ํ˜‘, ๊ทธ๋ฆฌ๊ณ  ์œค๋ฆฌ์  ๊ณ ๋ ค์‚ฌํ•ญ์„ ํƒ์ƒ‰ํ•˜๋Š” ๊ฒƒ์„ ๋ชฉ์ ์œผ๋กœ ํ•˜์˜€๋‹ค. Scopus 1๋ถ„์œ„ ํ•™์ˆ ์ง€์— ๊ฒŒ์žฌ๋œ ๋ฌธํ—Œ์„ ๋Œ€์ƒ์œผ๋กœ ์ค€-์งˆ์ (quasi-qualitative) ๋ฌธํ—Œ ๊ณ ์ฐฐ ๋ฐฉ๋ฒ•์„ ์ ์šฉํ•˜์—ฌ ๋ถ„์„์„ ์ˆ˜ํ–‰ํ•˜์˜€๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ChatGPT ๋“ฑ AI ๊ธฐ๋ฐ˜ ๋„๊ตฌ๋Š” ํ•™์ˆ ์  ํšจ์œจ์„ฑ์„ ํ–ฅ์ƒ์‹œํ‚ฌ ์ˆ˜ ์žˆ๋Š” ๊ฐ€๋Šฅ์„ฑ์„ ์ง€๋‹ˆ๋‚˜, ๋™์‹œ์— ํ•™์ˆ  ์ €์ž‘๋ฌผ์˜ ์ง„์œ„์„ฑ๊ณผ ์‹ ๋ขฐ์„ฑ์„ ํ›ผ์†ํ•  ์šฐ๋ ค๊ฐ€ ์žˆ์–ด ํˆฌ๋ช…ํ•œ ์‚ฌ์šฉ๊ณผ ์ธ๊ฐ„์˜ ๋น„ํŒ์  ์‚ฌ๊ณ ๋ฅผ ์šฐ์„ ์‹œํ•˜๋Š” ์œค๋ฆฌ์  ์›์น™์˜ ํ™•๋ฆฝ์ด ํ•„์š”ํ•˜๋‹ค๋Š” ๊ฒฐ๋ก ์„ ๋„์ถœํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

112Help, My Lung is Trapped! โ€“ A Case Report

2023Austin Journal of RadiologyDOI 10.26420/austinjradiol.2023.1210
Added: 2026-04-21 16:11View โ†—

113Performance of ChatGPT on USMLE: Potential for AI-assisted medical education using large language models

2023PLOS Digital Healthโญ Q1DOI 10.1371/journal.pdig.0000198

We evaluated the performance of a large language model called ChatGPT on the United States Medical Licensing Exam (USMLE), which consists of three exams: Step 1, Step 2CK, and Step 3. ChatGPT performed at or near the passing threshold for all three exams without any specialized training or reinforcement. Additionally, ChatGPT demonstrated a high level of concordance and insight in its explanations. These results suggest that large language models may have the potential to assist with medical education, and potentially, clinical decision-making.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์ธ ChatGPT๊ฐ€ ๋ฏธ๊ตญ ์˜์‚ฌ ๋ฉดํ—ˆ ์‹œํ—˜(USMLE)์˜ Step 1, Step 2CK, Step 3 ์ „ ๊ณผ์ •์—์„œ ์–ด๋– ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์ด๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€๋‹ค. ๋ณ„๋„์˜ ํŠนํ™” ํ›ˆ๋ จ์ด๋‚˜ ๊ฐ•ํ™” ํ•™์Šต ์—†์ด๋„ ChatGPT๋Š” ์„ธ ์‹œํ—˜ ๋ชจ๋‘์—์„œ ํ•ฉ๊ฒฉ ๊ธฐ์ค€์— ๊ทผ์ ‘ํ•˜๊ฑฐ๋‚˜ ์ด๋ฅผ ์ถฉ์กฑํ•˜๋Š” ์ ์ˆ˜๋ฅผ ๋‹ฌ์„ฑํ•˜์˜€์œผ๋ฉฐ, ๋‹ต๋ณ€ ์„ค๋ช…์—์„œ๋„ ๋†’์€ ์ผ๊ด€์„ฑ๊ณผ ํ†ต์ฐฐ๋ ฅ์„ ๋ณด์˜€๋‹ค. ์ด๋Ÿฌํ•œ ๊ฒฐ๊ณผ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์ด ์˜ํ•™ ๊ต์œก๋ฟ๋งŒ ์•„๋‹ˆ๋ผ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ๋„๊ตฌ๋กœ์„œ์˜ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ์‹œ์‚ฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

114Performance of Multimodal GPT-4V on USMLE with Image: Potential for Imaging Diagnostic Support with Explanations

2023medRxivDOI 10.1101/2023.10.26.23297629

Abstract Background Using artificial intelligence (AI) to help clinical diagnoses has been an active research topic for more than six decades. Past research, however, has not had the scale and accuracy for use in clinical decision making. The power of AI in large language model (LLM)-related technologies may be changing this. In this study, we evaluated the performance and interpretability of Generative Pre-trained Transformer 4 Vision (GPT-4V), a multimodal LLM, on medical licensing examination questions with images. Methods We used three sets of multiple-choice questions with images from the United States Medical Licensing Examination (USMLE), the USMLE question bank for medical students with different difficulty level (AMBOSS), and the Diagnostic Radiology Qualifying Core Exam (DRQCE) to test GPT-4Vโ€™s accuracy and explanation quality. We compared GPT-4V with two state-of-the-art LLMs, GPT-4 and ChatGPT. We also assessed the preference and feedback of healthcare professionals on GPT-4Vโ€™s explanations. We presented a case scenario on how GPT-4V can be used for clinical decision support. Results GPT-4V outperformed ChatGPT (58.4%) and GPT4 (83.6%) to pass the full USMLE exam with an overall accuracy of 90.7%. In comparison, the passing threshold was 60% for medical students. For questions with images, GPT-4V achieved a performance that was equivalent to the 70th - 80th percentile with AMBOSS medical students, with accuracies of 86.2%, 73.1%, and 62.0% on USMLE, DRQCE, and AMBOSS, respectively. While the accuracies decreased quickly among medical students when the difficulties of questions increased, the performance of GPT-4V remained relatively stable. On the other hand, GPT-4Vโ€™s performance varied across different medical subdomains, with the highest accuracy in immunology (100%) and otolaryngology (100%) and the lowest accuracy in anatomy (25%) and emergency medicine (25%). When GPT-4V answered correctly, its explanations were almost as good as those made by domain experts. However, when GPT-4V answered incorrectly, the quality of generated explanation was poor: 18.2% wrong answers had made-up text; 45.5% had inferencing errors; and 76.3% had image misunderstandings. Our results show that after experts gave GPT-4V a short hint about the image, it reduced 40.5% errors on average, and more difficult test questions had higher performance gains. Therefore, a hypothetical clinical decision support system as shown in our case scenario is a human-AI-in-the-loop system where a clinician can interact with GPT-4V with hints to maximize its clinical use. Conclusion GPT-4V outperformed other LLMs and typical medical student performance on results for medical licensing examination questions with images. However, uneven subdomain performance and inconsistent explanation quality may restrict its practical application in clinical settings. The observation that physiciansโ€™ hints significantly improved GPT-4Vโ€™s performance suggests that future research could focus on developing more effective human-AI collaborative systems. Such systems could potentially overcome current limitations and make GPT-4V more suitable for clinical use. 1-2 sentence description In this study the authors show that GPT-4V, a large multimodal chatbot, achieved accuracy on medical licensing exams with images equivalent to the 70th - 80th percentile with AMBOSS medical students. The authors also show issues with GPT-4V, including uneven performance in different clinical subdomains and explanation quality, which may hamper its clinical use.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋‹ค์ค‘๋ชจ๋‹ฌ ๋Œ€ํ˜•์–ธ์–ด๋ชจ๋ธ์ธ GPT-4V์˜ ์ž„์ƒ ์˜์ƒ ์ง„๋‹จ ๋ณด์กฐ ๊ฐ€๋Šฅ์„ฑ์„ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด, USMLEยทAMBOSSยท๋ฐฉ์‚ฌ์„ ๊ณผ ์ž๊ฒฉ์‹œํ—˜(DRQCE)์˜ ์ด๋ฏธ์ง€ ํฌํ•จ ๊ฐ๊ด€์‹ ๋ฌธํ•ญ์„ ํ™œ์šฉํ•˜์—ฌ GPT-4V์˜ ์ •๋‹ต๋ฅ  ๋ฐ ์„ค๋ช… ํ’ˆ์งˆ์„ GPT-4, ChatGPT์™€ ๋น„๊ต ๋ถ„์„ํ•˜์˜€๋‹ค. GPT-4V๋Š” USMLE ์ „์ฒด์—์„œ 90.7%์˜ ์ •ํ™•๋„๋กœ ๋‹ค๋ฅธ ๋ชจ๋ธ์„ ์ƒํšŒํ•˜์˜€์œผ๋ฉฐ ์ด๋ฏธ์ง€ ํฌํ•จ ๋ฌธํ•ญ์—์„œ AMBOSS ์˜๋Œ€์ƒ ๊ธฐ์ค€ 70~80๋ฐฑ๋ถ„์œ„์— ํ•ด๋‹นํ•˜๋Š” ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋‚˜, ๋ฉด์—ญํ•™ยท์ด๋น„์ธํ›„๊ณผ(๊ฐ 100%)์™€ ๋‹ฌ๋ฆฌ ํ•ด๋ถ€ํ•™ยท์‘๊ธ‰์˜ํ•™(๊ฐ 25%) ๋“ฑ ์ผ๋ถ€ ์„ธ๋ถ€ ๋ถ„์•ผ์—์„œ๋Š” ์„ฑ๋Šฅ์ด ํ˜„์ €ํžˆ ์ €ํ•˜๋˜์—ˆ๋‹ค. ์˜ค๋‹ต ์‹œ ํ™˜๊ฐ(hallucination), ์ถ”๋ก  ์˜ค๋ฅ˜, ์ด๋ฏธ์ง€ ์˜ค๋… ๋“ฑ ์„ค๋ช… ํ’ˆ์งˆ์ด ํฌ๊ฒŒ ๋–จ์–ด์ง€๋Š” ํ•œ๊ณ„๊ฐ€ ์žˆ์—ˆ์œผ๋‚˜, ์ž„์ƒ์˜์˜ ๊ฐ„๋žตํ•œ ํžŒํŠธ ์ œ๊ณต์ด ์˜ค๋ฅ˜๋ฅผ ํ‰๊ท  40.5% ๊ฐ์†Œ์‹œ์ผœ, ํ–ฅํ›„ ์ธ๊ฐ„-AI ํ˜‘์—… ๊ธฐ๋ฐ˜์˜ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ์‹œ์Šคํ…œ ๊ฐœ๋ฐœ์ด ํ˜„์žฌ์˜ ํ•œ๊ณ„๋ฅผ ๊ทน๋ณตํ•˜๋Š” ๋ฐ ํ•ต์‹ฌ์  ๋ฐฉํ–ฅ์ด ๋  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

115RadAdapt: Radiology Report Summarization via Lightweight Domain Adaptation of Large Language Models

2023arXiv (Cornell University)DOI 10.48550/arxiv.2305.01146

We systematically investigate lightweight strategies to adapt large language models (LLMs) for the task of radiology report summarization (RRS). Specifically, we focus on domain adaptation via pretraining (on natural language, biomedical text, or clinical text) and via discrete prompting or parameter-efficient fine-tuning. Our results consistently achieve best performance by maximally adapting to the task via pretraining on clinical text and fine-tuning on RRS examples. Importantly, this method fine-tunes a mere 0.32% of parameters throughout the model, in contrast to end-to-end fine-tuning (100% of parameters). Additionally, we study the effect of in-context examples and out-of-distribution (OOD) training before concluding with a radiologist reader study and qualitative analysis. Our findings highlight the importance of domain adaptation in RRS and provide valuable insights toward developing effective natural language processing solutions for clinical tasks.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ๋ฐฉ์‚ฌ์„ ๊ณผ ๋ณด๊ณ ์„œ ์š”์•ฝ(RRS) ๊ณผ์ œ์— ํšจ์œจ์ ์œผ๋กœ ์ ์šฉํ•˜๊ธฐ ์œ„ํ•œ ๊ฒฝ๋Ÿ‰ ๋„๋ฉ”์ธ ์ ์‘ ์ „๋žต์„ ์ฒด๊ณ„์ ์œผ๋กœ ๊ฒ€ํ† ํ•˜์˜€๋‹ค. ์ž„์ƒ ํ…์ŠคํŠธ ๊ธฐ๋ฐ˜ ์‚ฌ์ „ํ•™์Šต๊ณผ ๋งค๊ฐœ๋ณ€์ˆ˜ ํšจ์œจ์  ๋ฏธ์„ธ์กฐ์ •(์ „์ฒด ๋งค๊ฐœ๋ณ€์ˆ˜์˜ ๋‹จ 0.32% ์กฐ์ •)์„ ๊ฒฐํ•ฉํ•œ ๋ฐฉ๋ฒ•์ด ์ „์ฒด ๋งค๊ฐœ๋ณ€์ˆ˜ ๋ฏธ์„ธ์กฐ์ •(100%) ๋Œ€๋น„ ๋™๋“ฑํ•˜๊ฑฐ๋‚˜ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋‹ฌ์„ฑํ•จ์„ ํ™•์ธํ•˜์˜€๋‹ค. ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜ ํŒ๋… ์—ฐ๊ตฌ ๋ฐ ์ •์„ฑ์  ๋ถ„์„์„ ํ†ตํ•ด ๋„๋ฉ”์ธ ์ ์‘์ด ์ž„์ƒ ์ž์—ฐ์–ด ์ฒ˜๋ฆฌ ๊ณผ์ œ์˜ ํšจ๊ณผ์  ์ˆ˜ํ–‰์— ํ•ต์‹ฌ์  ์—ญํ• ์„ ํ•จ์„ ์ž…์ฆํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

116Revolutionizing healthcare: the role of artificial intelligence in clinical practice

2023BMC Medical Educationโญ Q1DOI 10.1186/s12909-023-04698-z

AI can be used to diagnose diseases, develop personalized treatment plans, and assist clinicians with decision-making. Rather than simply automating tasks, AI is about developing technologies that can enhance patient care across healthcare settings. However, challenges related to data privacy, bias, and the need for human expertise must be addressed for the responsible and effective implementation of AI in healthcare.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์ž„์ƒ ํ˜„์žฅ์—์„œ ์ธ๊ณต์ง€๋Šฅ(AI)์ด ์ˆ˜ํ–‰ํ•  ์ˆ˜ ์žˆ๋Š” ์—ญํ• , ์ฆ‰ ์งˆ๋ณ‘ ์ง„๋‹จ, ๊ฐœ์ธ ๋งž์ถคํ˜• ์น˜๋ฃŒ ๊ณ„ํš ์ˆ˜๋ฆฝ ๋ฐ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ๊ธฐ๋Šฅ์„ ํฌ๊ด„์ ์œผ๋กœ ๊ฒ€ํ† ํ•˜์˜€๋‹ค. AI๋Š” ๋‹จ์ˆœํ•œ ์—…๋ฌด ์ž๋™ํ™”๋ฅผ ๋„˜์–ด ๋‹ค์–‘ํ•œ ์˜๋ฃŒ ํ™˜๊ฒฝ์—์„œ ํ™˜์ž ์น˜๋ฃŒ์˜ ์งˆ์„ ํ–ฅ์ƒ์‹œํ‚ค๋Š” ๊ธฐ์ˆ ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์ œ์‹œํ•˜๋ฉฐ, ์ด๋ฅผ ์˜๋ฃŒ ์‹ค๋ฌด์— ํ†ตํ•ฉํ•˜๋Š” ๋ฐฉํ–ฅ์„ ๋…ผ์˜ํ•˜์˜€๋‹ค. ๋‹ค๋งŒ, AI์˜ ์ฑ…์ž„๊ฐ ์žˆ๊ณ  ํšจ๊ณผ์ ์ธ ๋„์ž…์„ ์œ„ํ•ด์„œ๋Š” ๋ฐ์ดํ„ฐ ํ”„๋ผ์ด๋ฒ„์‹œ, ์•Œ๊ณ ๋ฆฌ์ฆ˜ ํŽธํ–ฅ์„ฑ, ๊ทธ๋ฆฌ๊ณ  ์ธ๊ฐ„ ์ „๋ฌธ๊ฐ€์˜ ์—ญํ•  ์œ ์ง€ ๋“ฑ์˜ ๊ณผ์ œ๊ฐ€ ๋ฐ˜๋“œ์‹œ ํ•ด๊ฒฐ๋˜์–ด์•ผ ํ•จ์„ ๊ฐ•์กฐํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

117The future landscape of large language models in medicine

2023Communications Medicineโญ Q1DOI 10.1038/s43856-023-00370-1

Large language models (LLMs) are artificial intelligence (AI) tools specifically trained to process and generate text. LLMs attracted substantial public attention after OpenAI's ChatGPT was made publicly available in November 2022. LLMs can often answer questions, summarize, paraphrase and translate text on a level that is nearly indistinguishable from human capabilities. The possibility to actively interact with models like ChatGPT makes LLMs attractive tools in various fields, including medicine. While these models have the potential to democratize medical knowledge and facilitate access to healthcare, they could equally distribute misinformation and exacerbate scientific misconduct due to a lack of accountability and transparency. In this article, we provide a systematic and comprehensive overview of the potentials and limitations of LLMs in clinical practice, medical research and medical education.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ๋…ผ๋ฌธ์€ 2022๋…„ 11์›” ChatGPT ๊ณต๊ฐœ ์ดํ›„ ์ฃผ๋ชฉ๋ฐ›๊ณ  ์žˆ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์˜ํ•™์  ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ๊ณผ ํ•œ๊ณ„๋ฅผ ์ฒด๊ณ„์ ์œผ๋กœ ๊ณ ์ฐฐํ•˜๋Š” ๊ฒƒ์„ ๋ชฉ์ ์œผ๋กœ ํ•œ๋‹ค. ์ €์ž๋“ค์€ ๋ฌธํ—Œ ๊ธฐ๋ฐ˜์˜ ์ฒด๊ณ„์  ๊ฐœ์š” ๋ถ„์„์„ ํ†ตํ•ด ์ž„์ƒ ์ง„๋ฃŒ, ์˜ํ•™ ์—ฐ๊ตฌ, ์˜ํ•™ ๊ต์œก ๋ถ„์•ผ์—์„œ LLM์˜ ์ž ์žฌ์  ์—ญํ• ์„ ํฌ๊ด„์ ์œผ๋กœ ๊ฒ€ํ† ํ•˜์˜€๋‹ค. LLM์€ ์˜๋ฃŒ ์ง€์‹์˜ ๋Œ€์ค‘ํ™” ๋ฐ ์˜๋ฃŒ ์ ‘๊ทผ์„ฑ ํ–ฅ์ƒ์— ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ๋Š” ๋ฐ˜๋ฉด, ์ฑ…์ž„์„ฑ๊ณผ ํˆฌ๋ช…์„ฑ ๋ถ€์žฌ๋กœ ์ธํ•œ ์˜๋ฃŒ ์˜ค์ •๋ณด ํ™•์‚ฐ ๋ฐ ์—ฐ๊ตฌ ์œค๋ฆฌ ์œ„๋ฐ˜์˜ ์œ„ํ—˜์„ฑ๋„ ๋‚ดํฌํ•˜๊ณ  ์žˆ์Œ์„ ๊ฐ•์กฐํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

118The imperative for regulatory oversight of large language models (or generative AI) in healthcare

2023npj Digital Medicineโญ Q1DOI 10.1038/s41746-023-00873-0

The rapid advancements in artificial intelligence (AI) have led to the development of sophisticated large language models (LLMs) such as GPT-4 and Bard. The potential implementation of LLMs in healthcare settings has already garnered considerable attention because of their diverse applications that include facilitating clinical documentation, obtaining insurance pre-authorization, summarizing research papers, or working as a chatbot to answer questions for patients about their specific data and concerns. While offering transformative potential, LLMs warrant a very cautious approach since these models are trained differently from AI-based medical technologies that are regulated already, especially within the critical context of caring for patients. The newest version, GPT-4, that was released in March, 2023, brings the potentials of this technology to support multiple medical tasks; and risks from mishandling results it provides to varying reliability to a new level. Besides being an advanced LLM, it will be able to read texts on images and analyze the context of those images. The regulation of GPT-4 and generative AI in medicine and healthcare without damaging their exciting and transformative potential is a timely and critical challenge to ensure safety, maintain ethical standards, and protect patient privacy. We argue that regulatory oversight should assure medical professionals and patients can use LLMs without causing harm or compromising their data or privacy. This paper summarizes our practical recommendations for what we can expect from regulators to bring this vision to reality.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ๋…ผ๋ฌธ์€ GPT-4 ๋“ฑ ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์˜๋ฃŒ ํ™˜๊ฒฝ ๋„์ž…์ด ๊ธ‰์†ํžˆ ํ™•์‚ฐ๋จ์— ๋”ฐ๋ผ, ํ™˜์ž ์•ˆ์ „ยท์œค๋ฆฌ ๊ธฐ์ค€ยท๊ฐœ์ธ์ •๋ณด ๋ณดํ˜ธ๋ฅผ ๋‹ด๋ณดํ•˜๊ธฐ ์œ„ํ•œ ๊ทœ์ œ ๊ฐ๋…์˜ ํ•„์š”์„ฑ์„ ๋…ผ์˜ํ•œ๋‹ค. ์ €์ž๋“ค์€ LLM์ด ์ž„์ƒ ๋ฌธ์„œ ์ž‘์„ฑ, ๋ณดํ—˜ ์‚ฌ์ „ ์Šน์ธ, ์—ฐ๊ตฌ ๋…ผ๋ฌธ ์š”์•ฝ, ํ™˜์ž ๋Œ€์ƒ ์ฑ—๋ด‡ ๋“ฑ ๋‹ค์–‘ํ•œ ์˜๋ฃŒ ์—…๋ฌด์— ํ™œ์šฉ๋  ์ˆ˜ ์žˆ์œผ๋‚˜, ๊ธฐ์กด ๊ทœ์ œ ๋Œ€์ƒ AI ์˜๋ฃŒ๊ธฐ๊ธฐ์™€๋Š” ํ•™์Šต ๋ฐฉ์‹์ด ์ƒ์ดํ•˜๊ณ  ์‹ ๋ขฐ์„ฑ ๋ณ€๋™ ๋ฐ ์˜ค์šฉ ์œ„ํ—˜์ด ๋” ๋†’๋‹ค๋Š” ์ ์„ ์ง€์ ํ•œ๋‹ค. ์ด์— ๋”ฐ๋ผ ์˜๋ฃŒ ์ „๋ฌธ๊ฐ€์™€ ํ™˜์ž๊ฐ€ LLM์„ ์•ˆ์ „ํ•˜๊ฒŒ ํ™œ์šฉํ•  ์ˆ˜ ์žˆ๋„๋ก ๊ทœ์ œ ๊ธฐ๊ด€์ด ์ทจํ•ด์•ผ ํ•  ์‹ค์งˆ์  ๊ถŒ๊ณ  ์‚ฌํ•ญ์„ ์ œ์‹œํ•˜๋ฉฐ, ํ˜์‹ ์  ์ž ์žฌ๋ ฅ์„ ํ›ผ์†ํ•˜์ง€ ์•Š๋Š” ๊ท ํ˜• ์žกํžŒ ๊ทœ์ œ ํ”„๋ ˆ์ž„์›Œํฌ์˜ ์กฐ์†ํ•œ ๋งˆ๋ จ์„ ์ด‰๊ตฌํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

119Utility of ChatGPT in Clinical Practice

2023Journal of Medical Internet Researchโญ Q1DOI 10.2196/48568

ChatGPT is receiving increasing attention and has a variety of application scenarios in clinical practice. In clinical decision support, ChatGPT has been used to generate accurate differential diagnosis lists, support clinical decision-making, optimize clinical decision support, and provide insights for cancer screening decisions. In addition, ChatGPT has been used for intelligent question-answering to provide reliable information about diseases and medical queries. In terms of medical documentation, ChatGPT has proven effective in generating patient clinical letters, radiology reports, medical notes, and discharge summaries, improving efficiency and accuracy for health care providers. Future research directions include real-time monitoring and predictive analytics, precision medicine and personalized treatment, the role of ChatGPT in telemedicine and remote health care, and integration with existing health care systems. Overall, ChatGPT is a valuable tool that complements the expertise of health care providers and improves clinical decision-making and patient care. However, ChatGPT is a double-edged sword. We need to carefully consider and study the benefits and potential dangers of ChatGPT. In this viewpoint, we discuss recent advances in ChatGPT research in clinical practice and suggest possible risks and challenges of using ChatGPT in clinical practice. It will help guide and support future artificial intelligence research similar to ChatGPT in health.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ๋…ผ๋ฌธ์€ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์ธ ChatGPT์˜ ์ž„์ƒ ํ˜„์žฅ ์ ์šฉ ๊ฐ€๋Šฅ์„ฑ๊ณผ ์ตœ๊ทผ ์—ฐ๊ตฌ ๋™ํ–ฅ์„ ์ข…ํ•ฉ์ ์œผ๋กœ ๊ณ ์ฐฐํ•˜๋Š” ๊ฒƒ์„ ๋ชฉ์ ์œผ๋กœ ํ•œ๋‹ค. ChatGPT๋Š” ๊ฐ๋ณ„์ง„๋‹จ ์ƒ์„ฑ, ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์›, ์•” ๊ฒ€์ง„ ๊ด€๋ จ ์ธ์‚ฌ์ดํŠธ ์ œ๊ณต ๋“ฑ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ๋ณด์กฐ ๋ถ„์•ผ์™€ ๋”๋ถˆ์–ด, ์ง„๋ฃŒ ์„œํ•œยท๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œยทํ‡ด์› ์š”์•ฝ๋ฌธ ๋“ฑ ์˜๋ฃŒ ๋ฌธ์„œ ์ž‘์„ฑ์˜ ํšจ์œจ์„ฑ ๋ฐ ์ •ํ™•์„ฑ ํ–ฅ์ƒ์— ํšจ๊ณผ์ ์ž„์ด ๋ณด๊ณ ๋˜์—ˆ๋‹ค. ๊ทธ๋Ÿฌ๋‚˜ ChatGPT๋Š” ์–‘๋‚ ์˜ ๊ฒ€์œผ๋กœ์„œ, ์›๊ฒฉ์˜๋ฃŒยท์ •๋ฐ€์˜ํ•™ยท๊ธฐ์กด ์˜๋ฃŒ ์‹œ์Šคํ…œ๊ณผ์˜ ํ†ตํ•ฉ ๋“ฑ ํ–ฅํ›„ ์—ฐ๊ตฌ ๋ฐฉํ–ฅ์„ ์ œ์‹œํ•˜๋Š” ๋™์‹œ์—, ์ž„์ƒ ์ ์šฉ ์‹œ ๋ฐœ์ƒํ•  ์ˆ˜ ์žˆ๋Š” ์ž ์žฌ์  ์œ„ํ—˜๊ณผ ์œค๋ฆฌ์  ๊ณผ์ œ์— ๋Œ€ํ•œ ์‹ ์ค‘ํ•œ ๊ฒ€ํ† ์˜ ํ•„์š”์„ฑ์„ ๊ฐ•์กฐํ•œ๋‹ค.
Added: 2026-04-21 16:11View โ†—

120shs-nlp at RadSum23: Domain-Adaptive Pre-training of Instruction-tuned LLMs for Radiology Report Impression Generation

2023UnknownDOI 10.18653/v1/2023.bionlp-1.57

Instruction-tuned generative large language models (LLMs), such as ChatGPT and Bloomz, possess excellent generalization abilities. However, they face limitations in understanding radiology reports, particularly when generating the IMPRESSIONS section from the FINDINGS section. These models tend to produce either verbose or incomplete IMPRESSIONS, mainly due to insufficient exposure to medical text data during training. We present a system that leverages large-scale medical text data for domain-adaptive pre-training of instruction-tuned LLMs, enhancing their medical knowledge and performance on specific medical tasks. We demonstrate that this system performs better in a zero-shot setting compared to several pretrain-and-finetune adaptation methods on the IMPRESSIONS generation task. Furthermore, it ranks 1st among participating systems in Task 1B: Radiology Report Summarization.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ChatGPT ๋ฐ Bloomz์™€ ๊ฐ™์€ ๋ช…๋ น์–ด ์กฐ์ •(instruction-tuned) ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ์ด ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์˜ FINDINGS ์„น์…˜์œผ๋กœ๋ถ€ํ„ฐ IMPRESSIONS ์„น์…˜์„ ์ƒ์„ฑํ•˜๋Š” ๋ฐ ์žˆ์–ด ์˜๋ฃŒ ํ…์ŠคํŠธ ๋…ธ์ถœ ๋ถ€์กฑ์œผ๋กœ ์ธํ•ด ์žฅํ™ฉํ•˜๊ฑฐ๋‚˜ ๋ถˆ์™„์ „ํ•œ ๊ฒฐ๊ณผ๋ฅผ ์‚ฐ์ถœํ•˜๋Š” ํ•œ๊ณ„๋ฅผ ๊ทน๋ณตํ•˜๊ณ ์ž, ๋Œ€๊ทœ๋ชจ ์˜๋ฃŒ ํ…์ŠคํŠธ ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•œ ๋„๋ฉ”์ธ ์ ์‘ํ˜• ์‚ฌ์ „ํ•™์Šต(domain-adaptive pre-training) ๋ฐฉ๋ฒ•๋ก ์„ ์ œ์•ˆํ•˜์˜€๋‹ค. ์ œ์•ˆ๋œ ์‹œ์Šคํ…œ์€ ๊ธฐ์กด์˜ ์‚ฌ์ „ํ•™์Šต-๋ฏธ์„ธ์กฐ์ •(pretrain-and-finetune) ์ ์‘ ๋ฐฉ๋ฒ•๋“ค๊ณผ ๋น„๊ตํ•˜์—ฌ ์ œ๋กœ์ƒท(zero-shot) ํ™˜๊ฒฝ์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€๋‹ค. ๊ทธ ๊ฒฐ๊ณผ, RadSum23 ๊ณต์œ  ๊ณผ์ œ์˜ Task 1B(๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ์š”์•ฝ ๋ถ€๋ฌธ)์—์„œ ์ฐธ๊ฐ€ ์‹œ์Šคํ…œ ์ค‘ 1์œ„๋ฅผ ๋‹ฌ์„ฑํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—

121shs-nlp at RadSum23: Domain-Adaptive Pre-training of Instruction-tuned LLMs for Radiology Report Impression Generation

2023arXiv (Cornell University)DOI 10.48550/arxiv.2306.03264

Instruction-tuned generative Large language models (LLMs) like ChatGPT and Bloomz possess excellent generalization abilities, but they face limitations in understanding radiology reports, particularly in the task of generating the IMPRESSIONS section from the FINDINGS section. They tend to generate either verbose or incomplete IMPRESSIONS, mainly due to insufficient exposure to medical text data during training. We present a system which leverages large-scale medical text data for domain-adaptive pre-training of instruction-tuned LLMs to enhance its medical knowledge and performance on specific medical tasks. We show that this system performs better in a zero-shot setting than a number of pretrain-and-finetune adaptation methods on the IMPRESSIONS generation task, and ranks 1st among participating systems in Task 1B: Radiology Report Summarization at the BioNLP 2023 workshop.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ChatGPT ๋ฐ Bloomz์™€ ๊ฐ™์€ instruction-tuned ๋Œ€ํ˜• ์–ธ์–ด ๋ชจ๋ธ(LLM)์ด ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์˜ FINDINGS ์„น์…˜์œผ๋กœ๋ถ€ํ„ฐ IMPRESSIONS ์„น์…˜์„ ์ƒ์„ฑํ•˜๋Š” ๋ฐ ์žˆ์–ด ์˜๋ฃŒ ํ…์ŠคํŠธ ๋…ธ์ถœ ๋ถ€์กฑ์œผ๋กœ ์ธํ•ด ์„ฑ๋Šฅ์ด ์ œํ•œ๋œ๋‹ค๋Š” ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ์ด๋ฅผ ์œ„ํ•ด ๋Œ€๊ทœ๋ชจ ์˜๋ฃŒ ํ…์ŠคํŠธ ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•œ ๋„๋ฉ”์ธ ์ ์‘์  ์‚ฌ์ „ํ•™์Šต(domain-adaptive pre-training)์„ instruction-tuned LLM์— ์ ์šฉํ•˜์—ฌ ์˜๋ฃŒ ์ง€์‹ ๋ฐ ํŠน์ • ์ž„์ƒ ๊ณผ์ œ ์ˆ˜ํ–‰ ๋Šฅ๋ ฅ์„ ๊ฐ•ํ™”ํ•˜๋Š” ์‹œ์Šคํ…œ์„ ์ œ์•ˆํ•˜์˜€๋‹ค. ์ œ์•ˆ๋œ ์‹œ์Šคํ…œ์€ zero-shot ํ™˜๊ฒฝ์—์„œ ๊ธฐ์กด์˜ ์—ฌ๋Ÿฌ ์‚ฌ์ „ํ•™์Šต-๋ฏธ์„ธ์กฐ์ •(pretrain-and-finetune) ๋ฐฉ๋ฒ•๋ณด๋‹ค ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, BioNLP 2023 ์›Œํฌ์ˆ์˜ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ ์š”์•ฝ ๊ณผ์ œ(Task 1B)์—์„œ ์ฐธ๊ฐ€ ์‹œ์Šคํ…œ ์ค‘ 1์œ„๋ฅผ ๊ธฐ๋กํ•˜์˜€๋‹ค.
Added: 2026-04-21 16:11View โ†—
โ†‘ Back to top
0 selected