๐Ÿ“š Full Archive

Radiology LLM Personalized Reporting
โ† Dashboard ๐Ÿ“„ 59 papers ยท Generated 2026-08-23 00:10

1Extraction of distant recurrence sites for breast cancer patients from free-text clinical notes using large language models.

2026-06Journal of biomedical informaticsโญ Q1DOI 10.1016/j.jbi.2026.105032
OBJECTIVE

Accurate documentation of distant recurrence sites in breast cancer is essential for evaluating treatment effectiveness and outcomes research. However, such information is embedded in unstructured clinical notes, making manual abstraction labor-intensive. Large language models (LLMs) offer a scalable solution for extracting complex information from heterogeneous clinical narratives; however, generic LLMs often lack the specialized clinical reasoning needed for accurate interpretation of oncologic documentation. This study aims to develop an efficient LLM-based framework to automatically extract distant recurrence sites from free-text documentation. MATERIALS &

METHODS

We used clinical notes, pathology and radiology reports from recurrent breast cancer patients at Mayo Clinic (nย =ย 766) for model development and evaluated generalizability on internal hold-out samples (nย =ย 112) and an external Stanford Medicine cohort (nย =ย 110). For cross-disease domain adaptation, we further validated on prostate cancer patients (nย =ย 49). Our proposed framework employs BioLinkBERT, a pretrained language model (PLM) backbone, with weak supervision and an epoch-wise entropy optimization to address limited labeled data and class imbalance across recurrence sites. The fine-tuned model was compared against state-of-the-art models, including Llama2-7B, Llama-3-8B and MedAlpaca, using precision, recall, and F1-score.

RESULTS

The fine-tuned model outperformed generic and domain-specific LLM baselines, with notable gains in identifying multi-site distant recurrence. In-domain validation showed consistent F1-score improvement (average 0.78), particularly for rare recurrence sites. The model also demonstrated strong performance on the external Stanford cohort and on prostate cancer, achieving F1-score of 0.83 and 0.93, respectively.

CONCLUSION

This study presents an efficient, weakly supervised LLM framework that accurately extracts metastatic recurrence sites, reducing reliance on manual chart review. The results demonstrate that relatively small LLMs, optimized with domain-aware weak supervision, can outperform larger models for complex oncologic information extraction. The model is released as a platform-independent Docker image to support seamless cancer registry integration.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์œ ๋ฐฉ์•” ํ™˜์ž์˜ ๋น„์ •ํ˜• ์ž„์ƒ ๊ธฐ๋ก์—์„œ ์›๊ฒฉ ์ „์ด ๋ถ€์œ„๋ฅผ ์ž๋™์œผ๋กœ ์ถ”์ถœํ•˜๊ธฐ ์œ„ํ•ด BioLinkBERT ๊ธฐ๋ฐ˜์˜ ์•ฝ์ง€๋„ ํ•™์Šต(weakly supervised) LLM ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ๊ฐœ๋ฐœํ•˜์˜€์Šต๋‹ˆ๋‹ค. ํ•ด๋‹น ๋ชจ๋ธ์€ ๋‚ด๋ถ€ ๋ฐ ์™ธ๋ถ€ ๊ฒ€์ฆ ๋ฐ์ดํ„ฐ์…‹์—์„œ ๊ธฐ์กด LLM ๋Œ€๋น„ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ํŠนํžˆ ๋‹ค๋ฐœ์„ฑ ๋ฐ ํฌ๊ท€ ์ „์ด ๋ถ€์œ„ ์‹๋ณ„์—์„œ ๋†’์€ ์ •ํ™•๋„๋ฅผ ์ž…์ฆํ–ˆ์Šต๋‹ˆ๋‹ค. ์ด๋Š” ์ˆ˜๋™ ์ฐจํŠธ ๋ฆฌ๋ทฐ์˜ ๋ถ€๋‹ด์„ ์ค„์ด๊ณ  ์•” ๋“ฑ๋ก ์ฒด๊ณ„์˜ ํšจ์œจ์„ฑ์„ ๋†’์ด๋Š” ๋ฐ ๊ธฐ์—ฌํ•  ๊ฒƒ์œผ๋กœ ๊ธฐ๋Œ€๋ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 03:13View โ†—

2Exploring Full-cycle DeepSeek-assisted Case-based Learning in Undergraduate Radiology Education: A Respiratory System Example.

2026-04Academic radiologyโญ Q1DOI 10.1016/j.acra.2026.03.026

RATIONALE AND

OBJECTIVE

To evaluate the application of DeepSeek-assisted case-based learning (CBL) in respiratory radiology course across the full instructional cycle, including preparation, implementation, and evaluation.

METHODS

This prospective, single-center study was conducted in 2025 and involved third-year medical undergraduates. CBL Preparation: Six cases were retrieved from the Hospital Information System (HIS), and six generated via DeepSeek-R1. Preparation times were recorded and compared. CBL Implementation: Students were assigned to either a DeepSeek-assisted group or a traditional CBL group, with discussion time recorded for each subgroup. CBL Evaluation: Teaching effectiveness was evaluated through test scores and questionnaires. Subsequently, DeepSeek-R1 provided personalized feedback to students based on their individual scores.

RESULTS

A total of 200 students (mean age 21.02ย ยฑย 0.89 years, 94 males) participated. DeepSeek-generated cases required significantly less time than HIS-retrieved cases (p = 0.016). During implementation, the DeepSeek group spent less discussion time than traditional group (p = 0.026). The DeepSeek-assisted group achieved greater test score improvements compared to the traditional group (p < 0.05). Questionnaire responses indicated higher self-directed learning, greater interest in radiology, improved learning efficiency, and lower perceived learning burden in the DeepSeek-assisted group (p < 0.05). Additionally, personalized feedback generated by DeepSeek was qualitatively reviewed by the radiology teaching department and considered educationally useful.

CONCLUSION

This study demonstrates that DeepSeek-assisted CBL effectively supports respiratory radiology education throughout the entire course process-preparation, implementation, and evaluation-by enhancing efficiency, boosting student interest and engagement, improving performance, and providing valuable post-class feedback.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜๊ณผ๋Œ€ํ•™์ƒ์˜ ํ˜ธํก๊ธฐ ์˜์ƒ์˜ํ•™ ๊ต์œก์„ ์œ„ํ•ด DeepSeek๋ฅผ ํ™œ์šฉํ•œ ์‚ฌ๋ก€ ๊ธฐ๋ฐ˜ ํ•™์Šต(CBL)์˜ ์ „ ๊ณผ์ •์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, DeepSeek๋ฅผ ํ™œ์šฉํ•œ ๋ฐฉ์‹์€ ๊ธฐ์กด ๋ฐฉ์‹๋ณด๋‹ค ์‚ฌ๋ก€ ์ค€๋น„ ์‹œ๊ฐ„์„ ๋‹จ์ถ•์‹œ์ผฐ์„ ๋ฟ๋งŒ ์•„๋‹ˆ๋ผ, ํ•™์ƒ๋“ค์˜ ํ•™์Šต ํšจ์œจ์„ฑ์„ ๋†’์ด๊ณ  ์‹œํ—˜ ์„ฑ์  ํ–ฅ์ƒ ๋ฐ ํ•™์Šต ํฅ๋ฏธ ์œ ๋ฐœ์— ํšจ๊ณผ์ ์ž„์„ ํ™•์ธํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋˜ํ•œ, DeepSeek๊ฐ€ ์ œ๊ณตํ•˜๋Š” ๊ฐœ์ธ๋ณ„ ๋งž์ถคํ˜• ํ”ผ๋“œ๋ฐฑ์ด ๊ต์œก์ ์œผ๋กœ ์œ ์šฉํ•˜๋‹ค๋Š” ์ ์„ ์ž…์ฆํ•˜์—ฌ ์ฐจ์„ธ๋Œ€ ์˜ํ•™ ๊ต์œก ๋„๊ตฌ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ๋ณด์—ฌ์ฃผ์—ˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-12 00:00View โ†—

3Augmenting Large Language Model With Prompt Engineering and Supervised Fine-Tuning in Non-Small Cell Lung Cancer Tumor-Node-Metastasis Staging: Framework Development and Validation.

2026-04JMIR AIโญ Q1DOI 10.2196/77988
BACKGROUND

Accurate tumor node metastasis (TNM) staging is fundamental for treatment planning and prognosis in non-small cell lung cancer (NSCLC). However, its complexity poses significant challenges. Traditional rule-based natural language processing methods are constrained by their reliance on manually crafted rules and are susceptible to inconsistencies in clinical reporting.

OBJECTIVE

This study aimed to develop and validate a robust, accurate, and operationally efficient artificial intelligence framework for the TNM staging of NSCLC by strategically enhancing a large language model, GLM-4-Air (general language model), through advanced prompt engineering and supervised fine-tuning (SFT).

METHODS

We constructed a curated dataset of 492 deidentified real-world medical imaging reports, with TNM staging annotations rigorously validated by senior physicians according to the AJCC (American Joint Committee on Cancer) 8th edition guidelines. The GLM-4-Air model was systematically optimized via a multi-phase process: iterative prompt engineering incorporating chain-of-thought reasoning and domain knowledge injection for all staging tasks, followed by parameter-efficient SFT using low-rank adaptation for the reasoning-intensive primary tumor characteristics (T) and regional lymph node involvement (N) staging tasks. The final hybrid model was evaluated on a completely held-out test set (black-box) and benchmarked against GPT-4o using standard metrics, statistical tests, and a clinical impact analysis of staging errors.

RESULTS

The optimized hybrid GLM-4-Air model demonstrated reliable performance. It achieved higher staging accuracies on the black-box test set: 92% (95% CI 0.850-0.959) for T, 86% (95% CI 0.779-0.915) for N, 92% (95% CI 0.850-0.959) for distant metastasis status (M), and 90% for overall clinical staging; by comparison, GPT-4o attained 87% (95% CI 0.790-0.922), 70% (95% CI 0.604-0.781), 78% (95% CI 0.689-0.850), and 80%, respectively. The model's robustness was further evidenced by its macro-average F1-scores of 0.914 (T), 0.815 (N), and 0.831 (M), consistently surpassing those of GPT-4o (0.836, 0.620, and 0.698). Analysis of confusion matrices confirmed the model's proficiency in identifying critical staging features while effectively minimizing false negatives. Crucially, the clinical impact assessment showed a substantial reduction in severe category I errors, which are defined as misclassifications that could significantly influence subsequent clinical decisions. Our model committed 0 category I errors in M staging and fewer category I errors in T and N staging. Furthermore, the framework demonstrated practical deployability, achieving efficient inference on consumer-grade hardware (eg, 4 RTX 4090 GPUs) with latencies suitable and acceptable for clinical workflows.

CONCLUSION

The proposed hybrid framework, integrating structured prompt engineering and applying SFT to reasoning-heavy tasks (T/N), enables the GLM-4-Air model to serve as a highly accurate, clinically reliable, and cost-efficient solution for automated NSCLC TNM staging. This work demonstrates the efficacy and potential of a domain-optimized smaller model compared with an off-the-shelf generalist model, holding promise for enhancing diagnostic standardization in resource-aware health care environments.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” GLM-4-Air ๋ชจ๋ธ์— ํ”„๋กฌํ”„ํŠธ ์—”์ง€๋‹ˆ์–ด๋ง๊ณผ ์ง€๋„ ๋ฏธ์„ธ ์กฐ์ •(SFT)์„ ๊ฒฐํ•ฉํ•˜์—ฌ ๋น„์†Œ์„ธํฌํ์•”(NSCLC)์˜ TNM ๋ณ‘๊ธฐ ๊ฒฐ์ •์„ ์ž๋™ํ™”ํ•˜๋Š” ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ๊ฐœ๋ฐœํ•˜๊ณ  ๊ฒ€์ฆํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์‹ค์ œ ์ž„์ƒ ๋ณด๊ณ ์„œ ๋ฐ์ดํ„ฐ์…‹์„ ํ™œ์šฉํ•œ ํ‰๊ฐ€ ๊ฒฐ๊ณผ, ํ•ด๋‹น ๋ชจ๋ธ์€ GPT-4o ๋Œ€๋น„ ์šฐ์ˆ˜ํ•œ ๋ณ‘๊ธฐ ํŒ์ • ์ •ํ™•๋„์™€ F1-score๋ฅผ ๊ธฐ๋กํ•˜์˜€์œผ๋ฉฐ, ํŠนํžˆ ์ž„์ƒ์  ์˜์‚ฌ๊ฒฐ์ •์— ์ค‘๋Œ€ํ•œ ์˜ํ–ฅ์„ ๋ฏธ์น˜๋Š” ์˜ค๋ฅ˜๋ฅผ ์œ ์˜๋ฏธํ•˜๊ฒŒ ๊ฐ์†Œ์‹œ์ผฐ์Šต๋‹ˆ๋‹ค. ๋ณธ ํ”„๋ ˆ์ž„์›Œํฌ๋Š” ๋†’์€ ์ •ํ™•๋„์™€ ํšจ์œจ์ ์ธ ์ถ”๋ก  ์„ฑ๋Šฅ์„ ๋ฐ”ํƒ•์œผ๋กœ ์ž„์ƒ ํ˜„์žฅ์—์„œ ํ‘œ์ค€ํ™”๋œ ๋ณ‘๊ธฐ ๊ฒฐ์ •์„ ์ง€์›ํ•˜๋Š” ์‹ค์šฉ์ ์ธ ๋„๊ตฌ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ•˜์˜€์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

4Automating the segmentation, date extraction, and classification of multi-report PDFs in outside medical records using optical character recognition and generativeย artificial intelligence.

2026-04JAMIA openโญ Q1DOI 10.1093/jamiaopen/ooag027
OBJECTIVE

Patients referred for specialized care often arrive with outside medical records (OMRs) compiled into multi-report PDFs that include imaging, pathology, and clinical notes in unstructured formats. Reviewing these records is time consuming and mentally taxing, increasing the risk of delayed care, clinician frustration, and missed information affecting quality of care. This study aimed to automate the segmentation, classification, and date extraction of scanned OMRs, with a focus on records relevant to breast cancer care.

METHODS

We used optical character recognition (OCR) to extract machine-readable text from 1303 scanned PDF documents from 116 distinct external institutions. Gemini 1.5, a large language model (LLM), was then used to segment multi-report files into individual documents, classify them into clinically meaningful categories such as mammograms and pathology reports, and extract study dates to build diagnostic timelines. Document categories were informed by clinical workflows in a breast cancer center.

RESULTS

The system achieved an F1 score of 0.95 for segmentation, 0.96 for classification, and 0.90 for date extraction. In a pilot of 45 records reviewed by clinicians, only 2 classification errors and 1 date error were reported. Clinicians estimated that the tool reduced OMR review time by 40%, improved workflow efficiency, and increased satisfaction.

CONCLUSION

Our findings demonstrate that combining OCR with LLMs can significantly enhance the processing of unstructured medical records, reducing manual burden and supporting timely clinical decision-making.

CONCLUSION

This study demonstrates the successful application of OCR and LLMs for organizing scanned OMRs within a specialty clinic. By automating a previously manual process, the approach supports scalable review of incoming outside records and has potential for adaptation to other clinical workflows. Future work will focus on evaluating the system across additional specialties and institutions.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ด‘ํ•™ ๋ฌธ์ž ์ธ์‹(OCR)๊ณผ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์„ ๊ฒฐํ•ฉํ•˜์—ฌ ์™ธ๋ถ€ ์˜๋ฃŒ ๊ธฐ๋ก(OMR) PDF์˜ ๋ถ„ํ• , ๋ถ„๋ฅ˜ ๋ฐ ๋‚ ์งœ ์ถ”์ถœ์„ ์ž๋™ํ™”ํ•˜๋Š” ์‹œ์Šคํ…œ์„ ๊ฐœ๋ฐœํ•˜๊ณ  ๊ทธ ์œ ํšจ์„ฑ์„ ํ‰๊ฐ€ํ–ˆ์Šต๋‹ˆ๋‹ค. ์œ ๋ฐฉ์•” ์ง„๋ฃŒ ํ™˜๊ฒฝ์—์„œ ํ…Œ์ŠคํŠธํ•œ ๊ฒฐ๊ณผ, ํ•ด๋‹น ์‹œ์Šคํ…œ์€ ๋†’์€ ์ •ํ™•๋„(F1 ์ ์ˆ˜ 0.90~0.96)๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ์ž„์ƒ์˜์˜ ๊ธฐ๋ก ๊ฒ€ํ†  ์‹œ๊ฐ„์„ 40% ๋‹จ์ถ•ํ•˜๊ณ  ์—…๋ฌด ํšจ์œจ์„ฑ์„ ์œ ์˜๋ฏธํ•˜๊ฒŒ ๊ฐœ์„ ํ–ˆ์Šต๋‹ˆ๋‹ค. ์ด ์ ‘๊ทผ๋ฒ•์€ ๋น„์ •ํ˜• ์˜๋ฃŒ ๊ธฐ๋ก์˜ ์ฒ˜๋ฆฌ๋ฅผ ์ž๋™ํ™”ํ•จ์œผ๋กœ์จ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ •์˜ ์‹ ์†์„ฑ์„ ๋†’์ด๊ณ  ์˜๋ฃŒ์ง„์˜ ์—…๋ฌด ๋ถ€๋‹ด์„ ๊ฒฝ๊ฐํ•˜๋Š” ๋ฐ ๊ธฐ์—ฌํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 03:13View โ†—

5AI-assisted tumor board decision-making in pancreatic oncology.

2026-03-20BMC medical informatics and decision makingโญ Q1DOI 10.1186/s12911-026-03444-x
BACKGROUND

Pancreatic cancer requires nuanced, multidisciplinary treatment planning typically conducted within tumor boards. While Large Language Models (LLMs) have shown capabilities in medical reasoning, their ability to approximate complex, integrative decision-making in oncology remains underexplored.

METHODS

This study evaluated the performance of LLaMA 3.3 (70b) in predicting tumor board decisions for newly diagnosed pancreatic cancer patients. Clinical documentation (including free-text imaging reports, pathology findings, and patient history) from 42 first-diagnosis cases discussed in a real-world tumor board was collected. The model was tasked with predicting one of three treatment options: surgical resection (SURG), neoadjuvant chemotherapy (NEO), or palliative therapy (PALL). Four prompting strategies were evaluated: zero-shot, advanced (adv.) zero-shot, Chain-of-Thought (CoT), and few-shot prompting. Performance was assessed using accuracy, micro- and macro-averaged F1 scores, and category-specific recall.

RESULTS

The advanced zero-shot and CoT strategies achieved the highest overall accuracy of 78.6% and a micro-averaged F1 score of 0.786. However, this performance was driven primarily by the correct classification of majority classes (SURG and PALL). Crucially, both high-accuracy strategies failed to identify any of the neoadjuvant therapy candidates (Recall NEOโ€‰=โ€‰0.00; 0/7 cases), systematically misclassifying them as palliative or surgical. While few-shot prompting improved the detection of neoadjuvant cases (Recall NEOโ€‰=โ€‰1.00), it introduced substantial noise, reducing overall accuracy to 56.7%. LLaMA 3.3 (70b) demonstrates high concordance with tumor board decisions for clear-cut surgical or palliative cases but exhibits a critical systematic failure in identifying candidates for neoadjuvant therapy. The high global accuracy masks a significant safety limitation regarding the recognition of complex, intermediate-stage patients.

CONCLUSION

These findings suggest that current LLMs may approximate majority-class decisions but risk overlooking curative treatment pathways in nuanced scenarios, necessitating rigorous oversight and specific adaptation before clinical consideration.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์ทŒ์žฅ์•” ํ™˜์ž์˜ ๋‹คํ•™์ œ ์ง„๋ฃŒ(Tumor Board) ๊ฒฐ์ • ์˜ˆ์ธก์— ์žˆ์–ด LLaMA 3.3 ๋ชจ๋ธ์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์œผ๋ฉฐ, ์ˆ˜์ˆ  ๋ฐ ์™„ํ™” ์น˜๋ฃŒ ๊ฒฐ์ •์—์„œ๋Š” ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์œผ๋‚˜ ์„ ํ–‰ ํ•ญ์•”ํ™”ํ•™์š”๋ฒ• ๋Œ€์ƒ์ž๋ฅผ ์‹๋ณ„ํ•˜๋Š” ๋ฐ์—๋Š” ์ฒด๊ณ„์ ์ธ ํ•œ๊ณ„๋ฅผ ๋“œ๋Ÿฌ๋ƒˆ์Šต๋‹ˆ๋‹ค. ์ „๋ฐ˜์ ์ธ ์ •ํ™•๋„๊ฐ€ ๋†’๋”๋ผ๋„ ๋ณต์žกํ•œ ์ค‘๊ฐ„ ๋‹จ๊ณ„ ํ™˜์ž์˜ ์น˜๋ฃŒ ๊ฒฝ๋กœ๋ฅผ ๋†“์น  ์œ„ํ—˜์ด ํฌ๋ฏ€๋กœ, ์ž„์ƒ ์ ์šฉ์„ ์œ„ํ•ด์„œ๋Š” ๋ชจ๋ธ์˜ ์ •๊ตํ•œ ๋ณด์™„๊ณผ ์—„๊ฒฉํ•œ ์ „๋ฌธ์˜ ๊ฐ๋…์ด ํ•„์ˆ˜์ ์ž…๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

6Automated detection of primary soft tissue sarcomas of the extremities using artificial intelligence and ChatGPT.

2026-03-16Frontiers in oncology๐Ÿ”ท Q2DOI 10.3389/fonc.2026.1674509
OBJECTIVE

Developing effective Convolutional Neural Networks (CNN) for soft tissue sarcoma detection often requires numerous iterations and adjustments, demanding specialized IT (Information Technology) skills. This study aims to use ChatGPT 4 to simplify CNN adaptation, reducing the need for specialized IT skills while enabling efficient exploration of training configurations to enhance diagnostic accuracy.

METHODS

This study leveraged a preexisting Artificial Intelligence (AI) model adapted using a preexisting Convolutional Neural Network (CNN). The study involved 54 participants diagnosed with primary soft tissue sarcomas in the extremities and possessing complete Magnetic Resonance Imaging (MRI) datasets. AI adaptations and programming were conducted using TensorFlow and verified with ChatGPT. Model training involved a dataset split of 70% training, 15% validation and 15% test set on patient level split, processed over eight epochs.

RESULTS

The adapted CNN model demonstrated significant improvement across various MRI sequences, achieving high accuracy levels (up to 98.5%) and excellent sensitivity and specificity rates. The model performed robustly in differentiating tumor presence in MR images, with test accuracies as high as 93.9%. The inclusion of a Gradient-weighted Class Activation Mapping (Grad-CAM) heat map and probability scores in the diagnostic outputs further enhanced interpretative capabilities.

CONCLUSION

This study highlights the potential of AI, particularly CNNs, in the early and accurate detection of soft tissue sarcomas, underscoring the technology's adaptability across different imaging modalities. The integration of large language models like ChatGPT into the model adaptation process emphasizes the reduced need for specialized IT skills, making advanced diagnostic tools more accessible and potentially improving diagnostic accuracy and patient outcomes in radiology and oncology.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ChatGPT๋ฅผ ํ™œ์šฉํ•˜์—ฌ ์ „๋ฌธ์ ์ธ IT ๊ธฐ์ˆ  ์—†์ด๋„ ํ•ฉ์„ฑ๊ณฑ ์‹ ๊ฒฝ๋ง(CNN) ๋ชจ๋ธ์„ ํšจ์œจ์ ์œผ๋กœ ์ตœ์ ํ™”ํ•˜๊ณ , ์‚ฌ์ง€ ์—ฐ๋ถ€ ์กฐ์ง ์œก์ข…์˜ MRI ์ง„๋‹จ ์ •ํ™•๋„๋ฅผ ๋†’์ด๋Š” ๋ฐฉ๋ฒ•์„ ์ œ์‹œํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ์ตœ์ ํ™”๋œ CNN ๋ชจ๋ธ์€ ์ตœ๋Œ€ 98.5%์˜ ๋†’์€ ์ •ํ™•๋„์™€ ์šฐ์ˆ˜ํ•œ ๋ฏผ๊ฐ๋„ ๋ฐ ํŠน์ด๋„๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, Grad-CAM์„ ํ†ตํ•œ ์‹œ๊ฐ์  ํ•ด์„ ๊ฐ€๋Šฅ์„ฑ๊นŒ์ง€ ํ™•๋ณดํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ด๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์„ ํ™œ์šฉํ•œ AI ๋ชจ๋ธ ๊ตฌ์ถ•์ด ์˜์ƒ์˜ํ•™ ๋ฐ ์ข…์–‘ํ•™ ๋ถ„์•ผ์—์„œ ์ง„๋‹จ ํšจ์œจ์„ฑ์„ ๋†’์ด๊ณ  ์ž„์ƒ์  ์ ‘๊ทผ์„ฑ์„ ๊ฐœ์„ ํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

7Automated Report Generation in Ophthalmology: Integrating Artificial Intelligence, Multimodal Imaging, and Clinical Data.

2026-02-19Ophthalmology and therapyโญ Q1DOI 10.1007/s40123-026-01316-1

Artificial intelligence (AI) has emerged as a transformative force in ophthalmology, enabling automated, accurate, and efficient clinical reporting. This review summarizes recent advances in AI-driven report generation, emphasizing the integration of multimodal imaging and clinical data. Deep learning and natural language processing (NLP) models can synthesize information from diverse sources-including fundus photography, optical coherence tomography, fluorescein angiography, and patient records-to generate structured, interpretable, and personalized diagnostic reports. Such systems enhance diagnostic precision, streamline workflow, and reduce interobserver variability. We outline the technological foundations underlying these systems, including convolutional and transformer-based architectures, self-supervised and multimodal learning, and large language models. Representative applications in diabetic retinopathy, glaucoma, cataract, and age-related macular degeneration are discussed, highlighting their clinical value and emerging real-world deployment. Persistent challenges-including data heterogeneity, model interpretability, ethical governance, and clinical integration-are critically reviewed. Finally, we explore future directions such as real-time AI-assisted reporting, predictive and personalized analytics, and global scalability across healthcare ecosystems. Multimodal, explainable, and clinically integrated AI systems hold promise to redefine ophthalmic diagnostics and improve both clinician efficiency and patient outcomes.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์•ˆ๊ณผ ๋ถ„์•ผ์—์„œ ๋”ฅ๋Ÿฌ๋‹๊ณผ ์ž์—ฐ์–ด ์ฒ˜๋ฆฌ ๊ธฐ์ˆ ์„ ํ™œ์šฉํ•ด ๋‹ค์ค‘ ๋ชจ๋‹ฌ ์˜์ƒ ๋ฐ ์ž„์ƒ ๋ฐ์ดํ„ฐ๋ฅผ ํ†ตํ•ฉํ•œ ์ž๋™ ์ง„๋‹จ ๋ณด๊ณ ์„œ ์ƒ์„ฑ ์‹œ์Šคํ…œ์˜ ์ตœ์‹  ๋™ํ–ฅ์„ ๊ณ ์ฐฐํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ด๋Ÿฌํ•œ AI ๊ธฐ๋ฐ˜ ์‹œ์Šคํ…œ์€ ์ง„๋‹จ์˜ ์ •ํ™•๋„๋ฅผ ๋†’์ด๊ณ  ๊ด€์ฐฐ์ž ๊ฐ„ ๋ณ€์ด๋ฅผ ์ค„์—ฌ ์ž„์ƒ ์›Œํฌํ”Œ๋กœ์šฐ๋ฅผ ํšจ์œจํ™”ํ•  ์ˆ˜ ์žˆ์Œ์„ ํ™•์ธํ–ˆ์Šต๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ์ž„์ƒ ํ˜„์žฅ์— ์„ฑ๊ณต์ ์œผ๋กœ ๋„์ž…๋˜๊ธฐ ์œ„ํ•ด์„œ๋Š” ๋ฐ์ดํ„ฐ ์ด์งˆ์„ฑ ํ•ด๊ฒฐ, ๋ชจ๋ธ์˜ ํ•ด์„ ๊ฐ€๋Šฅ์„ฑ ํ™•๋ณด ๋ฐ ์œค๋ฆฌ์  ๊ฑฐ๋ฒ„๋„Œ์Šค ๊ตฌ์ถ•์ด ์„ ํ–‰๋˜์–ด์•ผ ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

8Differences and Trends of Artificial Intelligence in Medical Education: A Comparative Bibliometric Analysis Between China and the International Community.

2026-01-31Advances in medical education and practice๐Ÿ”ท Q2DOI 10.2147/amep.s573537
OBJECTIVE

This study aims to explore the application of artificial intelligence in medical education by comparing research hotspots and evolutionary trends between China and the international community, ultimately proposing informed educational practices and policy recommendations.

METHODS

Literature was retrieved from the core collections of CNKI and Web of Science for the period 2014-2024, limited to article and review publications. After applying a unified Boolean search strategy and deduplication, the data were analyzed using CiteSpace 6.4.R1 to examine publication trends, collaboration networks, keyword co-occurrence/clustering/burst detection, and co-citation patterns.

RESULTS

A total of 379 Chinese and 552 English records were included. Publications surged after 2018 and peaked during 2023-2024. International hotspots centered on machine learning, deep learning, and large language models for simulation-based training and clinical reasoning; Chinese studies focused on "New Medical Sciences", VR/AR, and medical imaging. The emergence of generative artificial intelligence and multimodal large models has become a new frontier in artificial intelligence research within global medical education from 2023 to 2024.

CONCLUSION

This study is based on a comparison of two databases to reveal the hotspots and differences in artificial intelligence and medical education research between China and the international research community. It not only compensates for the time lag of existing research, but also proposes three major trends driven by artificial intelligence in the development of medical education (generative AI, personalized learning, immersive experience). A complementary pattern exists between technology-driven and scenario-driven orientations. We recommend integrating AI literacy and ethics into curricula, establishing Generative-AI teaching/assessment guidelines, and building cross-institutional, yearly knowledge-map monitoring for sustainable innovation in medical education.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 2014๋…„๋ถ€ํ„ฐ 2024๋…„๊นŒ์ง€์˜ ๋ฌธํ—Œ์„ ๋ฐ”ํƒ•์œผ๋กœ ์ค‘๊ตญ๊ณผ ๊ตญ์ œ ์‚ฌํšŒ์˜ ์˜ํ•™ ๊ต์œก ๋‚ด ์ธ๊ณต์ง€๋Šฅ(AI) ์—ฐ๊ตฌ ๋™ํ–ฅ์„ ๋น„๊ต ๋ถ„์„ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ๊ตญ์ œ ์—ฐ๊ตฌ๋Š” ๋จธ์‹ ๋Ÿฌ๋‹๊ณผ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ์„ ํ™œ์šฉํ•œ ์ž„์ƒ ์ถ”๋ก  ๊ต์œก์— ์ง‘์ค‘๋œ ๋ฐ˜๋ฉด, ์ค‘๊ตญ์€ ๊ฐ€์ƒํ˜„์‹ค(VR/AR) ๋ฐ ์˜๋ฃŒ ์˜์ƒ ๊ธฐ์ˆ ์— ์ค‘์ ์„ ๋‘๋Š” ์ฐจ์ด๋ฅผ ๋ณด์˜€์Šต๋‹ˆ๋‹ค. ํ–ฅํ›„ ์˜ํ•™ ๊ต์œก์˜ ํ˜์‹ ์„ ์œ„ํ•ด ์ƒ์„ฑํ˜• AI์˜ ๊ต์œก์  ํ™œ์šฉ ๊ฐ€์ด๋“œ๋ผ์ธ ์ˆ˜๋ฆฝ๊ณผ AI ๋ฆฌํ„ฐ๋Ÿฌ์‹œ ๋ฐ ์œค๋ฆฌ ๊ต์œก์˜ ํ†ตํ•ฉ์ด ํ•„์š”ํ•จ์„ ์ œ์–ธํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

9Radiology Board-Style Examinations and LLMs: A Scoping Review of Model Performance.

2026-01-29Journal of the American College of Radiology : JACRโญ Q1DOI 10.1016/j.jacr.2026.01.017
BACKGROUND

Large language models (LLMs) are increasingly being evaluated for their ability to answer official radiology board-style examination questions. Understanding their accuracy, limitations, and potential applications in education is essential for assessing their utility in the field.

METHODS

A scoping review was conducted in October 2025 across PubMed, Scopus, and Web of Science, following Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Studies were included if they evaluated LLMs on official radiology board-style examination questions. After screening 205 unique records, 29 studies met the inclusion criteria. Data were extracted on study characteristics, including LLM type and version, input modality, language, examination type, answer format, comparison with humans, and reported outcomes.

RESULTS

The reviewed studies evaluated multiple LLMs, predominantly Chat Generative Pre-trained Transformer (GPT)-based models (GPT-3.5, GPT-4, GPT-4 Turbo, GPT-4o), as well as Claude, Gemini, Llama 3, and Mixtral. Text-only evaluations generally yielded higher accuracy (โ‰ˆ65%-90%) compared with multimodal tasks (45%-89%). GPT-4 and its variants consistently outperformed earlier versions, occasionally exceeding average human performance. Open-source models such as Llama 3 70B and Mixtral achieved comparable results to proprietary models, offering advantages in local deployment and privacy. Few studies directly compared LLM performance with human radiologists.

CONCLUSION

LLMs demonstrate promising performance in answering text-based radiology board-style examination questions, particularly GPT-4-based models. Nevertheless, significant limitations persist in multimodal tasks and complex reasoning scenarios.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜์ƒ์˜ํ•™ ์ „๋ฌธ์˜ ์‹œํ—˜ ๋ฌธํ•ญ์„ ํ™œ์šฉํ•˜์—ฌ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ๋ถ„์„ํ•œ ์Šค์ฝ”ํ•‘ ๋ฆฌ๋ทฐ๋กœ, 29๊ฐœ ์—ฐ๊ตฌ๋ฅผ ์ฒด๊ณ„์ ์œผ๋กœ ๊ฒ€ํ† ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, GPT-4 ๊ณ„์—ด ๋ชจ๋ธ์€ ํ…์ŠคํŠธ ๊ธฐ๋ฐ˜ ๋ฌธ์ œ์—์„œ ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋ณด์ด๋ฉฐ ์ธ๊ฐ„์˜ ํ‰๊ท  ์ ์ˆ˜๋ฅผ ์ƒํšŒํ•˜๊ธฐ๋„ ํ–ˆ์œผ๋‚˜, ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ ๊ณผ์ œ ๋ฐ ๋ณตํ•ฉ์  ์ถ”๋ก  ์ƒํ™ฉ์—์„œ๋Š” ์—ฌ์ „ํžˆ ์œ ์˜๋ฏธํ•œ ํ•œ๊ณ„๋ฅผ ๋‚˜ํƒ€๋ƒˆ์Šต๋‹ˆ๋‹ค. ํ–ฅํ›„ LLM์˜ ๊ต์œก์  ํ™œ์šฉ์„ ์œ„ํ•ด์„œ๋Š” ์ด๋Ÿฌํ•œ ๊ธฐ์ˆ ์  ์ œ์•ฝ์„ ๊ทน๋ณตํ•˜๊ณ  ์ž„์ƒ์  ํŒ๋‹จ ๋Šฅ๋ ฅ์„ ๋ณด์™„ํ•˜๋Š” ์ถ”๊ฐ€์ ์ธ ๊ฒ€์ฆ์ด ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

10Comparing the performance of radiomics, nomograms, machine learning, and large language models in predicting 28-day mortality in severe community-acquired pneumonia patients.

2026-01-19Frontiers in immunologyโญ Q1DOI 10.3389/fimmu.2025.1679496
BACKGROUND

Severe community-acquired pneumonia (SCAP) is a significant global health challenge due to its high mortality. Despite advances, early diagnosis and effective management remain critical. Tools like radiomics analyze imaging data for risk assessment, while machine learning and nomograms aid in personalized treatment. Large language models (LLMs) enhance clinical decision-making by analyzing data and supporting care strategies. This study integrates these methods to predict 28-day mortality in SCAP patients.

METHODS

A cohort of 599 patients diagnosed with severe community-acquired pneumonia (SCAP), including 316 males and 283 females, from Shanghai East Hospital and Xiamen Humanity Hospital were enrolled in this study. High-resolution lung CT scans were used to segment three-dimensional regions of interest, from which 1,050 radiomic features were extracted. The dataset was divided into a training set (80%) and an independent test set (20%), and k-fold cross-validation was applied to optimize model performance. To address class imbalance, the SMOTE oversampling technique was employed. The study integrated radiomics, nomograms, seven machine learning models, and five LLMs to predict the 28-day mortality risk in SCAP patients. SHAP values were utilized to enhance the interpretability of feature contributions. Not only that, this study integrates the prior knowledge provided by LLMs, processed through an embedding layer, with data-driven feature learning in the main network, and dynamically fuses their outputs using a bias network with a gating mechanism, thereby improving the accuracy and interpretability of LLMs in predicting 28-day mortality risk for SCAP patients.

RESULTS

Key predictors of 28-day mortality included inflammatory markers, cytokines, age, CRP, and oxygenation index. Clinical-Radiomics models achieved strong accuracy (AUC 0.92). Machine learning models, particularly XGBoost (AUC 0.90), were highly effective, with SHAP analysis emphasizing radscore's importance. LLMs like Chatgpt also performed well (AUC 0.78), showcasing the potential of integrating clinical, radiomic, and AI-driven approaches.

CONCLUSION

This study demonstrates the effectiveness of radiomics, machine learning, and LLMs to predict SCAP outcomes. Models like XGBoost achieved superior accuracy, while SHAP analysis improved interpretability. These advancements highlight the potential for enhanced SCAP prognosis and personalized care strategies.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์ค‘์ฆ ์ง€์—ญ์‚ฌํšŒ ํš๋“ ํ๋ ด(SCAP) ํ™˜์ž 599๋ช…์„ ๋Œ€์ƒ์œผ๋กœ ๋ฐฉ์‚ฌ์„  ํŠน์ง•, ๊ธฐ๊ณ„ ํ•™์Šต, ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์„ ํ†ตํ•ฉํ•˜์—ฌ 28์ผ ์‚ฌ๋ง๋ฅ ์„ ์˜ˆ์ธกํ•˜๋Š” ๋ชจ๋ธ์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ž„์ƒ ์ •๋ณด์™€ ๋ฐฉ์‚ฌ์„  ํŠน์ง•์„ ๊ฒฐํ•ฉํ•œ ๋ชจ๋ธ์ด 0.92์˜ ๋†’์€ AUC๋ฅผ ๊ธฐ๋กํ•˜์˜€์œผ๋ฉฐ, XGBoost์™€ ๊ฐ™์€ ๊ธฐ๊ณ„ ํ•™์Šต ๋ชจ๋ธ์ด ์šฐ์ˆ˜ํ•œ ์˜ˆ์ธก๋ ฅ์„ ๋ณด์˜€์Šต๋‹ˆ๋‹ค. ํŠนํžˆ LLM์˜ ์‚ฌ์ „ ์ง€์‹๊ณผ ๋ฐ์ดํ„ฐ ๊ธฐ๋ฐ˜ ํ•™์Šต์„ ์œตํ•ฉํ•œ ๋ฐฉ์‹์€ ์˜ˆ์ธก ์ •ํ™•๋„์™€ ํ•ด์„ ๊ฐ€๋Šฅ์„ฑ์„ ๋™์‹œ์— ํ–ฅ์ƒ์‹œ์ผœ ํ–ฅํ›„ SCAP ํ™˜์ž์˜ ๋งž์ถคํ˜• ์น˜๋ฃŒ ์ „๋žต ์ˆ˜๋ฆฝ์— ๊ธฐ์—ฌํ•  ๊ฒƒ์œผ๋กœ ๊ธฐ๋Œ€๋ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

11Artificial intelligence and machine learning-driven advancements in gastrointestinal cancer: Paving the way for precision medicine.

2026-01-07World journal of gastroenterologyโญ Q1DOI 10.3748/wjg.v32.i1.111428

Gastrointestinal (GI) cancers remain a leading cause of cancer-related morbidity and mortality worldwide. Artificial intelligence (AI), particularly machine learning and deep learning (DL), has shown promise in enhancing cancer detection, diagnosis, and prognostication. A narrative review of literature published from January 2015 to march 2025 was conducted using PubMed, Web of Science, and Scopus. Search terms included "gastrointestinal cancer", "artificial intelligence", "machine learning", "deep learning", "radiomics", "multimodal detection" and "predictive modeling". Studies were included if they focused on clinically relevant AI applications in GI oncology. AI algorithms for GI cancer detection have achieved high performance across imaging modalities, with endoscopic DL systems reporting accuracies of 85%-97% for polyp detection and segmentation. Radiomics-based models have predicted molecular biomarkers such as programmed cell death ligand 2 expression with area under the curves up to 0.92. Large language models applied to radiology reports demonstrated diagnostic accuracy comparable to junior radiologists (78.9% vs 80.0%), though without incremental value when combined with human interpretation. Multimodal AI approaches integrating imaging, pathology, and clinical data show emerging potential for precision oncology. AI in GI oncology has reached clinically relevant accuracy levels in multiple diagnostic tasks, with multimodal approaches and predictive biomarker modeling offering new opportunities for personalized care. However, broader validation, integration into clinical workflows, and attention to ethical, legal, and social implications remain critical for widespread adoption.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์†Œํ™”๊ธฐ ์•” ๋ถ„์•ผ์—์„œ ์ธ๊ณต์ง€๋Šฅ(AI) ๋ฐ ๋จธ์‹ ๋Ÿฌ๋‹ ๊ธฐ์ˆ ์ด ์ง„๋‹จ, ์˜ˆํ›„ ์˜ˆ์ธก, ์ •๋ฐ€ ์˜๋ฃŒ์— ๊ธฐ์—ฌํ•˜๋Š” ์ตœ์‹  ๋™ํ–ฅ์„ ๊ณ ์ฐฐํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋‚ด์‹œ๊ฒฝ ์˜์ƒ ๋ถ„์„๊ณผ ๋ฐฉ์‚ฌ์„ ํ•™์  ๋ชจ๋ธ์„ ํ†ตํ•ด ๋†’์€ ์ง„๋‹จ ์ •ํ™•๋„๋ฅผ ํ™•๋ณดํ•˜๊ณ  ๋ถ„์ž ์ƒ๋ฌผํ•™์  ์ง€ํ‘œ๋ฅผ ์˜ˆ์ธกํ•˜๋Š” ์„ฑ๊ณผ๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ํ–ฅํ›„ ์ž„์ƒ ํ˜„์žฅ ๋„์ž…์„ ์œ„ํ•ด์„œ๋Š” ๋‹คํ•™์ œ์  ๋ฐ์ดํ„ฐ ํ†ตํ•ฉ๊ณผ ์œค๋ฆฌ์ ยท๋ฒ•์  ๊ฒ€ํ† ๊ฐ€ ํ•„์ˆ˜์ ์ž„์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

12Evaluating Generative Artificial Intelligence as an Educational Tool for Radiology Resident Report Drafting.

2025-12-24Journal of the American College of Radiology : JACRโญ Q1DOI 10.1016/j.jacr.2025.12.024
OBJECTIVE

Radiology residents require timely, personalized feedback to develop accurate image analysis and reporting skills. Increasing clinical workload often limits attendings' ability to provide guidance. This study evaluates a HIPAA-compliant Generative Pretrained Transformer (GPT)-4o system that delivers automated feedback on breast imaging reports drafted by residents in real clinical settings.

METHODS

We analyzed 5,000 resident-attending report pairs from routine practice at a multisite US health system. GPT-4o was prompted with clinical instructions to identify common errors and provide feedback. A reader study using 100 report pairs was conducted. Four attending radiologists and four residents independently reviewed each pair, determined whether predefined error types were present, and rated GPT-4o's feedback as helpful or not. Agreement between GPT and readers was assessed using percent match. Interreader reliability was measured with Krippendorff's ฮฑ. Educational value was measured as the proportion of cases rated helpful.

RESULTS

Three common error types were identified: (1) omission or addition of key findings, (2) incorrect use or omission of technical descriptors, and (3) final assessment inconsistent with findings. GPT-4o showed strong agreement with attending consensus: 90.5%, 78.3%, and 90.4% (Cohen's ฮบ: 0.790, 0.550, and 0.615) across error types. Interreader reliability among all eight readers showed moderate to substantial variability (ฮฑ = 0.767, 0.595, 0.567). When each reader was individually replaced with GPT-4o and interreader agreement among seven readers and GPT was recalculated, the effect was not statistically significant (ฮ” = -0.004 to 0.002, all P > .05). GPT's feedback was rated helpful in most cases: 89.8%, 83.0%, and 92.0%.

CONCLUSION

ChatGPT-4o can reliably identify key educational errors. It may serve as a scalable tool to support radiology education.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜์ƒ์˜ํ•™๊ณผ ์ „๊ณต์˜์˜ ํŒ๋…๋ฌธ ์ž‘์„ฑ์„ ๋•๊ธฐ ์œ„ํ•ด HIPAA๋ฅผ ์ค€์ˆ˜ํ•˜๋Š” GPT-4o ๊ธฐ๋ฐ˜์˜ ์ž๋™ ํ”ผ๋“œ๋ฐฑ ์‹œ์Šคํ…œ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. 5,000๊ฑด์˜ ํŒ๋… ๋ฐ์ดํ„ฐ๋ฅผ ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, GPT-4o๋Š” ์ฃผ์š” ์˜ค๋ฅ˜ ์œ ํ˜•์„ ์ „๋ฌธ์˜ ์ˆ˜์ค€์œผ๋กœ ์ •ํ™•ํ•˜๊ฒŒ ์‹๋ณ„ํ•˜์˜€์œผ๋ฉฐ, 80% ์ด์ƒ์˜ ๋†’์€ ๋น„์œจ๋กœ ์ „๊ณต์˜๋“ค์—๊ฒŒ ์œ ์šฉํ•œ ๊ต์œก์  ํ”ผ๋“œ๋ฐฑ์„ ์ œ๊ณตํ•˜๋Š” ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค. ์ด๋Š” ์ƒ์„ฑํ˜• AI๊ฐ€ ์ž„์ƒ ํ˜„์žฅ์—์„œ ์ „๊ณต์˜ ๊ต์œก์„ ๋ณด์™„ํ•˜๊ณ  ํŒ๋… ์—ญ๋Ÿ‰์„ ๊ฐ•ํ™”ํ•˜๋Š” ํ™•์žฅ ๊ฐ€๋Šฅํ•œ ๋„๊ตฌ๋กœ ํ™œ์šฉ๋  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

13Is It Getting Better? An Evaluation of Two Successive Generations of ChatGPT in Answering Specialized Vascular Surgery Questions.

2025-12-10Vascular specialist internationalDOI 10.5758/vsi.250071
OBJECTIVE

Large language models (LLMs) can generate clinically relevant text; however, their performance in highly specialized medical domains remains uncertain. This study evaluated ChatGPT-3.5 and ChatGPT-4 (OpenAI) using vascular surgery board-style questions from the Vascular Education and Self-Assessment Program, version 4 (VESAP4) and compared the two public model versions (June and November 2023).

METHODS

All non-image VESAP4 questions (n=384) were presented independently three times to each model version (ChatGPT-3.5 June/November; ChatGPT-4, June/November). Outcomes included accuracy (proportion correct), consistency (same option letter across all three attempts and "consistently correct"), explanation length (word count), and modes of failure classified for a pre-specified index attempt using a multi-label taxonomy, with independent dual review and consensus. Accuracy and consistency were reported with 95% confidence intervals, and between-condition differences were compared using the chi-square (proportions) and Welch t-test (word count).

RESULTS

Accuracy was 47.9% and 46.5% for ChatGPT-3.5 (June/November) and 62.4% and 63.8% for ChatGPT-4, respectively. No significant within-model improvement occurred between June and November, whereas ChatGPT-4 outperformed ChatGPT-3.5 in both months (P<0.0001). For consistency, ChatGPT-3.5 increased in "same-letter" (55.5% to 65.6%; P=0.004) with no change in "consistently correct" (40.6% to 40.4%; P=0.94). ChatGPT-4 decreased in "same-letter" (90.1% to 79.7%; P<0.0001) with stable "consistently correct" (60.4% to 58.6%; P=0.61). Between the models, ChatGPT-4 exceeded ChatGPT-3.5 on both consistency metrics in June and November (all P<0.0001). Explanation length shifted within models: ChatGPT-3.5 produced shorter responses in November compared with June (105.4ยฑ1.0 vs. 164.2ยฑ1.5 words; P<0.0001), whereas ChatGPT-4 produced longer responses (282.0ยฑ1.3 vs. 120.0ยฑ1.2 words; P<0.0001). Performance across VESAP4 subsections was higher for Vascular Medicine and Radiological Imaging/Radiation Safety, and lower for Dialysis Access Management. The modes of failure were predominantly due to external information retrieval errors for ChatGPT-3.5 and logical errors for ChatGPT-4.

CONCLUSION

ChatGPT-4 outperformed ChatGPT-3.5 on vascular surgery board-style questions yet achieved only moderate accuracy. Version updates did not consistently improve the performance in this specialized domain, emphasizing the complexity of decision-making in vascular surgery and the current limitations of LLMs in surgical education.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ํ˜ˆ๊ด€์™ธ๊ณผ ์ „๋ฌธ์˜ ์‹œํ—˜ ๋ฌธํ•ญ(VESAP4)์„ ํ™œ์šฉํ•˜์—ฌ ChatGPT-3.5์™€ ChatGPT-4์˜ ์„ฑ๋Šฅ์„ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ChatGPT-4๊ฐ€ ChatGPT-3.5๋ณด๋‹ค ๋†’์€ ์ •ํ™•๋„์™€ ์ผ๊ด€์„ฑ์„ ๋ณด์˜€์œผ๋‚˜, ๋‘ ๋ชจ๋ธ ๋ชจ๋‘ ์ „๋ฌธ์ ์ธ ํ˜ˆ๊ด€์™ธ๊ณผ ์˜์—ญ์—์„œ ์ž„์ƒ์  ์˜์‚ฌ๊ฒฐ์ •์„ ๋Œ€์ฒดํ•˜๊ธฐ์—๋Š” ์—ฌ์ „ํžˆ ํ•œ๊ณ„๊ฐ€ ์žˆ๋Š” ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค. ๋˜ํ•œ ๋ชจ๋ธ์˜ ๋ฒ„์ „ ์—…๋ฐ์ดํŠธ๊ฐ€ ๋ฐ˜๋“œ์‹œ ์„ฑ๋Šฅ ํ–ฅ์ƒ์œผ๋กœ ์ด์–ด์ง€์ง€๋Š” ์•Š์•„, ํŠน์ˆ˜ ์˜๋ฃŒ ๋ถ„์•ผ์—์„œ์˜ LLM ํ™œ์šฉ์— ์‹ ์ค‘ํ•œ ์ ‘๊ทผ์ด ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

14Evaluating the role of large language models in supporting patient education during the informed consent process for routine radiology procedures.

2025-12-01The British journal of radiologyโญ Q1DOI 10.1093/bjr/tqaf225
OBJECTIVE

This study evaluated 3 LLM chatbots (GPT-3.5-turbo, GPT-4-turbo, and GPT-4o) on their effectiveness in supporting patient education by answering common patient questions for CT, MRI, and DSA informed consent, assessing their accuracy and clarity.

METHODS

Two radiologists formulated 90 questions categorized as general, clinical, or technical. Each LLM answered every question 5ร—. Radiologists then rated the responses for medical accuracy and clarity, while medical physicists assessed technical accuracy using a Likert scale. Semantic similarity was analyzed with SBERT and cosine similarity.

RESULTS

Ratings improved with newer model versions. Linear mixed-effects models revealed that GPT-4 models were rated significantly higher than GPT-3.5 (Pโ€‰<โ€‰.001) by both physicians and physicists. However, physicians' ratings for GPT-4 models showed a significant performance decrease for complex modalities like DSA and MRI (Pโ€‰<โ€‰.01), a pattern not observed in physicists' ratings. SBERT analysis revealed high internal consistency across all models. SBERT analysis revealed high internal consistency across all models.

CONCLUSION

Variability in ratings revealed that while models effectively handled general and technical questions, they struggled with contextually complex medical inquiries requiring personalized responses and nuanced understanding. Statistical analysis confirms that while newer models are superior, their performance is modality-dependent and perceived differently by clinical and technical experts. ADVANCES IN KNOWLEDGE: This study evaluates the potential of LLMs to enhance informed consent in radiology, highlighting strengths in general and technical questions while noting limitations with complex clinical inquiries, with performance varying significantly by model type and imaging modality.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜์ƒ์˜ํ•™ ๊ฒ€์‚ฌ ๋™์˜ ๊ณผ์ •์—์„œ ํ™˜์ž ๊ต์œก์„ ์ง€์›ํ•˜๊ธฐ ์œ„ํ•œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(GPT-3.5, GPT-4, GPT-4o)์˜ ์œ ์šฉ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ์ตœ์‹  ๋ชจ๋ธ์ผ์ˆ˜๋ก ์˜ํ•™์  ์ •ํ™•๋„์™€ ๋ช…ํ™•์„ฑ ์ธก๋ฉด์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋‚˜, ๋ณต์žกํ•œ ์ž„์ƒ์  ๋งฅ๋ฝ์ด ์š”๊ตฌ๋˜๋Š” ๊ฒ€์‚ฌ์—์„œ๋Š” ์—ฌ์ „ํžˆ ํ•œ๊ณ„๊ฐ€ ํ™•์ธ๋˜์—ˆ์Šต๋‹ˆ๋‹ค. ๋”ฐ๋ผ์„œ LLM์€ ์ผ๋ฐ˜์ ยท๊ธฐ์ˆ ์  ์งˆ๋ฌธ์— ํšจ๊ณผ์ ์ธ ๋ณด์กฐ ๋„๊ตฌ๋กœ ํ™œ์šฉ๋  ์ˆ˜ ์žˆ์œผ๋‚˜, ์ž„์ƒ์  ํŒ๋‹จ์ด ํ•„์š”ํ•œ ์˜์—ญ์—์„œ๋Š” ์ฃผ์˜ ๊นŠ์€ ๊ฒ€ํ† ๊ฐ€ ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

15FHIR-Former: enhancing clinical predictions through Fast Healthcare Interoperability Resources and large language models.

2025-12-01Journal of the American Medical Informatics Association : JAMIAโญ Q1DOI 10.1093/jamia/ocaf165
OBJECTIVE

To address the challenges of data heterogeneity and manual feature engineering in clinical predictive modeling, we introduce FHIR-Former, an open-source framework integrating Fast Healthcare Interoperability Resources (FHIR) with large language models (LLMs) to automate and standardize clinical prediction tasks.

METHODS

FHIR-Former dynamically processes structured (eg, lab results, medications) and unstructured (eg, clinical notes) data from FHIR resources. The pipeline supports multiple classification tasks, including 30-day readmission, imaging study prediction, and ICD code classification. Leveraging open-source LLMs (GeBERTa), we trained models on 1.1 million data points across ten FHIR resources using retrospective inpatient data (2018-2024). Hyperparameters were optimized via Bayesian methods, and outputs were mapped to FHIR RiskAssessment resources for interoperability.

RESULTS

FHIR-Former achieved an F1-score of 70.7% and accuracy of 72.9% for 30-day readmission, 51.8% F1-score (88.1% accuracy) for mortality prediction, and 61% macro F1-score for imaging study classification. The ICD code prediction model attained 94% accuracy. Performance demonstrated promising performance for readmission and showed scalability across tasks without manual feature engineering.

CONCLUSION

FHIR-Former eliminates institution-specific preprocessing by adapting to diverse FHIR implementations, enabling seamless integration of multimodal data. Its configurable architecture outperformed prior frameworks reliant on static inputs or limited to unstructured text. Real-time risk scores embedded in FHIR servers enhance clinical workflows without disrupting existing practices.

CONCLUSION

By harmonizing FHIR standardization with LLM flexibility, FHIR-Former advances scalable, interoperable predictive modeling in healthcare. The open-source framework facilitates automation, improves resource allocation, and supports personalized decision-making, bridging gaps between AI innovation and clinical practice.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
FHIR-Former๋Š” FHIR ํ‘œ์ค€ ๋ฐ์ดํ„ฐ์™€ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ๊ฒฐํ•ฉํ•˜์—ฌ ์ž„์ƒ ๋ฐ์ดํ„ฐ์˜ ์ด์งˆ์„ฑ์„ ๊ทน๋ณตํ•˜๊ณ  ์ˆ˜๋™ ํŠน์ง• ๊ณตํ•™ ์—†์ด๋„ ์˜ˆ์ธก ๋ชจ๋ธ๋ง์„ ์ž๋™ํ™”ํ•˜๋Š” ํ”„๋ ˆ์ž„์›Œํฌ์ž…๋‹ˆ๋‹ค. 110๋งŒ ๊ฑด์˜ ๋ฐ์ดํ„ฐ๋ฅผ ํ•™์Šตํ•œ ๊ฒฐ๊ณผ, 30์ผ ์žฌ์ž…์› ๋ฐ ์‚ฌ๋ง๋ฅ  ์˜ˆ์ธก ๋“ฑ ๋‹ค์–‘ํ•œ ์ž„์ƒ ๊ณผ์ œ์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ์ž…์ฆํ•˜๋ฉฐ FHIR ๊ธฐ๋ฐ˜์˜ ์ƒํ˜ธ์šด์šฉ ๊ฐ€๋Šฅํ•œ ์‹ค์‹œ๊ฐ„ ์œ„ํ—˜๋„ ํ‰๊ฐ€๋ฅผ ์ง€์›ํ•ฉ๋‹ˆ๋‹ค. ๋ณธ ์—ฐ๊ตฌ๋Š” ํ‘œ์ค€ํ™”๋œ ๋ฐ์ดํ„ฐ ๊ตฌ์กฐ์™€ LLM์˜ ์œ ์—ฐ์„ฑ์„ ํ†ตํ•ฉํ•จ์œผ๋กœ์จ ํ™•์žฅ ๊ฐ€๋Šฅํ•˜๊ณ  ์ž„์ƒ ํ˜„์žฅ์— ์ฆ‰์‹œ ์ ์šฉ ๊ฐ€๋Šฅํ•œ AI ์˜ˆ์ธก ๋ชจ๋ธ๋ง์˜ ์ƒˆ๋กœ์šด ๋ฐฉํ–ฅ์„ ์ œ์‹œํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

16Evaluation of AI models for radiology exam preparation: DeepSeek vs. ChatGPT-3.5.

2025-11-28Medical education onlineโญ Q1DOI 10.1080/10872981.2025.2589679

The rapid advancement of artificial intelligence (AI) chatbots has generated significant interest regarding their potential applications within medical education. This study sought to assess the performance of the open-source large language model DeepSeek-V3 in answering radiology board-style questions and to compare its accuracy with that of ChatGPT-3.5.A total of 161 questions (comprising 207 items) were randomly selected from the Exercise Book for the National Senior Health Professional Qualification Examination: Radiology. The question set included single-choice, multiple-choice, shared-stem, and case analysis questions. Both DeepSeek-V3 and ChatGPT-3.5 were evaluated using the same question set over a seven-day testing period. Response accuracy was systematically assessed, and statistical analyses were performed using Pearson's chi-square test and Fisher's exact test.DeepSeek-V3 achieved an overall accuracy of 72%, which was significantly higher than the 55.6% accuracy achieved by ChatGPT-3.5 (Pโ€‰<โ€‰0.001). Performance analysis by question type revealed DeepSeek's superior accuracy in single-choice questions (87.1%), though with comparatively lower performance in multiple-choice (55.7%) and case analysis questions (68.0%). Across clinical subspecialties, DeepSeek consistently outperformed ChatGPT, particularly in peripheral nervous system (Pโ€‰=โ€‰0.003), respiratory system (Pโ€‰=โ€‰0.008), circulatory system (Pโ€‰=โ€‰0.012), and musculoskeletal system (Pโ€‰=โ€‰0.021) domains.In conclusion, DeepSeek demonstrates considerable potential as an educational tool in radiology, particularly for knowledge recall and foundational learning applications. However, its relatively weaker performance on higher-order cognitive tasks and complex question formats suggests the need for further model refinement. Future research should investigate DeepSeek's capability in processing image-based questions and perform comparative analyses with more advanced models (e.g., GPT-5) to better evaluate its potential for medical education.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„ ๊ณผ ์ „๋ฌธ์˜ ์ž๊ฒฉ ์‹œํ—˜ ๋ฌธํ•ญ์„ ํ™œ์šฉํ•˜์—ฌ DeepSeek-V3์™€ ChatGPT-3.5์˜ ํ•™์ˆ ์  ์„ฑ๋Šฅ์„ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, DeepSeek-V3๋Š” 72%์˜ ์ •ํ™•๋„๋กœ ChatGPT-3.5(55.6%)๋ณด๋‹ค ์œ ์˜๋ฏธํ•˜๊ฒŒ ๋†’์€ ์„ฑ๊ณผ๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ํŠนํžˆ ๊ธฐ์ดˆ ์ง€์‹ ํšŒ์ƒ ๋ฐ ํŠน์ • ์ž„์ƒ ๋ถ„์•ผ์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋‚˜ํƒ€๋ƒˆ์Šต๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ๋ณตํ•ฉ์ ์ธ ์‚ฌ๊ณ ๋ฅผ ์š”ํ•˜๋Š” ๋ฌธํ•ญ์—์„œ๋Š” ์„ฑ๋Šฅ ์ œํ•œ์ด ํ™•์ธ๋˜์–ด ํ–ฅํ›„ ์˜ํ•™ ๊ต์œก ๋„๊ตฌ๋กœ์„œ์˜ ํ™œ์šฉ์„ ์œ„ํ•œ ์ถ”๊ฐ€์ ์ธ ๋ชจ๋ธ ๊ณ ๋„ํ™”๊ฐ€ ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

17Artificial intelligence in polycystic ovary syndrome: a systematic review of diagnostic and predictive applications.

2025-11-24BMC medical informatics and decision makingโญ Q1DOI 10.1186/s12911-025-03255-6
BACKGROUND

Polycystic ovary syndrome (PCOS) is one of the most common endocrine disorders, affecting 8โ€“13% of women of reproductive age. Its heterogeneous presentation and the variability of diagnostic criteria make accurate diagnosis and effective management challenging. Artificial intelligence (AI) methods, including machine learning (ML), deep learning (DL), explainable AI (XAI), and large language models (LLMs), have recently emerged as promising approaches to address these gaps.

OBJECTIVE

This systematic review aimed to provide a comprehensive synthesis of AI applications in PCOS, with emphasis on diagnostic performance, biomarker discovery, risk prediction, clinical decision support, model interpretability, and the emerging use of generative AI.

METHODS

Following PRISMA 2020 guidelines, PubMed, Scopus, and Web of Science were searched from inception to March 2025. Eligible studies applied AI techniques to PCOS and reported at least one performance metric. Two reviewers independently screened and extracted data, with quality appraisal conducted using QUADAS-2 and ROBIS. Given the heterogeneity of designs and outcomes, findings were narratively synthesized across imaging, clinical/EHR, and biomarker/-omics domains.

RESULTS

From 662 retrieved records, 80 studies met the inclusion criteria. CNN-based models dominated imaging applications, with accuracies often exceeding 95% and occasionally reaching 98โ€“99%. Supervised ML approaches, particularly random forests and support vector machines, achieved consistent high performance in clinical and biochemical datasets. Omics-based studies revealed novel biomarkers such as HDDC3, SDC2, MAP1LC3A, and OVGP1. However, only about one-quarter of studies applied XAI methods, limiting transparency and clinical trust. Early evaluations of LLMs suggested potential for patient education and decision support but highlighted risks of bias, hallucination, and lack of domain-specific training. Key limitations across studies included small sample sizes, class imbalance, methodological heterogeneity, and limited external validation.

CONCLUSION

AI offers substantial opportunities to advance PCOS diagnosis and prediction by integrating multimodal data and reducing diagnostic subjectivity. Yet its clinical adoption is constrained by interpretability gaps and insufficient validation. Future priorities include large multicenter studies, standardized reporting, systematic use of XAI, and careful evaluation of LLMs to ensure safe, equitable, and clinically meaningful integration into PCOS care. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1186/s12911-025-03255-6.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์ฒด๊ณ„์  ๋ฌธํ—Œ๊ณ ์ฐฐ์€ ๋‹ค๋‚ญ์„ฑ ๋‚œ์†Œ ์ฆํ›„๊ตฐ(PCOS) ์ง„๋‹จ ๋ฐ ์˜ˆ์ธก์„ ์œ„ํ•œ ์ธ๊ณต์ง€๋Šฅ(AI) ๊ธฐ์ˆ ์˜ ํ™œ์šฉ ํ˜„ํ™ฉ์„ ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, ์˜์ƒ ๋ถ„์„๊ณผ ๋จธ์‹ ๋Ÿฌ๋‹ ๋ชจ๋ธ์ด ๋†’์€ ์ง„๋‹จ ์ •ํ™•๋„๋ฅผ ๋ณด์ด๋ฉฐ ์ƒˆ๋กœ์šด ๋ฐ”์ด์˜ค๋งˆ์ปค ๋ฐœ๊ตด์—๋„ ๊ธฐ์—ฌํ•˜๊ณ  ์žˆ์Œ์„ ํ™•์ธํ–ˆ์Šต๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ํ˜„์žฌ ์—ฐ๊ตฌ๋“ค์€ ์„ค๋ช… ๊ฐ€๋Šฅํ•œ AI(XAI) ์ ์šฉ ๋ถ€์กฑ๊ณผ ์™ธ๋ถ€ ๊ฒ€์ฆ ๋ฏธํก ๋“ฑ์˜ ํ•œ๊ณ„๋ฅผ ๋ณด์ด๊ณ  ์žˆ์–ด, ํ–ฅํ›„ ์ž„์ƒ ํ˜„์žฅ ๋„์ž…์„ ์œ„ํ•ด์„œ๋Š” ๋Œ€๊ทœ๋ชจ ๋‹ค๊ธฐ๊ด€ ์—ฐ๊ตฌ์™€ ํ‘œ์ค€ํ™”๋œ ๊ฒ€์ฆ ์ฒด๊ณ„ ๋งˆ๋ จ์ด ํ•„์ˆ˜์ ์ž…๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

18Clinical applications of large language models in knee osteoarthritis: a systematic review.

2025-11-19Frontiers in medicineโญ Q1DOI 10.3389/fmed.2025.1670824

BACKGROUND AND

OBJECTIVE

Knee osteoarthritis (KOA) is a common chronic degenerative disease that significantly impacts patients' quality of life. With the rapid advancement of artificial intelligence, large language models (LLMs) have demonstrated potential in supporting medical information extraction, clinical decision-making, and patient education through their natural language processing capabilities. However, the current landscape of LLM applications in the KOA domain, along with their methodological quality, has yet to be systematically reviewed. Therefore, this systematic review aims to comprehensively summarize existing clinical studies on LLMs in KOA, evaluate their performance and methodological rigor, and identify current challenges and future research directions.

METHODS

Following the PRISMA guidelines, a systematic search was conducted in PubMed, Cochrane Library, Embase databases and Web of science for literature published up to June 2025. The protocol was preregistered on the OSF platform. Studies were screened using standardized inclusion and exclusion criteria. Key study characteristics and performance evaluation metrics were extracted. Methodological quality was assessed using tools such as Cochrane RoB, STROBE, STARD, and DISCERN. Additionally, the CLEAR-LLM and CliMA-10 frameworks were applied to provide complementary evaluations of quality and performance.

RESULTS

A total of 16 studies were included, covering various LLMs such as ChatGPT, Gemini, and Claude. Application scenarios encompassed text generation, imaging diagnostics, and patient education. Most studies were observational in nature, and overall methodological quality ranged from moderate to high. Based on CliMA-10 scores, LLMs exhibited upper-moderate performance in KOA-related tasks. The ChatGPT-4 series consistently outperformed other models, especially in structured output generation, interpretation of clinical terminology, and content accuracy. Key limitations included insufficient sample representativeness, inconsistent control over hallucinated content, and the lack of standardized evaluation tools.

CONCLUSION

Large language models show notable potential in the KOA field, but their clinical application is still exploratory and limited by issues such as sample bias and methodological heterogeneity. Model performance varies across tasks, underscoring the need for improved prompt design and standardized evaluation frameworks. With real-world data and ethical oversight, LLMs may contribute more significantly to personalized KOA management. SYSTEMATIC REVIEW REGISTRATION: https://osf.io/jy4kz, identifier 10.17605/OSF.IO/479R8.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์ฒด๊ณ„์  ๋ฌธํ—Œ๊ณ ์ฐฐ์€ ๋ฌด๋ฆŽ ๊ณจ๊ด€์ ˆ์—ผ(KOA) ๋ถ„์•ผ์—์„œ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ์ž„์ƒ์  ํ™œ์šฉ ํ˜„ํ™ฉ๊ณผ ์„ฑ๋Šฅ์„ ๋ถ„์„ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, LLM์€ ์ง„๋‹จ ๋ณด์กฐ, ์ •๋ณด ์ถ”์ถœ ๋ฐ ํ™˜์ž ๊ต์œก์—์„œ ์ค‘์ƒ์œ„ ์ˆ˜์ค€์˜ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋‚˜, ํ™˜๊ฐ ํ˜„์ƒ๊ณผ ๋ฐฉ๋ฒ•๋ก ์  ์ด์งˆ์„ฑ ๋“ฑ์˜ ํ•œ๊ณ„๊ฐ€ ํ™•์ธ๋˜์—ˆ์Šต๋‹ˆ๋‹ค. ๋”ฐ๋ผ์„œ ํ–ฅํ›„ ์ž„์ƒ ์ ์šฉ์„ ์œ„ํ•ด์„œ๋Š” ํ‘œ์ค€ํ™”๋œ ํ‰๊ฐ€ ํ”„๋ ˆ์ž„์›Œํฌ ๊ตฌ์ถ•๊ณผ ํ”„๋กฌํ”„ํŠธ ์„ค๊ณ„์˜ ๊ณ ๋„ํ™”๊ฐ€ ํ•„์ˆ˜์ ์ž…๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

19Open LLM-based actionable incidental finding extraction from [18F]fluorodeoxyglucose PET-CT radiology reports.

2025-11-14Frontiers in digital healthโญ Q1DOI 10.3389/fdgth.2025.1702082
BACKGROUND

We developed an open, large language model (LLM)-based pipeline to extract actionable incidental findings (AIFs) from [18F]fluorodeoxyglucose positron emission tomography-computed tomography ([18F]FDG PET-CT) reports. This imaging modality often uncovers AIFs, which can affect a patient's treatment. The pipeline classifies reports for the presence of AIFs, extracts the relevant sentences, and stores the results in structured JavaScript Object Notation format, enabling use in both short- and long-term applications.

METHODS

Training, validation, and test datasets of 1,999, 248, and 250 lung cancer [18F]FDG PET-CT reports, respectively, were annotated by a nuclear medicine physician. An external test dataset of 460 reports was annotated by two nuclear medicine physicians. The training dataset was used to fine-tune an LLM using QLoRA and chain-of-thought (CoT) prompting. This was evaluated quantitatively and qualitatively on both test datasets.

RESULTS

The pipeline achieved document-level F1 scores of 0.917โ€‰ยฑโ€‰0.016 and 0.79โ€‰ยฑโ€‰0.025 on the internal and external test datasets. At the sentence-level, F1 scores of 0.754โ€‰ยฑโ€‰0.011 and 0.522โ€‰ยฑโ€‰0.012 were recorded, and qualitative analysis demonstrated even higher practical utility. This qualitative analysis revealed how sentence-level performance is better in practice.

CONCLUSION

Llama-3.1-8B Instruct was the base LLM that provided the best combination of performance and computational efficiency. The utilisation of CoT prompting improved performance further. Radiology reporting characteristics such as length and style affect model generalisation.

CONCLUSION

We find that a QLoRA-adapted LLM utilising CoT prompting successfully extracts AIF information at both document- and sentence-level from both internal and external PET-CT reports. We believe this model can assist with short-term clinical challenges like clinical alerts and reminders, and long-term tasks like investigating comorbidities.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” [18F]FDG PET-CT ํŒ๋…๋ฌธ์—์„œ ์ž„์ƒ์  ์กฐ์น˜๊ฐ€ ํ•„์š”ํ•œ ์šฐ์—ฐํ•œ ๋ฐœ๊ฒฌ(AIF)์„ ์ถ”์ถœํ•˜๊ธฐ ์œ„ํ•ด QLoRA ๊ธฐ๋ฐ˜์œผ๋กœ ๋ฏธ์„ธ ์กฐ์ •๋œ Llama-3.1-8B ๋ชจ๋ธ๊ณผ Chain-of-Thought(CoT) ํ”„๋กฌํ”„ํŒ…์„ ๊ฒฐํ•ฉํ•œ ํŒŒ์ดํ”„๋ผ์ธ์„ ๊ฐœ๋ฐœํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๊ฒ€์ฆ ๊ฒฐ๊ณผ, ํ•ด๋‹น ๋ชจ๋ธ์€ ๋‚ด๋ถ€ ๋ฐ ์™ธ๋ถ€ ํ…Œ์ŠคํŠธ ๋ฐ์ดํ„ฐ์…‹์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์œผ๋กœ AIF๋ฅผ ์‹๋ณ„ ๋ฐ ๊ตฌ์กฐํ™”ํ•˜์˜€์œผ๋ฉฐ, ํ–ฅํ›„ ์ž„์ƒ ์•Œ๋ฆผ ๋ฐ ๋™๋ฐ˜ ์งˆํ™˜ ์—ฐ๊ตฌ ๋“ฑ ๋‹ค์–‘ํ•œ ์˜๋ฃŒ ํ˜„์žฅ์—์„œ ํšจ์œจ์ ์ธ ๋ณด์กฐ ๋„๊ตฌ๋กœ ํ™œ์šฉ๋  ๊ฒƒ์œผ๋กœ ๊ธฐ๋Œ€๋ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

20Locally deployed context-aware chatbot outperforms generic large language models for guideline-concordant pediatric imaging recommendations.

2025-11-13Pediatric radiologyโญ Q1DOI 10.1007/s00247-025-06453-6
BACKGROUND

Accurate modality selection in pediatric imaging is critical, yet adherence to the American College of Radiology (ACR) Appropriateness Criteria remains limited. Large language models (LLMs) offer potential as decision support tools but often lack domain-specific accuracy.

OBJECTIVE

To evaluate the performance of a locally run, context-aware chatbot (ped-Llama) based on an open-source LLM for providing personalized pediatric imaging recommendations grounded in ACR guidelines.

METHODS

A simulation-based study was conducted using 50 pediatric clinical scenarios derived from ACR guideline variants. The ped-Llama chatbot, built using a retrieval-augmented generation (RAG) approach with a locally deployed Llama-3.1-8B model, was compared against three generic LLMs (Llama-3.1-8B without RAG, Generative pre-trained transformer-4o [GPT-4o], Claude Opus) and three radiologists (junior resident, senior resident, specialist pediatric radiologist). Each provided imaging recommendations, which were classified as "Usually appropriate," "May be appropriate," or "Usually not appropriate" per ACR guidelines. Modal outputs were analyzed, and consistency across three independent LLM runs was assessed.

RESULTS

ped-Llama achieved 80% (40/50) "Usually appropriate" recommendations, outperforming Llama-3.1-8B (46%), GPT-4o (54%), and Claude Opus (46%), and matching the specialist radiologist (76%). When both "Usually appropriate" and "May be appropriate" were accepted, ped-Llama achieved 90% accuracy. Consistency across runs was 72% for ped-Llama versus 44-50% for generic LLMs.

CONCLUSION

This study shows that a locally run, RAG-enabled chatbot based on an open-source LLM can provide guideline-concordant pediatric imaging recommendations with expert-level accuracy. Such systems offer a practical and scalable approach to artificial intelligence-assisted decision support in radiology.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG) ๊ธฐ์ˆ ์„ ์ ์šฉํ•œ ๋กœ์ปฌ ๊ธฐ๋ฐ˜ ์ฑ—๋ด‡(ped-Llama)์ด ์†Œ์•„ ์˜์ƒ ๊ฒ€์‚ฌ ๊ถŒ๊ณ ์•ˆ์„ ACR ๊ฐ€์ด๋“œ๋ผ์ธ์— ๋งž์ถฐ ์–ผ๋งˆ๋‚˜ ์ •ํ™•ํ•˜๊ฒŒ ์ œ์‹œํ•˜๋Š”์ง€ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. 50๊ฐœ์˜ ์ž„์ƒ ์‹œ๋‚˜๋ฆฌ์˜ค๋ฅผ ๋Œ€์ƒ์œผ๋กœ ํ•œ ๊ฒฐ๊ณผ, ped-Llama๋Š” ๋ฒ”์šฉ LLM๋ณด๋‹ค ์šฐ์ˆ˜ํ•œ 80%์˜ ์ ์ ˆ์„ฑ ์ ์ˆ˜๋ฅผ ๊ธฐ๋กํ•˜๋ฉฐ ์†Œ์•„ ์˜์ƒ ์ „๋ฌธ์˜์™€ ๋Œ€๋“ฑํ•œ ์ˆ˜์ค€์˜ ์ •ํ™•๋„์™€ ์ผ๊ด€์„ฑ์„ ๋ณด์˜€์Šต๋‹ˆ๋‹ค. ์ด๋Š” ๋กœ์ปฌ ํ™˜๊ฒฝ์—์„œ ๊ตฌ๋™๋˜๋Š” ํŠนํ™”๋œ LLM์ด ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ๋„๊ตฌ๋กœ์„œ ์‹ค์šฉ์ ์ด๊ณ  ์‹ ๋ขฐํ•  ์ˆ˜ ์žˆ๋Š” ๋Œ€์•ˆ์ด ๋  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

21Web based AI-driven framework combining multi-modal data with CNN and LLM for Parkinson's disease diagnosis.

2025-11-04Scientific reportsโญ Q1DOI 10.1038/s41598-025-22448-7

Parkinson's disease (PD) is a progressive neurodegenerative disorder characterized by a wide spectrum of motor and non-motor symptoms, often leading to delayed or inaccurate diagnosis. Conventional diagnostic methods frequently suffer from limited sensitivity, scalability, and interpretability, thereby restricting their utility in clinical settings. To address these limitations, this study presents a novel AI-driven diagnostic framework that integrates multimodal data fusion, deep learning-based classification, and generative language modeling to improve diagnostic accuracy and enable personalized reporting. The proposed framework leverages the Parkinson's Progression Marker Initiative (PPMI) dataset, incorporating structural Magnetic resonance imaging (MRI), Single-Photon Emission Computed Tomography (SPECT) imaging, cerebrospinal fluid (CSF) biomarkers, and clinical assessments. Statistical analysis was employed to select 14 key biomarkers-including dopamine transporter SBR values and CSF protein levels-from a total of 21 features identified as clinically relevant. A 1D Convolutional Neural Network (1D-CNN) was developed and trained using 121 engineered features, comprising radiomic descriptors and biologically derived metrics. Preprocessing and extensive feature engineering were conducted prior to a 70:30 train-test split, with data augmentation applied to the training set to enhance model generalization. The classifier achieved an accuracy of 93.7%, surpassing baseline approaches and emphasizing the value of domain-informed feature design. To improve interpretability and clinician usability, a Mini ChatGPT-4.0 Large Language Model (LLM) was fine-tuned using approximately 1,000 domain-specific prompt-response pairs generated from literature, classifier-derived eXplainable AI (XAI) feature scores, and expert annotations. The generated responses were evaluated using a custom scoring metric (0.0-5.0) based on their semantic alignment with ground truth completions. This LLM module produces patient-specific diagnostic summaries and treatment suggestions. Additionally, a cloud-based interface was developed to facilitate real-time MRI uploads, automated inference, and chatbot-driven consultations. Overall, the framework demonstrates high diagnostic performance, transparency, and user accessibility, offering significant potential for real-world clinical deployment in PD diagnosis and decision support.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” MRI, SPECT, ๋‡Œ์ฒ™์ˆ˜์•ก ๋ฐ”์ด์˜ค๋งˆ์ปค ๋“ฑ ๋‹ค์ค‘ ๋ชจ๋‹ฌ ๋ฐ์ดํ„ฐ๋ฅผ 1D-CNN์œผ๋กœ ํ†ตํ•ฉ ๋ถ„์„ํ•˜์—ฌ ํŒŒํ‚จ์Šจ๋ณ‘์„ 93.7%์˜ ์ •ํ™•๋„๋กœ ์ง„๋‹จํ•˜๋Š” AI ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ์ œ์•ˆํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋˜ํ•œ, ๋ฏธ์„ธ ์กฐ์ •๋œ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์„ ๊ฒฐํ•ฉํ•˜์—ฌ ์ง„๋‹จ ๊ฒฐ๊ณผ์— ๋Œ€ํ•œ ์„ค๋ช… ๊ฐ€๋Šฅํ•œ ์š”์•ฝ ๋ฐ ๋งž์ถคํ˜• ์น˜๋ฃŒ ์ œ์•ˆ์„ ์ œ๊ณตํ•จ์œผ๋กœ์จ ์ž„์ƒ์  ํ•ด์„ ๊ฐ€๋Šฅ์„ฑ๊ณผ ํ™œ์šฉ๋„๋ฅผ ํฌ๊ฒŒ ํ–ฅ์ƒ์‹œ์ผฐ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

22Minimum Reporting Items for Clear Evaluation of Accuracy Reports of Large Language Models in Healthcare (MI-CLEAR-LLM): 2025 Updates.

2025-11-03Korean journal of radiologyโญ Q1DOI 10.3348/kjr.2025.1522

Recent systematic reviews have raised concerns about the quality of reporting in studies evaluating the accuracy of large language models (LLMs) in medical applications. Incomplete and inconsistent reporting hampers the ability of reviewers and readers to assess study methodology, interpret results, and evaluate reproducibility. To address this issue, the MInimum reporting items for CLear Evaluation of Accuracy Reports of Large Language Models in healthcare (MI-CLEAR-LLM) checklist was developed. This article presents an extensively updated version. While the original version focused on proprietary LLMs accessed via web-based chatbot interfaces, the updated checklist incorporates considerations relevant to application programming interfaces and self-managed models, typically based on open-source LLMs. As before, the revised MI-CLEAR-LLM focuses on reporting practices specific to LLM accuracy evaluations: specifically, the reporting of how LLMs are specified, accessed, adapted, and applied in testing, with special attention to methodological factors that influence outputs. The checklist includes essential items across categories such as model identification, access mode, input data type, adaptation strategy, prompt optimization, prompt execution, stochasticity management, and test data independence. This article also presents reporting examples from the literature. Adoption of the updated MI-CLEAR-LLM can help ensure transparency in reporting and enable more accurate and meaningful evaluation of studies.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜๋ฃŒ ๋ถ„์•ผ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ์ •ํ™•๋„ ํ‰๊ฐ€ ์—ฐ๊ตฌ์—์„œ ๋ณด๊ณ ์˜ ๋ถˆ์™„์ „์„ฑ๊ณผ ์ผ๊ด€์„ฑ ๋ถ€์กฑ ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๊ธฐ ์œ„ํ•ด MI-CLEAR-LLM ์ฒดํฌ๋ฆฌ์ŠคํŠธ์˜ 2025๋…„ ๊ฐœ์ •ํŒ์„ ์ œ์‹œํ•ฉ๋‹ˆ๋‹ค. ์ด๋ฒˆ ๊ฐœ์ •์•ˆ์€ API ๋ฐ ์˜คํ”ˆ์†Œ์Šค ๊ธฐ๋ฐ˜ ๋ชจ๋ธ๊นŒ์ง€ ๋ฒ”์œ„๋ฅผ ํ™•์žฅํ•˜์—ฌ ๋ชจ๋ธ ์‹๋ณ„, ์ ‘๊ทผ ๋ฐฉ์‹, ํ”„๋กฌํ”„ํŠธ ์ตœ์ ํ™”, ํ™•๋ฅ ์„ฑ ๊ด€๋ฆฌ ๋“ฑ ์—ฐ๊ตฌ์˜ ์žฌํ˜„์„ฑ๊ณผ ํˆฌ๋ช…์„ฑ์„ ํ™•๋ณดํ•˜๊ธฐ ์œ„ํ•œ ํ•„์ˆ˜ ๋ณด๊ณ  ํ•ญ๋ชฉ๋“ค์„ ์ฒด๊ณ„ํ™”ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ด๋ฅผ ํ†ตํ•ด ์˜๋ฃŒ AI ์—ฐ๊ตฌ์˜ ๋ฐฉ๋ฒ•๋ก ์  ์—„๋ฐ€์„ฑ์„ ๋†’์ด๊ณ , ์—ฐ๊ตฌ ๊ฒฐ๊ณผ์˜ ํ•ด์„ ๋ฐ ํ‰๊ฐ€๋ฅผ ์œ„ํ•œ ํ‘œ์ค€ํ™”๋œ ๊ฐ€์ด๋“œ๋ผ์ธ์„ ์ œ๊ณตํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

23The performance of large language models in dentomaxillofacial radiology: a systematic review.

2025-11-01Dento maxillo facial radiologyโญ Q1DOI 10.1093/dmfr/twaf060
OBJECTIVE

This study aimed to systematically review the current performance of large language models (LLMs) in dento-maxillofacial radiology (DMFR).

METHODS

Five electronic databases were used to identify studies that developed, fine-tuned, or evaluated LLMs for DMFR-related tasks. Data extracted included study purpose, LLM type, images/text source, applied language, dataset characteristics, input and output, performance outcomes, evaluation methods, and reference standards. Customized assessment criteria adapted from the TRIPOD-LLM reporting guideline were used to evaluate the risk-of-bias in the included studies specifically regarding the clarity of dataset origin, the robustness of performance evaluation methods, and the validity of the reference standards.

RESULTS

The initial search yielded 1621 titles, and 19 studies were included. These studies investigated the use of LLMs for tasks including the production and answering of DMFR-related qualification exams and educational questions (nโ€‰=โ€‰8), diagnosis and treatment recommendations (nโ€‰=โ€‰7), and radiology report generation and patient communication (nโ€‰=โ€‰4). LLMs demonstrated varied performance in diagnosing dental conditions, with accuracy ranging from 37% to 92.5% and expert ratings for differential diagnosis and treatment planning between 3.6 and 4.7 on a 5-point scale. For DMFR-related qualification exams and board-style questions, LLMs achieved correctness rates between 33.3% and 86.1%. Automated radiology report generation showed moderate performance with accuracy ranging from 70.4% to 81.3%.

CONCLUSION

LLMs demonstrate promising potential in DMFR, particularly for diagnostic, educational, and report generation tasks. However, their current accuracy, completeness, and consistency remain variable. Further development, validation, and standardization are needed before LLMs can be reliably integrated as supportive tools in clinical workflows and educational settings.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์น˜๊ณผ์•…์•ˆ๋ฉด ๋ฐฉ์‚ฌ์„ ํ•™(DMFR) ๋ถ„์•ผ์—์„œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ์ฒด๊ณ„์ ์œผ๋กœ ๊ณ ์ฐฐํ•˜์˜€์œผ๋ฉฐ, ์ง„๋‹จ, ๊ต์œก, ํŒ๋…๋ฌธ ์ƒ์„ฑ ๋“ฑ ๋‹ค์–‘ํ•œ ์˜์—ญ์—์„œ์˜ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, LLM์€ ์ง„๋‹จ ๋ฐ ๊ต์œก์  ๊ณผ์ œ์—์„œ ์œ ๋งํ•œ ์„ฑ๊ณผ๋ฅผ ๋ณด์˜€์œผ๋‚˜ ์ •ํ™•๋„์™€ ์ผ๊ด€์„ฑ์— ๋ณ€๋™์„ฑ์ด ์ปค, ์ž„์ƒ ๋ฐ ๊ต์œก ํ˜„์žฅ์— ๋„์ž…ํ•˜๊ธฐ ์œ„ํ•ด์„œ๋Š” ํ‘œ์ค€ํ™”๋œ ๊ฒ€์ฆ๊ณผ ์ถ”๊ฐ€์ ์ธ ๊ธฐ์ˆ ์  ๋ณด์™„์ด ํ•„์š”ํ•œ ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

24A systematic review and meta-analysis of GPT-based differential diagnostic accuracy in radiological cases: 2023-2025.

2025-10-28Frontiers in radiologyโญ Q1DOI 10.3389/fradi.2025.1670517
OBJECTIVE

To systematically evaluate the diagnostic accuracy of various GPT models in radiology, focusing on differential diagnosis performance across textual and visual input modalities, model versions, and clinical contexts.

METHODS

A systematic review and meta-analysis were conducted using PubMed and SCOPUS databases on March 24, 2025, retrieving 639 articles. Studies were eligible if they evaluated GPT model diagnostic accuracy on radiology cases. Non-radiology applications, fine-tuned/custom models, board-style multiple-choice questions, or studies lacking accuracy data were excluded. After screening, 28 studies were included. Risk of bias was assessed using the Newcastle-Ottawa Scale (NOS). Diagnostic accuracy was assessed as top diagnosis accuracy (correct diagnosis listed first) and differential accuracy (correct diagnosis listed anywhere). Statistical analysis involved Mann-Whitney U tests using study-level median (median) accuracy with interquartile ranges (IQR), and a generalized linear mixed-effects model (GLMM) to evaluate predictors influencing model performance.

RESULTS

Analysis included 8,852 radiological cases across multiple radiology subspecialties. Differential accuracy varied significantly among GPT models, with newer models (GPT-4T: 72.00%, median 82.32%; GPT-4o: 57.23%, median 53.75%; GPT-4: 56.46%, median 56.65%) outperforming earlier versions (GPT-3.5: 37.87%, median 36.33%). Textual inputs demonstrated higher accuracy (GPT-4: 56.46%, median 58.23%) compared to visual inputs (GPT-4V: 42.32%, median 41.41%). The provision of clinical history was associated with improved diagnostic accuracy in the GLMM (ORโ€‰=โ€‰1.27, pโ€‰=โ€‰.001), despite unadjusted medians showing lower performance when history was provided (61.74% vs. 52.28%). Private data (86.51%, median 94.00%) yielded higher accuracy than public data (47.62%, median 46.45%). Accuracy trends indicated improvement in newer models over time, while GPT-3.5's accuracy declined. GLMM results showed higher odds of accuracy for advanced models (ORโ€‰=โ€‰1.84), and lower odds for visual inputs (ORโ€‰=โ€‰0.29) and public datasets (ORโ€‰=โ€‰0.34), while accuracy showed no significant trend over successive study years (pโ€‰=โ€‰0.57). Egger's test found no significant publication bias, though considerable methodological heterogeneity was observed.

CONCLUSION

This meta-analysis highlights significant variability in GPT model performance influenced by input modality, data source, and model version. High methodological heterogeneity across studies emphasizes the need for standardized protocols in future research, and readers should interpret pooled estimates and medians with this variability in mind.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 2023๋…„๋ถ€ํ„ฐ 2025๋…„๊นŒ์ง€ ๋ฐœํ‘œ๋œ 28ํŽธ์˜ ์—ฐ๊ตฌ๋ฅผ ๋ฉ”ํƒ€๋ถ„์„ํ•˜์—ฌ ์˜์ƒ์˜ํ•™ ๋ถ„์•ผ์—์„œ GPT ๋ชจ๋ธ์˜ ๊ฐ๋ณ„ ์ง„๋‹จ ์ •ํ™•๋„๋ฅผ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ์ตœ์‹  ๋ชจ๋ธ์ผ์ˆ˜๋ก, ์‹œ๊ฐ ์ •๋ณด๋ณด๋‹ค ํ…์ŠคํŠธ ๊ธฐ๋ฐ˜ ์ž…๋ ฅ์ผ์ˆ˜๋ก, ๊ทธ๋ฆฌ๊ณ  ๊ณต๊ฐœ ๋ฐ์ดํ„ฐ์…‹๋ณด๋‹ค ์‚ฌ์„ค ๋ฐ์ดํ„ฐ์…‹์—์„œ ๋” ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ GPT ๋ชจ๋ธ์˜ ์ง„๋‹จ ์„ฑ๋Šฅ์€ ์ž…๋ ฅ ๋ฐฉ์‹๊ณผ ๋ชจ๋ธ ๋ฒ„์ „์— ๋”ฐ๋ผ ์œ ์˜๋ฏธํ•œ ์ฐจ์ด๋ฅผ ๋ณด์ด๋ฉฐ, ํ–ฅํ›„ ์—ฐ๊ตฌ์—์„œ๋Š” ํ‘œ์ค€ํ™”๋œ ํ‰๊ฐ€ ํ”„๋กœํ† ์ฝœ ํ™•๋ฆฝ์ด ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

25Exploring multimodal large language models on transthoracic Echocardiogram (TTE) tasks for cardiovascular decision support.

2025-10-23Journal of biomedical informaticsโญ Q1DOI 10.1016/j.jbi.2025.104930
OBJECTIVE

Multimodal large language models (LLMs) offer new potential for enhancing cardiovascular decision support, particularly in interpreting echocardiographic data. This study systematically evaluates and benchmarks foundation models from diverse domains on echocardiogram-based tasks to assess their effectiveness, limitations and potential in clinical cardiovascular applications.

METHODS

We curated three cardiovascular imaging datasets-EchoNet-Dynamic, TMED2, and an expert-annotated echocardiogram (TTE) dataset-to evaluate performance on four critical tasks: (1) cardiac function evaluation through ejection fraction (EF) prediction, (2) cardiac view classification, (3) aortic stenosis (AS) severity assessment, and (4) cardiovascular disease classification. We evaluated six multimodal LLMs: EchoClip (cardiovascular-specific), BiomedGPT and LLaVA-Med (medical-domain), and MiniCPM-V 2.6, LLaMA-3-Vision-Alpha, and Gemini-1.5 (general-domain). Models were assessed using zero-shot, few-shot, and fine-tuning strategies, where applicable. Performance was measured using mean absolute error (MAE) and root mean squared error (RMSE) for EF prediction, and accuracy, precision, recall, and F1 score for classification tasks.

RESULTS

Domain-specific models such as EchoClip demonstrated the strongest zero-shot performance in EF prediction, achieving an MAE of 10.34. General-domain models showed limited effectiveness without adaptation, with MiniCPM-V 2.6 reporting an MAE of 251.92. Fine-tuning significantly improved outcomes; for example, MiniCPM-V 2.6's MAE decreased to 31.93, and view classification accuracy increased from 20ย % to 63.05ย %. In classification tasks, EchoClip achieved F1 scores of 0.2716 for AS severity and 0.4919 for disease classification but exhibited limited performance in view classification (F1ย =ย 0.1457). Few-shot learning yielded modest gains but was generally less effective than fine-tuning.

CONCLUSION

This evaluation and benchmarking study demonstrated the importance of domain-specific pretraining and model adaptation in cardiovascular decision support tasks. Cardiovascular-focused models and fine-tuned general-domain models achieved superior performance, especially for complex assessments such as EF estimation. These findings offer critical insights into the current capabilities and future directions for clinically meaningful AI integration in cardiovascular medicine.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์‹ฌ์žฅ์ดˆ์ŒํŒŒ(TTE) ๋ฐ์ดํ„ฐ ํ•ด์„์„ ์œ„ํ•œ ๋‹ค์ค‘ ๋ชจ๋‹ฌ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ์ž„์ƒ์  ์œ ์šฉ์„ฑ์„ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด ์‹ฌ์žฅ ํŠนํ™” ๋ชจ๋ธ๊ณผ ๋ฒ”์šฉ ๋ชจ๋ธ์˜ ์„ฑ๋Šฅ์„ ๋น„๊ต ๋ถ„์„ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ์‹ฌ์žฅ ํŠนํ™” ๋ชจ๋ธ์ธ EchoClip์ด ์ œ๋กœ์ƒท ํ•™์Šต์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ๋ฒ”์šฉ ๋ชจ๋ธ์€ ๋ฏธ์„ธ ์กฐ์ •(fine-tuning)์„ ๊ฑฐ์ณค์„ ๋•Œ ์‹ฌ๋ฐ•์ถœ๋ฅ  ์˜ˆ์ธก ๋ฐ ์˜์ƒ ๋ถ„๋ฅ˜ ์ •ํ™•๋„๊ฐ€ ํฌ๊ฒŒ ํ–ฅ์ƒ๋˜์—ˆ์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ ์‹ฌํ˜ˆ๊ด€ ๋ถ„์•ผ์˜ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์›์„ ์œ„ํ•ด์„œ๋Š” ๋„๋ฉ”์ธ ํŠนํ™” ์‚ฌ์ „ ํ•™์Šต๊ณผ ์ ์ ˆํ•œ ๋ชจ๋ธ ๋ฏธ์„ธ ์กฐ์ •์ด ํ•„์ˆ˜์ ์ž„์„ ํ™•์ธํ•˜์˜€์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

26Performance of GPT-4o and o1-Pro on United Kingdom Medical Licensing Assessment-style items: a comparative study.

2025-10-10Journal of educational evaluation for health professionsโญ Q1DOI 10.3352/jeehp.2025.22.30
OBJECTIVE

Large language models (LLMs) such as ChatGPT, and their potential to support autonomous learning for licensing exams like the UK Medical Licensing Assessment (UKMLA), are of growing interest. However, empirical evaluations of artificial intelligence (AI) performance against the UKMLA standard remain limited.

METHODS

We evaluated the performance of 2 recent ChatGPT versions, GPT-4o and o1-Pro, on a curated set of 374 UKMLA-style single-best-answer items spanning diverse medical specialties. Statistical comparisons using McNemar's test assessed the significance of differences between the 2 models. Specialties were analyzed to identify domain-specific variation. In addition, 20 image-based items were evaluated.

RESULTS

GPT-4o achieved an accuracy of 88.8%, while o1-Pro achieved 93.0%. McNemar's test revealed a statistically significant difference in favor of o1-Pro. Across specialties, both models demonstrated excellent performance in surgery, psychiatry, and infectious diseases. Notable differences arose in dermatology, respiratory medicine, and imaging, where o1-Pro consistently outperformed GPT-4o. Nevertheless, isolated weaknesses in general practice were observed. The analysis of image-based items showed 75% accuracy for GPT-4o and 90% for o1-Pro (P=0.25).

CONCLUSION

ChatGPT shows strong potential as an adjunct learning tool for UKMLA preparation, with both models achieving scores above the calculated pass mark. This underscores the promise of advanced AI models in medical education. However, specialty-specific inconsistencies suggest AI tools should complement, rather than replace, traditional study methods.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜๊ตญ ์˜์‚ฌ๋ฉดํ—ˆ์‹œํ—˜(UKMLA) ํ˜•์‹์˜ ๋ฌธํ•ญ 374๊ฐœ๋ฅผ ํ™œ์šฉํ•˜์—ฌ GPT-4o์™€ o1-Pro์˜ ์„ฑ๋Šฅ์„ ๋น„๊ต ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, o1-Pro(93.0%)๊ฐ€ GPT-4o(88.8%)๋ณด๋‹ค ํ†ต๊ณ„์ ์œผ๋กœ ์œ ์˜ํ•˜๊ฒŒ ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ํŠนํžˆ ์˜์ƒ ๊ธฐ๋ฐ˜ ๋ฌธํ•ญ ๋ฐ ํŠน์ • ์ „๋ฌธ ๋ถ„์•ผ์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋‚˜ํƒ€๋ƒˆ์Šต๋‹ˆ๋‹ค. ๋‘ ๋ชจ๋ธ ๋ชจ๋‘ ํ•ฉ๊ฒฉ ๊ธฐ์ค€์„ ์ƒํšŒํ•˜๋Š” ์„ฑ์ ์„ ๊ฑฐ๋‘์—ˆ์œผ๋‚˜, ์ผ๋ถ€ ๋ถ„์•ผ์—์„œ์˜ ๋ถˆ๊ท ์ผํ•œ ์„ฑ๋Šฅ์„ ๊ณ ๋ คํ•  ๋•Œ AI๋Š” ํ•™์Šต ๋ณด์กฐ ๋„๊ตฌ๋กœ ํ™œ์šฉํ•˜๋Š” ๊ฒƒ์ด ์ ์ ˆํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

27Performance Comparison of Cutting-Edge Large Language Models on the ACR In-Training Examination: An Update for 2025.

2025-09-24Academic radiologyโญ Q1DOI 10.1016/j.acra.2025.09.008
OBJECTIVE

This study represents a continuation of prior work by Payne et al. evaluating large language model (LLM) performance on radiology board-style assessments, specifically the ACR diagnostic radiology in-training examination (DXIT). Building upon earlier findings with GPT-4, we assess the performance of newer, cutting-edge models, such as GPT-4o, GPT-o1, GPT-o3, Claude, Gemini, and Grok on standardized DXIT questions. In addition to overall performance, we compare model accuracy on text-based versus image-based questions to assess multi-modal reasoning capabilities. As a secondary aim, we investigate the potential impact of data contamination by comparing model performance on original versus revised image-based questions.

METHODS

Seven LLMs - GPT-4, GPT-4o, GPT-o1, GPT-o3, Claude 3.5 Sonnet, Gemini 1.5 Pro, and Grok 2.0-were evaluated using 106 publicly available DXIT questions. Each model was prompted using a standardized instruction set to simulate a radiology resident answering board-style questions. For each question, the model's selected answer, rationale, and confidence score were recorded. Unadjusted accuracy (based on correct answer selection) and logic-adjusted accuracy (based on clinical reasoning pathways) were calculated. Subgroup analysis compared model performance on text-based versus image-based questions. Additionally, 63 image-based questions were revised to test novel reasoning while preserving the original diagnostic image to assess the impact of potential training data contamination.

RESULTS

Across 106 DXIT questions, GPT-o1 demonstrated the highest unadjusted accuracy (71.7%), followed closely by GPT-4o (69.8%) and GPT-o3 (68.9%). GPT-4 (59.4%) and Grok 2.0 exhibited similar scores (59.4% and 52.8%). Claude 3.5 Sonnet had the lowest unadjusted accuracies (34.9%). Similar trends were observed for logic-adjusted accuracy, with GPT-o1 (60.4%), GPT-4o (59.4%), and GPT-o3 (59.4%) again outperforming other models, while Grok 2.0 and Claude 3.5 Sonnet lagged behind (34.0% and 30.2%, respectively). GPT-4o's performance was significantly higher on text-based questions compared to image-based ones. Unadjusted accuracy for the revised DXIT questions was 49.2%, compared to 56.1% on matched original DXIT questions. Logic-adjusted accuracy for the revised DXIT questions was 40.0% compared to 44.4% on matched original DXIT questions. No significant difference in performance was observed between original and revised questions.

CONCLUSION

Modern LLMs, especially those from OpenAI, demonstrate strong and improved performance on board-style radiology assessments. Comparable performance on revised prompts suggests that data contamination may have played a limited role. As LLMs improve, they hold strong potential to support radiology resident learning through personalized feedback and board-style question review.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 7์ข…์˜ ์ตœ์‹  ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์„ ๋Œ€์ƒ์œผ๋กœ ๋ฏธ๊ตญ ์˜์ƒ์˜ํ•™ํšŒ(ACR) ์ˆ˜๋ จ์˜ ์‹œํ—˜(DXIT) ๋ฌธ์ œ ํ’€์ด ๋Šฅ๋ ฅ์„ ํ‰๊ฐ€ํ•˜์˜€์œผ๋ฉฐ, ํŠนํžˆ ํ…์ŠคํŠธ ๊ธฐ๋ฐ˜ ๋ฌธ์ œ์™€ ์˜์ƒ ๊ธฐ๋ฐ˜ ๋ฌธ์ œ ๊ฐ„์˜ ์ถ”๋ก  ์„ฑ๋Šฅ ๋ฐ ๋ฐ์ดํ„ฐ ์˜ค์—ผ ๊ฐ€๋Šฅ์„ฑ์„ ๋ถ„์„ํ–ˆ์Šต๋‹ˆ๋‹ค. ํ‰๊ฐ€ ๊ฒฐ๊ณผ GPT-o1, GPT-4o, GPT-o3 ๋ชจ๋ธ์ด ์šฐ์ˆ˜ํ•œ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ์›๋ณธ ๋ฌธ์ œ์™€ ์ˆ˜์ •๋œ ๋ฌธ์ œ ๊ฐ„์˜ ์„ฑ๋Šฅ ์ฐจ์ด๊ฐ€ ํฌ์ง€ ์•Š์•„ ๋ฐ์ดํ„ฐ ์˜ค์—ผ์˜ ์˜ํ–ฅ์€ ์ œํ•œ์ ์ธ ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ ์ตœ์‹  LLM์€ ์˜์ƒ์˜ํ•™ ์ˆ˜๋ จ์˜์˜ ํ•™์Šต ๋ณด์กฐ ๋ฐ ์‹œํ—˜ ๋Œ€๋น„๋ฅผ ์œ„ํ•œ ๋„๊ตฌ๋กœ์„œ ๋†’์€ ์ž ์žฌ๋ ฅ์„ ๋ณด์œ ํ•˜๊ณ  ์žˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

28The top 100 most-cited articles on large language models in medicine: A bibliometric analysis.

2025-09-12Digital healthโญ Q1DOI 10.1177/20552076251365059
OBJECTIVE

Large language models (LLMs) are revolutionizing medical research. However, there is a lack of bibliometric analysis that identifies citation trends shaping the history of this field. This study analyzes the top 100 (T100) most-cited articles on LLMs in medicine to assess their impact and characteristics.

METHODS

A bibliometric analysis of top-cited articles in the Web of Science database using search terms like "LLMs, generative artificial intelligence, GPT" from 2022 to 2025. Two reviewers identified the T100 papers, extracting publication details, citations, and research themes, adhering to BIBLIO reporting guidelines.

RESULTS

The T100 articles had contributed from 655 authors, and 92 articles were published in 2023. Original research constituted the majority of publications (60 articles). Collectively, these works accumulated 14,847 citations, with individual citations ranging from 50 to 1057 (average 148.47). The U.S. led global contributions with 56 articles, Stanford University emerging as the most prolific institution (8 articles). The top seven journals contributed to 31% of the T100, and Journal of Medical Internet Research published the largest share (8 articles) in 70 peer-reviewed journals. The most-cited article is "Evolutionary-scale prediction of atomic-level protein structure with a language model" (Lin et al., Science 2023; 1057 citations). The research themes centered on evaluating LLMs' performance in exam-style evaluations, medical knowledge synthesis, and question-answering tasks in medicine.

CONCLUSION

This analysis provides a core overview of high-impact LLMs research in medicine, guiding future applications. The findings highlighted the remarkable progress in clinical decision support, drug discovery, multimodal medical imaging analysis, and personalized medical information-answering. They also stress the need for prospective trials to assess real-world clinical impacts, boost the reliability of LLMs-generated medical info, develop consensus-driven solutions to address ethical challenges, and launch global initiatives to democratize LLMs tools.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 2022๋…„๋ถ€ํ„ฐ 2025๋…„๊นŒ์ง€ ๋ฐœํ‘œ๋œ ์˜ํ•™ ๋ถ„์•ผ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM) ๊ด€๋ จ ๋…ผ๋ฌธ ์ค‘ ํ”ผ์ธ์šฉ ์ƒ์œ„ 100ํŽธ์„ ๋ถ„์„ํ•˜์—ฌ ์—ฐ๊ตฌ ๋™ํ–ฅ๊ณผ ์˜ํ–ฅ๋ ฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ํ•ด๋‹น ๋…ผ๋ฌธ๋“ค์€ ์ฃผ๋กœ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์›, ์•ฝ๋ฌผ ๋ฐœ๊ฒฌ, ์˜๋ฃŒ ์˜์ƒ ๋ถ„์„ ๋ฐ ์ง€์‹ ํ•ฉ์„ฑ ๋ถ„์•ผ์— ์ง‘์ค‘๋˜์–ด ์žˆ์œผ๋ฉฐ, ๋ฏธ๊ตญ๊ณผ ์Šคํƒ ํผ๋“œ ๋Œ€ํ•™์ด ์—ฐ๊ตฌ๋ฅผ ์ฃผ๋„ํ•˜๋Š” ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค. ํ–ฅํ›„ LLM์˜ ์ž„์ƒ์  ์‹ ๋ขฐ์„ฑ ํ™•๋ณด์™€ ์œค๋ฆฌ์  ๋ฌธ์ œ ํ•ด๊ฒฐ์„ ์œ„ํ•œ ์ „ํ–ฅ์  ์—ฐ๊ตฌ ๋ฐ ๊ตญ์ œ์  ํ˜‘๋ ฅ์˜ ํ•„์š”์„ฑ์ด ๊ฐ•์กฐ๋˜์—ˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

29Implementing a Resource-Light and Low-Code Large Language Model System for Information Extraction from Mammography Reports: A Pilot Study.

2025-09-10Journal of imaging informatics in medicineโญ Q1DOI 10.1007/s10278-025-01659-4

Large language models (LLMs) have been successfully used for data extraction from free-text radiology reports. Most current studies were conducted with LLMs accessed via an application programming interface (API). We evaluated the feasibility of using open-source LLMs, deployed on limited local hardware resources for data extraction from free-text mammography reports, using a common data element (CDE)-based structure. Seventy-nine CDEs were defined by an interdisciplinary expert panel, reflecting real-world reporting practice. Sixty-one reports were classified by two independent researchers to establish ground truth. Five different open-source LLMs deployable on a single GPU were used for data extraction using the general-classifier Python package. Extractions were performed for five different prompt approaches with calculation of overall accuracy, micro-recall and micro-F1. Additional analyses were conducted using thresholds for the relative probability of classifications. High inter-rater agreement was observed between manual classifiers (Cohen's kappa 0.83). Using default prompts, the LLMs achieved accuracies of 59.2-72.9%. Chain-of-thought prompting yielded mixed results, while few-shot prompting led to decreased accuracy. Adaptation of the default prompts to precisely define classification tasks improved performance for all models, with accuracies of 64.7-85.3%. Setting certainty thresholds further improved accuracies toโ€‰>โ€‰90% but reduced the coverage rate toโ€‰<โ€‰50%. Locally deployed open-source LLMs can effectively extract information from mammography reports, maintaining compatibility with limited computational resources. Selection and evaluation of the model and prompting strategy are critical. Clear, task-specific instructions appear crucial for high performance. Using a CDE-based framework provides clear semantics and structure for the data extraction.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์ œํ•œ๋œ ๋กœ์ปฌ ํ•˜๋“œ์›จ์–ด ํ™˜๊ฒฝ์—์„œ ์˜คํ”ˆ์†Œ์Šค ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์„ ํ™œ์šฉํ•˜์—ฌ ์œ ๋ฐฉ์ดฌ์˜์ˆ  ๋ณด๊ณ ์„œ๋กœ๋ถ€ํ„ฐ ๊ณตํ†ต ๋ฐ์ดํ„ฐ ์š”์†Œ(CDE)๋ฅผ ์ถ”์ถœํ•˜๋Š” ์‹œ์Šคํ…œ์˜ ํƒ€๋‹น์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ์ž‘์—… ํŠนํ™”ํ˜• ํ”„๋กฌํ”„ํŠธ ์„ค๊ณ„์™€ ํ™•์‹ค์„ฑ ์ž„๊ณ„๊ฐ’ ์„ค์ •์„ ํ†ตํ•ด 90% ์ด์ƒ์˜ ๋†’์€ ์ •ํ™•๋„๋กœ ์ •๋ณด ์ถ”์ถœ์ด ๊ฐ€๋Šฅํ•จ์„ ํ™•์ธํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ด๋Š” ๊ณ ๊ฐ€์˜ API ์—†์ด๋„ ๋กœ์ปฌ ํ™˜๊ฒฝ์—์„œ ํšจ์œจ์ ์ธ ๋ฐ์ดํ„ฐ ๊ตฌ์กฐํ™”๊ฐ€ ๊ฐ€๋Šฅํ•จ์„ ์‹œ์‚ฌํ•˜๋ฉฐ, ์ ์ ˆํ•œ ๋ชจ๋ธ ์„ ํƒ๊ณผ ํ”„๋กฌํ”„ํŠธ ์ „๋žต์ด ์„ฑ๋Šฅ ์ตœ์ ํ™”์˜ ํ•ต์‹ฌ์ž„์„ ๋ณด์—ฌ์ค๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

30Comparison of Large Language Models' Performance on 600 Nuclear Medicine Technology Board Examination-Style Questions.

2025-09-05Journal of nuclear medicine technologyDOI 10.2967/jnmt.124.269335

This study investigated the application of large language models (LLMs) with and without retrieval-augmented generation (RAG) in nuclear medicine, particularly their performance across various topics relevant to the field, to evaluate their potential use as reliable tools for professional education and clinical decision-making.

METHODS

We evaluated the performance of LLMs, including the OpenAI GPT-4o series, Google Gemini, Cohere, Anthropic, and Meta Llama3, across 15 nuclear medicine topics. The models' accuracy was assessed using a set of 600 sample questions, covering a range of clinical and technical domains in nuclear medicine. Overall accuracy was measured by averaging performance across these topics. Additional performance comparisons were conducted across individual models.

RESULTS

OpenAI's models, particularly openai_nvidia_gpt-4o_final and openai_mxbai_gpt-4o_final, demonstrated the highest overall accuracy, achieving scores of 0.787 and 0.783, respectively, when RAG was implemented. Anthropic Opus and Google Gemini 1.5 Pro followed closely, with competitive overall accuracy scores of 0.773 and 0.750 with RAG. Cohere and Llama3 models showed more variability in performance, with the Llama3 ollama_llama3 model (without RAG) achieving the lowest accuracy. Discrepancies were noted in question interpretation, particularly in complex clinical guidelines and imaging-based queries.

CONCLUSION

LLMs show promising potential in nuclear medicine, improving diagnostic accuracy, especially in areas like radiation safety and skeletal system scintigraphy. This study also demonstrates that adding a RAG workflow can increase the accuracy of an off-the-shelf model. However, challenges persist in handling nuanced guidelines and visual data, emphasizing the need for further optimization in LLMs for medical applications.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 600๊ฐœ์˜ ํ•ต์˜ํ•™ ์ „๋ฌธ์˜ ์‹œํ—˜ ๋ฌธํ•ญ์„ ํ™œ์šฉํ•˜์—ฌ ๋‹ค์–‘ํ•œ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜๊ณ  ๊ฒ€์ƒ‰ ์ฆ๊ฐ• ์ƒ์„ฑ(RAG) ๋„์ž…์˜ ํšจ๊ณผ๋ฅผ ๋ถ„์„ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ํ‰๊ฐ€ ๊ฒฐ๊ณผ, GPT-4o ๋ชจ๋ธ์ด RAG ์ ์šฉ ์‹œ ๊ฐ€์žฅ ๋†’์€ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ์ „๋ฐ˜์ ์œผ๋กœ RAG ๋„์ž…์ด ๋ชจ๋ธ์˜ ์ž„์ƒ์  ์ถ”๋ก  ๋Šฅ๋ ฅ์„ ํ–ฅ์ƒ์‹œํ‚ค๋Š” ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ๋ณต์žกํ•œ ์ž„์ƒ ๊ฐ€์ด๋“œ๋ผ์ธ ํ•ด์„ ๋ฐ ์˜์ƒ ๊ธฐ๋ฐ˜ ์งˆ์˜์—๋Š” ์—ฌ์ „ํžˆ ํ•œ๊ณ„๊ฐ€ ์žˆ์–ด ํ–ฅํ›„ ์˜๋ฃŒ ๋ถ„์•ผ ํŠนํ™” ์ตœ์ ํ™”๊ฐ€ ํ•„์š”ํ•œ ๊ฒƒ์œผ๋กœ ํ™•์ธ๋˜์—ˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

31Leveraging ChatGPT to strengthen pediatric healthcare systems: a systematic review.

2025-07-12European journal of pediatricsโญ Q1DOI 10.1007/s00431-025-06320-4

UNLABELLED: This systematic review is the first to investigate ChatGPT's applications in pediatric healthcare systems by assessing its accuracy and readability, with a focus on its impact across key areas such as clinical decision-making, clinical documentation, patient education, and training. The primary question guiding this review is: How does ChatGPT impact pediatric healthcare systems, for example, in terms of improving patient education, enhancing providers' efficiency, and assisting with clinical decision-making? A systematic review was conducted using PubMed, EMBASE, and Web of Science (February 16, 2025). Inclusion criteria encompassed peer-reviewed studies evaluating ChatGPT in pediatric healthcare (ages 0-18 and guardians). Of 475 screened articles, 58 met eligibility criteria. Two independent reviewers extracted data on study characteristics, intervention types, outcomes, and results. ChatGPT's primary applications were patient education (nโ€‰=โ€‰38), clinical decision-making (nโ€‰=โ€‰12), and clinical documentation (nโ€‰=โ€‰5). Accuracy was highest in patient education, where it generated educational materials and answered FAQs, though readability was often at a high school level, necessitating adaptation. Clinical documentation benefits included improved efficiency in drafting notes and discharge instructions. However, clinical decision-making and training (nโ€‰=โ€‰3) showed mixed accuracy, particularly in management recommendations and patient care plans.

CONCLUSION

ChatGPT demonstrates potential in enhancing physician efficiency and tailoring patient education in pediatric healthcare. However, most studies relied on observational designs, with only one quasi-experimental study. Further experimental research is required to evaluate AI's impact on pediatric care system effectiveness and patient outcomes. WHAT IS KNOWN: โ€ข AI has been increasingly integrated into pediatric healthcare, particularly in imaging, diagnostics, and decision support. โ€ข Prior studies have explored LLMs like ChatGPT in medical education and training. WHAT IS NEW: โ€ข This is the first systematic review assessing ChatGPT's broader applications in pediatric healthcare, including decision-making, patient education, clinical documentation, and training. โ€ข ChatGPT shows high accuracy in patient education and documentation but variable performance in decision-making and training, emphasizing the need for medical supervision.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์ฒด๊ณ„์  ๋ฌธํ—Œ๊ณ ์ฐฐ์€ ์†Œ์•„์ฒญ์†Œ๋…„๊ณผ ์˜์—ญ์—์„œ ChatGPT์˜ ์ž„์ƒ์  ํ™œ์šฉ๋„์™€ ์ •ํ™•๋„๋ฅผ ํ‰๊ฐ€ํ•œ ๊ฒฐ๊ณผ, ํ™˜์ž ๊ต์œก ์ž๋ฃŒ ์ƒ์„ฑ ๋ฐ ์ž„์ƒ ๋ฌธ์„œ ์ž‘์„ฑ ํšจ์œจํ™” ์ธก๋ฉด์—์„œ ๋†’์€ ์ž ์žฌ๋ ฅ์„ ํ™•์ธํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ๋ฐ ๊ต์œก ๋ถ„์•ผ์—์„œ๋Š” ์ •ํ™•๋„๊ฐ€ ๊ฐ€๋ณ€์ ์ด๋ฏ€๋กœ ์‹ค์ œ ์ ์šฉ ์‹œ ์˜๋ฃŒ์ง„์˜ ์ฒ ์ €ํ•œ ๊ฐ๋…์ด ํ•„์ˆ˜์ ์ž…๋‹ˆ๋‹ค. ํ–ฅํ›„ AI ๋„์ž…์ด ์†Œ์•„ ์ง„๋ฃŒ์˜ ์‹ค์งˆ์  ํšจ์œจ์„ฑ๊ณผ ํ™˜์ž ์˜ˆํ›„์— ๋ฏธ์น˜๋Š” ์˜ํ–ฅ์„ ๊ทœ๋ช…ํ•˜๊ธฐ ์œ„ํ•œ ์ถ”๊ฐ€์ ์ธ ์‹คํ—˜์  ์—ฐ๊ตฌ๊ฐ€ ์š”๊ตฌ๋ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

32Leveraging ChatGPT for Enhancing Learning in Radiology Resident Education.

2025-07-07Academic radiologyโญ Q1DOI 10.1016/j.acra.2025.06.019

RATIONALE AND

OBJECTIVE

Chat generative pre-trained transformer (ChatGPT) is a generative artificial intelligence chatbot based on a LLM at the forefront of technological development with promising applications in medical education. This study aims to evaluate the use of ChatGPT in generating board-style practice questions for radiology resident education.

METHODS

Multiple-choice questions (MCQs) were generated by ChatGPT from resident lecture transcripts using a custom prompt. 17 of the ChatGPT-generated MCQs were selected for inclusion in the study and randomly combined with 11 attending radiologist-written MCQs. For each MCQ, the 21 participating radiology residents answered the MCQ, rated the MCQ from 1-10 on effectiveness in reinforcing lecture material, and responded whether they thought an attending radiologist at their institution wrote the MCQ versus an alternative source.

RESULTS

Perceived MCQ quality was not significantly different between ChatGPT-generated (M=6.93, SD=0.29) and attending radiologist-written MCQs (M=7.08, SD=0.51) (p=0.15). MCQ correct answer percentages did not significantly differ between ChatGPT-generated (M=57%, SD=20%) and attending radiologist-written MCQs (M=59%, SD=25%) (p=0.78). The percentage of MCQs thought to be written by an attending radiologist was significantly different between ChatGPT-generated (M=57%, SD=13%) and attending radiologist-written MCQs (M=71%, SD=20%) (p=0.04).

CONCLUSION

LLMs such as ChatGPT demonstrate potential in generating and presenting educational material for radiology education, and their use should be explored further on a larger scale.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜์ƒ์˜ํ•™๊ณผ ์ „๊ณต์˜ ๊ต์œก์„ ์œ„ํ•œ ๋ณด๋“œ ์‹œํ—˜ ํ˜•์‹์˜ ๊ฐ๊ด€์‹ ๋ฌธํ•ญ ์ƒ์„ฑ์— ChatGPT์˜ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ChatGPT๊ฐ€ ์ƒ์„ฑํ•œ ๋ฌธํ•ญ์˜ ๊ต์œก์  ํšจ๊ณผ์™€ ์ •๋‹ต๋ฅ ์€ ์ „๋ฌธ์˜๊ฐ€ ์ž‘์„ฑํ•œ ๋ฌธํ•ญ๊ณผ ์œ ์˜๋ฏธํ•œ ์ฐจ์ด๊ฐ€ ์—†์—ˆ์œผ๋ฉฐ, ์ „๊ณต์˜๋“ค ๋˜ํ•œ ๋‘ ๋ฌธํ•ญ ๊ฐ„์˜ ํ’ˆ์งˆ ์ฐจ์ด๋ฅผ ํฌ๊ฒŒ ๋А๋ผ์ง€ ๋ชปํ–ˆ์Šต๋‹ˆ๋‹ค. ๋”ฐ๋ผ์„œ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์€ ์˜์ƒ์˜ํ•™ ๊ต์œก ์ž๋ฃŒ๋ฅผ ์ƒ์„ฑํ•˜๋Š” ๋ฐ ์žˆ์–ด ํšจ๊ณผ์ ์ธ ๋ณด์กฐ ๋„๊ตฌ๋กœ ํ™œ์šฉ๋  ์ž ์žฌ๋ ฅ์ด ์ถฉ๋ถ„ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

33Performance of open-source and proprietary large language models in generating patient-friendly radiology chest CT reports.

2025-07-05Clinical imaging๐Ÿ”ท Q2DOI 10.1016/j.clinimag.2025.110557

RATIONALE AND

OBJECTIVE

Large Language Models (LLMs) show promise for generating patient-friendly radiology reports, but the performance of open-source versus proprietary LLMs needs assessment. To compare open-source and proprietary LLMs in generating patient-friendly radiology reports from chest CTs using quantitative readability metrics and qualitative assessments by radiologists.

METHODS

Fifty chest CT reports were processed by seven LLMs: three open-source models (Llama-3-70b, Mistral-7b, Mixtral-8x7b) and four proprietary models (GPT-4, GPT-3.5-Turbo, Claude-3-Opus, Gemini-Ultra). Simplification was evaluated using five quantitative readability metrics. Three radiologists rated patient-friendliness on a five-point Likert scale across five criteria. Content and coherence errors were counted. Inter-rater reliability and differences among models were statistically assessed.

RESULTS

Inter-rater reliability was substantial to near perfect (ฮบย =ย 0.76-0.86). Qualitatively, Llama-3-70b was non-inferior to leading proprietary models in 4/5 categories. GPT-3.5-Turbo showed the best overall readability, outperforming GPT-4 in two metrics. Llama-3-70b outperformed GPT-3.5-Turbo on the CLI (pย =ย 0.006). Claude-3-Opus and Gemini-Ultra scored lower on readability but were rated highly in qualitative assessments. Claude-3-Opus maintained perfect factual accuracy. Claude-3-Opus and GPT-4 outperformed Llama-3-70b in emotional sensitivity (90.0ย % vs 46.0ย %, pย <ย 0.001).

CONCLUSION

Llama-3-70b shows strong potential in generating quality, patient-friendly radiology reports, challenging proprietary models. With further adaptation, open-source LLMs could advance patient-friendly reporting technology.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ํ‰๋ถ€ CT ํŒ๋…๋ฌธ์„ ํ™˜์ž ์นœํ™”์ ์ธ ์–ธ์–ด๋กœ ๋ณ€ํ™˜ํ•˜๋Š” ๋ฐ ์žˆ์–ด ์˜คํ”ˆ์†Œ์Šค ๋ฐ ์ƒ์šฉ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ์ •๋Ÿ‰์  ๊ฐ€๋…์„ฑ ์ง€ํ‘œ์™€ ์˜์ƒ์˜ํ•™ ์ „๋ฌธ์˜ ํ‰๊ฐ€๋ฅผ ํ†ตํ•ด ๋น„๊ต ๋ถ„์„ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ์˜คํ”ˆ์†Œ์Šค ๋ชจ๋ธ์ธ Llama-3-70b๊ฐ€ ์ฃผ์š” ์ƒ์šฉ ๋ชจ๋ธ๋“ค๊ณผ ๋Œ€๋“ฑํ•œ ์ˆ˜์ค€์˜ ํŒ๋…๋ฌธ ์ƒ์„ฑ ๋Šฅ๋ ฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ํŠนํžˆ Claude-3-Opus๋Š” ๋†’์€ ์‚ฌ์‹ค์  ์ •ํ™•๋„๋ฅผ ๋‚˜ํƒ€๋ƒˆ์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ ์˜คํ”ˆ์†Œ์Šค LLM์€ ํ–ฅํ›„ ์ถ”๊ฐ€์ ์ธ ์ตœ์ ํ™”๋ฅผ ํ†ตํ•ด ํ™˜์ž ์ค‘์‹ฌ์˜ ์˜๋ฃŒ ๋ณด๊ณ ์„œ ์ž‘์„ฑ ๊ธฐ์ˆ ์„ ๋ฐœ์ „์‹œํ‚ฌ ์ˆ˜ ์žˆ๋Š” ์œ ๋งํ•œ ๋Œ€์•ˆ์œผ๋กœ ํ™•์ธ๋˜์—ˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

34Radiology report generation using automatic keyword adaptation, frequency-based multi-label classification and text-to-text large language models.

2025-07-03Computers in biology and medicineโญ Q1DOI 10.1016/j.compbiomed.2025.110625
BACKGROUND

Radiology reports are essential in medical imaging, providing critical insights for diagnosis, treatment, and patient management by bridging the gap between radiologists and referring physicians. However, the manual generation of radiology reports is time-consuming and labor-intensive, leading to inefficiencies and delays in clinical workflows, particularly as case volumes increase. Although deep learning approaches have shown promise in automating radiology report generation, existing methods, particularly those based on the encoder-decoder framework, suffer from significant limitations. These include a lack of explainability due to black-box features generated by encoder and limited adaptability to diverse clinical settings.

METHODS

In this study, we address these challenges by proposing a novel deep learning framework for radiology report generation that enhances explainability, accuracy, and adaptability. Our approach replaces traditional black-box features in computer vision with transparent keyword lists, improving the interpretability of the feature extraction process. To generate these keyword lists, we apply a multi-label classification technique, which is further enhanced by an automatic keyword adaptation mechanism. This adaptation dynamically configures the multi-label classification to better adapt specific clinical environments, reducing the reliance on manually curated reference keyword lists and improving model adaptability across diverse datasets. We also introduce a frequency-based multi-label classification strategy to address the issue of keyword imbalance, ensuring that rare but clinically significant terms are accurately identified. Finally, we leverage a pre-trained text-to-text large language model (LLM) to generate human-like, clinically relevant radiology reports from the extracted keyword lists, ensuring linguistic quality and clinical coherence.

RESULTS

We evaluate our method using two public datasets, IU-XRay and MIMIC-CXR, demonstrating superior performance over state-of-the-art methods. Our framework not only improves the accuracy and reliability of radiology report generation but also enhances the explainability of the process, fostering greater trust and adoption of AI-driven solutions in clinical practice. Comprehensive ablation studies confirm the robustness and effectiveness of each component, highlighting the significant contributions of our framework to advancing automated radiology reporting.

CONCLUSION

In conclusion, we developed a novel deep-learning based radiology report generation method for preparing high-quality and explainable radiology report for chest X-ray images using the multi-label classification and a text-to-text large language model. Our method could address the lack of explainability in the current workflow and provide a clear and flexible automated pipeline to reduce the workload of radiologists and support the further applications related to Human-AI interactive communications.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜์ƒ ์˜ํ•™ ๋ณด๊ณ ์„œ ์ƒ์„ฑ์˜ ์„ค๋ช… ๊ฐ€๋Šฅ์„ฑ๊ณผ ์ •ํ™•๋„๋ฅผ ๋†’์ด๊ธฐ ์œ„ํ•ด ํˆฌ๋ช…ํ•œ ํ‚ค์›Œ๋“œ ์ถ”์ถœ ๊ธฐ๋ฐ˜์˜ ๋‹ค์ค‘ ๋ ˆ์ด๋ธ” ๋ถ„๋ฅ˜์™€ ์‚ฌ์ „ ํ•™์Šต๋œ ๊ฑฐ๋Œ€ ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ๊ฒฐํ•ฉํ•œ ์ƒˆ๋กœ์šด ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ์ œ์•ˆํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋นˆ๋„ ๊ธฐ๋ฐ˜์˜ ํ‚ค์›Œ๋“œ ์ ์‘ํ˜• ๋ฉ”์ปค๋‹ˆ์ฆ˜์„ ํ†ตํ•ด ์ž„์ƒ์  ๋ถˆ๊ท ํ˜• ๋ฌธ์ œ๋ฅผ ํ•ด๊ฒฐํ•˜๊ณ  ์ž„์ƒ ํ™˜๊ฒฝ์— ์œ ์—ฐํ•˜๊ฒŒ ๋Œ€์‘ํ•˜๋„๋ก ์„ค๊ณ„๋˜์—ˆ์Šต๋‹ˆ๋‹ค. IU-XRay ๋ฐ MIMIC-CXR ๋ฐ์ดํ„ฐ์…‹์„ ํ†ตํ•œ ๊ฒ€์ฆ ๊ฒฐ๊ณผ, ๊ธฐ์กด ๋ชจ๋ธ ๋Œ€๋น„ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ๊ณผ ๋†’์€ ์„ค๋ช… ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ•˜์—ฌ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜์˜ ์—…๋ฌด ํšจ์œจ์„ฑ์„ ๊ฐœ์„ ํ•  ์ˆ˜ ์žˆ๋Š” ๊ฐ€๋Šฅ์„ฑ์„ ํ™•์ธํ•˜์˜€์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

35RadGPT: A System Based on a Large Language Model That Generates Sets of Patient-Centered Materials to Explain Radiology Report Information.

2025-06-10Journal of the American College of Radiology : JACRโญ Q1DOI 10.1016/j.jacr.2025.06.013
OBJECTIVE

The 21st Century Cures Act final rule requires that patients have real-time access to their radiology reports, which contain technical language. The objective of this study to was to use a novel system called RadGPT, which integrates concept extraction and a large language model (LLM), to help patients understand their radiology reports.

METHODS

RadGPT generated 150 concept explanations and 390 question-and-answer pairs from 30 radiology report impressions from between 2012 and 2020. The extracted concepts were used to create concept-based explanations, as well as concept-based question-and-answer pairs for which questions were generated using either a fixed template or an LLM. Additionally, report-based question-and-answer pairs were generated directly from the impression using an LLM without concept extraction. One board-certified radiologist and four radiology residents rated the material quality using a standardized rubric.

RESULTS

Concept-based LLM-generated questions were of significantly higher quality than concept-based template-generated questions (P < .001). Excluding those template-based question-and-answer pairs from further analysis, nearly all (>95%) of RadGPT-generated materials were rated highly, with at least 50% receiving the highest possible ranking from all five raters. No answers or explanations were rated as likely to affect the safety or effectiveness of patient care. Report-level LLM-based questions and answers were rated particularly highly, with 92% of report-level LLM-based questions and 61% of the corresponding report-level answers receiving the highest rating from all raters.

CONCLUSION

The educational tool RadGPT generated high-quality explanations and question-and-answer pairs that were personalized for each radiology report, unlikely to produce harmful explanations, and likely to enhance patient understanding of radiology information.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
RadGPT๋Š” ๊ฐœ๋… ์ถ”์ถœ๊ณผ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์„ ํ†ตํ•ฉํ•˜์—ฌ ์ „๋ฌธ ์šฉ์–ด๊ฐ€ ํฌํ•จ๋œ ์˜์ƒ์˜ํ•™ ํŒ๋…๋ฌธ์„ ํ™˜์ž๊ฐ€ ์ดํ•ดํ•˜๊ธฐ ์‰ฌ์šด ์„ค๋ช…๊ณผ ์งˆ์˜์‘๋‹ต ํ˜•ํƒœ๋กœ ๋ณ€ํ™˜ํ•˜๋Š” ์‹œ์Šคํ…œ์ž…๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, RadGPT๊ฐ€ ์ƒ์„ฑํ•œ ์ฝ˜ํ…์ธ ๋Š” ์˜๋ฃŒ์ง„์œผ๋กœ๋ถ€ํ„ฐ ๋†’์€ ํ’ˆ์งˆ๊ณผ ์•ˆ์ „์„ฑ์„ ์ธ์ •๋ฐ›์•˜์œผ๋ฉฐ, ํŠนํžˆ LLM์„ ํ™œ์šฉํ•œ ๋ณด๊ณ ์„œ ๋‹จ์œ„์˜ ์งˆ์˜์‘๋‹ต์ด ํ™˜์ž์˜ ํŒ๋…๋ฌธ ์ดํ•ด๋„๋ฅผ ๋†’์ด๋Š” ๋ฐ ํšจ๊ณผ์ ์ธ ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

36Integrating Large language models into radiology workflow: Impact of generating personalized report templates from summary.

2025-05-25European journal of radiologyโญ Q1DOI 10.1016/j.ejrad.2025.112198
OBJECTIVE

To evaluate feasibility of large language models (LLMs) to convert radiologist-generated report summaries into personalized report templates, and assess its impact on scan reporting time and quality.

METHODS

In this retrospective study, 100 CT scans from oncology patients were randomly divided into two equal sets. Two radiologists generated conventional reports for one set and summary reports for the other, and vice versa. Three LLMs - GPT-4, Google Gemini, and Claude Opus - generated complete reports from the summaries using institution-specific generic templates. Two expert radiologists qualitatively evaluated the radiologist summaries and LLM-generated reports using the ACR RADPEER scoring system, using conventional radiologist reports as reference. Reporting time for conventional versus summary-based reports was compared, and LLM-generated reports were analyzed for errors. Quantitative similarity and linguistic metrics were computed to assess report alignment across models with the original radiologist-generated report summaries. Statistical analyses were performed using Python 3.0 to identify significant differences in reporting times, error rates and quantitative metrics.

RESULTS

The average reporting time was significantly shorter for summary method (6.76ย min) compared to conventional method (8.95ย min) (pย <ย 0.005). Among the 100 radiologist summaries, 10 received RADPEER scores worse than 1, with three deemed to have clinically significant discrepancies. Only one LLM-generated report received a worse RADPEER score than its corresponding summary. Error frequencies among LLM-generated reports showed no significant differences across models, with template-related errors being most common (ฯ‡2ย =ย 1.146, pย =ย 0.564). Quantitative analysis indicated significant differences in similarity and linguistic metrics among the three LLMs (pย <ย 0.05), reflecting unique generation patterns.

CONCLUSION

Summary-based scan reporting along with use of LLMs to generate complete personalized report templates can shorten reporting time while maintaining the report quality. However, there remains a need for human oversight to address errors in the generated reports. RELEVANCE STATEMENT: Summary-based reporting of radiological studies along with the use of large language models to generate tailored reports using generic templates has the potential to make the workflow more efficient by shortening the reporting time while maintaining the quality of reporting.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜๊ฐ€ ์ž‘์„ฑํ•œ ์š”์•ฝ๋ฌธ์„ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ํ†ตํ•ด ์ •์‹ ํŒ๋…๋ฌธ์œผ๋กœ ๋ณ€ํ™˜ํ•˜๋Š” ๋ฐฉ์‹์˜ ํšจ์œจ์„ฑ๊ณผ ์ •ํ™•์„ฑ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ์š”์•ฝ ๊ธฐ๋ฐ˜ ํŒ๋… ๋ฐฉ์‹์€ ๊ธฐ์กด ๋ฐฉ์‹ ๋Œ€๋น„ ํŒ๋… ์‹œ๊ฐ„์„ ์œ ์˜๋ฏธํ•˜๊ฒŒ ๋‹จ์ถ•ํ•˜๋ฉด์„œ๋„ ํŒ๋… ํ’ˆ์งˆ์„ ์•ˆ์ •์ ์œผ๋กœ ์œ ์ง€ํ•˜๋Š” ๊ฒƒ์œผ๋กœ ๋‚˜ํƒ€๋‚ฌ์Šต๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ์ƒ์„ฑ๋œ ํŒ๋…๋ฌธ์˜ ์˜ค๋ฅ˜๋ฅผ ๋ณด์™„ํ•˜๊ธฐ ์œ„ํ•ด ์ „๋ฌธ์˜์˜ ์ตœ์ข… ๊ฒ€์ˆ˜์™€ ์ธ๊ฐ„์˜ ๊ฐœ์ž…์ด ํ•„์ˆ˜์ ์ž„์„ ํ™•์ธํ•˜์˜€์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

37Precision health monitoring in spaceflight with integration of lower body negative pressure and advanced large language model artificial intelligence.

2025-05-20Life sciences in space researchโญ Q1DOI 10.1016/j.lssr.2025.05.010

Long-term exposure to microgravity influences musculoskeletal health and enhances the likelihood of sustaining orthopedic injuries while on a microgravity mission and upon return to Earth. Although countermeasures are being investigated to alleviate some risks of injury, such as resistive (or weight) exercise and Lower Body Negative Pressure (LBNP), evidence is accumulating that current paradigms do not ensure the safety or health of astronauts because of a lack of in-flight diagnostic methods, in which load/diagnostic metrics can be assessed over time. Here, we suggest the integration of a new vision-language large language model (DeepSeek-VL) as a potential autonomous diagnostic agent for monitoring musculoskeletal health in a microgravity environment. DeepSeek-VL will autonomously analyze radiographic data and biomechanical data streamed from a LBNP device. Determinations will be made based on lost or compromised density in bone, lost joint-centered stability, or ineffective loading patterns - providing personalized and specific feedback regarding musculoskeletal health with the astronaut as the primary user. Unlike conventional reporting approaches that rely on cross-institutional analysis by household experts, DeepSeek-VL allows for real-time, and autonomous interpretation of musculoskeletal imaging metrics (and physiological metrics) for on-time personalized countermeasure development. Here, we review architectural adaptations including microgravity specific samplings of data, training protocols and implications of deployment in the ISS. We anticipate DeepSeek's timely development of flight-ready diagnostic reporting will facilitate in-flight/systematic monitoring of musculoskeletal health and safety, especially for astronauts undergoing load management training (e.g., LBNP) and ensure effectiveness of countermeasures, their outputs. We will address methods to circumvent limitations and barriers to risk, and establish the importance of a federated, adaptive, and resilient AI-based platform to mitigate risk for astronaut musculoskeletal health during extended missions. Finally, we address some considerations for terrestrial model and a healthcare authority within a current context of growing importance for effective orthopedic healthcare.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์žฅ๊ธฐ ์šฐ์ฃผ ๋น„ํ–‰ ์ค‘ ๋ฐœ์ƒํ•˜๋Š” ๊ทผ๊ณจ๊ฒฉ๊ณ„ ์†์ƒ์„ ๋ฐฉ์ง€ํ•˜๊ธฐ ์œ„ํ•ด ํ•˜์ฒด ์Œ์•• ์žฅ์น˜(LBNP)์™€ ๊ฑฐ๋Œ€ ์‹œ๊ฐ-์–ธ์–ด ๋ชจ๋ธ(DeepSeek-VL)์„ ํ†ตํ•ฉํ•œ ์ž์œจ ์ง„๋‹จ ์‹œ์Šคํ…œ์„ ์ œ์•ˆํ•ฉ๋‹ˆ๋‹ค. ์ด ์‹œ์Šคํ…œ์€ ๋ฐฉ์‚ฌ์„  ๋ฐ ์ƒ์ฒด์—ญํ•™ ๋ฐ์ดํ„ฐ๋ฅผ ์‹ค์‹œ๊ฐ„์œผ๋กœ ๋ถ„์„ํ•˜์—ฌ ์šฐ์ฃผ๋น„ํ–‰์‚ฌ์˜ ๊ทผ๊ณจ๊ฒฉ๊ณ„ ์ƒํƒœ๋ฅผ ํ‰๊ฐ€ํ•˜๊ณ  ๋งž์ถคํ˜• ๋Œ€์‘์ฑ…์„ ์ œ๊ณตํ•จ์œผ๋กœ์จ, ๊ธฐ์กด์˜ ์‚ฌํ›„ ๋ณด๊ณ  ๋ฐฉ์‹๋ณด๋‹ค ํšจ์œจ์ ์ธ ์‹ค์‹œ๊ฐ„ ๊ฑด๊ฐ• ๋ชจ๋‹ˆํ„ฐ๋ง์„ ๊ฐ€๋Šฅํ•˜๊ฒŒ ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

38Enhancing Bidirectional Encoder Representations From Transformers (BERT) With Frame Semantics to Extract Clinically Relevant Information From German Mammography Reports: Algorithm Development and Validation.

2025-04-25Journal of medical Internet researchโญ Q1DOI 10.2196/68427
BACKGROUND

Structured reporting is essential for improving the clarity and accuracy of radiological information. Despite its benefits, the European Society of Radiology notes that it is not widely adopted. For example, while structured reporting frameworks such as the Breast Imaging Reporting and Data System provide standardized terminology and classification for mammography findings, radiology reports still mostly comprise free-text sections. This variability complicates the systematic extraction of key clinical data. Moreover, manual structuring of reports is time-consuming and prone to inconsistencies. Recent advancements in large language models have shown promise for clinical information extraction by enabling models to understand contextual nuances in medical text. However, challenges such as domain adaptation, privacy concerns, and generalizability remain. To address these limitations,ย frame semantics offers an approach to information extraction grounded in computational linguistics, allowing a structured representation of clinically relevant concepts.

OBJECTIVE

This study explores the combination of Bidirectional Encoder Representations from Transformers (BERT) architecture with the linguistic concept of frame semantics to extract and normalize information from free-text mammography reports.

METHODS

After creating an annotated corpus of 210 German reports for fine-tuning, we generate several BERT model variants by applying 3 pretraining strategies to hospital data. Afterward, a fact extraction pipeline is built, comprising an extractive question-answering model and a sequence labeling model. We quantitatively evaluate all model variants using common evaluation metrics (model perplexity, Stanford Question Answering Dataset 2.0 [SQuAD_v2], seqeval) and perform a qualitative clinician evaluation of the entire pipeline on a manually generated synthetic dataset of 21 reports, as well as a comparison with a generative approach following best practice prompting techniques using the open-source Llama 3.3 model (Meta).

RESULTS

Our system is capable of extracting 14 fact types and 40 entities from the clinical findings section of mammography reports. Further pretraining on hospital data reduced model perplexity, although it did not significantly impact the 2 downstream tasks. We achieved average F1-scores of 90.4% and 81% for question answering and sequence labeling, respectively (best pretraining strategy). Qualitative evaluation of the pipeline based on synthetic data shows an overall precision of 96.1% and 99.6% for facts and entities, respectively. In contrast, generative extraction shows an overall precision of 91.2% and 87.3% for facts and entities, respectively. Hallucinations and extraction inconsistencies were observed.

CONCLUSION

This study demonstrates that frame semantics provides a robust and interpretable framework for automating structured reporting. By leveraging frame semantics, the approach enables customizable information extraction and supports generalization to diverse radiological domains and clinical contexts with additional annotation efforts. Furthermore, the BERT-based model architecture allows for efficient, on-premise deployment, ensuring data privacy. Future research should focus on validating the model's generalizability across external datasets and different report types to ensure its broader applicability in clinical practice.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” BERT ๋ชจ๋ธ์— ํ”„๋ ˆ์ž„ ์˜๋ฏธ๋ก (frame semantics)์„ ๊ฒฐํ•ฉํ•˜์—ฌ ๋…์ผ์–ด ์œ ๋ฐฉ ์ดฌ์˜์ˆ  ์ž์œ  ํ…์ŠคํŠธ ๋ณด๊ณ ์„œ์—์„œ ์ž„์ƒ ์ •๋ณด๋ฅผ ์ถ”์ถœํ•˜๊ณ  ๊ตฌ์กฐํ™”ํ•˜๋Š” ์•Œ๊ณ ๋ฆฌ์ฆ˜์„ ๊ฐœ๋ฐœ ๋ฐ ๊ฒ€์ฆํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์งˆ์˜์‘๋‹ต ๋ฐ ์‹œํ€€์Šค ๋ผ๋ฒจ๋ง ํŒŒ์ดํ”„๋ผ์ธ์„ ํ†ตํ•ด 14๊ฐœ ์‚ฌ์‹ค ์œ ํ˜•๊ณผ 40๊ฐœ ์—”ํ‹ฐํ‹ฐ๋ฅผ ๋†’์€ ์ •ํ™•๋„(F1-score 81~90.4%)๋กœ ์ถ”์ถœํ•˜์˜€์œผ๋ฉฐ, ์ƒ์„ฑํ˜• ๋ชจ๋ธ ๋Œ€๋น„ ํ™˜๊ฐ ํ˜„์ƒ์ด ์ ๊ณ  ๋ฐ์ดํ„ฐ ๋ณด์•ˆ์ด ๋ณด์žฅ๋˜๋Š” ์˜จํ”„๋ ˆ๋ฏธ์Šค ํ™˜๊ฒฝ์— ์ตœ์ ํ™”๋œ ์„ฑ๋Šฅ์„ ์ž…์ฆํ–ˆ์Šต๋‹ˆ๋‹ค. ์ด๋Š” ํ”„๋ ˆ์ž„ ์˜๋ฏธ๋ก ์ด ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์˜ ๊ตฌ์กฐํ™”๋ฅผ ์ž๋™ํ™”ํ•˜๊ณ  ์ž„์ƒ์  ํ•ด์„ ๊ฐ€๋Šฅ์„ฑ์„ ๋†’์ด๋Š” ๋ฐ ํšจ๊ณผ์ ์ธ ํ”„๋ ˆ์ž„์›Œํฌ์ž„์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

39A scoping review of advancements in machine learning for glaucoma: current trends and future direction.

2025-04-24Frontiers in medicineโญ Q1DOI 10.3389/fmed.2025.1573329
BACKGROUND

Machine learning technology has demonstrated significant potential in glaucoma research, particularly in early diagnosis, predicting disease progression, evaluating treatment responses, and developing personalized treatment strategies. The application of machine learning not only enhances the understanding of the pathological mechanism of glaucoma and optimizes the diagnostic process but also provides patients with accurate medical services.

METHODS

This study aimed to describe the current state of research, highlight directions for further development, and identify potential trends for improvement. This review was conducted following the scoping review of the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) extension to showcase advancements in the application of machine learning in glaucoma research and treatment.

RESULTS

We employed a comprehensive search strategy to retrieve literature from the Web of Science Core Collection database, ultimately including 3,581 articles in the analysis. Through data analysis, we identified current research hotspots, noted differences in researchers' attitudes and opinions, and predicted potential future development trends.

CONCLUSION

We divided the research topics into six categories, clearly identifying "eye diseases", "retinal fundus imaging" and "risk factors" as the key terms for the development of this field. These findings signify the promising prospects of machine learning, particularly when integrated with multimodal technologies and large language models, to enhance the diagnosis and treatment of glaucoma.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋…น๋‚ด์žฅ ๋ถ„์•ผ์—์„œ ๋จธ์‹ ๋Ÿฌ๋‹์˜ ์ตœ์‹  ์—ฐ๊ตฌ ๋™ํ–ฅ๊ณผ ๋ฐœ์ „ ๋ฐฉํ–ฅ์„ ํŒŒ์•…ํ•˜๊ธฐ ์œ„ํ•ด 3,581ํŽธ์˜ ๋ฌธํ—Œ์„ ๋Œ€์ƒ์œผ๋กœ ์Šค์ฝ”ํ•‘ ๋ฆฌ๋ทฐ๋ฅผ ์ˆ˜ํ–‰ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ๋ง๋ง‰ ์•ˆ์ € ์˜์ƒ๊ณผ ์œ„ํ—˜ ์ธ์ž ๋ถ„์„์ด ํ•ต์‹ฌ ์—ฐ๊ตฌ ๋ถ„์•ผ๋กœ ํ™•์ธ๋˜์—ˆ์œผ๋ฉฐ, ํ–ฅํ›„ ๋‹ค์ค‘ ๋ชจ๋‹ฌ ๊ธฐ์ˆ  ๋ฐ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ๊ณผ์˜ ํ†ตํ•ฉ์ด ๋…น๋‚ด์žฅ์˜ ์กฐ๊ธฐ ์ง„๋‹จ๊ณผ ์ •๋ฐ€ ์น˜๋ฃŒ๋ฅผ ๊ณ ๋„ํ™”ํ•  ๊ฒƒ์œผ๋กœ ์ „๋ง๋ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

40Evaluating large language models in echocardiography reporting: opportunities and challenges.

2025-03-31European heart journal. Digital healthโญ Q1DOI 10.1093/ehjdh/ztae086
OBJECTIVE

The increasing need for diagnostic echocardiography tests presents challenges in preserving the quality and promptness of reports. While Large Language Models (LLMs) have proven effective in summarizing clinical texts, their application in echo remains underexplored. METHODS AND

RESULTS

Adult echocardiography studies, conducted at the Mayo Clinic from 1 January 2017 to 31 December 2017, were categorized into two groups: development (all Mayo locations except Arizona) and Arizona validation sets. We adapted open-source LLMs (Llama-2, MedAlpaca, Zephyr, and Flan-T5) using In-Context Learning and Quantized Low-Rank Adaptation fine-tuning (FT) for echo report summarization from 'Findings' to 'Impressions.' Against cardiologist-generated Impressions, the models' performance was assessed both quantitatively with automatic metrics and qualitatively by cardiologists. The development dataset included 97 506 reports from 71 717 unique patients, predominantly male (55.4%), with an average age of 64.3 ยฑ 15.8 years. EchoGPT, a fine-tuned Llama-2 model, outperformed other models with win rates ranging from 87% to 99% in various automatic metrics, and produced reports comparable to cardiologists in qualitative review (significantly preferred in conciseness (P < 0.001), with no significant preference in completeness, correctness, and clinical utility). Correlations between automatic and human metrics were fair to modest, with the best being RadGraph F1 scores vs. clinical utility (r = 0.42) and automatic metrics showed insensitivity (0-5% drop) to changes in measurement numbers.

CONCLUSION

EchoGPT can generate draft reports for human review and approval, helping to streamline the workflow. However, scalable evaluation approaches dedicated to echo reports remains necessary.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์‹ฌ์žฅ์ดˆ์ŒํŒŒ ๊ฒฐ๊ณผ ๋ณด๊ณ ์„œ์˜ ํšจ์œจ์ ์ธ ์ž‘์„ฑ์„ ์œ„ํ•ด ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์„ ํ™œ์šฉํ•œ 'EchoGPT'๋ฅผ ๊ฐœ๋ฐœํ•˜๊ณ  ๊ทธ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. Llama-2 ๊ธฐ๋ฐ˜์˜ EchoGPT๋Š” ์ „๋ฌธ์˜๊ฐ€ ์ž‘์„ฑํ•œ ๊ฒฐ๊ณผ์ง€์™€ ๋น„๊ตํ–ˆ์„ ๋•Œ ์ž„์ƒ์  ์œ ์šฉ์„ฑ๊ณผ ์ •ํ™•๋„ ์ธก๋ฉด์—์„œ ๋Œ€๋“ฑํ•œ ์ˆ˜์ค€์„ ๋ณด์˜€์œผ๋ฉฐ, ํŠนํžˆ ๋ณด๊ณ ์„œ์˜ ๊ฐ„๊ฒฐ์„ฑ ๋ฉด์—์„œ ์šฐ์ˆ˜ํ•œ ํ‰๊ฐ€๋ฅผ ๋ฐ›์•˜์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ EchoGPT๋Š” ์ดˆ์ŒํŒŒ ํŒ๋… ์›Œํฌํ”Œ๋กœ์šฐ๋ฅผ ๊ฐœ์„ ํ•  ์ˆ˜ ์žˆ๋Š” ์ž ์žฌ๋ ฅ์„ ์ง€๋…”์œผ๋‚˜, ์ž„์ƒ ์ ์šฉ์„ ์œ„ํ•ด์„œ๋Š” ์ •๊ตํ•œ ํ‰๊ฐ€ ์ฒด๊ณ„ ๋งˆ๋ จ์ด ์„ ํ–‰๋˜์–ด์•ผ ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

41Impact of hospital-specific domain adaptation on BERT-based models to classify neuroradiology reports.

2025-03-17European radiologyโญ Q1DOI 10.1007/s00330-025-11500-9
OBJECTIVE

To determine the effectiveness of hospital-specific domain adaptation through masked language modelling (MLM) on BERT-based models' performance in classifying neuroradiology reports, and to compare these models with open-source large language models (LLMs).

METHODS

This retrospective study (2008-2019) utilised 126,556 and 86,032 MRI brain reports from two tertiary hospitals-King's College Hospital (KCH) and Guys and St Thomas' Trust (GSTT). Various BERT-based models, including RoBERTa, BioBERT and RadBERT, underwent MLM on unlabelled reports from these centres. The downstream tasks were binary abnormality classification and multi-label classification. Performances of models with and without hospital-specific domain adaptation were compared against each other and LLMs on internal (KCH) and external (GSTT) hold-out test sets. Model performances for binary classification were compared using 2-way and 1-way ANOVA.

RESULTS

All models that underwent hospital-specific domain adaptation performed better than their baseline counterparts (all p-valuesโ€‰<โ€‰0.001). For binary classification, MLM on all available unlabelled reports (194,467 reports) yielded the highest balanced accuracies (KCH: mean 97.0โ€‰ยฑโ€‰0.4% (standard deviation), GSTT: 95.5โ€‰ยฑโ€‰1.0%), after which no differences between BERT-based models remained (1-way ANOVA, p-valuesโ€‰>โ€‰0.05). There was a log-linear relationship between the number of reports and performance. LLama-3.0 70B was the best-performing LLM (KCH: 97.1%, GSTT: 94.0%). Multi-label classification demonstrated consistent performance improvements from MLM for all abnormality categories.

CONCLUSION

Hospital-specific domain adaptation should be considered best practice when deploying BERT-based models in new clinical settings. When labelled data is scarce or unavailable, LLMs can serve as a viable alternative, assuming adequate computational power is accessible. KEY POINTS: Question BERT-based models can classify radiology reports, but it is unclear if there is any incremental benefit from additional hospital-specific domain adaptation. Findings Hospital-specific domain adaptation resulted in the highest BERT-based model accuracies and performance scaled log-linearly with the number of reports. Clinical relevance BERT-based models after hospital-specific domain adaptation achieve the best classification results provided sufficient high-quality training labels. When labelled data is scarce, LLMs such as Llama-3.0 70B are a viable alternative provided there are sufficient computational resources.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์‹ ๊ฒฝ์˜์ƒ์˜ํ•™ ํŒ๋…๋ฌธ ๋ถ„๋ฅ˜๋ฅผ ์œ„ํ•ด BERT ๊ธฐ๋ฐ˜ ๋ชจ๋ธ์— ๋ณ‘์›๋ณ„ ๋„๋ฉ”์ธ ์ ์‘(MLM)์„ ์ ์šฉํ•œ ํšจ๊ณผ๋ฅผ ํ‰๊ฐ€ํ•˜๊ณ , ์ด๋ฅผ ์˜คํ”ˆ์†Œ์Šค ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)๊ณผ ๋น„๊ตํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ๋ณ‘์›๋ณ„ ๋„๋ฉ”์ธ ์ ์‘์„ ๊ฑฐ์นœ BERT ๋ชจ๋ธ์ด ๋ฒ ์ด์Šค๋ผ์ธ ๋ชจ๋ธ๋ณด๋‹ค ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ๋ฐ์ดํ„ฐ ์–‘๊ณผ ์„ฑ๋Šฅ ๊ฐ„์˜ ๋กœ๊ทธ ์„ ํ˜• ๊ด€๊ณ„๊ฐ€ ํ™•์ธ๋˜์—ˆ์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ ์ถฉ๋ถ„ํ•œ ๋ผ๋ฒจ๋ง ๋ฐ์ดํ„ฐ๊ฐ€ ์žˆ์„ ๋•Œ๋Š” ๋„๋ฉ”์ธ ์ ์‘๋œ BERT ๋ชจ๋ธ์ด, ๋ฐ์ดํ„ฐ๊ฐ€ ๋ถ€์กฑํ•  ๋•Œ๋Š” LLM์ด ์ž„์ƒ ํ˜„์žฅ์—์„œ ํšจ๊ณผ์ ์ธ ๋Œ€์•ˆ์ด ๋  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

42Artificial Intelligence in Relation to Accurate Information and Tasks in Gynecologic Oncology and Clinical Medicine-Dunning-Kruger Effects and Ultracrepidarianism.

2025-03-15Diagnostics (Basel, Switzerland)๐Ÿ”ท Q2DOI 10.3390/diagnostics15060735

Publications on the application of artificial intelligence (AI) to many situations, including those in clinical medicine, created in 2023-2024 are reviewed here. Because of the short time frame covered, here, it is not possible to conduct exhaustive analysis as would be the case in meta-analyses or systematic reviews. Consequently, this literature review presents an examination of narrative AI's application in relation to contemporary topics related to clinical medicine. The landscape of the findings reviewed here span 254 papers published in 2024 topically reporting on AI in medicine, of which 83 articles are considered in the present review because they contain evidence-based findings. In particular, the types of cases considered deal with AI accuracy in initial differential diagnoses, cancer treatment recommendations, board-style exams, and performance in various clinical tasks, including clinical imaging. Importantly, summaries of the validation techniques used to evaluate AI findings are presented. This review focuses on AIs that have a clinical relevancy evidenced by application and evaluation in clinical publications. This relevancy speaks to both what has been promised and what has been delivered by various AI systems. Readers will be able to understand when generative AI may be expressing views without having the necessary information (ultracrepidarianism) or is responding as if the generative AI had expert knowledge when it does not. A lack of awareness that AIs may deliver inadequate or confabulated information can result in incorrect medical decisions and inappropriate clinical applications (Dunning-Kruger effect). As a result, in certain cases, a generative AI system might underperform and provide results which greatly overestimate any medical or clinical validity.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 2024๋…„ ๋ฐœํ‘œ๋œ 83ํŽธ์˜ ์ž„์ƒ ์˜ํ•™ ๊ด€๋ จ AI ๋…ผ๋ฌธ์„ ๊ฒ€ํ† ํ•˜์—ฌ ์ง„๋‹จ, ์น˜๋ฃŒ ๊ถŒ๊ณ  ๋ฐ ์˜์ƒ ๋ถ„์„ ๋ถ„์•ผ์—์„œ์˜ AI ์„ฑ๋Šฅ๊ณผ ๊ฒ€์ฆ ๋ฐฉ์‹์„ ๋ถ„์„ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ƒ์„ฑํ˜• AI๊ฐ€ ์ „๋ฌธ ์ง€์‹ ์—†์ด ์ž˜๋ชป๋œ ์ •๋ณด๋ฅผ ์ œ๊ณตํ•˜๋Š” '์šธํŠธ๋ผํฌ๋ ˆํ”ผ๋‹ค๋ฆฌ์•„๋‹ˆ์ฆ˜(ultracrepidarianism)'๊ณผ AI์˜ ์„ฑ๋Šฅ์„ ๊ณผ๋Œ€ํ‰๊ฐ€ํ•˜๋Š” '๋”๋‹-ํฌ๋ฃจ๊ฑฐ ํšจ๊ณผ'๊ฐ€ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ •์— ์‹ฌ๊ฐํ•œ ์˜ค๋ฅ˜๋ฅผ ์ดˆ๋ž˜ํ•  ์ˆ˜ ์žˆ์Œ์„ ๊ฒฝ๊ณ ํ•ฉ๋‹ˆ๋‹ค. ๋”ฐ๋ผ์„œ ์˜๋ฃŒ ํ˜„์žฅ์—์„œ AI ํ™œ์šฉ ์‹œ ๊ฒฐ๊ณผ์˜ ์ •ํ™•์„ฑ์— ๋Œ€ํ•œ ๋น„ํŒ์  ๊ฒ€์ฆ๊ณผ ์ฃผ์˜ ๊นŠ์€ ํ•ด์„์ด ํ•„์ˆ˜์ ์ž„์„ ๊ฐ•์กฐํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

43Intuitive Human-Artificial Intelligence Theranostic Complementarity.

2025-02-20Cancer biotherapy & radiopharmaceuticals๐Ÿ”ท Q2DOI 10.1089/cbr.2025.0021

Deep learning artificial intelligence (AI) algorithms are poised to subsume diagnostic imaging specialists in radiology and nuclear medicine, where radiomics can consistently outperform human analysis and reporting capability, and do it faster, with greater accuracy and reliability. However, claims made for generative AI in respect of decision-making in the clinical practice of theranostic nuclear medicine are highly contentious. Statistical computer algorithms cannot emulate human emotion, reason, instinct, intuition, or empathy. AI simulates intelligence without possessing it. AI has no understanding of the meaning of its outputs. The unique statistical probability attributes of large language models of AI must be complemented by the innate human intuitive capabilities of nuclear physicians who accept the responsibility and accountability for direct clinical care of each individual patient referred for theranostic management of specified cancers. Complementarity envisions synergistic engagement with AI radiomics, genomics, radiobiology, dosimetry, and data collation from multidimensional sources, including the electronic medical record, to enable the nuclear physician to spend informed face time with their patient. Together with physician discernment, application of the technical insights from AI will facilitate optimal formulation of a personalized precision theranostic strategy for empathic, efficacious, targeted treatment of the patient with cancer in accordance with their wishes.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ํ…Œ๋ผ๋…ธ์Šคํ‹ฑ์Šค(theranostics) ๋ถ„์•ผ์—์„œ AI์˜ ๋ฐฉ๋Œ€ํ•œ ๋ฐ์ดํ„ฐ ๋ถ„์„ ๋ฐ ํ†ต๊ณ„์  ์˜ˆ์ธก ๋Šฅ๋ ฅ๊ณผ ํ•ต์˜ํ•™ ์ „๋ฌธ์˜์˜ ์ง๊ด€, ๊ณต๊ฐ, ์ž„์ƒ์  ์ฑ…์ž„๊ฐ์ด ์ƒํ˜ธ ๋ณด์™„์ ์œผ๋กœ ๊ฒฐํ•ฉํ•ด์•ผ ํ•จ์„ ๊ฐ•์กฐํ•ฉ๋‹ˆ๋‹ค. AI๊ฐ€ ์ œ๊ณตํ•˜๋Š” ์ •๋ฐ€ํ•œ ๊ธฐ์ˆ ์  ํ†ต์ฐฐ์„ ์˜์‚ฌ์˜ ์ž„์ƒ์  ํŒ๋‹จ๊ณผ ํ†ตํ•ฉํ•จ์œผ๋กœ์จ, ํ™˜์ž ๊ฐœ๊ฐœ์ธ์˜ ์š”๊ตฌ๋ฅผ ๋ฐ˜์˜ํ•œ ์ตœ์ ์˜ ๋งž์ถคํ˜• ์ •๋ฐ€ ์น˜๋ฃŒ ์ „๋žต์„ ์ˆ˜๋ฆฝํ•  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

44Large Language Models in Healthcare: A Bibliometric Analysis and Examination of Research Trends.

2025-01-17Journal of multidisciplinary healthcareโญ Q1DOI 10.2147/jmdh.s502351
BACKGROUND

The integration of large language models (LLMs) in healthcare has generated significant interest due to their potential to improve diagnostic accuracy, personalization of treatment, and patient care efficiency.

OBJECTIVE

This study aims to conduct a comprehensive bibliometric analysis to identify current research trends, main themes and future directions regarding applications in the healthcare sector.

METHODS

A systematic scan of publications until 08.05.2024 was carried out from an important database such as Web of Science.Using bibliometric tools such as VOSviewer and CiteSpace, we analyzed data covering publication counts, citation analysis, co-authorship, co- occurrence of keywords and thematic development to map the intellectual landscape and collaborative networks in this field.

RESULTS

The analysis included more than 500 articles published between 2021 and 2024. The United States, Germany and the United Kingdom were the top contributors to this field. The study highlights that neural network applications in diagnostic imaging, natural language processing for clinical documentation, and patient data in the field of general internal medicine, radiology, medical informatics, health care services, surgery, oncology, ophthalmology, neurology, orthopedics and psychiatry have seen significant growth in publications over the past two years. Keyword trend analysis revealed emerging sub-themes such as clinical research, artificial intelligence, ChatGPT, education, natural language processing, clinical management, virtual reality, chatbot, indicating a shift towards addressing the broader implications of LLM application in healthcare.

CONCLUSION

The use of LLM in healthcare is an expanding field with significant academic and clinical interest. This bibliometric analysis not only maps the current state of the research, but also identifies important areas that require further research and development. Continued advances in this field are expected to significantly impact future healthcare applications, with a focus on increasing the accuracy and personalization of patient care through advanced data analytics.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 2021๋…„๋ถ€ํ„ฐ 2024๋…„๊นŒ์ง€์˜ ๋ฌธํ—Œ์„ ๋Œ€์ƒ์œผ๋กœ ์„œ์ง€ํ•™์  ๋ถ„์„์„ ์ˆ˜ํ–‰ํ•˜์—ฌ ์˜๋ฃŒ ๋ถ„์•ผ ๋‚ด ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ์—ฐ๊ตฌ ๋™ํ–ฅ๊ณผ ํ•ต์‹ฌ ์ฃผ์ œ๋ฅผ ํŒŒ์•…ํ•˜๊ณ ์ž ํ•˜์˜€๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ์ง„๋‹จ ์˜์ƒ, ์ž„์ƒ ๋ฌธ์„œํ™”, ํ™˜์ž ๋ฐ์ดํ„ฐ ๊ด€๋ฆฌ ๋“ฑ ๋‹ค์–‘ํ•œ ์ž„์ƒ ์˜์—ญ์—์„œ LLM ํ™œ์šฉ์ด ๊ธ‰์ฆํ•˜๊ณ  ์žˆ์œผ๋ฉฐ, ํŠนํžˆ ๋ฏธ๊ตญ๊ณผ ์œ ๋Ÿฝ์„ ์ค‘์‹ฌ์œผ๋กœ ์ž„์ƒ ์—ฐ๊ตฌ ๋ฐ ์ธ๊ณต์ง€๋Šฅ ๊ธฐ๋ฐ˜์˜ ํ™˜์ž ๋งž์ถคํ˜• ์ง„๋ฃŒ ์ฒด๊ณ„๊ฐ€ ์ฃผ์š” ์—ฐ๊ตฌ ์ฃผ์ œ๋กœ ๋ถ€์ƒํ•˜๊ณ  ์žˆ์Œ์„ ํ™•์ธํ•˜์˜€๋‹ค. ๋ณธ ์—ฐ๊ตฌ๋Š” LLM์ด ํ–ฅํ›„ ์˜๋ฃŒ ์„œ๋น„์Šค์˜ ์ •ํ™•๋„์™€ ๊ฐœ์ธํ™”๋œ ์น˜๋ฃŒ๋ฅผ ๊ฐœ์„ ํ•˜๋Š” ๋ฐ ์ค‘์š”ํ•œ ์—ญํ• ์„ ํ•  ๊ฒƒ์œผ๋กœ ์ „๋งํ•˜๋ฉฐ, ํ–ฅํ›„ ์ง€์†์ ์ธ ๊ธฐ์ˆ  ๋ฐœ์ „๊ณผ ์—ฐ๊ตฌ์˜ ํ•„์š”์„ฑ์„ ์ œ์‹œํ•œ๋‹ค.
Added: 2026-04-04 06:48View โ†—

45Privacy-ensuring Open-weights Large Language Models Are Competitive with Closed-weights GPT-4o in Extracting Chest Radiography Findings from Free-Text Reports.

2025-01Radiologyโญ Q1DOI 10.1148/radiol.240895

Background Large-scale secondary use of clinical databases requires automated tools for retrospective extraction of structured content from free-text radiology reports. Purpose To share data and insights on the application of privacy-preserving open-weights large language models (LLMs) for reporting content extraction with comparison to standard rule-based systems and the closed-weights LLMs from OpenAI. Materials and Methods In this retrospective exploratory study conducted between May 2024 and September 2024, zero-shot prompting of 17 open-weights LLMs was preformed. These LLMs with model weights released under open licenses were compared with rule-based annotation and with OpenAI's GPT-4o, GPT-4o-mini, GPT-4-turbo, and GPT-3.5-turbo on a manually annotated public English chest radiography dataset (Indiana University, 3927 patients and reports). An annotated nonpublic German chest radiography dataset (18โ€‰500 reports, 16โ€‰844 patients [10โ€‰340 male; mean age, 62.6 years ยฑ 21.5 {SD}]) was used to compare local fine-tuning of all open-weights LLMs via low-rank adaptation and 4-bit quantization to bidirectional encoder representations from transformers (BERT) with different subsets of reports (from 10 to 14โ€‰580). Nonoverlapping 95% CIs of macro-averaged F1 scores were defined as relevant differences. Results For the English reports, the highest zero-shot macro-averaged F1 score was observed for GPT-4o (92.4% [95% CI: 87.9, 95.9]); GPT-4o outperformed the rule-based CheXpert [Stanford University] (73.1% [95% CI: 65.1, 79.7]) but was comparable in performance to several open-weights LLMs (top three: Mistral-Large [Mistral AI], 92.6% [95% CI: 88.2, 96.0]; Llama-3.1-70b [Meta AI], 92.2% [95% CI: 87.1, 95.8]; and Llama-3.1-405b [Meta AI]: 90.3% [95% CI: 84.6, 94.5]). For the German reports, Mistral-Large (91.6% [95% CI: 90.5, 92.7]) had the highest zero-shot macro-averaged F1 score among the six other open-weights LLMs and outperformed the rule-based annotation (74.8% [95% CI: 73.3, 76.1]). Using 1000 reports for fine-tuning, all LLMs (top three: Mistral-Large, 94.3% [95% CI: 93.5, 95.2]; OpenBioLLM-70b [Saama]: 93.9% [95% CI: 92.9, 94.8]; and Mixtral-8ร—22b [Mistral AI]: 93.8% [95% CI: 92.8, 94.7]) achieved significantly higher macro-averaged F1 score than did BERT (86.7% [95% CI: 85.0, 88.3]); however, the differences were not relevant when 2000 or more reports were used for fine-tuning. Conclusion LLMs have the potential to outperform rule-based systems for zero-shot "out-of-the-box" structuring of report databases, with privacy-ensuring open-weights LLMs being competitive with closed-weights GPT-4o. Additionally, the open-weights LLM outperformed BERT when moderate numbers of reports were used for fine-tuning. Published under a CC BY 4.0 license. Supplemental material is available for this article. See also the editorial by Gee and Yao in this issue.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ํ‰๋ถ€ ๋ฐฉ์‚ฌ์„  ๋ณด๊ณ ์„œ์—์„œ ์ž„์ƒ ์ •๋ณด๋ฅผ ์ถ”์ถœํ•˜๊ธฐ ์œ„ํ•ด ์˜คํ”ˆ ์›จ์ดํŠธ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜๊ณ , ์ด๋ฅผ ๊ธฐ์กด ๊ทœ์น™ ๊ธฐ๋ฐ˜ ์‹œ์Šคํ…œ ๋ฐ ํ์‡„ํ˜• ๋ชจ๋ธ์ธ GPT-4o์™€ ๋น„๊ตํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ์˜คํ”ˆ ์›จ์ดํŠธ LLM์€ ์ œ๋กœ์ƒท ํ™˜๊ฒฝ์—์„œ GPT-4o์™€ ๋Œ€๋“ฑํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ์†Œ๋Ÿ‰์˜ ๋ฐ์ดํ„ฐ๋กœ ๋ฏธ์„ธ ์กฐ์ •์„ ์ˆ˜ํ–‰ํ•  ๊ฒฝ์šฐ BERT ๋ชจ๋ธ๋ณด๋‹ค ์šฐ์ˆ˜ํ•œ ์ถ”์ถœ ์„ฑ๋Šฅ์„ ๋‚˜ํƒ€๋ƒˆ์Šต๋‹ˆ๋‹ค. ์ด๋Š” ๋ฐ์ดํ„ฐ ํ”„๋ผ์ด๋ฒ„์‹œ๋ฅผ ๋ณดํ˜ธํ•˜๋ฉด์„œ๋„ ์ž„์ƒ ๋ฐ์ดํ„ฐ์˜ ๊ตฌ์กฐํ™”๋œ ์ถ”์ถœ์„ ์œ„ํ•ด ์˜คํ”ˆ ์›จ์ดํŠธ LLM์„ ํšจ๊ณผ์ ์œผ๋กœ ํ™œ์šฉํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

46ReXamine-Global: A Framework for Uncovering Inconsistencies in Radiology Report Generation Metrics.

2025Pacific Symposium on Biocomputing. Pacific Symposium on Biocomputing๐Ÿ”ท Q2DOI 10.1142/9789819807024_0014

Given the rapidly expanding capabilities of generative AI models for radiology, there is a need for robust metrics that can accurately measure the quality of AI-generated radiology reports across diverse hospitals. We develop ReXamine-Global, a LLM-powered, multi-site framework that tests metrics across different writing styles and patient populations, exposing gaps in their generalization. First, our method tests whether a metric is undesirably sensitive to reporting style, providing different scores depending on whether AI-generated reports are stylistically similar to ground-truth reports or not. Second, our method measures whether a metric reliably agrees with experts, or whether metric and expert scores of AI-generated report quality diverge for some sites. Using 240 reports from 6 hospitals around the world, we apply ReXamine-Global to 7 established report evaluation metrics and uncover serious gaps in their generalizability. Developers can apply ReXamine-Global when designing new report evaluation metrics, ensuring their robustness across sites. Additionally, our analysis of existing metrics can guide users of those metrics towards evaluation procedures that work reliably at their sites of interest.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
ReXamine-Global์€ ๋‹ค์–‘ํ•œ ์˜๋ฃŒ๊ธฐ๊ด€์˜ ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ ์ƒ์„ฑ ๋ชจ๋ธ์„ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด LLM ๊ธฐ๋ฐ˜์˜ ๋‹ค๊ธฐ๊ด€ ํ”„๋ ˆ์ž„์›Œํฌ๋ฅผ ์ œ์•ˆํ•˜์—ฌ ๊ธฐ์กด ํ‰๊ฐ€ ์ง€ํ‘œ๋“ค์˜ ์ผ๋ฐ˜ํ™” ์„ฑ๋Šฅ๊ณผ ์Šคํƒ€์ผ ๋ฏผ๊ฐ๋„๋ฅผ ๊ฒ€์ฆํ•ฉ๋‹ˆ๋‹ค. 6๊ฐœ๊ตญ 6๊ฐœ ๋ณ‘์›์˜ ๋ฐ์ดํ„ฐ๋ฅผ ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, ๊ธฐ์กด 7๊ฐœ ํ‰๊ฐ€ ์ง€ํ‘œ์—์„œ ๋ณด๊ณ ์„œ ์ž‘์„ฑ ์Šคํƒ€์ผ ๋ฐ ๊ธฐ๊ด€ ๊ฐ„ ์ฐจ์ด์— ๋”ฐ๋ฅธ ์„ฑ๋Šฅ ํŽธ์ฐจ๊ฐ€ ํ™•์ธ๋˜์—ˆ์Šต๋‹ˆ๋‹ค. ๋ณธ ์—ฐ๊ตฌ๋Š” ํ–ฅํ›„ ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ ํ‰๊ฐ€ ์ง€ํ‘œ์˜ ๊ฒฌ๊ณ ์„ฑ์„ ํ™•๋ณดํ•˜๊ณ , ์ž„์ƒ ํ˜„์žฅ์—์„œ ์‹ ๋ขฐํ•  ์ˆ˜ ์žˆ๋Š” ํ‰๊ฐ€ ์ ˆ์ฐจ๋ฅผ ์ˆ˜๋ฆฝํ•˜๋Š” ๋ฐ ๊ธฐ์—ฌํ•  ๊ฒƒ์œผ๋กœ ๊ธฐ๋Œ€๋ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

47Large Language Models with Vision on Diagnostic Radiology Board Exam Style Questions.

2024-12-04Academic radiologyโญ Q1DOI 10.1016/j.acra.2024.11.028

RATIONALE AND

OBJECTIVE

The expansion of large language models to process images offers new avenues for application in radiology. This study aims to assess the multimodal capabilities of contemporary large language models, which allow analysis of image inputs in addition to textual data, on radiology board-style examination questions with images.

METHODS

280 questions were retrospectively selected from the AuntMinnie public test bank. The test questions were converted into three formats of prompts; (1) Multimodal, (2) Image-only, and (3) Text-only input. Three models, GPT-4V, Gemini 1.5 Pro, and Claude 3.5 Sonnet, were evaluated using these prompts. The Cochran Q test and pairwise McNemar test were used to compare performances between prompt formats and models.

RESULTS

No difference was found for the performance in terms of % correct answers between the text, image, and multimodal prompt formats for GPT-4V (54%, 52%, and 57%, respectively; pย =ย .31) and Gemini 1.5 Pro (53%, 54%, and 57%, respectively; pย =ย .53). For Claude 3.5 Sonnet, the image input (48%) significantly underperformed compared to the text input (63%, pย <ย .001) and the multimodal input (66%, pย <ย .001), but no difference was found between the text and multimodal inputs (pย =ย .29). Claude significantly outperformed GPT and Gemini in the text and multimodal formats (pย <ย .01).

CONCLUSION

Vision-capable large language models cannot effectively use images to increase performance on radiology board-style examination questions. When using textual data alone, Claude 3.5 Sonnet outperforms GPT-4V and Gemini 1.5 Pro, highlighting the advancements in the field and its potential for use in further research.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜์ƒ ์˜ํ•™ ๋ณด๋“œ ์‹œํ—˜ ๋ฌธํ•ญ์„ ํ™œ์šฉํ•˜์—ฌ ์ตœ์‹  ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(GPT-4V, Gemini 1.5 Pro, Claude 3.5 Sonnet)์˜ ์˜์ƒ ๋ถ„์„ ๋ฐ ์ง„๋‹จ ๋Šฅ๋ ฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, ๋ชจ๋“  ๋ชจ๋ธ์—์„œ ์˜์ƒ ์ •๋ณด์˜ ์ถ”๊ฐ€๊ฐ€ ์ •๋‹ต๋ฅ  ํ–ฅ์ƒ์œผ๋กœ ์ด์–ด์ง€์ง€ ์•Š์•˜์œผ๋ฉฐ, Claude 3.5 Sonnet์ด ํ…์ŠคํŠธ ๊ธฐ๋ฐ˜ ๋ฌธํ•ญ์—์„œ ํƒ€ ๋ชจ๋ธ ๋Œ€๋น„ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ ํ˜„์žฌ์˜ ๋ฉ€ํ‹ฐ๋ชจ๋‹ฌ ๋ชจ๋ธ๋“ค์€ ์˜์ƒ ์ •๋ณด๋ฅผ ํšจ๊ณผ์ ์œผ๋กœ ํ™œ์šฉํ•˜์—ฌ ์ง„๋‹จ ์ •ํ™•๋„๋ฅผ ๋†’์ด๋Š” ๋ฐ ํ•œ๊ณ„๊ฐ€ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

48Performance of a Generative Pre-Trained Transformer in Generating Scientific Abstracts in Dentistry: A Comparative Observational Study.

2024-11-19European journal of dental education : official journal of the Association for Dental Education in Europeโญ Q1DOI 10.1111/eje.13057
OBJECTIVE

To evaluate the performance of a Generative Pre-trained Transformer (GPT) in generating scientific abstracts in dentistry.

METHODS

Ten scientific articles in dental radiology had their original abstracts collected, while another 10 articles had their methodology and results added to a ChatGPT prompt to generate an abstract. All abstracts were randomised and compiled into a single file for subsequent assessment. Five evaluators classified whether the abstract was generated by a human using a 5-point scale and provided justifications within seven aspects: formatting, information accuracy, orthography, punctuation, terminology, text fluency, and writing style. Furthermore, an online GPT detector provided "Human Score" values, and a plagiarism detector assessed similarity with existing literature.

RESULTS

Sensitivity values for detecting human writing ranged from 0.20 to 0.70, with a mean of 0.58; specificity values ranged from 0.40 to 0.90, with a mean of 0.62; and accuracy values ranged from 0.50 to 0.80, with a mean of 0.60. Orthography and Punctuation were the most indicated aspects for the abstract generated by ChatGPT. The GPT detector revealed confidence levels for a "Human Score" of 16.9% for the AI-generated texts and plagiarism levels averaging 35%.

CONCLUSION

The GPT exhibited commendable performance in generating scientific abstracts when evaluated by humans, as the generated abstracts were indistinguishable from those generated by humans. When evaluated by an online GPT detector, the use of GPT became apparent.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์น˜์˜ํ•™ ๋ถ„์•ผ์—์„œ ์ƒ์„ฑํ˜• ์‚ฌ์ „ ํ•™์Šต ํŠธ๋žœ์Šคํฌ๋จธ(GPT)๊ฐ€ ์ž‘์„ฑํ•œ ์ดˆ๋ก์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜๊ธฐ ์œ„ํ•ด ์ธ๊ฐ„ ํ‰๊ฐ€์ž์™€ AI ํƒ์ง€๊ธฐ๋ฅผ ํ™œ์šฉํ•˜์—ฌ ๊ธฐ์กด ์ดˆ๋ก๊ณผ ๋น„๊ต ๋ถ„์„ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ํ‰๊ฐ€ ๊ฒฐ๊ณผ, ์ธ๊ฐ„ ํ‰๊ฐ€์ž๋Š” GPT๊ฐ€ ์ž‘์„ฑํ•œ ์ดˆ๋ก์„ ์‹ค์ œ ์—ฐ๊ตฌ์ž๊ฐ€ ์ž‘์„ฑํ•œ ๊ฒƒ๊ณผ ๊ตฌ๋ณ„ํ•˜๊ธฐ ์–ด๋ ค์šธ ์ •๋„๋กœ ๋†’์€ ์™„์„ฑ๋„๋ฅผ ๋ณด์˜€์œผ๋‚˜, ์˜จ๋ผ์ธ AI ํƒ์ง€ ๋„๊ตฌ๋Š” ํ•ด๋‹น ์ดˆ๋ก์„ ๋†’์€ ํ™•๋ฅ ๋กœ ์‹๋ณ„ํ•ด ๋‚ผ ์ˆ˜ ์žˆ์—ˆ์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ GPT๋Š” ํ•™์ˆ ์  ์ดˆ๋ก ์ž‘์„ฑ์— ์žˆ์–ด ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์ด๋‚˜, ํ•™์ˆ ์  ์ง„์‹ค์„ฑ ๊ฒ€์ฆ์„ ์œ„ํ•œ ๋ณด์™„์  ๋„๊ตฌ์˜ ํ™œ์šฉ์ด ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

49Enhancing chatbot performance for imaging recommendations: Leveraging GPT-4 and context-awareness for trustworthy clinical guidance.

2024-09-24European journal of radiologyโญ Q1DOI 10.1016/j.ejrad.2024.111756
OBJECTIVE

To investigate if GPT-4 improves the accuracy, consistency, and trustworthiness of a context-aware chatbot to provide personalized imaging recommendations from American College of Radiology (ACR) appropriateness criteria documents using semantic similarity processing: In addition, we sought to enable auditability of the output by revealing the information source the decision relies on.

METHODS

We refined an existing chatbot that incorporated specialized knowledge of the ACR guidelines by upgrading GPT-3.5-Turbo to its successor GPT-4 by OpenAI, using the latest version of LlamaIndex, and improving the prompting strategy. This chatbot was compared to the previous version, generic GPT-3.5-Turbo and GPT-4, and general radiologists regarding the performance in applying the ACR appropriateness guidelines.

RESULTS

The refined context-aware chatbot performed superior to the previous version using GPT-3.5-Turbo, generic chatbots GPT-3.5-Turbo and GPT-4, and general radiologists in providing "usually or may be appropriate" recommendations according to the ACR guidelines (all pย <ย 0.001). It also outperformed GPT-3.5-Turbo and general radiologists in respect to "usually appropriate" recommendations (both pย <ย 0.001). Moreover, the consistency in correct answers was higher with 78ย % consistent correct "usually appropriate" answers and 94ย % for "usually or may be appropriate" recommendations. In all cases, the same source documents were chosen, ensuring transparency.

CONCLUSION

Our study demonstrates the significance of context awareness in ensuring the use of appropriate knowledge and proposes a strategy to enhance trust in chatbot-based outputs to provide transparency. The improvements in accuracy, consistency, and source transparency address trust issues and enhance the clinical decision support process. ABBREVIATIONS: ACR, American College of Radiology; accGPT, appropriateness criteria context aware GPT; accGPT-4, appropriateness criteria context aware GPT using GPT-4; GPT, generative pre-trained transformer; LLM, Large Language Model.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฏธ๊ตญ์˜์ƒ์˜ํ•™ํšŒ(ACR) ์ ์ ˆ์„ฑ ๊ธฐ์ค€์„ ๊ธฐ๋ฐ˜์œผ๋กœ GPT-4์™€ ๋ฌธ๋งฅ ์ธ์‹ ๊ธฐ์ˆ ์„ ๊ฒฐํ•ฉํ•œ ์ฑ—๋ด‡์˜ ์˜์ƒ ๊ฒ€์‚ฌ ๊ถŒ๊ณ  ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๊ฐœ์„ ๋œ ์ฑ—๋ด‡์€ ๊ธฐ์กด ๋ชจ๋ธ ๋ฐ ์ผ๋ฐ˜ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜๋ณด๋‹ค ๋†’์€ ์ •ํ™•๋„์™€ ์ผ๊ด€์„ฑ์„ ๋ณด์˜€์œผ๋ฉฐ, ๊ทผ๊ฑฐ ๋ฌธ์„œ๋ฅผ ๋ช…์‹œํ•˜์—ฌ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ๋„๊ตฌ๋กœ์„œ์˜ ์‹ ๋ขฐ์„ฑ๊ณผ ํˆฌ๋ช…์„ฑ์„ ํ™•๋ณดํ•˜์˜€์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

50Based on Medicine, The Now and Future of Large Language Models.

2024-09-16Cellular and molecular bioengineering๐Ÿ”ท Q2DOI 10.1007/s12195-024-00820-3
OBJECTIVE

This review explores the potential applications of large language models (LLMs) such as ChatGPT, GPT-3.5, and GPT-4 in the medical field, aiming to encourage their prudent use, provide professional support, and develop accessible medical AI tools that adhere to healthcare standards.

METHODS

This paper examines the impact of technologies such as OpenAI's Generative Pre-trained Transformers (GPT) series, including GPT-3.5 and GPT-4, and other large language models (LLMs) in medical education, scientific research, clinical practice, and nursing. Specifically, it includes supporting curriculum design, acting as personalized learning assistants, creating standardized simulated patient scenarios in education; assisting with writing papers, data analysis, and optimizing experimental designs in scientific research; aiding in medical imaging analysis, decision-making, patient education, and communication in clinical practice; and reducing repetitive tasks, promoting personalized care and self-care, providing psychological support, and enhancing management efficiency in nursing.

RESULTS

LLMs, including ChatGPT, have demonstrated significant potential and effectiveness in the aforementioned areas, yet their deployment in healthcare settings is fraught with ethical complexities, potential lack of empathy, and risks of biased responses.

CONCLUSION

Despite these challenges, significant medical advancements can be expected through the proper use of LLMs and appropriate policy guidance. Future research should focus on overcoming these barriers to ensure the effective and ethical application of LLMs in the medical field.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜๋ฃŒ ๊ต์œก, ์—ฐ๊ตฌ, ์ž„์ƒ ๋ฐ ๊ฐ„ํ˜ธ ๋ถ„์•ผ์—์„œ ๊ฑฐ๋Œ€์–ธ์–ด๋ชจ๋ธ(LLM)์˜ ํ™œ์šฉ ๊ฐ€๋Šฅ์„ฑ์„ ๋ถ„์„ํ•˜๊ณ , ์ด๋ฅผ ์˜๋ฃŒ ํ˜„์žฅ์— ํšจ๊ณผ์ ์œผ๋กœ ๋„์ž…ํ•˜๊ธฐ ์œ„ํ•œ ๋ฐฉ์•ˆ์„ ๊ณ ์ฐฐํ•˜์˜€์Šต๋‹ˆ๋‹ค. LLM์€ ๋‹ค์–‘ํ•œ ์˜๋ฃŒ ์˜์—ญ์—์„œ ์—…๋ฌด ํšจ์œจ์„ฑ๊ณผ ์ž„์ƒ์  ์ง€์› ๋Šฅ๋ ฅ์„ ์ž…์ฆํ–ˆ์œผ๋‚˜, ์œค๋ฆฌ์  ๋ฌธ์ œ์™€ ํŽธํ–ฅ์„ฑ ๋“ฑ์˜ ํ•œ๊ณ„์  ๋˜ํ•œ ๋ช…ํ™•ํžˆ ๋“œ๋Ÿฌ๋ƒˆ์Šต๋‹ˆ๋‹ค. ๋”ฐ๋ผ์„œ ํ–ฅํ›„ ์˜๋ฃŒ ํ˜„์žฅ์—์„œ LLM์„ ์•ˆ์ „ํ•˜๊ณ  ํšจ๊ณผ์ ์œผ๋กœ ํ™œ์šฉํ•˜๊ธฐ ์œ„ํ•ด์„œ๋Š” ๊ธฐ์ˆ ์  ํ•œ๊ณ„ ๊ทน๋ณต๊ณผ ํ•จ๊ป˜ ์ ์ ˆํ•œ ์ •์ฑ…์  ๊ฐ€์ด๋“œ๋ผ์ธ ๋งˆ๋ จ์ด ํ•„์ˆ˜์ ์ž…๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

51Performance of GPT-4 with Vision on Text- and Image-based ACR Diagnostic Radiology In-Training Examination Questions.

2024-09Radiologyโญ Q1DOI 10.1148/radiol.240153

Background Recent advancements, including image processing capabilities, present new potential applications of large language models such as ChatGPT (OpenAI), a generative pretrained transformer, in radiology. However, baseline performance of ChatGPT in radiology-related tasks is understudied. Purpose To evaluate the performance of GPT-4 with vision (GPT-4V) on radiology in-training examination questions, including those with images, to gauge the model's baseline knowledge in radiology. Materials and Methods In this prospective study, conducted between September 2023 and March 2024, the September 2023 release of GPT-4V was assessed using 386 retired questions (189 image-based and 197 text-only questions) from the American College of Radiology Diagnostic Radiology In-Training Examinations. Nine question pairs were identified as duplicates; only the first instance of each duplicate was considered in ChatGPT's assessment. A subanalysis assessed the impact of different zero-shot prompts on performance. Statistical analysis included ฯ‡2 tests of independence to ascertain whether the performance of GPT-4V varied between question types or subspecialty. The McNemar test was used to evaluate performance differences between the prompts, with Benjamin-Hochberg adjustment of the P values conducted to control the false discovery rate (FDR). A P value threshold of less than.05 denoted statistical significance. Results GPT-4V correctly answered 246 (65.3%) of the 377 unique questions, with significantly higher accuracy on text-only questions (81.5%, 159 of 195) than on image-based questions (47.8%, 87 of 182) (ฯ‡2 test, P < .001). Subanalysis revealed differences between prompts on text-based questions, where chain-of-thought prompting outperformed long instruction by 6.1% (McNemar, P = .02; FDR = 0.063), basic prompting by 6.8% (P = .009, FDR = 0.044), and the original prompting style by 8.9% (P = .001, FDR = 0.014). No differences were observed between prompts on image-based questions with P values of .27 to >.99. Conclusion While GPT-4V demonstrated a level of competence in text-based questions, it showed deficits interpreting radiologic images. ยฉ RSNA, 2024 See also the editorial by Deng in this issue.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฏธ๊ตญ ์˜์ƒ์˜ํ•™ํšŒ(ACR) ์ˆ˜๋ จ์˜ ์‹œํ—˜ ๋ฌธํ•ญ 377๊ฐœ๋ฅผ ํ™œ์šฉํ•˜์—ฌ GPT-4V์˜ ์˜์ƒ์˜ํ•™ ์ง€์‹ ์ˆ˜์ค€์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, GPT-4V๋Š” ํ…์ŠคํŠธ ๊ธฐ๋ฐ˜ ๋ฌธํ•ญ์—์„œ๋Š” 81.5%์˜ ๋†’์€ ์ •๋‹ต๋ฅ ์„ ๋ณด์˜€์œผ๋‚˜, ์˜์ƒ ๊ธฐ๋ฐ˜ ๋ฌธํ•ญ์—์„œ๋Š” 47.8%์˜ ๋‚ฎ์€ ์ •๋‹ต๋ฅ ์„ ๊ธฐ๋กํ•˜๋ฉฐ ์˜์ƒ ํ•ด์„ ๋Šฅ๋ ฅ์— ํ•œ๊ณ„๋ฅผ ๋“œ๋Ÿฌ๋ƒˆ์Šต๋‹ˆ๋‹ค. ์ด๋Š” GPT-4V๊ฐ€ ํ…์ŠคํŠธ ์ •๋ณด ์ฒ˜๋ฆฌ์—๋Š” ์ƒ๋‹นํ•œ ์—ญ๋Ÿ‰์„ ๊ฐ–์ถ”์—ˆ์œผ๋‚˜, ์ž„์ƒ ์˜์ƒ ํŒ๋… ๋ฐ ์ง„๋‹จ ๋ณด์กฐ ๋„๊ตฌ๋กœ ํ™œ์šฉ๋˜๊ธฐ์—๋Š” ์•„์ง ๋ณด์™„์ด ํ•„์š”ํ•จ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

52Large Language Models as Tools to Generate Radiology Board-Style Multiple-Choice Questions.

2024-07-15Academic radiologyโญ Q1DOI 10.1016/j.acra.2024.06.046

RATIONALE AND

OBJECTIVE

To determine the potential of large language models (LLMs) to be used as tools by radiology educators to create radiology board-style multiple choice questions (MCQs), answers, and rationales.

METHODS

Two LLMs (Llama 2 and GPT-4) were used to develop 104 MCQs based on the American Board of Radiology exam blueprint. Two board-certified radiologists assessed each MCQ using a 10-point Likert scale across five criteria-clarity, relevance, suitability for a board exam based on level of difficulty, quality of distractors, and adequacy of rationale. For comparison, MCQs from prior American College of Radiology (ACR) Diagnostic Radiology In-Training (DXIT) exams were also assessed using these criteria, with radiologists blinded to the question source.

RESULTS

Mean scores (ยฑstandard deviation) for clarity, relevance, suitability, quality of distractors, and adequacy of rationale were 8.7 (ยฑ1.4), 9.2 (ยฑ1.3), 9.0 (ยฑ1.2), 8.4 (ยฑ1.9), and 7.2 (ยฑ2.2), respectively, for Llama 2; 9.9 (ยฑ0.4), 9.9 (ยฑ0.5), 9.9 (ยฑ0.4), 9.8 (ยฑ0.5), and 9.9 (ยฑ0.3), respectively, for GPT-4; and 9.9 (ยฑ0.3), 9.9 (ยฑ0.2), 9.9 (ยฑ0.2), 9.9 (ยฑ0.4), and 9.8 (ยฑ0.6), respectively, for ACR DXIT items (pย <ย 0.001 for Llama 2 vs. ACR DXIT across all criteria; no statistically significant difference for GPT-4 vs. ACR DXIT). The accuracy of model-generated answers was 69% for Llama 2 and 100% for GPT-4.

CONCLUSION

A state-of-the art LLM such as GPT-4 may be used to develop radiology board-style MCQs and rationales to enhance exam preparation materials and expand exam banks, and may allow radiology educators to further use MCQs as teaching and learning tools.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์ธ Llama 2์™€ GPT-4๋ฅผ ํ™œ์šฉํ•˜์—ฌ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜ ์‹œํ—˜ ์ˆ˜์ค€์˜ ๊ฐ๊ด€์‹ ๋ฌธํ•ญ(MCQ)์„ ์ƒ์„ฑํ•˜๊ณ  ๊ทธ ํ’ˆ์งˆ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ํ‰๊ฐ€ ๊ฒฐ๊ณผ, GPT-4๊ฐ€ ์ƒ์„ฑํ•œ ๋ฌธํ•ญ์€ ์ „๋ฌธ์˜๊ฐ€ ํ‰๊ฐ€ํ•œ ๋ชจ๋“  ํ•ญ๋ชฉ์—์„œ ๊ธฐ์กด ACR DXIT ๋ฌธํ•ญ๊ณผ ํ†ต๊ณ„์ ์œผ๋กœ ์œ ์˜๋ฏธํ•œ ์ฐจ์ด ์—†์ด ๋†’์€ ํ’ˆ์งˆ๊ณผ ์ •ํ™•๋„๋ฅผ ๋ณด์˜€์Šต๋‹ˆ๋‹ค. ๋”ฐ๋ผ์„œ GPT-4์™€ ๊ฐ™์€ ์ตœ์‹  LLM์€ ์˜์ƒ์˜ํ•™ ๊ต์œก์ž์˜ ๋ฌธํ•ญ ๊ฐœ๋ฐœ ๋ฐ ํ•™์Šต ์ž๋ฃŒ ๊ตฌ์ถ•์„ ์œ„ํ•œ ํšจ๊ณผ์ ์ธ ๋„๊ตฌ๋กœ ํ™œ์šฉ๋  ์ˆ˜ ์žˆ์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

53GPT-Driven Radiology Report Generation with Fine-Tuned Llama 3

2024Bioengineering๐Ÿ”ท Q2DOI 10.3390/bioengineering11101043

The integration of deep learning into radiology has the potential to enhance diagnostic processes, yet its acceptance in clinical practice remains limited due to various challenges. This study aimed to develop and evaluate a fine-tuned large language model (LLM), based on Llama 3-8B, to automate the generation of accurate and concise conclusions in magnetic resonance imaging (MRI) and computed tomography (CT) radiology reports, thereby assisting radiologists and improving reporting efficiency. A dataset comprising 15,000 radiology reports was collected from the University of Medicine and Pharmacy of Craiova's Imaging Center, covering a diverse range of MRI and CT examinations made by four experienced radiologists. The Llama 3-8B model was fine-tuned using transfer-learning techniques, incorporating parameter quantization to 4-bit precision and low-rank adaptation (LoRA) with a rank of 16 to optimize computational efficiency on consumer-grade GPUs. The model was trained over five epochs using an NVIDIA RTX 3090 GPU, with intermediary checkpoints saved for monitoring. Performance was evaluated quantitatively using Bidirectional Encoder Representations from Transformers Score (BERTScore), Recall-Oriented Understudy for Gisting Evaluation (ROUGE), Bilingual Evaluation Understudy (BLEU), and Metric for Evaluation of Translation with Explicit Ordering (METEOR) metrics on a held-out test set. Additionally, a qualitative assessment was conducted, involving 13 independent radiologists who participated in a Turing-like test and provided ratings for the AI-generated conclusions. The fine-tuned model demonstrated strong quantitative performance, achieving a BERTScore F1 of 0.8054, a ROUGE-1 F1 of 0.4998, a ROUGE-L F1 of 0.4628, and a METEOR score of 0.4282. In the human evaluation, the artificial intelligence (AI)-generated conclusions were preferred over human-written ones in approximately 21.8% of cases, indicating that the model's outputs were competitive with those of experienced radiologists. The average rating of the AI-generated conclusions was 3.65 out of 5, reflecting a generally favorable assessment. Notably, the model maintained its consistency across various types of reports and demonstrated the ability to generalize to unseen data. The fine-tuned Llama 3-8B model effectively generates accurate and coherent conclusions for MRI and CT radiology reports. By automating the conclusion-writing process, this approach can assist radiologists in reducing their workload and enhancing report consistency, potentially addressing some barriers to the adoption of deep learning in clinical practice. The positive evaluations from independent radiologists underscore the model's potential utility. While the model demonstrated strong performance, limitations such as dataset bias, limited sample diversity, a lack of clinical judgment, and the need for large computational resources require further refinement and real-world validation. Future work should explore the integration of such models into clinical workflows, address ethical and legal considerations, and extend this approach to generate complete radiology reports.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” Llama 3-8B ๋ชจ๋ธ์„ ๋ฏธ์„ธ ์กฐ์ •(fine-tuning)ํ•˜์—ฌ MRI ๋ฐ CT ์˜์ƒ์˜ ํŒ๋…๋ฌธ ๊ฒฐ๋ก ์„ ์ž๋™์œผ๋กœ ์ƒ์„ฑํ•˜๋Š” ์‹œ์Šคํ…œ์„ ๊ฐœ๋ฐœํ•˜๊ณ  ๊ทธ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. 15,000๊ฑด์˜ ๋ฐ์ดํ„ฐ๋ฅผ ํ•™์Šตํ•œ ๋ชจ๋ธ์€ ์ •๋Ÿ‰์  ์ง€ํ‘œ์—์„œ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ๋…๋ฆฝ์ ์ธ ์˜์ƒ์˜ํ•™ ์ „๋ฌธ์˜ ํ‰๊ฐ€์—์„œ๋„ ์ธ๊ฐ„์ด ์ž‘์„ฑํ•œ ํŒ๋…๋ฌธ๊ณผ ๋Œ€๋“ฑํ•œ ์ˆ˜์ค€์˜ ์ •ํ™•๋„์™€ ์ผ๊ด€์„ฑ์„ ์ž…์ฆํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ด ๋ชจ๋ธ์€ ํŒ๋… ์—…๋ฌด์˜ ํšจ์œจ์„ฑ์„ ๋†’์ด๊ณ  ์˜์ƒ์˜ํ•™ ๋ณด๊ณ ์„œ์˜ ํ‘œ์ค€ํ™”๋ฅผ ์ง€์›ํ•˜๋Š” ์ž„์ƒ ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์ œ์‹œํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

54Large language models (LLMs) in radiology exams for medical students: Performance and consequences

2024RรถFo - Fortschritte auf dem Gebiet der Rรถntgenstrahlen und der bildgebenden VerfahrenDOI 10.1055/a-2437-2067

The evolving field of medical education is being shaped by technological advancements, including the integration of Large Language Models (LLMs) like ChatGPT. These models could be invaluable resources for medical students, by simplifying complex concepts and enhancing interactive learning by providing personalized support. LLMs have shown impressive performance in professional examinations, even without specific domain training, making them particularly relevant in the medical field. This study aims to assess the performance of LLMs in radiology examinations for medical students, thereby shedding light on their current capabilities and implications.This study was conducted using 151 multiple-choice questions, which were used for radiology exams for medical students. The questions were categorized by type and topic and were then processed using OpenAI's GPT-3.5 and GPT- 4 via their API, or manually put into Perplexity AI with GPT-3.5 and Bing. LLM performance was evaluated overall, by question type and by topic.GPT-3.5 achieved a 67.6% overall accuracy on all 151 questions, while GPT-4 outperformed it significantly with an 88.1% overall accuracy (p<0.001). GPT-4 demonstrated superior performance in both lower-order and higher-order questions compared to GPT-3.5, Perplexity AI, and medical students, with GPT-4 particularly excelling in higher-order questions. All GPT models would have successfully passed the radiology exam for medical students at our university.In conclusion, our study highlights the potential of LLMs as accessible knowledge resources for medical students. GPT-4 performed well on lower-order as well as higher-order questions, making ChatGPT-4 a potentially very useful tool for reviewing radiology exam questions. Radiologists should be aware of ChatGPT's limitations, including its tendency to confidently provide incorrect responses. ยท ChatGPT demonstrated remarkable performance, achieving a passing grade on a radiology examination for medical students that did not include image questions.. ยท GPT-4 exhibits significantly improved performance compared to its predecessors GPT-3.5 and Perplexity AI with 88% of questions answered correctly.. ยท Radiologists as well as medical students should be aware of ChatGPT's limitations, including its tendency to confidently provide incorrect responses.. ยท Gotta J, Le Hong QA, Koch V et al. Large language models (LLMs) in radiology exams for medical students: Performance and consequences. Rofo 2025; 197: 1057-1067.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜๋Œ€์ƒ ๋Œ€์ƒ ์˜์ƒ์˜ํ•™๊ณผ ์‹œํ—˜ ๋ฌธ์ œ 151๊ฐœ๋ฅผ ํ™œ์šฉํ•˜์—ฌ GPT-3.5์™€ GPT-4 ๋“ฑ ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ(LLM)์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ๋ถ„์„ ๊ฒฐ๊ณผ, GPT-4๋Š” 88.1%์˜ ์ •ํ™•๋„๋กœ GPT-3.5 ๋ฐ Perplexity AI๋ณด๋‹ค ์šฐ์ˆ˜ํ•œ ์„ฑ์ ์„ ๊ฑฐ๋‘๋ฉฐ ๋ชจ๋“  ๋ชจ๋ธ์ด ํ•ฉ๊ฒฉ ๊ธฐ์ค€์„ ์ƒํšŒํ•˜์˜€์Šต๋‹ˆ๋‹ค. LLM์€ ์˜์ƒ์˜ํ•™๊ณผ ํ•™์Šต ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ ๋†’์€ ์ž ์žฌ๋ ฅ์„ ๋ณด์˜€์œผ๋‚˜, ์ž˜๋ชป๋œ ์ •๋ณด๋ฅผ ํ™•์‹ ์— ์ฐจ์„œ ๋‹ต๋ณ€ํ•  ์ˆ˜ ์žˆ๋Š” ํ•œ๊ณ„์ ์— ๋Œ€ํ•œ ์ฃผ์˜๊ฐ€ ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

55Personalized Impression Generation for PET Reports Using Large Language Models

2024Journal of Imaging Informatics in Medicineโญ Q1DOI 10.1007/s10278-024-00985-3

Large language models (LLMs) have shown promise in accelerating radiology reporting by summarizing clinical findings into impressions. However, automatic impression generation for whole-body PET reports presents unique challenges and has received little attention. Our study aimed to evaluate whether LLMs can create clinically useful impressions for PET reporting. To this end, we fine-tuned twelve open-source language models on a corpus of 37,370 retrospective PET reports collected from our institution. All models were trained using the teacher-forcing algorithm, with the report findings and patient information as input and the original clinical impressions as reference. An extra input token encoded the reading physician's identity, allowing models to learn physician-specific reporting styles. To compare the performances of different models, we computed various automatic evaluation metrics and benchmarked them against physician preferences, ultimately selecting PEGASUS as the top LLM. To evaluate its clinical utility, three nuclear medicine physicians assessed the PEGASUS-generated impressions and original clinical impressions across 6 quality dimensions (3-point scales) and an overall utility score (5-point scale). Each physician reviewed 12 of their own reports and 12 reports from other physicians. When physicians assessed LLM impressions generated in their own style, 89% were considered clinically acceptable, with a mean utility score of 4.08/5. On average, physicians rated these personalized impressions as comparable in overall utility to the impressions dictated by other physicians (4.03, Pโ€‰=โ€‰0.41). In summary, our study demonstrated that personalized impressions generated by PEGASUS were clinically useful in most cases, highlighting its potential to expedite PET reporting by automatically drafting impressions.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” 37,370๊ฑด์˜ PET ํŒ๋… ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•˜์—ฌ ํŒ๋…์˜๋ณ„ ๊ณ ์œ  ์Šคํƒ€์ผ์„ ํ•™์Šตํ•œ ๋งž์ถคํ˜• ์ธ์ƒ(impression) ์ƒ์„ฑ ๋ชจ๋ธ์ธ PEGASUS๋ฅผ ๊ฐœ๋ฐœํ•˜๊ณ  ๊ทธ ์ž„์ƒ์  ์œ ์šฉ์„ฑ์„ ํ‰๊ฐ€ํ–ˆ์Šต๋‹ˆ๋‹ค. ํ‰๊ฐ€ ๊ฒฐ๊ณผ, ๋ชจ๋ธ์ด ์ƒ์„ฑํ•œ ๋งž์ถคํ˜• ์ธ์ƒ์€ 89%์˜ ์ž„์ƒ์  ์ˆ˜์šฉ๋„๋ฅผ ๋ณด์˜€์œผ๋ฉฐ, ํ•ต์˜ํ•™๊ณผ ์ „๋ฌธ์˜๊ฐ€ ์ง์ ‘ ์ž‘์„ฑํ•œ ํŒ๋…๋ฌธ๊ณผ ๋Œ€๋“ฑํ•œ ์ˆ˜์ค€์˜ ์œ ์šฉ์„ฑ์„ ๋‚˜ํƒ€๋ƒˆ์Šต๋‹ˆ๋‹ค. ์ด๋Š” ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์ด PET ํŒ๋… ๊ณผ์ •์—์„œ ํšจ์œจ์ ์ด๊ณ  ๊ฐœ์ธํ™”๋œ ์ดˆ์•ˆ์„ ์ œ๊ณตํ•จ์œผ๋กœ์จ ์—…๋ฌด ๋ถ€๋‹ด์„ ๊ฒฝ๊ฐํ•  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

56Domain-adapted Large Language Models for Classifying Nuclear Medicine Reports.

2023-09-27Radiology. Artificial intelligenceโญ Q1DOI 10.1148/ryai.220281
OBJECTIVE

To evaluate the impact of domain adaptation on the performance of language models in predicting five-point Deauville scores on the basis of clinical fluorine 18 fluorodeoxyglucose PET/CT reports.

METHODS

The authors retrospectively retrieved 4542 text reports and images for fluorodeoxyglucose PET/CT lymphoma examinations from 2008 to 2018 in the University of Wisconsin-Madison institutional clinical imaging database. Of these total reports, 1664 had Deauville scores that were extracted from the reports and served as training labels. The bidirectional encoder representations from transformers (BERT) model and initialized BERT models BioClinicalBERT, RadBERT, and RoBERTa were adapted to the nuclear medicine domain by pretraining using masked language modeling. These domain-adapted models were then compared with the non-domain-adapted versions on the task of five-point Deauville score prediction. The language models were compared against vision models, multimodal vision-language models, and a nuclear medicine physician, with sevenfold Monte Carlo cross-validation. Means and SDs for accuracy are reported, with P values from paired t testing.

RESULTS

Domain adaptation improved the performance of all language models (P = .01). For example, BERT improved from 61.3% ยฑ 2.9 (SD) five-class accuracy to 65.7% ยฑ 2.2 (P = .01) following domain adaptation. Domain-adapted RoBERTa (named DA RoBERTa) performed best, achieving 77.4% ยฑ 3.4 five-class accuracy; this model performed similarly to its multimodal counterpart (named Multimodal DA RoBERTa) (77.2% ยฑ 3.2) and outperformed the best vision-only model (48.1% ยฑ 3.5, P โ‰ค .001). A physician given the task on a subset of the data had a five-class accuracy of 66%.

CONCLUSION

Domain adaptation improved the performance of large language models in predicting Deauville scores in PET/CT reports.Keywords Lymphoma, PET, PET/CT, Transfer Learning, Unsupervised Learning, Convolutional Neural Network (CNN), Nuclear Medicine, Deauville, Natural Language Processing, Multimodal Learning, Artificial Intelligence, Machine Learning, Language Modeling Supplemental material is available for this article. ยฉ RSNA, 2023See also the commentary by Abajian in this issue.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฆผํ”„์ข… ํ™˜์ž์˜ 18F-FDG PET/CT ํŒ๋…๋ฌธ์—์„œ ๋„๋นŒ ์ ์ˆ˜(Deauville score)๋ฅผ ์˜ˆ์ธกํ•˜๊ธฐ ์œ„ํ•ด ๋Œ€๊ทœ๋ชจ ์–ธ์–ด ๋ชจ๋ธ์— ํ•ต์˜ํ•™ ๋„๋ฉ”์ธ ์ ์‘(domain adaptation)์„ ์ ์šฉํ•œ ํšจ๊ณผ๋ฅผ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์—ฐ๊ตฌ ๊ฒฐ๊ณผ, ๋„๋ฉ”์ธ ์ ์‘์„ ๊ฑฐ์นœ RoBERTa ๋ชจ๋ธ์ด 77.4%์˜ ์ •ํ™•๋„๋กœ ๊ฐ€์žฅ ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋ฉฐ, ์ด๋Š” ๊ธฐ์กด์˜ ๋น„์ ์‘ ๋ชจ๋ธ ๋ฐ ์˜์ƒ ๊ธฐ๋ฐ˜ ๋ชจ๋ธ๋ณด๋‹ค ๋›ฐ์–ด๋‚œ ์˜ˆ์ธก๋ ฅ์„ ๋‚˜ํƒ€๋ƒˆ์Šต๋‹ˆ๋‹ค. ๊ฒฐ๋ก ์ ์œผ๋กœ ๋„๋ฉ”์ธ ์ ์‘์€ ํ•ต์˜ํ•™ ํŒ๋…๋ฌธ ๋ถ„์„์„ ์œ„ํ•œ ์–ธ์–ด ๋ชจ๋ธ์˜ ์„ฑ๋Šฅ์„ ์œ ์˜๋ฏธํ•˜๊ฒŒ ํ–ฅ์ƒ์‹œํ‚ค๋Š” ํšจ๊ณผ์ ์ธ ์ „๋žต์ž„์„ ํ™•์ธํ•˜์˜€์Šต๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

57Benchmarking ChatGPT-4 on a radiation oncology in-training exam and Red Journal Gray Zone cases: potentials and challenges for ai-assisted medical education and decision making in radiation oncology.

2023-09-14Frontiers in oncology๐Ÿ”ท Q2DOI 10.3389/fonc.2023.1265024
OBJECTIVE

The potential of large language models in medicine for education and decision-making purposes has been demonstrated as they have achieved decent scores on medical exams such as the United States Medical Licensing Exam (USMLE) and the MedQA exam. This work aims to evaluate the performance of ChatGPT-4 in the specialized field of radiation oncology.

METHODS

The 38th American College of Radiology (ACR) radiation oncology in-training (TXIT) exam and the 2022 Red Journal Gray Zone cases are used to benchmark the performance of ChatGPT-4. The TXIT exam contains 300 questions covering various topics of radiation oncology. The 2022 Gray Zone collection contains 15 complex clinical cases.

RESULTS

For the TXIT exam, ChatGPT-3.5 and ChatGPT-4 have achieved the scores of 62.05% and 78.77%, respectively, highlighting the advantage of the latest ChatGPT-4 model. Based on the TXIT exam, ChatGPT-4's strong and weak areas in radiation oncology are identified to some extent. Specifically, ChatGPT-4 demonstrates better knowledge of statistics, CNS & eye, pediatrics, biology, and physics than knowledge of bone & soft tissue and gynecology, as per the ACR knowledge domain. Regarding clinical care paths, ChatGPT-4 performs better in diagnosis, prognosis, and toxicity than brachytherapy and dosimetry. It lacks proficiency in in-depth details of clinical trials. For the Gray Zone cases, ChatGPT-4 is able to suggest a personalized treatment approach to each case with high correctness and comprehensiveness. Importantly, it provides novel treatment aspects for many cases, which are not suggested by any human experts.

CONCLUSION

Both evaluations demonstrate the potential of ChatGPT-4 in medical education for the general public and cancer patients, as well as the potential to aid clinical decision-making, while acknowledging its limitations in certain domains. Owing to the risk of hallucinations, it is essential to verify the content generated by models such as ChatGPT for accuracy.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฐฉ์‚ฌ์„  ์ข…์–‘ํ•™ ์ „๋ฌธ์˜ ์ˆ˜๋ จ ์‹œํ—˜(TXIT)๊ณผ ๋ณตํ•ฉ ์ž„์ƒ ์‚ฌ๋ก€(Gray Zone cases)๋ฅผ ํ†ตํ•ด ChatGPT-4์˜ ์˜ํ•™์  ์ง€์‹ ๋ฐ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ๋Šฅ๋ ฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. ํ‰๊ฐ€ ๊ฒฐ๊ณผ, ChatGPT-4๋Š” ์ˆ˜๋ จ ์‹œํ—˜์—์„œ 78.77%์˜ ๋†’์€ ์ •๋‹ต๋ฅ ์„ ๋ณด์˜€์œผ๋ฉฐ ๋ณตํ•ฉ ์ž„์ƒ ์‚ฌ๋ก€์—์„œ๋„ ํฌ๊ด„์ ์ธ ์น˜๋ฃŒ ์ „๋žต์„ ์ œ์‹œํ•˜๋Š” ๋“ฑ ๊ต์œก ๋ฐ ์ž„์ƒ ๋ณด์กฐ ๋„๊ตฌ๋กœ์„œ์˜ ๊ฐ€๋Šฅ์„ฑ์„ ์ž…์ฆํ–ˆ์Šต๋‹ˆ๋‹ค. ๋‹ค๋งŒ, ํŠน์ • ์„ธ๋ถ€ ๋ถ„์•ผ์˜ ์ง€์‹ ๋ถ€์กฑ๊ณผ ํ™˜๊ฐ ํ˜„์ƒ(hallucination)์˜ ์œ„ํ—˜์ด ์กด์žฌํ•˜๋ฏ€๋กœ ์ž„์ƒ ์ ์šฉ ์‹œ ์ƒ์„ฑ๋œ ์ •๋ณด์˜ ์ •ํ™•์„ฑ์— ๋Œ€ํ•œ ์ „๋ฌธ๊ฐ€์˜ ๊ฒ€์ฆ์ด ํ•„์ˆ˜์ ์ž…๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

58Performance of ChatGPT on a Radiology Board-style Examination: Insights into Current Strengths and Limitations.

2023-05-16Radiologyโญ Q1DOI 10.1148/radiol.230582

Background ChatGPT is a powerful artificial intelligence large language model with great potential as a tool in medical practice and education, but its performance in radiology remains unclear. Purpose To assess the performance of ChatGPT on radiology board-style examination questions without images and to explore its strengths and limitations. Materials and Methods In this exploratory prospective study performed from February 25 to March 3, 2023, 150 multiple-choice questions designed to match the style, content, and difficulty of the Canadian Royal College and American Board of Radiology examinations were grouped by question type (lower-order [recall, understanding] and higher-order [apply, analyze, synthesize] thinking) and topic (physics, clinical). The higher-order thinking questions were further subclassified by type (description of imaging findings, clinical management, application of concepts, calculation and classification, disease associations). ChatGPT performance was evaluated overall, by question type, and by topic. Confidence of language in responses was assessed. Univariable analysis was performed. Results ChatGPT answered 69% of questions correctly (104 of 150). The model performed better on questions requiring lower-order thinking (84%, 51 of 61) than on those requiring higher-order thinking (60%, 53 of 89) (P = .002). When compared with lower-order questions, the model performed worse on questions involving description of imaging findings (61%, 28 of 46; P = .04), calculation and classification (25%, two of eight; P = .01), and application of concepts (30%, three of 10; P = .01). ChatGPT performed as well on higher-order clinical management questions (89%, 16 of 18) as on lower-order questions (P = .88). It performed worse on physics questions (40%, six of 15) than on clinical questions (73%, 98 of 135) (P = .02). ChatGPT used confident language consistently, even when incorrect (100%, 46 of 46). Conclusion Despite no radiology-specific pretraining, ChatGPT nearly passed a radiology board-style examination without images; it performed well on lower-order thinking questions and clinical management questions but struggled with higher-order thinking questions involving description of imaging findings, calculation and classification, and application of concepts. ยฉ RSNA, 2023 See also the editorial by Lourenco et al and the article by Bhayana et al in this issue.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ์˜์ƒ์˜ํ•™ ์ „๋ฌธ์˜ ์‹œํ—˜ ์Šคํƒ€์ผ์˜ ๋ฌธํ•ญ์„ ํ†ตํ•ด ChatGPT์˜ ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•œ ๊ฒฐ๊ณผ, ์ „์ฒด ๋ฌธํ•ญ์˜ 69% ์ •๋‹ต๋ฅ ์„ ๋ณด์ด๋ฉฐ ํ•ฉ๊ฒฉ๊ถŒ์— ๊ทผ์ ‘ํ•œ ์„ฑ๊ณผ๋ฅผ ํ™•์ธํ–ˆ์Šต๋‹ˆ๋‹ค. ChatGPT๋Š” ๋‹จ์ˆœ ์ง€์‹ ๊ธฐ๋ฐ˜์˜ ํ•˜์œ„ ์‚ฌ๊ณ ๋ ฅ ๋ฌธํ•ญ๊ณผ ์ž„์ƒ ๊ด€๋ฆฌ ๋ฌธ์ œ์—์„œ๋Š” ์šฐ์ˆ˜ํ•œ ์„ฑ๋Šฅ์„ ๋ณด์˜€์œผ๋‚˜, ์˜์ƒ ์†Œ๊ฒฌ ๊ธฐ์ˆ , ๊ณ„์‚ฐ ๋ฐ ๊ฐœ๋… ์ ์šฉ์ด ํ•„์š”ํ•œ ๊ณ ์ฐจ์›์  ์‚ฌ๊ณ  ๋ฌธํ•ญ๊ณผ ๋ฌผ๋ฆฌํ•™ ๋ถ„์•ผ์—์„œ๋Š” ์ƒ๋Œ€์ ์œผ๋กœ ์ทจ์•ฝํ–ˆ์Šต๋‹ˆ๋‹ค. ํŠนํžˆ ์˜ค๋‹ต์„ ์ œ์‹œํ•  ๋•Œ๋„ ์ผ๊ด€๋˜๊ฒŒ ํ™•์‹ ์— ์ฐฌ ์–ด์กฐ๋ฅผ ๋ณด์ด๋Š” ๊ฒฝํ–ฅ์ด ์žˆ์–ด, ์˜๋ฃŒ ๊ต์œก ๋ฐ ์‹ค๋ฌด ํ™œ์šฉ ์‹œ ์ฃผ์˜๊ฐ€ ํ•„์š”ํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—

59A Context-based Chatbot Surpasses Radiologists and Generic ChatGPT in Following the ACR Appropriateness Guidelines

2023Radiologyโญ Q1DOI 10.1148/radiol.230970

Background Radiological imaging guidelines are crucial for accurate diagnosis and optimal patient care as they result in standardized decisions and thus reduce inappropriate imaging studies. Purpose In the present study, we investigated the potential to support clinical decision-making using an interactive chatbot designed to provide personalized imaging recommendations from American College of Radiology (ACR) appropriateness criteria documents using semantic similarity processing. Methods We utilized 209 ACR appropriateness criteria documents as specialized knowledge base and employed LlamaIndex, a framework that allows to connect large language models with external data, and the ChatGPT 3.5-Turbo to create an appropriateness criteria contexted chatbot (accGPT). Fifty clinical case files were used to compare the accGPT's performance against general radiologists at varying experience levels and to generic ChatGPT 3.5 and 4.0. Results All chatbots reached at least human performance level. For the 50 case files, the accGPT performed best in providing correct recommendations that were "usually appropriate" according to the ACR criteria and also did provide the highest proportion of consistently correct answers in comparison with generic chatbots and radiologists. Further, the chatbots provided substantial time and cost savings, with an average decision time of 5 minutes and a cost of 0.19 โ‚ฌ for all cases, compared to 50 minutes and 29.99 โ‚ฌ for radiologists (both p < 0.01). Conclusion ChatGPT-based algorithms have the potential to substantially improve the decision-making for clinical imaging studies in accordance with ACR guidelines. Specifically, a context-based algorithm performed superior to its generic counterpart, demonstrating the value of tailoring AI solutions to specific healthcare applications.

๐Ÿ‡ฐ๐Ÿ‡ท ํ•ต์‹ฌ ์š”์•ฝ
๋ณธ ์—ฐ๊ตฌ๋Š” ๋ฏธ๊ตญ์˜์ƒ์˜ํ•™ํšŒ(ACR) ์ ์ •์„ฑ ๊ธฐ์ค€์„ ํ•™์Šต์‹œํ‚จ ๋ฌธ๋งฅ ๊ธฐ๋ฐ˜ ์ฑ—๋ด‡(accGPT)์˜ ์ž„์ƒ ์˜์‚ฌ๊ฒฐ์ • ์ง€์› ์„ฑ๋Šฅ์„ ํ‰๊ฐ€ํ•˜์˜€์Šต๋‹ˆ๋‹ค. 50๊ฐœ์˜ ์ž„์ƒ ์‚ฌ๋ก€๋ฅผ ๋ถ„์„ํ•œ ๊ฒฐ๊ณผ, accGPT๋Š” ์ผ๋ฐ˜ ์˜์ƒ์˜ํ•™๊ณผ ์ „๋ฌธ์˜ ๋ฐ ๋ฒ”์šฉ ChatGPT ๋ชจ๋ธ๋ณด๋‹ค ACR ๊ฐ€์ด๋“œ๋ผ์ธ์— ๋ถ€ํ•ฉํ•˜๋Š” ์ •ํ™•ํ•œ ๊ถŒ๊ณ ์•ˆ์„ ๋” ์ผ๊ด€๋˜๊ฒŒ ์ œ์‹œํ•˜์˜€์œผ๋ฉฐ, ์˜์‚ฌ๊ฒฐ์ •์— ์†Œ์š”๋˜๋Š” ์‹œ๊ฐ„๊ณผ ๋น„์šฉ์„ ์œ ์˜๋ฏธํ•˜๊ฒŒ ์ ˆ๊ฐํ•˜์˜€์Šต๋‹ˆ๋‹ค. ์ด๋Š” ํŠนํ™”๋œ ์˜๋ฃŒ ๋ฐ์ดํ„ฐ๋ฅผ ํ™œ์šฉํ•œ ๋ฌธ๋งฅ ๊ธฐ๋ฐ˜ AI ๋ชจ๋ธ์ด ์˜์ƒ ๊ฒ€์‚ฌ ์ ์ •์„ฑ ํŒ๋‹จ์˜ ํšจ์œจ์„ฑ๊ณผ ์ •ํ™•๋„๋ฅผ ๋†’์ด๋Š” ๋ฐ ํšจ๊ณผ์ ์ธ ๋„๊ตฌ๊ฐ€ ๋  ์ˆ˜ ์žˆ์Œ์„ ์‹œ์‚ฌํ•ฉ๋‹ˆ๋‹ค.
Added: 2026-04-04 06:48View โ†—
โ†‘ Back to top
0 selected