(radiology[tiab] OR "radiology report*"[tiab]) AND ("large language model*"[tiab] OR LLM[tiab] OR GPT[tiab]) AND (report*[tiab] OR impression[tiab] OR conclusion[tiab]) AND (style[tiab] OR preference*[tiab] OR personalized[tiab] OR personalization[tiab] OR imitation[tiab] OR adaptation[tiab]) AND (radiologist*[tiab] OR reader*[tiab] OR physician*[tiab] OR "individual radiologist"[tiab] OR "expert feedback"[tiab] OR "human feedback"[tiab] OR fine-tun*[tiab] OR finetun*[tiab] OR LoRA[tiab] OR PEFT[tiab])
Artificial intelligence (AI), particularly large language models (LLMs), is increasingly applied to radiographic interpretation in healthcare. In dentistry, radiographic imaging is essential for diagnosis and treatment planning, yet remains subject to variability and human error. AI may enhance diagnostic accuracy and consistency.
To evaluate and compare the diagnostic accuracy, consistency, and interpretive performance of multimodal AI models-ChatGPT, Grok, and MANUS-with expert radiologists in dental radiograph interpretation.
A total of 120 anonymised radiographs (40 OPGs, 40 periapical, 40 CT slices) were selected from validated academic sources. Two board-certified oral and maxillofacial radiologists established gold standard diagnoses. Each image was independently assessed by the three AI models under standardised prompting. Diagnostic accuracy and intra-model consistency were analysed using descriptive statistics, Cohen's kappa, McNemar's test, and logistic regression.
In the first assessment, MANUS and ChatGPT achieved 92.5% accuracy (111/120), while Grok reached 88.3% (106/120). Performance improved in the second round: MANUS 95.0%, ChatGPT 93.3%, and Grok 90.8%, compared with 96.7% for radiologists. ChatGPT showed the highest reproducibility (ฮบ = 0.937), whereas MANUS demonstrated the highest overall accuracy. Strong agreement was observed between ChatGPT and MANUS, with greater variability in Grok. No significant systematic bias was detected between AI outputs and radiologist benchmarks.
The evaluated LLMs demonstrated diagnostic performance comparable to expert radiologists. MANUS excelled in accuracy and ChatGPT in reproducibility, supporting their potential as adjunct tools in dental radiology, while maintaining the need for expert oversight.Clinical trial number: Not applicable.
Large language models (LLMs) show promise in automatically detecting errors in radiology reports, but their performance remains insufficiently validated in large-scale, real-world clinical datasets.
This study aimed to systematically evaluate the performance of LLMs in detecting and correcting errors in Chinese radiology reports derived from authentic clinical data.
A large-scale dataset of 4480 Chinese radiology reports with modification records containing real clinical practice-generated errors was retrospectively collected between January 2023 and June 2024 at a single institution. After exclusions, 1363 reports containing 1551 errors were included. The dataset covers various anatomical parts of the body from different imaging modalities and was randomly divided into a test set (n=1263) and an internal validation set (n=100). Additionally, 100 error-free reports were added to the internal validation set. An additional 200 English-language reports from the Medical Information Mart for Intensive Care (MIMIC-III) were used for external validation. Eight human readers and 8 widely adopted LLMs, enhanced by prompt engineering, were tasked with error detection. Overall and subgroup detection performance and reading time were evaluated. Correction suggestions from the 2 best-performing LLMs were reviewed by a senior radiologist.
On the test set, DeepSeek-R1 achieved the highest overall detection rate at 89% (95% CI 87%-90%), significantly better than the other 7 models (P=.001-.007). On the internal validation set, DeepSeek-R1 and Claude-3.5-Sonnet achieved detection rates of 83% (100/120; 95% CI 76%-89%) and 80% (96/120; 95% CI 72%-86%), respectively. DeepSeek-R1 showed performance comparable to radiologists (83%, 95% CI 76%-89% vs 80%, 95% CI 72%-86% for junior radiologists and 78%, 95% CI 70%-85% for senior radiologists; P=.39 and P=.19, respectively) and significantly better performance than that of nonradiologists and nonphysicians (83%, 95% CI 76%-89% vs 66%, 95% CI 57%-74% and 38%, 95% CI 30%-47%; P<.001, respectively). DeepSeek-R1 showed a false-positive rate comparable to radiologists (DeepSeek-R1 vs senior radiologists and junior radiologists, 3% vs 0% and 1%; P=.25 and P=.61, respectively) and a significantly lower rate than nonradiologists and nonphysicians (3% vs 13% and 17%; P=.02 and P=.002, respectively). On the external validation set, DeepSeek-R1 and Claude-3.5-Sonnet achieved detection rates of 94% (95% CI 89%-97%) and 93% (95% CI 88%-97%), respectively. The correction accuracy of DeepSeek-R1 and Claude-3.5-Sonnet was 95% and 91%, respectively.
Enhanced LLMs, particularly DeepSeek-R1, demonstrated robust performance in error detection and correction within real-world Chinese radiology reports, supporting their clinical use for automated quality assurance and integration into workflows to improve reporting accuracy and efficiency.
Information extraction (IE) from clinical texts has advanced rapidly with recent advances in natural language processing, particularly the advent of large language models (LLMs). However, inconsistent and incomplete reporting of methodologies limits reproducibility, comparability, and clinical translation. We aimed to develop a consensus-based reporting guideline tailored to clinical IE studies.
We developed the Clinical Information Extraction Reporting Guideline (CINEX) through a multi-phase process. The initiative was prospectively registered on the EQUATOR Network as a reporting guideline under development, and a detailed Delphi study protocol was published in advance. A scoping review informed an initial set of items, which was refined through a 3-round electronic Delphi study with 20 international experts, followed by a final consensus meeting. Items were iteratively refined based on predefined inclusion criteria and expert feedback.
The final CINEX guideline comprises 29 checklist items grouped into 5 domains: information model, architecture, data, annotation, and outcomes. The 3 eDelphi rounds included 20, 15, and 12 experts, respectively. Two items were added after round one. Consensus for inclusion was reached for a 21 and additional 7 items after rounds 2 and 3, respectively.
CINEX provides a structured framework to improve transparency, reproducibility, and interpretability in clinical IE research. By standardizing reporting of key methodological components, such as data provenance, annotation processes, and evaluation strategies, it facilitates meaningful comparison across studies and supports safer clinical implementation.
CINEX complements existing AI reporting standards by addressing domain-specific challenges in clinical IE.
Radiology reports are written for clinicians, leaving the majority of patients unable to understand their own imaging findings. Large language models (LLMs) offer a means to simplify reports for patient-facing use; however, most implementations rely on cloud-based platforms that raise data privacy and regulatory concerns. Arabic-speaking populations - over 400 million people worldwide - remain substantially underserved in medical AI research, with no prospectively evaluated AI Arabic radiology report simplification system previously reported.
This study aimed to conduct a preliminary prospective feasibility and safety evaluation of patient-perceived understandability outcomes and radiologist-assessed clinical safety of an on-site, institutionally governed AI system that generates simplified radiology reports in plain English and Arabic for Arabic-speaking outpatients.
In this prospective single-center observational study conducted at King Faisal Specialist Hospital and Research Centre - Jeddah (KFSHRC-J) in January 2026, 98 adult outpatients (mean age 50.0ย ยฑย 14.9ย years; 50 men, 48 women; IRB #2251488) reviewed three versions of their own radiology report in randomized order: a traditional radiologist report, an AI-simplified English report, and an AI-simplified Arabic report. The system used a two-stage pipeline: Qwen3-14B-FP8 with structured prompt engineering for English simplification (Stage 1), and a fine-tuned Hala-1.2B model for Arabic translation (Stage 2; 8,319 training samples). Patient-perceived understandability was measured on a 5-point Likert scale (1ย =ย Very Easy; 5ย =ย Very Difficult). Three board-certified radiologists independently assessed clinical safety using a standardized seven-domain rubric; inter-rater agreement was quantified using ICC(2,k).
AI-simplified Arabic reports produced a median perceived understandability score of 1 (IQR 1-2) versus 5 (IQR 3-5) for traditional reports (pย <ย 0.001; rank-biserial rย =ย 0.97), with 94.9% (93/98) of patients preferring the Arabic simplified version. Safety evaluation demonstrated 96.9% (95/98) of reports safe for patient release, with a hallucination rate of 3.1% (nย =ย 3; 1 unsafe, 2 safe by consensus). Inter-rater reliability was moderate by conventional thresholds (ICC 0.48-0.55) due to a ceiling effect from uniformly high safety scores (mean fidelity 4.8/5.0; 90.8% highest rating; SDย =ย 0.3), with 91.2% absolute agreement on binary safety classification and Fleiss' kappaย =ย 0.71 for safety.
In this preliminary prospective feasibility evaluation, a locally deployed, clinician- and AI team-built system significantly improved patient-perceived understandability of radiology reports for Arabic-speaking patients, with a radiologist-assessed safety profile that is promising but requires further validation before broad clinical deployment. The hallucination rate of 3.1% (1.0% unsafe) is lower than published pooled benchmarks, though direct comparisons are limited by methodological heterogeneity across studies (pooled 7.2% across 38 simplification studies; 1.12% for fine-tuned models). To our knowledge, this is among the first prospectively evaluated AI systems for Arabic radiology report simplification, highlighting the potential of on-site AI deployment as a safe, privacy-preserving approach for multilingual radiology communication, warranting further investigation of its contribution to health equity.
RATIONALE AND
Photon -Counting CT (PCCT) provides higher resolution, reduced dose, and valuable spectral data, but generates a large number of images per study, underlining radiologist and Picture Archiving and Communication Systemsย (PACS) capacity. While no guidelines defining when PCCT should be preferred over conventional Energy Integrating Detector (EID) CT in neuroradiology, and manual decision-making being impractical, Natural Language Processing (NLP)-based large language models (LLMs) may help automate routing of CT requests to the most appropriate CT technology.
We conducted a retrospective study using Spanish-language neuroradiology CT requests retrieved from the Radiology Information System (RIS) (January 2012-October 2025). A random sample of 800 requests was independently labeled by two neuroradiologists as Basic Protocol (BP); feasible on EID-CT or basic PCCT) or Advanced Protocol (AP); requiring full-advanced PCCT), considering the gold standard an expert consensus (Cohen's k = 0.74). For discriminative modeling, transformer classifiers (Spanish bidirectional encoder representations from transformers (BERT), biomedical RoBERTa, Llama-3.1-8B, and Mistral-7B) were fine-tuned for binary BP/AP classification and evaluated with stratified five-fold cross-validation. For generative modeling, ChatGPT 5.2 and Gemini 3 were applied in a zero-shot setting using structured prompting and post-processed into binary outputs. Performance was assessed using accuracy, precision, recall, specificity, F1-score, and AUC.
Discriminative transformers achieved the best performance for the classification, with RoBERTa leading (accuracy 0.925, precision 0.906, recall 0.967, F1 0.935, AUC 0.920) and the lowest total error, without a significant difference versus Spanish BERT. In contrast, RoBERTa outperformed Llama, Mistral, Gemini, and ChatGPT with statistically significant improvements, while the generative LLMs yielded the poorest overall performance.
Discriminative transformer models, particularly RoBERTa, enabled highly accurate automatic identification of CT requests requiring AP to be performed at PCCT, substantially outperforming general-purpose generative LLMs. These findings support this exploratory proof-of-concept study of AI-driven eligibility routing to optimize PCCT utilization and streamline neuroradiology workflows.
Large language models (LLMs) show promise for converting complex radiology reports into patient-centric language, but inherent output instability may limit clinical application.
To quantitatively assess the translational accuracy, error rates, and instability of various LLMs when generating patient-centric radiology reports, and evaluate demographic influences on report readability.
This retrospective study evaluated 320 de-identified radiology reports processed by three LLMs using a two-stage (baseline and optimized) prompt engineering strategy. Two senior radiologists evaluated medical accuracy, completeness, and recommendation suitability. Readability was evaluated by 16 non-medical participants stratified by age and education.
Professional radiological evaluation revealed that all tested models exhibited inherent instability, omitted information, and tended to generate risk-averse, generalized clinical recommendations. To address these limitations, optimized structured prompts significantly reduced model output variance and improved translational accuracy, with particularly prominent effects observed in DeepSeek-R1 and ChatGPT-4.0. Overall, large language models significantly enhanced the readability of radiology reports (P < 0.05), with DeepSeek-R1 achieving the best performance. However, patients' self-reported comprehension of the reports was affected by demographic characteristics.
Large language models can effectively improve the readability of radiology reports, yet all such models inherently suffer from output instability and information omission. Optimized structured prompting can substantially reduce the variability of model outputs and improve the accuracy of medical text translation. Nevertheless, LLMs should currently be strictly confined to human-supervised auxiliary tools rather than applied as standalone clinical solutions.
Extracting disease labels from radiology reports is essential for developing deep learning-based diagnostic models and enabling large-scale retrospective clinical research. Classification of usual interstitial pneumonia (UIP) patterns from high-resolution computed tomography (HRCT) reports according to Fleischner Society guidelines is a particularly demanding task, requiring synthesis of spatial distribution, fibrotic features, and exclusion criteria. As open-source large language models (LLMs) are released at an accelerating pace with steadily improving general benchmarks, a practical question arises: Do these improvements translate to better performance on complex, real-world clinical classification, and does the optimal prompting strategy differ across model architectures? While prior studies have evaluated LLMs for radiology report labeling, none have compared how prompting strategies interact with the native reasoning capabilities of newer model architectures. We evaluated 10 open-source LLMs from three architecture families (Llama, Qwen, Gemma) spanning 8 to 405 billion parameters, each tested with three prompting strategies on 270 HRCT reports classified by expert consensus of two senior thoracic radiologists. Four models with native reasoning ("thinking") capability were additionally tested in thinking mode. The best configuration achieved a Cohen's kappa (ฮบ) of 0.70 and 82% four-class accuracy. Structured reasoning prompting improved all Llama models but degraded all models with native reasoning capability (Qwen 3.5 and Gemma 4), revealing an architecture-dependent interaction. Thinking mode hurt performance on criteria-based prompts and never yielded the best configuration. Larger models did not consistently outperform smaller ones: Llama 3.1 405B offered no advantage over Llama 3.3 70B, and the Qwen 397B model underperformed the dense Qwen 27B. These findings demonstrate that newer model generations with improved general benchmarks and larger parameter counts do not guarantee better performance on specialized medical classification tasks and that prompt design must be matched to model architecture.
Background Clinical histories accompanying imaging orders guide protocol selection and diagnostic focus. However, they are often incomplete, potentially compromising diagnostic accuracy and workflow efficiency. Purpose To evaluate whether large language models (LLMs) can improve the clinical utility of provided imaging indications by leveraging clinical notes. Materials and Methods This retrospective study curated a dataset from deidentified electronic health records at the University of California San Francisco (January 2012 to August 2024), consisting of radiology reports with paired referring clinician-provided and radiologist-curated indications linked to clinical notes. The dataset was stratified across five body systems and five pathophysiologic categories to derive LLM selection and reader study internal test sets. For the reader study, 20 radiologists with 2-25 years of experience compared indications from the referring clinician, radiologist, and best-performing LLMs. Readers scored comprehensiveness, factuality, and conciseness and ranked indications for usefulness in protocoling, usefulness in interpretation, and overall ranking. Models and clinicians were compared using cumulative link mixed models with Tukey-adjusted post hoc comparisons. Results From 28โ313 patients (mean age, 59 years ยฑ 20.6 [SD]; 14โ912 women), 250 examinations from 247 patients were sampled for the reader study. After nine exclusions, 241 examinations were analyzed, yielding 482 reader-examination evaluations. Indications from the best-performing proprietary (Claude 3.5 Sonnet; Anthropic) and open-source (Qwen 2.5-7B Instruct; Alibaba) LLM were rated as more comprehensive (Likert rating of 5: 37.14% and 28.42%, respectively; both P < .001) and factual (68.05% and 59.75%; both P < .001) than referring clinician indications. The proprietary LLM ranked most useful in protocoling (rank 1: 40.87%; all P < .001), useful in interpretation (44.61%; all P < .001), and overall ranking (44.19%, all P < .001). Comprehensiveness (65.77% of ratings; both P < .001) most strongly influenced overall rankings. Conclusion LLMs generated radiology-relevant indications from clinical notes that were more comprehensive and factual than clinician indications, and when generated by the proprietary LLM, were ranked most useful in protocoling and imaging interpretation. ยฉ RSNA, 2026 Supplemental material is available for this article. See also the editorial by Yilmaz and Cardoza-Ochoa in this issue.
To evaluate text-only versus vision-enabled performance of late-2025 large language models (LLMs) on the Japan Diagnostic Radiology Board Examination (JDRBE) and compare model performance with newly board-certified radiologists.
Image-based questions from the JDRBE 2021 and 2023โ2025 were collected, and ground truth answers were determined by expert consensus. Four commercial multimodal LLMs were evaluated: Gemini 2.5 Pro (March 2025, baseline), Gemini 3 Pro, GPT-5.1, and Claude Opus 4.5 (all released in November 2025). Each question was answered with image input (โvisionโ) and without images (โtext-onlyโ). For the JDRBE 2025, subjective legitimacy of responses was independently rated by two radiologists using a five-point Likert scale, and low-rated responses were further analyzed by error type. Additional analyses on the JDRBE 2025 subset included image-shuffling and multi-run variability assessment (five runs). Model accuracies were also compared with those of five newly board-certified radiologists who passed the JDRBE 2025.
Gemini 3 Pro achieved the highest accuracy among all models, scoring 85.3% (279/327) in the vision condition and significantly outperforming its text-only accuracy (74.3%, Pโ<โ0.001). Gemini 2.5 Pro and Claude Opus 4.5 also improved with image input, whereas GPT-5.1 did not. For the JDRBE 2025, Gemini 3 Pro in the vision condition received the highest legitimacy ratings, and its accuracy (88%) was above the range observed in a reference group of five newly board-certified radiologists (65%โ83%), but hallucination was still the most common error type. Image-shuffling analysis using the 2025 subset showed no performance gain in all models, supporting reliance on visual input. Multi-run variability analysis showed high agreement across runs.
Among late-2025 commercial LLMs, Gemini 3 Pro demonstrated board-level performance on the JDRBE through direct medical image interpretation. SECONDARY ABSTRACT: The performance of vision-enabled large language models on the Japan Diagnostic Radiology Board Examination was evaluated. Among the models released in November 2025, Gemini 3 Pro demonstrated significant capabilities in direct medical image interpretation, achieving accuracy above that of a reference group of five newly board-certified radiologists. SUPPLEMENTARY INFORMATION: The online version contains supplementary material available at 10.1007/s11604-026-01983-x.
Poor cardiac MR image quality can prompt repeat examinations and hinder clinical decision-making.
To evaluate whether pre-imaging clinical information, extracted using a large language model (LLM), is independently associated with cardiac MR image quality. STUDY TYPE: Retrospective. POPULATION: 1006 adults undergoing clinical cardiac MR examinations. FIELD STRENGTH/SEQUENCE: 1.5โT and 3โT scanners with cine, black blood, MR angiogram, or late gadolinium enhancement protocols. ASSESSMENT: Image quality was categorized per study as excellent, slightly limited, severely limited, or nondiagnostic using institutional reporting conventions finalized by radiologists and cardiologists. A HIPAA-compliant LLM assigned image quality labels based on radiology reports through an iteratively refined prompt, with reliability confirmed by two radiologists. Labels were binarized as Good (excellent and slightly limited) versus Poor (severely limited and nondiagnostic). A repeat-imaging-adjusted image quality label was used in a sensitivity analysis. Pre-imaging clinical and patient variables were extracted from electronic health records. Associations between variables and image quality labels were investigated. STATISTICAL TESTS: Cohen's kappa (ฮบ) for label agreement. Chi-square and t-tests for univariate analysis. Variance inflation factor (VIF) and multivariable logistic regression. Significance level: pโ<โ0.05.
Binarized image quality labels showed substantial agreement with interpreters' assessments for both the primary dataset (ฮบโ=โ0.689) and the repeat-adjusted dataset (ฮบโ=โ0.879). There was no significant multicollinearity (VIFโ=โ1.01-1.39). Cognitive and communication impairment (OR 1.81, 95% CI [1.30-2.54], pโ<โ0.001) and respiratory issues (1.57 [1.14-2.17], pโ=โ0.006) were independently associated with poor image quality. These associations remained significant in the repeat-adjusted sensitivity analysis (cognitive and communication impairment (OR 1.75, 95% CI [1.27-2.44], pโ<โ0.001) and respiratory compromise (OR 1.37, 95% CI [1.04-1.82], pโ=โ0.027)). Other clinical variables were not independently associated after adjustment. DATA
Cognitive/communication impairment and respiratory compromise were independently associated with poor cardiac MR image quality. TECHNICAL EFFICACY: Stage 2. This study examined why some cardiac magnetic resonance scans have poor image quality. Poorโquality scans can limit diagnosis and may require repeat imaging. It analyzed adult examinations and used artificial intelligence (AI) to extract clinical information from the medical records available before the scan. Cognitive or communication difficulties and breathingโrelated problems were associated with worse image quality. These findings might suggest that these patients may benefit from additional preparation before scanning. More broadly, this study shows that information in the medical record can be efficiently extracted using AI to support largeโscale clinical research.
| Run at | Source | Hits | New | Status |
|---|---|---|---|---|
| 2026-08-23 00:00 | LitReview | 4 | 1 | completed |
| 2026-08-16 00:00 | LitReview | 3 | 1 | completed |
| 2026-08-09 00:00 | LitReview | 11 | 5 | completed |
| 2026-08-02 00:00 | LitReview | 10 | 6 | completed |
| 2026-07-26 00:00 | LitReview | 2 | completed | |
| 2026-07-19 00:00 | LitReview | 6 | 4 | completed |
| 2026-07-12 00:00 | LitReview | 4 | 2 | completed |
| 2026-07-05 00:00 | LitReview | 13 | 6 | completed |
"prostatic neoplasms"[mesh] AND "magnetic resonance imaging"[mesh] AND "clinical trial"[pt]
To evaluate the feasibility and clinical value of a prostate cancer screening model incorporating total prostate-specific antigen (tPSA) and biparametric magnetic resonance imaging (bpMRI) in a Chinese population.
Between May 2024 and May 2025, 2 251 men aged 50 years or older from three community health service centers in Beijing were enrolled, who were randomly assigned to the precise screening group and the standard screening group in a 2 โถ 1 ratio in the community health service centers. In the precision screening group, participants with tPSAโฅ4 ฮผg/L underwent bpMRI; those with a prostate imaging reporting and data system (PI-RADS) scoreโฅ3 were recommended to undergo systematic combined with targeted biopsy, while those with tPSAโฅ10 ฮผg/L and PI-RADS < 3 were recommended to undergo systematic biopsy alone. In the standard screening group, participants with tPSAโฅ4 ฮผg/L were recommended to undergo systematic biopsy. The number of participants who actually underwent biopsy, biopsy positivity rate, and Gleason score concordance rate was recorded. The primary outcome measure was the detection rate of clinically significant prostate cancer (csPCa), and the intergroup comparisons were performed using the Mann-Whitney U test.
A total of 2 251 men were included in the analysis, with a median age of 68 years (range: 57-90 years). In the precision screening group (n=985), 115 (11.7%) participants had tPSAโฅ4 ฮผg/L, of whom 30 underwent bpMRI. The recommended biopsy rate was 2.9% (29/985), and 20 participants actually underwent prostate biopsy. A total of 15 prostate cancer cases were detected, including 14 csPCa and one clinically insignificant prostate cancer. In the standard screening group (n=1 266), 111 (8.8%) participants had tPSA โฅ4 ฮผg/L, with a recommended biopsy rate of 8.8% (111/1 266); 26 of them underwent systematic biopsy, and 14 prostate cancer cases were detected, including 7 csPCa and 7 clinically insignificant prostate cancer. The overall positive biopsy rate in the precision screening group was 75.0%, which was not statistically significant compared with that in the standard screening group (58.3%, P=0.141). However, the csPCa detection rate in the precision screening group reached 70.0%, which was significantly higher than that in the standard screening group (26.9%), and the difference was statistically significant (P=0.004). Among the 29 diagnosed prostate cancer patients, 17 (58.6%) had localized disease, 10 (34.5%) had locally advanced disease, and 2 (6.9%) had metastatic disease. Among the 21 patients who underwent radical surgery, the concordance rate between biopsy and post-operative pathological Gleason scores was 70.0% in the precision screening group, higher than that in the standard screening group (36.4%), although the difference was not statistically significant (P=0.198).
The prostate cancer screening model based on serum tPSA combined with bpMRI demonstrates good feasibility, reduces unnecessary biopsies, optimizes the allocation of screening resources, and provides a basis for precision screening strategies. ็ฎ็: ๆข่ฎจๆปๅๅ่ บ็นๅผๆงๆๅ(total prostate-specific antigen, tPSA)่ๅๅๅๆฐ็ฃๅ ฑๆฏๆๅ(biparametric magnetic resonance imaging, bpMRI)็ๅๅ่ บ็็ญๆฅๆจกๅผๅจไธญๅฝไบบ็พคไธญ็ๅฏ่กๆงไธไธดๅบไปทๅผใ ๆนๆณ: ็บณๅ ฅๅไบฌๅธ3ๅฎถ็คพๅบๅซ็ๆๅกไธญๅฟ2024ๅนด5ๆ่ณ2025ๅนด5ๆ50ๅฒๅไปฅไธ็ทๆงไฝไธบ็ ็ฉถๅฏน่ฑก๏ผๆ2 โถ 1ๆฏไพ้ๆบๅไธบ็ฒพๅ็ญๆฅ็ปไธๆ ๅ็ญๆฅ็ปใ็ฒพๅ็ญๆฅ็ป๏ผtPSAโฅ4 ฮผg/L่ ่ฟ่กbpMRIๆฃๆฅ๏ผ่ฅๅๅ่ บๅฝฑๅๆฅๅๅๆฐๆฎ็ณป็ป(prostate imaging reporting and data system, PI-RADS)่ฏๅโฅ3ๅ๏ผๅปบ่ฎฎ่ก็ณป็ป+้ถๅ็ฉฟๅบ๏ผ่ฅtPSAโฅ10 ฮผg/LไธPI-RADS < 3ๅ๏ผๅปบ่ฎฎ่ก็ณป็ป็ฉฟๅบใๆ ๅ็ญๆฅ็ป๏ผtPSAโฅ4 ฮผg/L่ ๅปบ่ฎฎ่ก็ณป็ป็ฉฟๅบใ่ฎฐๅฝไธค็ปๅฎ้ ๆฅๅ็ฉฟๅบไบบๆฐ๏ผ็ฉฟๅบ้ณๆง็ๅGleason่ฏๅไธ่ด็ใไธป่ฆ็ปๅฑๆๆ ไธบไธดๅบๆๆไนๅๅ่ บ็(clinically significant prostate cancer, csPCa)ๆฃๅบ็๏ผ็ป้ดๆฏ่พ้็จๅกๆนๆฃ้ชใ ็ปๆ: ๅ ฑ2 251ๅ็ทๆง็บณๅ ฅๅๆ๏ผไธญไฝๅนด้พ68ๅฒ(57~90ๅฒ)ใ็ฒพๅ็ญๆฅ็ป985ไพ๏ผ115ไพ(11.7%)ๅ่ฏ่ tPSAโฅ4 ฮผg/L๏ผๅ ถไธญ30ไพๆฅๅไบbpMRIๆฃๆฅ๏ผๅปบ่ฎฎ็ฉฟๅบ็ไธบ2.9%(29/985)๏ผๅฎ้ ๆฅๅๅๅ่ บ็ฉฟๅบ่ 20ไพ๏ผๆ็ปๆฃๅบๅๅ่ บ็15ไพ๏ผๅ ๆฌ14ไพไธดๅบๆๆไนๅๅ่ บ็(csPCa)ๅ1ไพไธดๅบๆ ๆไนๅๅ่ บ็ใๆ ๅ็ญๆฅ็ป1 266ไพ๏ผ111ไพๆฃ่ tPSAโฅ4 ฮผg/L๏ผๅปบ่ฎฎ็ฉฟๅบ็ไธบ8.8%๏ผๅ ถไธญ26ไพๆฅๅ็ณป็ป็ฉฟๅบ๏ผๅ ฑๆฃๅบ14ไพๅๅ่ บ็๏ผๅ ๆฌ7ไพcsPCaๅ7ไพไธดๅบๆ ๆไนๅๅ่ บ็ใ็ฒพๅ็ญๆฅ็ป็ๆปไฝ็ฉฟๅบ้ณๆง็ไธบ75.0%๏ผไธๆ ๅ็ญๆฅ็ป(58.3%)็ธๆฏ๏ผๅทฎๅผๆ ็ป่ฎกๅญฆๆไน(P=0.141)๏ผ่็ฒพๅ็ญๆฅ็ปcsPCa็ฉฟๅบ้ณๆง็(70.0%)ๆพ่้ซไบๆ ๅ็ญๆฅ็ป(26.9%)๏ผๅทฎๅผๆ็ป่ฎกๅญฆๆไน(P=0.004)ใ็กฎ่ฏ็29ไพๅๅ่ บ็ๆฃ่ ไธญ๏ผๅฑ็ถๆ17ไพ(58.6%)ใๅฑ้จ่ฟๅฑๆ10ไพ(34.5%)ใ่ฝฌ็งปๆง2ไพ(6.9%)ใๅจๆฅๅๆ นๆฒปๆๆฏ็21ไพๆฃ่ ไธญ๏ผ็ฒพๅ็ญๆฅ็ป็ฉฟๅบๆ ๆฌไธๆฏๅ็ ็Gleason่ฏๅไธ่ด็ไธบ70.0%๏ผ้ซไบๆ ๅ็ญๆฅ็ป(36.4%)๏ผไฝๅทฎๅผๆ ็ป่ฎกๅญฆๆไน(P=0.198)ใ ็ป่ฎบ: ๅบไบ่กๆธ tPSA่ๅbpMRI็ๅๅ่ บ็็ญๆฅๆจกๅผๅ ทๆๅฏ่กๆง๏ผๅฏๅๅฐไธๅฟ ่ฆ็็ฉฟๅบ๏ผไผๅ็ญๆฅ่ตๆบ้ ็ฝฎ๏ผๅฏไธบ็ฒพๅ็ญๆฅๆไพไพๆฎใ
To evaluate prostate-specific antigen (PSA) and prostate-specific membrane antigen (PSMA) kinetics following PSMA-positron emission tomography (PET) and magnetic resonance (MR) guided stereotactic body radiation therapy (SBRT) and short-term androgen deprivation therapy with dominant intraprostatic lesion (DIL) boost in localized prostate cancer. METHODS AND MATERIALS: This prespecified analysis from the phase 2 PROBE trial (IRB approval: IEC 900917) included patients with intermediate- or high-risk prostate cancer treated with 5-fraction SBRT (36.25 Gy to prostate; 40 to 42.5 Gy to PSMA/MRI-defined DIL) and 6 months of androgen deprivation therapy. PSA kinetics relied on 3-monthly PSA levels. Longitudinal metabolic response, based on the adapted PET Response Criteria in Solid Tumors, was assessed using 68Ga-PSMA-PET/computed tomography done at SBRT completion, 3 to 6 months, and 9 to 12 months after SBRT completion. Cumulative genitourinary/gastrointestinal toxicities were reported per National Cancer Institute Common Terminology Criteria for Adverse Events v5.0.
Thirty enrolled patients (54% intermediate risk, 46% high risk) were included. A PSA <0.1 ng/mL was achieved in 79% of patients at 6 months and 50% at 12 months. The mean maximum standardized uptake value (SUVmax) decreased from 12.2 (range, 4.0-39.6) pretreatment to 3.2 at 3 to 6 months and 1.8 at 9 to 12 months. Complete response (CR) rates increased from 17% at SBRT completion to 48% at 3 to 6 months and 65% at 9 to 12 months. Patients achieving CR at 3 to 6 months reached a lower mean PSA (0.04 vs 0.3 ng/mL, P = .009). Those with PSA <0.1 ng/mL at 6 months were more likely to attain CR at 12 months (76% vs 28%, P = .003). An early standardized uptake value (SUV) "flare" was observed in 4 (14%) patients, resolving by 9 to 12 months. No National Cancer Institute Common Terminology Criteria for Adverse Events grade โฅ3 genitourinary/gastrointestinal toxicity was seen.
PSMA-PET and MR guided prostate SBRT with DIL boost achieved deep PSA nadirs, gradual PSMA kinetics, and a favorable safety profile. About 80% of patients achieved a PSA <0.1 ng/mL at 6 months, with two-thirds achieving CR by 12 months post-SBRT. Achieving a PSA <0.1 ng/mL could potentially predict PSMA-CR. Early PSMA-CR at 3 to 6 months may predict durable PSA response and warrants further evaluation as a biomarker for risk-adapted strategies in localized prostate cancer.
Prostate cancer (PC) screening reduces PC mortality but also causes burden such as overdiagnosis and unnecessary biopsies. The European Association of Urology (EAU) recently proposed a risk-adapted screening protocol incorporating prostate-specific antigen (PSA)-based intervals, a risk calculator (RC), and magnetic resonance imaging (MRI), though its long-term outcomes remain unquantified. We constructed the MIcrosimulation SCreening Analysis-PSA model by adapting the existing MIcrosimulation SCreening Analysis-PROstate microsimulation framework to simulate individual PSA trajectories. The model parameters were calibrated to outcomes of the European Randomized Study of Screening for Prostate Cancer (ERSPC). Using this model, we simulated five screening protocols for men aged 55-69: fixed 4-year intervals between PSA tests and a biopsy when PSA โฅ3.0โng/mL (ERSPC protocol); PSA-based intervals; MRI prior to biopsy; the RC and MRI prior to biopsy; and the full EAU protocol (PSA-based intervals, RC, and MRI). Outcomes included PC mortality, overdiagnoses, the number of PSA tests, biopsies, and MRIs. Compared to the ERSPC protocol, PSA-based intervals reduced PSA tests by 21%. The MRI-only protocol decreased overdiagnosis by 6% but also required many MRIs. Incorporating the RC further reduced overdiagnosis to 10% and required 36% fewer MRIs than the MRI-only protocol. The EAU combines the best of all these approaches while maintaining equal PC mortality (200 deaths per 10,000 men). The EAU protocol optimizes long-term screening efficiency, significantly reducing biopsies and overdiagnosis with minimal mortality trade-offs. MRI and RC integration enhance resource allocation.
Non-contrast MRI (bi-parametric MRI-bpMRI) has been investigated as a potential tool to be integrated in clinically significant prostate cancer (csPCa) screening. Moreover, artificial intelligence (AI) is emerging too as a potential support, especially for less-experienced radiologists. Therefore, the aim of this study was to evaluate the effectiveness of an AI-based software in csPCa screening using bpMRI, with a focus on supporting less-experienced radiologists.
A retrospective analysis was conducted within the PROSA-trial, a randomized, single-center study involving 759 men eligible for PCa screening. BpMRI were acquired using prostate imaging reporting and data system (PI-RADS) v2.1-compliant protocols and evaluated independently by an expert radiologist, a less-experienced reader, AI-based software, and the less-experienced reader with AI support. Diagnostic performance was assessed using ROC curves and inter-reader agreement (Cohen's kappa), using expert interpretation as the reference standard.
Four hundred ninety-nine bpMRI were analyzed. The AI-assisted less-experienced reader achieved the highest diagnostic performance (sensitivity 76.5%, specificity 97.2%, accuracy 95.8%, AUC 0.969), surpassing both AI-alone (sensitivity 58.8%, specificity 96.6%, accuracy 94.0%, AUC 0.952) and unaided less-experienced reader (sensitivity 67.6%, specificity 95.1%, accuracy 93.2%, AUC 0.932). Inter-reader agreement improved with AI assistance (from ฮบโ=โ0.64 to 0.84). AI assistance reduced equivocal PI-RADS 3 cases (from 77 to 53) and improved exact agreement with the expert from 32.5% to 54.5%, while also reducing diagnostic discordance.
AI can support less experienced radiologists and enhance consistency in bpMRI interpretation, especially considering equivocal cases. Moreover, integrating AI into radiology workflows can alleviate reporting burden and help prioritize suspicious cases, offering critical advantages in high-volume PCa screening settings. KEY POINTS: Question Can AI improve the diagnostic performance and consistency of less experienced radiologists interpreting non-contrast prostate MRI in a screening setting? Findings AI assistance significantly improved effectiveness and inter-reader agreement in prostate MRI interpretation, particularly for less experienced radiologists within a screening population. Clinical relevance Integrating AI into prostate MRI workflows may enhance screening efficiency, reduce variability, and support equitable early detection of csPCa, with a positive impact on prioritization of reporting.
MRI is recommended for men with clinical suspicion of significant prostate cancer. Those with high clinical risk but non-suspicious or equivocal MRI often undergo prostate biopsy, but have a low likelihood of clinically significant prostate cancer, and a high incidence of clinically insignificant prostate cancer. We aimed to investigate whether gallium-68 ([68Ga]Ga)-prostate-specific membrane antigen (PSMA)-11 PET-CT could reduce the number of people requiring prostate biopsy and limit biopsy to targeted cores, without compromising clinically significant prostate cancer diagnosis.
In this multicentre, non-inferiority, phase 3, randomised controlled trial, done at at seven Australian hospitals, we recruited biopsy-naive participants with clinical suspicion of significant prostate cancer, equivocal (Prostate Imaging-Reporting and Data System [PI-RADS] 3) or non-suspicious (PI-RADS 2) MRI but high clinical risk (eg, prostate-specific antigen [PSA] density of >0ยท1 ng/mL/mL, strong family history of prostate cancer, abnormal digital rectal examination, BRCA mutation, PSA >10 ng/mL, PSA doubling time <36 months, or PSA velocity >0ยท75 ng/mL per year), PSA of 20 ng/mL or less, and clinical T2 disease or less. Participants were randomly assigned (1:1) using a centralised web-based system to undergo [68Ga]Ga-PSMA-11 PET-CT (experimental group) or systematic transperineal prostate biopsy (control group), using block sizes of two or four and stratification by study site. There was no masking for participants or investigators. Participants with positive [68Ga]Ga-PSMA-11 PET-CT (PRIMARY score 3-5) underwent PSMA-PET-targeted transperineal prostate biopsies, whereas those with a negative result (PRIMARY score 1-2) avoided biopsy. The co-primary outcomes were the proportion of participants with clinically significant prostate cancer, defined as a Gleason score of 3โ+โ4 (โฅ10% pattern 4) or higher, and the proportion of participants in the [68Ga]Ga-PSMA-11 PET-CT group who avoided biopsy within 6 months of random assignment. A two-sided 95% Wald CI based on a binomial model was used to estimate the risk difference in the proportion of participants with clinically significant prostate cancer (non-inferiority margin 10%) and to estimate the proportion of participants in the experimental group who had avoided biopsy 6 months after random assignment (20% threshold), analysed based on intention to treat. This trial is registered with ClinicalTrials.gov, NCT05154162, and participant follow-up is ongoing.
Between March 2, 2022, and Aug 24, 2025, 660 eligible male participants were enrolled and had a median age of 61 years (IQR 56-66), a median PSA of 5ยท2 ng/mL (4ยท0-7ยท0), and a median PSA density of 0ยท13 ng/mL/mL (0ยท09-0ยท17). There were PI-RADS 2 in 335 (51%) participants and PI-RADS 3 in 325 (49%) participants. Ethnicity data were not collected. 329 (50%) were assigned to the control group with systematic transperineal prostate biopsy, and 331 (50%) were assigned to the experimental group with [68Ga]Ga-PSMA-11 PET-CT. The proportion of participants with clinically significant prostate cancer in the experimental group (39 [12%] of 331) was non-inferior to the control (51 [16%] of 329; difference -3ยท7% [95% CI -8ยท9 to 1ยท5%]; p=0ยท0093). Use of [68Ga]Ga-PSMA-11 PET-CT avoided biopsy in 163 (49%) of 331 participants (95% CI 44 to 55%; p <0ยท0001). After prostate biopsy, participants reported similar proportions of pain (33 [21%] in the experimental group vs 62 [21%] in the control group), haematuria (60 [38%] vs 126 [43%]), and haematospermia (77 [48%] vs 133 [45%]).
[68Ga]Ga-PSMA-11 PET-CT could have the potential to improve the diagnostic pathway of patients with a high clinical risk but non-suspicious or equivocal prostate MRI. Further research, including health-economic analyses and validation with other PSMA radiopharmaceuticals, are needed to confirm the clinical implementation and generalisability of this approach. FUNDING: Prostate Cancer Foundation, National Health and Medical Research Council, St Vincent's Curran Foundation, and Peter MacCallum Cancer Foundation.
Index lesion-focused ipsilateral systematic biopsy (iSB) has been proposed as a core-reduction alternative to MRI-targeted biopsy combined with systematic biopsy (TBย +ย SB) for prostate cancer, but prospective multicenter validation is limited.
This prospective multicenter trial (NCT06584279) enrolled biopsy-naive men with Prostate Imaging Reporting and Data System (PI-RADS) โฅ4 lesions or PI-RADS 3 plus prostate-specific antigen (PSA) density โฅ0.15 ng/mL/cm3. All underwent MRI-targeted biopsy (โฅ2 cores/lesion) with 12-core SB. Through core-level reclassification, the authors simulated TBย +ย iSB (targeted biopsy combined with index lesion-ipsilateral systematic biopsy). The primary outcome was cancer detection rate (CDR) of clinically significant prostate cancer (csPCa; Grade Group โฅ2) with a noninferiority margin of -3%. Secondary outcomes included clinically insignificant prostate cancer (cisPCa; Grade Group 1) detection and pathological concordance after radical prostatectomy.
Among 564 men (median age, 69 years; median PSA, 7.3 ng/mL), TBย +ย iSB demonstrated a noninferior csPCa CDR versus TBย +ย SB (40.78% vs. 42.38%; difference, -1.6 percentage points; 95% CI, -2.69 to -0.50), with the lower bound above the -3% noninferiority margin. The cisPCa CDR was slightly lower with TBย +ย iSB than with TBย +ย SB (13.83% vs. 14.36%; difference, -0.53 percentage points; 95% CI, -2.01 to 0.92). Among 189 surgical patients, pathological concordance was similar, whereas upgrading was numerically more frequent and downgrading numerically less frequent under simulated TBย +ย iSB.
TBย +ย iSB met the prespecified noninferiority criterion for csPCa detection relative to TBย +ย SB and may represent a promising core-reduction strategy in transperineal MRI-guided biopsy, although the full protocol still provided additional contralateral information.
PSMA-PET offers an opportunity to reduce the mischaracterization of disease in active surveillance (AS) and focal therapy (FT) candidates. We describe the results of a pilot clinical trial evaluating 18F-radiohybrid(rh)PSMA-7.3-PET/MRI to detect occult adverse pathology among potential AS and FT candidates (NCT05852041).
We enrolled 20 men with low risk or favorable intermediate risk prostate cancer and Decipher score โฅโ0.45 diagnosed after an MRI-informed prostate biopsy. All patients underwent PSMA-PET/MRI followed by either PET/MRI-guided biopsy or radical prostatectomy within 90 days. The outcome of interest was detection of grade group (GG) 3-5 disease, seminal vesicle invasion (pT3b), or lymph node involvement (pN1). A detection rate of 15% for this outcome was considered clinically significant. Management decisions and confidence in those decisions were recorded before and after the scan using a 3-point scale. A paired t-test was performed to compare the change in confidence decisions.
At enrollment, 17 patients (85%) had favorable intermediate risk and 3 (15%) had low risk prostate cancer. The median Decipher score was 0.58 (IQR 0.49-0.65). Five patients (25%) demonstrated the outcome of interest based on upgrading alone. None had upstaging to pT3b or pN1. Major changes in management plan occurred in 7 patients (35%). Average confidence in decisions improved from moderate (2.05) to high (2.80) after the scan (pโ<โ0.001).
18F-rhPSMA-7.3-PET/MRI can detect occult higher risk disease in men who are otherwise candidates for AS or FT. The scan prompted major changes in management and increased confidence in the final treatment strategy.
Non-White patients are poorly represented in prostate cancer trials. MRI PI-RADS scoring was developed in primarily White populations, but prostate cancer differs in non-White men. We aimed to explore differences in PI-RADS calibration for Asian and Black men.
This is a secondary analysis of PREVENT, a multi-institutional study of infection rates for transrectal vs. transperineal biopsy. We compared cancer detection for self-identifying Asian and Black men. We compared detection rates on a per-person basis, stratified by index PI-RADS lesion, to White men, using Fisher's exact and logistic regression.
Of 665/752 trial patients with PI-RADS 3-5 lesions, 88 (13%) were Black and 36 (6%) were Asian. Black men were younger at diagnosis with increased rates of overall (70% vs. 43%%, Pโ=โ0.004) and clinically significant prostate cancer (60% vs. 27%, Pโ=โ0.003) and Asian men had decreased rates of overall (0% vs. 47%, Pโ=โ0.004) and clinically significant prostate cancer (0% vs. 27%, Pโ=โ0.003) in PI-RADS 3 lesions compared to White men. On multivariable regression, Black men with PI-RADS 3/4 lesions had higher odds of overall (OR 1.17, Pโ=โ0.009) and clinically significant prostate cancer (OR 1.20, Pโ=โ0.004) and Asian men had lower odds of overall (OR 0.79, Pโ=โ0.01) but not clinically significant prostate cancer (OR 0.94, Pโ=โ0.5).
Black men with PI-RADS 3/4 lesions had 20% higher odds of clinically significant prostate cancer than White men while all PI-RADS 3 lesions in Asian men were negative. These findings suggest PI-RADS may require differential interpretation when assessing prostate cancer risk in non-White men. TRIAL REGISTRATION: Registered at ClinicalTrials.gov ( NCT04843566 , https://clinicaltrials.gov/study/NCT04843566 ).
To determine which treatment parameters optimize focal therapy for intermediate-risk prostate cancer by balancing oncologic control with healthy tissue preservation, in a phase 2b multicenter trial of MRI-guided Focused Ultrasound (MRgFUS). Additionally, to assess the relationship of ablation volume relative to lesion volume with oncologic outcomes, urinary, and erectile function.
In this retrospective interpretation of prospectively acquired data, the non-perfused volume (NPV) of prostate tissue encompassing the MRI-visible lesion volume defined the ablation-volume-to-lesion-volume ratio (ALVR). Oncologic efficacy was assessed as the absence of clinically significant (GGGโโฅโ2) cancer in the treatment zone at 24-month biopsy. Associations between ALVR and outcomes were assessed using Student's t-tests. Baseline characteristics were compared using Kruskal-Wallis tests.
Eighty-nine men (mean age,ย 63 yearsโยฑโ7) had MRI-visible lesions with a volume of 0.47โmL (IQR: 0.20-0.95), with a surrounding NPV of 6.9โmL (IQR: 5.2-10.4). Men achieving oncologic efficacy had twice the ALVR compared to those with recurrence at the treatment site (17 vs 8, mean difference 8.8, 95% CI: 2.1, 16, pโ=โ0.013). Increasing NPV relative to total prostate volume did not improve oncologic outcomes. Baseline characteristics did not significantly differ between men with and without GGGโโฅโ2 at 24-month biopsy. ALVR did not differ in men with new erectile dysfunction (mean difference in ALVR: 2.1, 95% CI: -12, 16, pโ=โ0.8) or urinary symptoms (mean difference in ALVR 4.0, 95% CI: -21, 29, pโ=โ0.71).
In patients with intermediate-risk prostate cancer, higher ALVR was associated with superior 2-year oncologic outcomes without increased risk of urinary or erectile dysfunction. KEY POINTS: Question What treatment parameters optimize focal therapy for prostate cancer by balancing healthy tissue preservation with favorable oncologic outcomes? Findings Patients without residual cancer at 24-month biopsy had twice the ALVR of those with recurrence, with no adverse impact on erectile or urinary function. Clinical relevance While fixed intra-prostatic margins (e.g., 5โmm or 10โmm) are commonly prescribed in focal therapy, this study highlights the importance of scaling the ALVR in the treatment plan to achieve sufficient oncologic coverage.
To assess the oncological outcomes of targeted microwave ablation (TMA) using organ-based tracking (OBT) Fusionยฎ via KOELIS Trinityยฎ (KOELIS, Meylan, France) in men with intermediate-risk prostate cancer (PCa): the VIOLETTE trial (ClinicalTrials.gov identifier: NCT04582656) PATIENTS AND
In this prospective phase II, multicentre European study, men with a prostate-specific antigen (PSA) level <20โng/mL, a single magnetic resonance imaging (MRI)-visible lesion โค15โmm, International Society of Urological Pathology (ISUP) Grade Group 2 on MRI-targeted biopsy, and clinical T stage โค2, were enrolled. The microwave applicator was placed using OBT Fusion guidance, either transperineally or transrectally. The primary endpoint was the absence of clinically significant PCa (csPCa), defined as ISUP Grade Group โฅ2, within the treated area at 12โmonths. Secondary endpoints included safety, functional outcomes using validated measures, and the need for subsequent radical treatment.
A total of 76 patients were treated across six centres with 66 (87%) completing the 12-month follow-up. At 6โmonths, six patients had csPCa after positive MRI control, including four within the treated area. At 12โmonths, csPCa was detected in 15 additional patients, including nine in-field recurrences, yielding an 81% in-field csPCa-free rate. Five serious adverse events in three patients were reported. Sexual (-2.5 points; Pโ<โ0.001) and ejaculatory (-1 points; Pโ<โ0.001) scores decreased significantly, whereas urinary function remained stable. Radical treatment was required in four (5.2%) patients at 12โmonths.
Targeted microwave ablation using OBT Fusion technology appears to be a safe and effective focal therapy procedure for localised intermediate-risk PCa. The VIOLETTE trial achieved its primary endpoint, with 81% patients free of in-field csPCa at 12โmonths.
| Run at | Source | Hits | New | Status |
|---|---|---|---|---|
| 2026-08-23 00:00 | LitReview | 1 | 1 | completed |
| 2026-08-16 00:00 | LitReview | completed | ||
| 2026-08-09 00:00 | LitReview | 1 | completed | |
| 2026-08-02 00:00 | LitReview | 1 | 1 | completed |
| 2026-07-26 00:00 | LitReview | completed | ||
| 2026-07-19 00:00 | LitReview | 1 | 1 | completed |
| 2026-07-12 00:00 | LitReview | completed | ||
| 2026-07-05 00:00 | LitReview | 3 | 3 | completed |
(radiology[tiab] OR imaging[tiab]) AND ("large language model*"[tiab] OR LLM[tiab]) AND (personalized[tiab] OR personalization[tiab] OR style[tiab] OR adaptation[tiab]) AND (reporting[tiab] OR impression[tiab] OR conclusion[tiab])
Accurate documentation of distant recurrence sites in breast cancer is essential for evaluating treatment effectiveness and outcomes research. However, such information is embedded in unstructured clinical notes, making manual abstraction labor-intensive. Large language models (LLMs) offer a scalable solution for extracting complex information from heterogeneous clinical narratives; however, generic LLMs often lack the specialized clinical reasoning needed for accurate interpretation of oncologic documentation. This study aims to develop an efficient LLM-based framework to automatically extract distant recurrence sites from free-text documentation. MATERIALS &
We used clinical notes, pathology and radiology reports from recurrent breast cancer patients at Mayo Clinic (nย =ย 766) for model development and evaluated generalizability on internal hold-out samples (nย =ย 112) and an external Stanford Medicine cohort (nย =ย 110). For cross-disease domain adaptation, we further validated on prostate cancer patients (nย =ย 49). Our proposed framework employs BioLinkBERT, a pretrained language model (PLM) backbone, with weak supervision and an epoch-wise entropy optimization to address limited labeled data and class imbalance across recurrence sites. The fine-tuned model was compared against state-of-the-art models, including Llama2-7B, Llama-3-8B and MedAlpaca, using precision, recall, and F1-score.
The fine-tuned model outperformed generic and domain-specific LLM baselines, with notable gains in identifying multi-site distant recurrence. In-domain validation showed consistent F1-score improvement (average 0.78), particularly for rare recurrence sites. The model also demonstrated strong performance on the external Stanford cohort and on prostate cancer, achieving F1-score of 0.83 and 0.93, respectively.
This study presents an efficient, weakly supervised LLM framework that accurately extracts metastatic recurrence sites, reducing reliance on manual chart review. The results demonstrate that relatively small LLMs, optimized with domain-aware weak supervision, can outperform larger models for complex oncologic information extraction. The model is released as a platform-independent Docker image to support seamless cancer registry integration.
RATIONALE AND
To evaluate the application of DeepSeek-assisted case-based learning (CBL) in respiratory radiology course across the full instructional cycle, including preparation, implementation, and evaluation.
This prospective, single-center study was conducted in 2025 and involved third-year medical undergraduates. CBL Preparation: Six cases were retrieved from the Hospital Information System (HIS), and six generated via DeepSeek-R1. Preparation times were recorded and compared. CBL Implementation: Students were assigned to either a DeepSeek-assisted group or a traditional CBL group, with discussion time recorded for each subgroup. CBL Evaluation: Teaching effectiveness was evaluated through test scores and questionnaires. Subsequently, DeepSeek-R1 provided personalized feedback to students based on their individual scores.
A total of 200 students (mean age 21.02ย ยฑย 0.89 years, 94 males) participated. DeepSeek-generated cases required significantly less time than HIS-retrieved cases (p = 0.016). During implementation, the DeepSeek group spent less discussion time than traditional group (p = 0.026). The DeepSeek-assisted group achieved greater test score improvements compared to the traditional group (p < 0.05). Questionnaire responses indicated higher self-directed learning, greater interest in radiology, improved learning efficiency, and lower perceived learning burden in the DeepSeek-assisted group (p < 0.05). Additionally, personalized feedback generated by DeepSeek was qualitatively reviewed by the radiology teaching department and considered educationally useful.
This study demonstrates that DeepSeek-assisted CBL effectively supports respiratory radiology education throughout the entire course process-preparation, implementation, and evaluation-by enhancing efficiency, boosting student interest and engagement, improving performance, and providing valuable post-class feedback.
Accurate tumor node metastasis (TNM) staging is fundamental for treatment planning and prognosis in non-small cell lung cancer (NSCLC). However, its complexity poses significant challenges. Traditional rule-based natural language processing methods are constrained by their reliance on manually crafted rules and are susceptible to inconsistencies in clinical reporting.
This study aimed to develop and validate a robust, accurate, and operationally efficient artificial intelligence framework for the TNM staging of NSCLC by strategically enhancing a large language model, GLM-4-Air (general language model), through advanced prompt engineering and supervised fine-tuning (SFT).
We constructed a curated dataset of 492 deidentified real-world medical imaging reports, with TNM staging annotations rigorously validated by senior physicians according to the AJCC (American Joint Committee on Cancer) 8th edition guidelines. The GLM-4-Air model was systematically optimized via a multi-phase process: iterative prompt engineering incorporating chain-of-thought reasoning and domain knowledge injection for all staging tasks, followed by parameter-efficient SFT using low-rank adaptation for the reasoning-intensive primary tumor characteristics (T) and regional lymph node involvement (N) staging tasks. The final hybrid model was evaluated on a completely held-out test set (black-box) and benchmarked against GPT-4o using standard metrics, statistical tests, and a clinical impact analysis of staging errors.
The optimized hybrid GLM-4-Air model demonstrated reliable performance. It achieved higher staging accuracies on the black-box test set: 92% (95% CI 0.850-0.959) for T, 86% (95% CI 0.779-0.915) for N, 92% (95% CI 0.850-0.959) for distant metastasis status (M), and 90% for overall clinical staging; by comparison, GPT-4o attained 87% (95% CI 0.790-0.922), 70% (95% CI 0.604-0.781), 78% (95% CI 0.689-0.850), and 80%, respectively. The model's robustness was further evidenced by its macro-average F1-scores of 0.914 (T), 0.815 (N), and 0.831 (M), consistently surpassing those of GPT-4o (0.836, 0.620, and 0.698). Analysis of confusion matrices confirmed the model's proficiency in identifying critical staging features while effectively minimizing false negatives. Crucially, the clinical impact assessment showed a substantial reduction in severe category I errors, which are defined as misclassifications that could significantly influence subsequent clinical decisions. Our model committed 0 category I errors in M staging and fewer category I errors in T and N staging. Furthermore, the framework demonstrated practical deployability, achieving efficient inference on consumer-grade hardware (eg, 4 RTX 4090 GPUs) with latencies suitable and acceptable for clinical workflows.
The proposed hybrid framework, integrating structured prompt engineering and applying SFT to reasoning-heavy tasks (T/N), enables the GLM-4-Air model to serve as a highly accurate, clinically reliable, and cost-efficient solution for automated NSCLC TNM staging. This work demonstrates the efficacy and potential of a domain-optimized smaller model compared with an off-the-shelf generalist model, holding promise for enhancing diagnostic standardization in resource-aware health care environments.
Patients referred for specialized care often arrive with outside medical records (OMRs) compiled into multi-report PDFs that include imaging, pathology, and clinical notes in unstructured formats. Reviewing these records is time consuming and mentally taxing, increasing the risk of delayed care, clinician frustration, and missed information affecting quality of care. This study aimed to automate the segmentation, classification, and date extraction of scanned OMRs, with a focus on records relevant to breast cancer care.
We used optical character recognition (OCR) to extract machine-readable text from 1303 scanned PDF documents from 116 distinct external institutions. Gemini 1.5, a large language model (LLM), was then used to segment multi-report files into individual documents, classify them into clinically meaningful categories such as mammograms and pathology reports, and extract study dates to build diagnostic timelines. Document categories were informed by clinical workflows in a breast cancer center.
The system achieved an F1 score of 0.95 for segmentation, 0.96 for classification, and 0.90 for date extraction. In a pilot of 45 records reviewed by clinicians, only 2 classification errors and 1 date error were reported. Clinicians estimated that the tool reduced OMR review time by 40%, improved workflow efficiency, and increased satisfaction.
Our findings demonstrate that combining OCR with LLMs can significantly enhance the processing of unstructured medical records, reducing manual burden and supporting timely clinical decision-making.
This study demonstrates the successful application of OCR and LLMs for organizing scanned OMRs within a specialty clinic. By automating a previously manual process, the approach supports scalable review of incoming outside records and has potential for adaptation to other clinical workflows. Future work will focus on evaluating the system across additional specialties and institutions.
Pancreatic cancer requires nuanced, multidisciplinary treatment planning typically conducted within tumor boards. While Large Language Models (LLMs) have shown capabilities in medical reasoning, their ability to approximate complex, integrative decision-making in oncology remains underexplored.
This study evaluated the performance of LLaMA 3.3 (70b) in predicting tumor board decisions for newly diagnosed pancreatic cancer patients. Clinical documentation (including free-text imaging reports, pathology findings, and patient history) from 42 first-diagnosis cases discussed in a real-world tumor board was collected. The model was tasked with predicting one of three treatment options: surgical resection (SURG), neoadjuvant chemotherapy (NEO), or palliative therapy (PALL). Four prompting strategies were evaluated: zero-shot, advanced (adv.) zero-shot, Chain-of-Thought (CoT), and few-shot prompting. Performance was assessed using accuracy, micro- and macro-averaged F1 scores, and category-specific recall.
The advanced zero-shot and CoT strategies achieved the highest overall accuracy of 78.6% and a micro-averaged F1 score of 0.786. However, this performance was driven primarily by the correct classification of majority classes (SURG and PALL). Crucially, both high-accuracy strategies failed to identify any of the neoadjuvant therapy candidates (Recall NEOโ=โ0.00; 0/7 cases), systematically misclassifying them as palliative or surgical. While few-shot prompting improved the detection of neoadjuvant cases (Recall NEOโ=โ1.00), it introduced substantial noise, reducing overall accuracy to 56.7%. LLaMA 3.3 (70b) demonstrates high concordance with tumor board decisions for clear-cut surgical or palliative cases but exhibits a critical systematic failure in identifying candidates for neoadjuvant therapy. The high global accuracy masks a significant safety limitation regarding the recognition of complex, intermediate-stage patients.
These findings suggest that current LLMs may approximate majority-class decisions but risk overlooking curative treatment pathways in nuanced scenarios, necessitating rigorous oversight and specific adaptation before clinical consideration.
Developing effective Convolutional Neural Networks (CNN) for soft tissue sarcoma detection often requires numerous iterations and adjustments, demanding specialized IT (Information Technology) skills. This study aims to use ChatGPT 4 to simplify CNN adaptation, reducing the need for specialized IT skills while enabling efficient exploration of training configurations to enhance diagnostic accuracy.
This study leveraged a preexisting Artificial Intelligence (AI) model adapted using a preexisting Convolutional Neural Network (CNN). The study involved 54 participants diagnosed with primary soft tissue sarcomas in the extremities and possessing complete Magnetic Resonance Imaging (MRI) datasets. AI adaptations and programming were conducted using TensorFlow and verified with ChatGPT. Model training involved a dataset split of 70% training, 15% validation and 15% test set on patient level split, processed over eight epochs.
The adapted CNN model demonstrated significant improvement across various MRI sequences, achieving high accuracy levels (up to 98.5%) and excellent sensitivity and specificity rates. The model performed robustly in differentiating tumor presence in MR images, with test accuracies as high as 93.9%. The inclusion of a Gradient-weighted Class Activation Mapping (Grad-CAM) heat map and probability scores in the diagnostic outputs further enhanced interpretative capabilities.
This study highlights the potential of AI, particularly CNNs, in the early and accurate detection of soft tissue sarcomas, underscoring the technology's adaptability across different imaging modalities. The integration of large language models like ChatGPT into the model adaptation process emphasizes the reduced need for specialized IT skills, making advanced diagnostic tools more accessible and potentially improving diagnostic accuracy and patient outcomes in radiology and oncology.
Artificial intelligence (AI) has emerged as a transformative force in ophthalmology, enabling automated, accurate, and efficient clinical reporting. This review summarizes recent advances in AI-driven report generation, emphasizing the integration of multimodal imaging and clinical data. Deep learning and natural language processing (NLP) models can synthesize information from diverse sources-including fundus photography, optical coherence tomography, fluorescein angiography, and patient records-to generate structured, interpretable, and personalized diagnostic reports. Such systems enhance diagnostic precision, streamline workflow, and reduce interobserver variability. We outline the technological foundations underlying these systems, including convolutional and transformer-based architectures, self-supervised and multimodal learning, and large language models. Representative applications in diabetic retinopathy, glaucoma, cataract, and age-related macular degeneration are discussed, highlighting their clinical value and emerging real-world deployment. Persistent challenges-including data heterogeneity, model interpretability, ethical governance, and clinical integration-are critically reviewed. Finally, we explore future directions such as real-time AI-assisted reporting, predictive and personalized analytics, and global scalability across healthcare ecosystems. Multimodal, explainable, and clinically integrated AI systems hold promise to redefine ophthalmic diagnostics and improve both clinician efficiency and patient outcomes.
This study aims to explore the application of artificial intelligence in medical education by comparing research hotspots and evolutionary trends between China and the international community, ultimately proposing informed educational practices and policy recommendations.
Literature was retrieved from the core collections of CNKI and Web of Science for the period 2014-2024, limited to article and review publications. After applying a unified Boolean search strategy and deduplication, the data were analyzed using CiteSpace 6.4.R1 to examine publication trends, collaboration networks, keyword co-occurrence/clustering/burst detection, and co-citation patterns.
A total of 379 Chinese and 552 English records were included. Publications surged after 2018 and peaked during 2023-2024. International hotspots centered on machine learning, deep learning, and large language models for simulation-based training and clinical reasoning; Chinese studies focused on "New Medical Sciences", VR/AR, and medical imaging. The emergence of generative artificial intelligence and multimodal large models has become a new frontier in artificial intelligence research within global medical education from 2023 to 2024.
This study is based on a comparison of two databases to reveal the hotspots and differences in artificial intelligence and medical education research between China and the international research community. It not only compensates for the time lag of existing research, but also proposes three major trends driven by artificial intelligence in the development of medical education (generative AI, personalized learning, immersive experience). A complementary pattern exists between technology-driven and scenario-driven orientations. We recommend integrating AI literacy and ethics into curricula, establishing Generative-AI teaching/assessment guidelines, and building cross-institutional, yearly knowledge-map monitoring for sustainable innovation in medical education.
Large language models (LLMs) are increasingly being evaluated for their ability to answer official radiology board-style examination questions. Understanding their accuracy, limitations, and potential applications in education is essential for assessing their utility in the field.
A scoping review was conducted in October 2025 across PubMed, Scopus, and Web of Science, following Preferred Reporting Items for Systematic Reviews and Meta-Analyses guidelines. Studies were included if they evaluated LLMs on official radiology board-style examination questions. After screening 205 unique records, 29 studies met the inclusion criteria. Data were extracted on study characteristics, including LLM type and version, input modality, language, examination type, answer format, comparison with humans, and reported outcomes.
The reviewed studies evaluated multiple LLMs, predominantly Chat Generative Pre-trained Transformer (GPT)-based models (GPT-3.5, GPT-4, GPT-4 Turbo, GPT-4o), as well as Claude, Gemini, Llama 3, and Mixtral. Text-only evaluations generally yielded higher accuracy (โ65%-90%) compared with multimodal tasks (45%-89%). GPT-4 and its variants consistently outperformed earlier versions, occasionally exceeding average human performance. Open-source models such as Llama 3 70B and Mixtral achieved comparable results to proprietary models, offering advantages in local deployment and privacy. Few studies directly compared LLM performance with human radiologists.
LLMs demonstrate promising performance in answering text-based radiology board-style examination questions, particularly GPT-4-based models. Nevertheless, significant limitations persist in multimodal tasks and complex reasoning scenarios.
Severe community-acquired pneumonia (SCAP) is a significant global health challenge due to its high mortality. Despite advances, early diagnosis and effective management remain critical. Tools like radiomics analyze imaging data for risk assessment, while machine learning and nomograms aid in personalized treatment. Large language models (LLMs) enhance clinical decision-making by analyzing data and supporting care strategies. This study integrates these methods to predict 28-day mortality in SCAP patients.
A cohort of 599 patients diagnosed with severe community-acquired pneumonia (SCAP), including 316 males and 283 females, from Shanghai East Hospital and Xiamen Humanity Hospital were enrolled in this study. High-resolution lung CT scans were used to segment three-dimensional regions of interest, from which 1,050 radiomic features were extracted. The dataset was divided into a training set (80%) and an independent test set (20%), and k-fold cross-validation was applied to optimize model performance. To address class imbalance, the SMOTE oversampling technique was employed. The study integrated radiomics, nomograms, seven machine learning models, and five LLMs to predict the 28-day mortality risk in SCAP patients. SHAP values were utilized to enhance the interpretability of feature contributions. Not only that, this study integrates the prior knowledge provided by LLMs, processed through an embedding layer, with data-driven feature learning in the main network, and dynamically fuses their outputs using a bias network with a gating mechanism, thereby improving the accuracy and interpretability of LLMs in predicting 28-day mortality risk for SCAP patients.
Key predictors of 28-day mortality included inflammatory markers, cytokines, age, CRP, and oxygenation index. Clinical-Radiomics models achieved strong accuracy (AUC 0.92). Machine learning models, particularly XGBoost (AUC 0.90), were highly effective, with SHAP analysis emphasizing radscore's importance. LLMs like Chatgpt also performed well (AUC 0.78), showcasing the potential of integrating clinical, radiomic, and AI-driven approaches.
This study demonstrates the effectiveness of radiomics, machine learning, and LLMs to predict SCAP outcomes. Models like XGBoost achieved superior accuracy, while SHAP analysis improved interpretability. These advancements highlight the potential for enhanced SCAP prognosis and personalized care strategies.
| Run at | Source | Hits | New | Status |
|---|---|---|---|---|
| 2026-04-19 00:00 | LitReview | 1 | completed | |
| 2026-04-14 14:29 | LitReview | completed | ||
| 2026-04-14 14:28 | LitReview | completed | ||
| 2026-04-14 14:28 | LitReview | error | ||
| 2026-04-14 11:57 | LitReview | error | ||
| 2026-04-12 00:00 | LitReview | 1 | 1 | completed |
| 2026-04-05 00:00 | LitReview | 2 | completed | |
| 2026-04-04 07:56 | pubmed:seed | 58 | seed-completed |
("ureteral stone*"[tiab] OR "ureteral calculus"[tiab] OR "ureteral calculi"[tiab] OR urolithiasis[tiab]) AND (KUB[tiab] OR "kidney, ureter, and bladder"[tiab] OR "kidney ureter bladder"[tiab] OR "abdominal radiograph*"[tiab] OR "plain radiograph*"[tiab] OR radiograph*[tiab] OR x-ray[tiab]) AND ("artificial intelligence"[tiab] OR "deep learning"[tiab] OR "machine learning"[tiab] OR "neural network*"[tiab] OR "computer-aided diagnosis"[tiab] OR AI[tiab] OR CNN[tiab])
Rapid and reliable detection of kidney stones on non-contrast abdominal CT is essential for timely decision-making in emergency radiology. However, rising imaging volumes and workflow pressures continue to limit reporting capacity, creating a need for AI systems capable of supporting routine diagnostic practice. Although many AI-based stone detection models have been proposed, most rely on retrospective datasets, and few have been evaluated prospectively within environments that reflect real radiology workflow conditions. This study prospectively evaluates the performance, usability, and workflow compatibility of a deep learning-based kidney stone detection model deployed within a web-based platform designed to emulate key components of routine radiology practice, enabling forward-in-time evaluation without direct integration into routine clinical operations such as PACS/RIS or clinical reporting. A dual-stage convolutional neural network was developed using an internal dataset of 235 cases (3,452 slices) and validated through five-fold patient-level cross-validation. An independent set of 732 slices served as an independent hold-out set. For prospective evaluation, the trained model was integrated into a secure, browser-based interface supporting case upload, slice-level review, independent radiologist labeling, and visualization of AI-generated predictions. Over a six-month period, three radiologists uploaded and annotated a total of 5,152 anonymized CT slices. The platform dynamically calculated diagnostic metrics and logged human-AI interactions to assess performance stability and concordance. The pilot deployment demonstrated strong diagnostic performance under real-world variability, achieving 97.83% accuracy, 94.64% sensitivity, 98.27% specificity, 88.50% precision, and a Cohen's kappa of 0.90. Concordance between radiologists and the model exhibited increasing stability across sequential pilot stages. These findings present a reproducible framework for transitioning radiological AI systems from retrospective validation toward workflow-aligned, prospective pilot deployment. Although full PACS/RIS integration was not attempted, the results underscore the importance of pilot-stage evaluation as a critical intermediary step toward clinical implementation and regulatory approval.
Axial gout is an underrecognized manifestation of monosodium urate (MSU) crystal deposition and frequently mimics inflammatory, infectious, or mechanical spinal disorders, particularly without overt hyperuricaemia. We report a 29-year-old man with recurrent low back pain for over two years and recurrent nephrolithiasis who had undergone three extracorporeal shock wave lithotripsy sessions with only transient relief. Admission laboratory investigations showed normal serum urate (262 ยตmol/L, ~โ4.4ย mg/dL) but markedly elevated Cโreactive protein (119ย mg/L) and erythrocyte sedimentation rate (53ย mm/h). Retrospective review of prior hospitalizations for ureteral stones revealed that routine serum urate measurements had never been elevated. Conventional imaging, including plain radiography and nonโcontrast computed tomography, was inconclusive. Dualโenergy computed tomography (DECT) revealed MSU deposition in the L4/L5 and L5/S1 facet joints and sacral foramina, supporting the diagnosis of axial gout. The patient was treated with etoricoxib (60ย mg daily) and febuxostat (20ย mg daily). On telephone followโup on 27 May 2026 (the only followโup to date), he reported that the J stent had been removed and low back pain had completely resolved without recurrence. A single repeat laboratory test performed locally was reported as normal, though specific values were unavailable. A systematic literature review of 58 eligible articles revealed that the lumbar spine was most frequently involved (โโ60%), DECT sensitivity ranged from 78% to 100%, and 77.5% of historically reported cases required surgical diagnosis, underscoring the underrecognition of axial gout in non-surgical settings. This caseโbased review highlights that axial gout should be considered in young nonโhyperuricaemic patients with persistent axial pain, and that DECT is a valuable nonโinvasive diagnostic tool when interpreted alongside clinical and laboratory features.
CONTEXT: Experimental evidence supporting the existence of the viscerosomatic reflex highlights an involvement of multiple vertebral levels when renal pathology is present. Further exploration of this reflex, particularly in the context of nephrolithiasis, could offer valuable insights for osteopathic treatments related to this pathology. Open-sourced machine learning datasets provide a valuable source of imaging data for investigating osteopathic phenomena including the viscerosomatic reflex.
This study aimed to compare the rotation of vertebrae at levels associated with the viscerosomatic reflex in renal pathology in patients with nephrolithiasis vs. those without kidney stones.
A total of 210 unenhanced computed tomography (CT) scans were examined from an open-sourced dataset designed for kidney and kidney stone segmentation. Among these, 166 scans were excluded due to pathologies that could affect analysis (osteophytes, renal masses, etc.). The 44 scans included in the analysis encompassed 292 relevant vertebrae. Of those, 15 scans were of patients with kidney stones in the right kidney, 13 in the left kidney, 7 bilaterally, and 11 without kidney stones. These scans included vertebral levels from T5-L5, with the majority falling within T10-L5. An open-sourced algorithm was employed to segment individual vertebrae, generating models that maintained their orientation in three-dimensional (3D) space. A self-coded 3D slicer module utilizing vertebral symmetry for rotation detection was then applied. Two-way analysis of variance (ANOVA) testing was conducted to assess differences in vertebral rotation between the four possible combinations of kidney stone location (left-sided, right-sided, bilateral, or none) and vertebral levels (T10-L4). Subsequently, the two-way ANOVA analysis was narrowed down to include various combinations of three vertebral levels (T10-L4) to identify the most significant levels.
We observed a statistically significant difference in average vertebral rotation (p=0.0038) dependent on kidney stone location. Post-hoc analysis showed an average difference in rotation ofย -1.38ยฐ leftward between scans that contained left kidney stones compared to no kidney stones (p=0.027), as well as an average difference ofย -1.72ยฐ leftward in the scans containing right kidney stones compared to no kidney stone (p=0.0037). The average differences in rotation between the remaining stone location combinations were not statistically significant. Narrowed analysis of three vertebral level combinations showed a single statistically significant combination (T10, T12, and L4) out of a total of 35 combinations (p=0.028). A subsequent post-hoc procedure showed that angular rotation at these levels had the only statistically significant contribution to the difference between scans containing right kidney stones and no kidney stones (p=0.046).
This study observed a statistically significant difference in the rotation of vertebrae at the levels associated with the viscerosomatic reflex between patients with unilateral kidney stones and those without kidney stones. The vertebral levels with the highest significance of association with this finding, particularly in right kidney stones, were T10, T12, andย L4.
Urolithiasis is a prevalent urological condition, and Non-Contrast Computed Tomography (NCCT) is the gold standard for diagnosis. In recent years, there has been growing interest in investigating machine learning (ML)- based detection of urolithiasis and the wider potential of AI in urology.
To synthesise the diagnostic accuracy of ML-based UTS detection on NCCT and in externally validated cohorts.
We performed a systematic review and bivariate meta-analysis of studies evaluating ML for detecting urinary stones. We used QUADAS-2 to assess the risk of bias. Subgroup analyses examined performance by model type, classification task, stone site, dataset source, and CT orientation. Bivariate meta-regression was performed to further explore heterogeneity. Publication bias was assessed using Deeks' test. The study was prospectively registered in Prospero (CRD42024542409).
Forty-five studies were included qualitatively. 24 studies (49,277 test images) provided extractable 2ย รย 2 data for meta-analysis. For NCCT (10 studies), pooled sensitivity was 96% (95% CI 92-98%) and pooled specificity was 98% (95% CI 97-99%). In externally validated NCCT cohorts (4 studies; 1,056 images), pooled sensitivity was 95% (95% CI 92-97%) and pooled specificity was 96% (95% CI 70-100%). Subgroup performance remained high, but heterogeneity persisted; meta-regression found stone site contributed to variability (pย =ย 0.014), while other moderators were not significant. Deeks' test showed no small-study effects (pย =ย 0.571).
ML models show high image-level diagnostic performance for stone detection on NCCT and may support radiologists as decision support tools. Translation is limited by heterogeneity and limited external validation. Future studies should move beyond detection-alone tasks towards clinically meaningful outputs that are actionable for radiologists and downstream clinicians, including urologists and nephrologists.
Three-dimensional (3D) rendering of urologic pathology plays an important role in simulation-based education, surgical training, and computer vision research; however, a standardized, open-access repository of high-fidelity kidney stone models stratified by chemical composition is lacking. We developed and validated a reproducible photogrammetry-based pipeline to generate realistic 3D kidney stone renderings. Chemically characterized human stones composed of calcium oxalate monohydrate (COM) (nโ=โ11), uric acid (UA) (nโ=โ5), cystine (nโ=โ4), magnesium ammonium phosphate hexahydrate/carbonate apatite (MAPH/CA) (nโ=โ2), and calcium hydrogen phosphate dihydrate (CHPD) (nโ=โ3) were photographed using a custom-built rotating stage and dual fixed 4ย K cameras. Rendered models were sent to 25 endourologists using a 5-point Likert-scale survey assessing geometric and surface texture fidelity. Successful 3D renderings were obtained for 8/11 COM stones, 5/5 UA stones, 2/2 MAPH/CA fragments, and 3/3 CHPD fragments, while all cystine stones failed to render. Across stone types, mean fidelity scores were highest for UA and COM stones (mean 3.8-3.9), intermediate for calcium phosphate stones (mean 3.6-3.8), and lowest for struvite stones (mean 3.0-3.3). Geometry scores were higher than texture scores overall, though this difference was not significant. Significant differences in geometric fidelity were observed across stone compositions (ฯยฒ = 9.30, pโ=โ0.026). Inter-rater reliability was poor for individual evaluators (ICCโ=โ0.10) but moderate for aggregated mean ratings (ICCโ=โ0.67). This validated workflow enables the creation of generally realistic, open-access 3D kidney stone models (github.com/uro-glidar/3d-rendering-diverse-stones) for simulation, education, and future machine learning applications in endourology.
Visceral adiposity has been implicated in metabolic dysregulation and chronic inflammation, both of which may contribute to kidney stone recurrence. However, accurate and reproducible quantification of visceral fat in routine clinical practice remains challenging. This study aimed to investigate the association between visceral fat area (VFA) quantified from computed tomography (CT) images and the risk of kidney stone recurrence. We retrospectively analyzed patients with urolithiasis who underwent abdominal CT imaging. Visceral fat area was automatically quantified using a previously validated artificial intelligence (AI)-based CT segmentation system. Clinical characteristics and stone recurrence outcomes were collected. Multivariable regression models were applied to assess the association between VFA and stone recurrence after adjustment for relevant confounders. A total of 131 patients were included, of whom 73 (48%) experienced stone recurrence during a mean follow-up of 47 weeks. Patients with recurrence had significantly higher visceral fat area. High VFA was independently associated with increased recurrence risk (adjusted hazard ratio [HR] 1.71, 95% confidence interval [CI] 1.03 to 2.82). Subgroup analyses demonstrated a stronger association in younger patients (HR 2.45, 95% CI 1.23 to 4.89), while no significant association was observed in older patients or across sexes. CT-derived visceral fat area was independently associated with kidney stone recurrence in this retrospective cohort. These findings suggest that visceral adiposity may serve as a useful imaging biomarker for risk stratification. Further prospective studies are warranted to validate its clinical utility.
Ceftriaxone-associated biliary pseudolithiasis is well reported, but the true risk of urolithiasis remains less clearly defined. Prior reviews frequently combine biliary and urinary tract calcifications or rely mainly on descriptive findings, limiting accurate incidence estimates.
This systematic review and meta-analysis aimed to provide the first pooled frequency estimate specifically for ceftriaxone-induced urolithiasis in children. DATA SOURCES: A systematic search of PubMed, Google Scholar, Web of Science, and Scopus was conducted up to November 2025. STUDY ELIGIBILITY CRITERIA: Studies reporting the cases of urolithiasis in pediatric patients receiving ceftriaxone were included. STUDY APPRAISAL AND DATA SYNTHESIS: Data extraction and synthesis were performed by two independent reviewers and analyzed by Stata 19.5. Pooled proportions were calculated using a random-effects model with REML estimation to determine overall frequency, between-study variance (ฯ2), and 95% prediction intervals (PI).
Eight studies met the inclusion criteria. The pooled frequency of ceftriaxone-induced urolithiasis was 7% (95% CI: 2-12%). Heterogeneity was substantial (I2โ=โ89.8%; ฯ2โ=โ0.0034), and the 95% PI ranged from 2.7 to 15.8%. Subgroup analyses showed that retrospective studies from Asian regions reported markedly higher rates (up to 34%), whereas prospective Western studies consistently demonstrated lower frequencies. Sensitivity analysis excluding the main outlier reduced the pooled estimate to 4% (pโ=โ0.004). Publication bias was detected (Egger's test, pโ=โ0.006), indicating underreporting of studies with low event rates. CONCLUSIONS AND IMPLICATIONS OF KEY
This meta-analysis suggests that ceftriaxone-associated urinary tract lithiasis occurs in approximately 7% of pediatric patients, with substantial variability driven by study design and diagnostic approach. Well-powered, prospective studies with standardized imaging protocols are required to define the true incidence and modifiable risk factors. SYSTEMATIC REVIEW REGISTRATION NUMBER: (CRD42024598001).
Accurate assessment of stone burden is fundamental in urolithiasis, as it directly influences treatment selection and prognostic evaluation. Although maximum stone diameter on non-contrast computed tomography remains the most widely used parameter, it does not adequately reflect the three-dimensional complexity of urinary calculi. This review aimed to summarize the evolution of stone burden assessment from conventional imaging-based methods to software-assisted volumetry and artificial intelligence (AI)-driven tools, with emphasis on their accuracy, reproducibility, and clinical utility.
A narrative review of the literature was performed using PubMed/MEDLINE, Scopus, and Google Scholar for English-language studies published up to March 2026. Original studies, validation studies, technical reports, reviews, and guideline-related papers addressing conventional CT-based measurement, software-assisted volumetry, AI-based stone segmentation, and the clinical significance of stone volume were included. Due to heterogeneity in study design and reported outcomes, the evidence was synthesized narratively.
Maximum stone diameter remains simple and widely available, but it incompletely represents true stone burden, particularly in larger or irregular stones. Formula-based ellipsoid calculations are practical yet show limited accuracy in complex geometries. Semi-automated CT-based segmentation software provides more reliable volumetric assessment, with excellent agreement with reference standards and reduced interobserver variability. AI-based approaches have further improved efficiency by enabling rapid and highly accurate automated stone detection and volume calculation. Across the reviewed literature, stone volume was consistently shown to be more clinically informative than linear dimensions for predicting spontaneous passage, stone-free rates after shockwave lithotripsy and ureteroscopy, and future symptomatic events during surveillance.
Stone volume offers a more accurate and clinically meaningful estimate of stone burden than maximum stone diameter alone. The transition from formula-based methods to software-assisted and AI-driven volumetry represents an important advance in urolithiasis imaging. Wider adoption will depend on standardized imaging protocols, improved software accessibility, and validation of volume-based thresholds for routine clinical practice.
PURPOSE OF REVIEW: Stone volume represents the most accurate measure of urolithiasis burden. While this may be obvious to all, this stone metric has not yet found its way into standard clinical practice or guidelines algorithms guiding treatment strategies. Although linear measurements have their obvious limitations, they are still most commonly used in practice and reported in literature as well as guidelines. This review evaluates the current available evidence supporting stone volume as the most important stone metric. RECENT
By now, literature has been able to confidently demonstrate that stone volume is a more accurate predictor of stone free status in shockwave lithotripsy and ureteroscopy. Recent advances in three-dimensional imaging reconstruction, automated segmentation and artificial intelligence have lowered thresholds for obtaining stone volume from computed tomography imaging. SUMMARY: Today, historical barriers have been overcome, and fast, reproducible stone volume assessment is at our fingertips. And yet, we do not use volume in our daily practice, as we ourselves hold back evolution. While future efforts should be put towards embracing volume as a stone metric, more studies are also needed to identify volume-based thresholds for treatment success aiding in the development of new treatment algorithms.
In low-invasive surgical treatment of urolithiasis, there is a need for an analytical method to determine the chemical composition of urinary stones in real-time mode, i.e., intraoperatively. While a thorough phase analysis can be done after the surgery, preliminary information about a target stone would be helpful for the specialists for choosing an optimal strategy of treatment and giving some immediate dietary or drug prescriptions to a patient. Near-infrared spectroscopy (NIRS) is a good candidate for such a method that can provide immediate results without obligatory sample preparation. Fiber optic probes, often used for acquiring near-infrared spectra, are compatible with surgical instrumentation. Chemometric algorithms can successfully resolve the complexity of NIR spectra, which consist of overlapped signals. For the first time, we applied NIRS in diffuse reflectance mode to classify three major types of urinary stones: oxalates, urates, and phosphates. To imitate the real conditions of a surgery, the NIR spectra were acquired not only under ambient conditions but also in saline medium. A trained and optimized multinomial classifier (Error Correcting Output Codes) showed an acceptable precision and recall for an independent validation dataset. Even considering the strong absorbance of saline, the calculated geometric mean was 94ย %, 87ย %, and 71ย % for oxalates, urates, and phosphates, respectively. A first real-time approbation during a real surgery (percutaneous nephrolithotomy) demonstrated a compatibility of the suggested approach with the surgical protocols and a good agreement of the acquired NIR spectra and the results of reference X-ray phase analysis.
| Run at | Source | Hits | New | Status |
|---|---|---|---|---|
| 2026-08-23 00:00 | LitReview | completed | ||
| 2026-08-16 00:00 | LitReview | completed | ||
| 2026-08-09 00:00 | LitReview | 1 | completed | |
| 2026-08-02 00:00 | LitReview | 1 | 1 | completed |
| 2026-07-26 00:00 | LitReview | completed | ||
| 2026-07-19 00:00 | LitReview | completed | ||
| 2026-07-12 00:00 | LitReview | 1 | 1 | completed |
| 2026-07-05 00:00 | LitReview | 2 | 1 | completed |