22nd AIAI 2026, 16 - 19 July 2026, Chania, Crete, Greece

A Fine-Tuned Vision Language Model Framework for Medical Forensic Image Analysis

Kallipolitis Athanasios, Tziomaka Melina , Zafeiriou Argyrios, Koulouris Dionysios, Maglogiannis Ilias

Abstract:

  Medical forensics investigations put significant burden on image anal-ysis for reaching conclusions concerning injury findings, characteristics, and causes. When conducted manually, the analysis is time-consuming, subjective, and prone to bias. To solve this issue, an AI-based forensic Image Description System is proposed as a methodology for systematically collecting and standard-izing expert knowledge in autopsy injury descriptions. Zero-shot approaches re-lying on general-purpose visual-language models present a baseline against our system that employs a fine-tuned BLIP-2 (Bootstrapping Language-Image Pre-training 2) model. Trained on a custom dataset of 3,113 forensic image-caption pairs and implemented using LoRA (Low-Rank Adaptation) for parameter-effi-cient fine-tuning, the model generates contextually accurate forensic descrip-tions. These descriptions are parsed into structured classifications across seven specialized tasks, including finding type identification, mechanism inference, and anatomical localization. The experimental results show that fine-tuning the BLIP-2 model on forensic data improves classification accuracy by an average of 21.31% across seven tasks compared to zero-shot BLIP-2. Moreover, the cap-tion generation quality metrics (BLEU-4: 22.1%, ROUGE-L: 43.3%) confirm the system's practical utility. Finally, an integrated Human-in-the-Loop (HITL) ar-chitecture is proposed as future work to enable continuous expert consensus and iterative dataset refinement.  

*** Title, author list and abstract as submitted during Camera-Ready version delivery. Small changes that may have occurred during processing by Springer may not appear in this window.