| Large‑scale datasets underpin modern vision systems, yet many images are redundant or detrimental to model generalisation. We introduce a Deep Q‑Network (DQN) driven sample‑selection policy that adaptively curates training data for semantic segmentation on the LaRS maritime benchmark. Based on the sample selection trajectory during training, we demonstrate the DQN agent's proficiency in selecting salient samples and highlighting those with little to no salient information. We use this analysis to retrain our model on a curated reduced dataset, to determine the efficacy of our approach to reducing dataset size through saliency quantification by the DQN agent. Our findings demonstrate the efficacy of reinforcement‑based dataset distillation for dense prediction and highlight the advantage of learned saliency quantification over hand‑crafted heuristics. As a result, we train on 40\% of the original dataset and maintain comparable performance on the validation dataset. We further discuss the implications of dataset bias and propose extensions and limitations to our work. |
*** Title, author list and abstract as submitted during Camera-Ready version delivery. Small changes that may have occurred during processing by Springer may not appear in this window.