Abstract ID: 26-295

Automated Phase Recognition and Anatomical Landmark Segmentation in Endonasal Dacryocystorhinostomy Using Video Foundation Models

Author: Nir Mathur
Base Hospital / Institution: Guy’s, King’s and St. Thomas’

Presentation Type: Rapid Fire Presentation
Session: Trauma / War / Miscellaneous
Date: 12th September
Time: 17.05PM

Purpose

Endonasal DCR has a recognised learning curve. Trainees struggle to identify key anatomical landmarks, particularly during bone removal and sac identification. No automated surgical workflow analysis tools currently exist for this procedure. We assess whether video foundation models can recognise surgical phases and segment relevant anatomy during endonasal DCR, with the goal of providing objective training feedback.


Methods

Endoscopic video from endonasal DCR cases performed at a single centre was analysed retrospectively. Four surgical phases were defined with the operating surgeon: mucosal flap creation, bone removal, sac identification, and sac opening. Phase recognition was performed using a vision-language model with procedure-specific prompting and temporal context across sequential frames. Anatomical segmentation used the Segment Anything Model 3 (SAM3) in video propagation mode, with surgeon-placed point prompts on index frames. Target structures included the maxillary line, middle turbinate, lacrimal bone, and lacrimal sac. Phase accuracy was evaluated against surgeon annotations. Segmentation quality was assessed by intersection-over-union and surgeon review of clinical usefulness.


Results

Phase recognition correctly identified surgical step transitions across analysed cases. SAM3 maintained segmentation masks across consecutive frames, performing best on structures with distinct visual boundaries. Structures with subtle tissue interfaces required more frequent re-prompting.


Conclusion

Video foundation models show feasibility for both phase recognition and anatomical segmentation in endonasal DCR. Combining these capabilities allows relevant structures to be highlighted only at the surgical steps where trainees need them most, without cluttering the operative view. To our knowledge, this represents the first application of automated workflow analysis to DCR.


Additional Authors

First name Last name Base Hospital / Institution
Robert Thomas Moorfields Eye Hospital
Alejandro Granados King’s College London
Hannah Timlin Moorfields Eye Hospital
Mohsan Malik King’s College London

↑ Back to top