Unsupervised multi-latent space reinforcement learning framework for video summarization in ultrasound imaging

Mathews, Roshan P; Panicker, Mahesh Raveendranatha; Hareendranathan, Abhilash R; Chen, Yale Tung; Jaremko, Jacob L; Buchanan, Brian; Narayan, Kiran Vishnu; C, Kesavadas; Mathews, Greeta

Electrical Engineering and Systems Science > Image and Video Processing

arXiv:2109.01309 (eess)

COVID-19 e-print

Important: e-prints posted on arXiv are not peer-reviewed by arXiv; they should not be relied upon without context to guide clinical practice or health-related behavior and should not be reported in news media as established information without consulting multiple experts in the field.

[Submitted on 3 Sep 2021]

Title:Unsupervised multi-latent space reinforcement learning framework for video summarization in ultrasound imaging

Authors:Roshan P Mathews, Mahesh Raveendranatha Panicker, Abhilash R Hareendranathan, Yale Tung Chen, Jacob L Jaremko, Brian Buchanan, Kiran Vishnu Narayan, Kesavadas C, Greeta Mathews

View PDF

Abstract:The COVID-19 pandemic has highlighted the need for a tool to speed up triage in ultrasound scans and provide clinicians with fast access to relevant information. The proposed video-summarization technique is a step in this direction that provides clinicians access to relevant key-frames from a given ultrasound scan (such as lung ultrasound) while reducing resource, storage and bandwidth requirements. We propose a new unsupervised reinforcement learning (RL) framework with novel rewards that facilitates unsupervised learning avoiding tedious and impractical manual labelling for summarizing ultrasound videos to enhance its utility as a triage tool in the emergency department (ED) and for use in telemedicine. Using an attention ensemble of encoders, the high dimensional image is projected into a low dimensional latent space in terms of: a) reduced distance with a normal or abnormal class (classifier encoder), b) following a topology of landmarks (segmentation encoder), and c) the distance or topology agnostic latent representation (convolutional autoencoders). The decoder is implemented using a bi-directional long-short term memory (Bi-LSTM) which utilizes the latent space representation from the encoder. Our new paradigm for video summarization is capable of delivering classification labels and segmentation of key landmarks for each of the summarized keyframes. Validation is performed on lung ultrasound (LUS) dataset, that typically represent potential use cases in telemedicine and ED triage acquired from different medical centers across geographies (India, Spain and Canada).

Comments:	24 pages, submitted to Elsevier Medical Image Analysis for review
Subjects:	Image and Video Processing (eess.IV); Computer Vision and Pattern Recognition (cs.CV)
Cite as:	arXiv:2109.01309 [eess.IV]
	(or arXiv:2109.01309v1 [eess.IV] for this version)
	https://doi.org/10.48550/arXiv.2109.01309

Submission history

From: Mahesh Raveendranatha Panicker [view email]
[v1] Fri, 3 Sep 2021 04:50:35 UTC (1,806 KB)

Electrical Engineering and Systems Science > Image and Video Processing

Title:Unsupervised multi-latent space reinforcement learning framework for video summarization in ultrasound imaging

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Electrical Engineering and Systems Science > Image and Video Processing

Title:Unsupervised multi-latent space reinforcement learning framework for video summarization in ultrasound imaging

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators