aalto1 untyped-item.component.html

Joint modeling of X-ray images of different dimensions with native resolution vision transformer

Loading...
Thumbnail Image

URL

Journal Title

Journal ISSN

Volume Title

School of Science | Master's thesis
Electronic archive copy is available via Aalto Thesis Database.

Department

Mcode

Language

en

Pages

77

Series

Abstract

Medical images play a crucial role in detecting potential sicknesses in hard-to-reach parts of the human body. Fortunately, for the time being, computer vision methods are helpful in distinguishing between ailing organs and healthy ones. It’s well known how effectively medical image processing can improve clinical verdicts and sometimes prolong or even save patients’ lives. Current work demonstrates the precision with which computer vision methods can predict diseases from Chest X-ray images, including their original sizes and aspect ratios. In addition, some performance will be shown regarding training time, which saves computer resources and delivers results at a fast pace. Depending on the model, a transfer learning paradigm with pre-trained models on ImageNet-1k was used to achieve reasonable results. The output result shows how likely patients are to be diagnosed with all 14 diseases described in the CheXpert X-Ray dataset. The fine-tuning process was applied to all models. A baseline was established with ResNet-18, and the results were improved by applying more complex deep learning models such as Vision Transformers. Ultimately, the NaViT model demonstrates promising results when using different resolutions. This study could aid in understanding the rationale behind selecting optimal image resolution and experimenting with different image resolutions to efficiently train the NaViT model with reduced training duration.

Description

Supervisor

Marttinen, Pekka

Thesis advisor

Kumar, Yogesh

Other note

Citation

Endorsement

Review

Supplemented By

Referenced By