학술논문

Home

자료검색

학술논문

검색결과 돌아가기

검색화면

내보내기 프린트

Vision Transformer Hashing for Image Retrieval

Resource Type: Conference
Authors: Dubey, Shiv Ram; Singh, Satish Kumar; Chu, Wei-Ta
Source: 2022 IEEE International Conference on Multimedia and Expo (ICME) Multimedia and Expo (ICME), 2022 IEEE International Conference on. :1-6 Jul, 2022
Subject: Communication, Networking and Broadcast Technologies
Components, Circuits, Devices and Systems
Computing and Processing
Signal Processing and Analysis
Deep learning
Visualization
Codes
Image recognition
Head
Convolution
Image retrieval
Image Retrieval
Transformer
Deep Learning
Hashing
Multimedia Retrieval
Language
ISSN: 1945-788X

Online Access

Full Text (IEEE)

초록

Recently, Transformer has emerged as a new architecture in deep learning by utilizing self-attention without convolution. Transformer is also extended to Vision Transformer (ViT) for the visual recognition with a promising performance on ImageNet. In this paper, we propose a Vision Transformer Hashing (VTS) for image retrieval. We utilize the pre-trained ViT on ImageNet as the backbone network and add the hashing head. The proposed VTS model is fine tuned for hashing under six different image retrieval frameworks with their objective functions. We perform the extensive experiments on CIFAR10, ImageNet, NUS-Wide, and COCO datasets. The proposed VTS based image retrieval outperforms the recent state-of-the-art hashing techniques with a significant margin. We also find the proposed VTS model as the backbone network is better than the existing networks, such as AlexNet and ResNet. The code is released at https://github.com/shivram1987/VisionTransformerHashing.

공지

DAU Library

학술논문

요약정보

Vision Transformer Hashing for Image Retrieval

Online Access

초록