Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

CNN and ViT Efficiency Study on Tiny ImageNet and DermaMNIST Datasets

About

This study evaluates the trade-offs between convolutional and transformer-based architectures on both medical and general-purpose image classification benchmarks. We use ResNet-18 as our baseline and introduce a fine-tuning strategy applied to four Vision Transformer variants (Tiny, Small, Base, Large) on DermatologyMNIST and TinyImageNet. Our goal is to reduce inference latency and model complexity with acceptable accuracy degradation. Through systematic hyperparameter variations, we demonstrate that appropriately fine-tuned Vision Transformers can match or exceed the baseline's performance, achieve faster inference, and operate with fewer parameters, highlighting their viability for deployment in resource-constrained environments.

Aidar Amangeldi, Angsar Taigonyrov, Muhammad Huzaifa Jawad, Chinedu Emmanuel Mbonu• 2025

Related benchmarks

TaskDatasetResultRank
Order-level classificationBIOSCAN-5M 1.0 (test)
Accuracy75.88
11
Order-level classificationInsects-1M 1.0 (test)
Accuracy74
11
Showing 2 of 2 rows

Other info

Follow for update