arXiv (Cornell University)· 2020· Preprint
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- 21,699citations
- 2020year
Short summary
A pure Transformer architecture, Vision Transformer (ViT), directly applied to sequences of image patches achieves state-of-the-art results in image classification, outperforming convolutional networks with fewer computational resources for training.
AI-generated from the title and abstract; the full text is not read.
TakeawaysIn the app
Key pointsIn the app
Ask the paperIn the app
The rest is in the Pofolia app
Takeaways, key points and questions to the paper; new summaries every day for your field. Free.
Sign in on the web to openField: Computer Vision and Pattern Recognition
Computer Vision and Pattern RecognitionComputer Science