Computer Vision and Pattern Recognition papers
Pofolia’s corpus holds 99 papers from Computer Vision and Pattern Recognition (2002–2026), each with a short summary. Below, the 20 most-cited, then the most recently added.
At a glance
- Most cited: Deep Residual Learning for Image Recognition (2016, 223,583 citations).
- The 20 papers listed have 910,043 citations between them.
- 3 of the 3 with a known journal quartile appeared in a Q1 journal.
- 2 have a free full text (open access).
- Most frequent journals: IEEE Transactions on Pattern Analysis and Machine Intelligence, 2021 IEEE/CVF International Conference on Computer Vision (ICCV), arXiv (Cornell University).
- Published between 2014 and 2024.
Most cited
Ranked by citation count. Because citations accumulate over time, this list naturally leans towards work published a few years ago; for where the field is now, see “recently added”. How to read the signals
Deep Residual Learning for Image Recognition
2016 · 223,583 citations
A new residual learning framework enables training significantly deeper neural networks, achieving state-of-the-art image recognition performance.
ImageNet classification with deep convolutional neural networks
Communications of the ACM · 2017 · Q1 · SJR 1.00 · FWCI 4435.36 · 75,715 citations · Open access
A deep convolutional neural network achieved state-of-the-art ImageNet classification, reducing top-5 error rates to 17.0% in 2010 and a winning 15.3% in 2012.
Very Deep Convolutional Networks for Large-Scale Image Recognition
arXiv (Cornell University) · 2014 · 75,542 citations · Open access
Increasing convolutional network depth to 16-19 weight layers, using small 3x3 filters, significantly improves accuracy in large-scale image recognition over prior configurations.
Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2016 · Q1 · SJR 4.00 · FWCI 1672.08 · 54,589 citations
This paper introduces a Region Proposal Network (RPN) that integrates seamlessly with object detection networks, enabling real-time performance by generating region proposals at nearly no computational cost.
Going deeper with convolutions
2015 · 46,948 citations
A novel deep convolutional neural network architecture, Inception, has set a new state-of-the-art for image classification and detection on the ILSVRC14 challenge.
Densely Connected Convolutional Networks
2017 · 44,970 citations
DenseNets introduce a new convolutional network architecture where each layer connects to every other layer in a feed-forward manner, significantly increasing the number of direct connections compared to traditional networks.
Fully convolutional networks for semantic segmentation
2015 · 37,045 citations
This paper introduces fully convolutional networks (FCNs) that can perform semantic segmentation on images of any size, achieving state-of-the-art results with efficient inference.
Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation
2014 · 31,870 citations
A new algorithm, R-CNN, significantly boosts object detection accuracy by over 30% on the PASCAL VOC 2012 dataset, achieving a 53.3% mAP.
Rethinking the Inception Architecture for Computer Vision
2016 · 31,095 citations
A new network design, Inception, efficiently scales up convolutional networks by factorizing convolutions and using aggressive regularization, achieving state-of-the-art results on the ILSVRC 2012 classification challenge.
Swin Transformer: Hierarchical Vision Transformer using Shifted Windows
2021 IEEE/CVF International Conference on Computer Vision (ICCV) · 2021 · 30,585 citations
The Swin Transformer introduces a hierarchical architecture with shifted windows, enabling efficient self-attention computation for general-purpose computer vision tasks.
Mask R-CNN
2017 · 28,989 citations
Mask R-CNN is a new framework that efficiently detects objects and generates high-quality segmentation masks for each instance in an image, extending Faster R-CNN with a mask prediction branch.
Feature Pyramid Networks for Object Detection
2017 · 28,919 citations
Feature Pyramid Networks (FPNs) efficiently construct a rich multi-scale feature representation from deep convolutional networks with marginal computational cost.
Squeeze-and-Excitation Networks
2018 · 28,761 citations
A novel 'Squeeze-and-Excitation' (SE) block adaptively recalibrates channel-wise feature responses by modeling interdependencies between channels, significantly improving deep learning model performance.
Fast R-CNN
2015 · 28,041 citations
Fast R-CNN is a novel object detection method that significantly improves training and testing speeds while increasing accuracy by efficiently classifying object proposals using deep convolutional networks.
Focal Loss for Dense Object Detection
2017 · 26,050 citations
A novel 'Focal Loss' function effectively addresses the extreme class imbalance in training dense object detectors, enabling them to surpass the accuracy of slower, two-stage detectors while maintaining speed.
MobileNetV2: Inverted Residuals and Linear Bottlenecks
2018 · 25,715 citations
MobileNetV2 introduces a novel mobile architecture featuring inverted residuals and linear bottlenecks, significantly improving state-of-the-art performance on mobile tasks like image classification, object detection, and semantic segmentation across various model sizes.
Learning Multiple Layers of Features from Tiny Images
2024 · FWCI 670.28 · 25,502 citations
Researchers developed a multi-layer generative model that learns meaningful visual features from millions of tiny web images, mimicking aspects of the human visual cortex.
DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs
IEEE Transactions on Pattern Analysis and Machine Intelligence · 2017 · Q1 · SJR 4.00 · FWCI 632.10 · 22,241 citations
DeepLab introduces 'atrous convolution' to control feature resolution and enlarge filter receptive fields in deep networks, enabling better semantic image segmentation.
Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks
2017 · 22,010 citations
This paper introduces a novel method for image-to-image translation that does not require paired training data, a common limitation in existing techniques.
Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization
2017 · 21,873 citations
Grad-CAM generates visual explanations for Convolutional Neural Network (CNN) decisions by highlighting important image regions, making complex models more transparent.
Recently added
Constellation Dataset: Benchmarking High-Altitude Object Detection for an Urban Intersection
International Journal of Computer Vision · 2026 · Q1 · SJR 3.00 · 1 citations · Open access
The Constellation dataset, comprising 13K images from high-altitude urban cameras, enables benchmarking object detection for privacy-preserving edge processing, revealing a 10% lower AP for small pedestrians vs. vehicles.
Performance analysis and deployment considerations of post-quantum cryptography for consumer Electronics
Scientific Reports · 2026 · Q1 · FWCI 11.11 · 3 citations · Open access
Lattice-based post-quantum cryptography (PQC) schemes, specifically ML-KEM and ML-DSA, offer a promising balance of execution speed and object size for gateway-class consumer electronics, outperforming RSA/ECC baselines in key encapsulation and digital signature tasks.
Causal Cohesion and Story Coherence
2026 · 360 citations
Children's ability to understand and remember stories hinges on the causal links between events, with more cohesive stories leading to easier comprehension and memory representation.
CP-Diffusion: Conditional Prompt-Based Diffusion Models for Video Generation
ACM Transactions on Multimedia Computing Communications and Applications · 2026 · Q1 · FWCI 11.61 · 3 citations
CP-Diffusion introduces a few-shot learning approach using a Multi-Head Temporal Attention (MHTA) module for motion customization in text-to-video diffusion models, significantly reducing computational needs while improving motion quality.
Abstract Functional Language Logic: A Competitive Mixture of Experts Architecture for Paradox-Free Reasoning and Adaptive Intelligence
Zenodo (CERN European Organization for Nuclear Research) · 2025 · FWCI 44.00 · 595 citations
This paper introduces a Competitive Mixture of Experts (CMoE) architecture based on Functional Language Logic (FLL) that replaces probabilistic LLM reasoning with mathematical functional approximators for paradox-free, efficient deduction.
Reference-Vector Removed Product Quantization for Approximate Nearest Neighbor Search
Applied Sciences · 2025 · Q2 · FWCI 0.46 · 1 citations
Reference-Vector Removed Product Quantization (RvRPQ) improves approximate nearest neighbor (ANN) search accuracy by subtracting encoded reference-vectors from database vectors, yielding residuals that are then quantized.
VM-UNet: Vision Mamba UNet for Medical Image Segmentation
ACM Transactions on Multimedia Computing Communications and Applications · 2025 · Q1 · FWCI 226.04 · 541 citations
VM-UNet, a novel U-shaped architecture for medical image segmentation, leverages State Space Models (SSMs) like Mamba to achieve linear computational complexity and effective long-range modeling, outperforming previous CNN and Transformer models.
V0.1IC
Zenodo (CERN European Organization for Nuclear Research) · 2025 · 499 citations
A new deep learning framework, Hamming Cube, jointly learns compact binary codes and continuous embeddings while preserving Hamming distance structure.
Journals in this field
The journals that publish most of this field’s papers. Quartile (Q1–Q4), SJR and h-index are from SCImago Journal Rank; “in this field” is how many of the journal’s pooled papers belong here. What is a Q1 journal? · What is the h-index?
| Journal | Quartile | SJR | h-index | In this field | Summaries |
|---|---|---|---|---|---|
| IEEE Transactions on Pattern Analysis and Machine Intelligence | Q1 | 4.00 | 460 | 8 | 27 |
| Proceedings of the VLDB Endowment | Q1 | 1.00 | 174 | 3 | 3 |
| IEEE Transactions on Circuits and Systems for Video Technology | Q1 | 2.00 | 206 | 2 | 7 |
| Communications of the ACM | Q1 | 1.00 | 259 | 2 | 14 |
| Journal of Big Data | Q1 | 1.00 | 108 | 2 | 5 |
| ACM Transactions on Multimedia Computing, Communications and Applications | Q1 | — | 81 | 2 | 4 |
| Artificial Intelligence Review | Q1 | 3.00 | 169 | 1 | 3 |
| IEEE Transactions on Medical Imaging | Q1 | 2.00 | 283 | 1 | 9 |
Every journal in Computer Vision and Pattern Recognition (Q1–Q4) →
Add this field to your daily feed
Pick your interests and new work in your area arrives every day, summarised. Full summaries live in the app.
Open the appOther subfields in Computer Science
- Computer Networks and Communications
- Computer Science Applications
- Computer Graphics and Computer-Aided Design
- Information Systems
- Hardware and Architecture
- Computational Theory and Mathematics
- Human-Computer Interaction
- Signal Processing
- Artificial Intelligence
- Software