Key papers in Computer Vision and Pattern Recognition

Pofolia’s corpus holds 85 papers from the Computer Vision and Pattern Recognition subfield (2002–2025). The list below starts with the most cited.

Most cited

Ranked by citation count. Because citations accumulate over time, this list naturally leans towards work published a few years ago; for where the field is now, see “recently added”.

  • Deep Residual Learning for Image Recognition

    2016 · 223,583 citations

    A new residual learning framework enables training significantly deeper neural networks, achieving state-of-the-art image recognition performance.

    Go to source

  • ImageNet classification with deep convolutional neural networks

    Communications of the ACM · 2017 · Q1 · SJR 1.00 · FWCI 4435.36 · 75,715 citations · Open access

    A deep convolutional neural network achieved state-of-the-art ImageNet classification, reducing top-5 error rates to 17.0% in 2010 and a winning 15.3% in 2012.

    Go to source

  • Very Deep Convolutional Networks for Large-Scale Image Recognition

    arXiv (Cornell University) · 2014 · 75,542 citations · Open access

    Increasing convolutional network depth to 16-19 weight layers, using small 3x3 filters, significantly improves accuracy in large-scale image recognition over prior configurations.

    Go to source

  • Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2016 · Q1 · SJR 4.00 · FWCI 1672.08 · 54,589 citations

    This paper introduces a Region Proposal Network (RPN) that integrates seamlessly with object detection networks, enabling real-time performance by generating region proposals at nearly no computational cost.

    Go to source

  • Going deeper with convolutions

    2015 · 46,948 citations

    A novel deep convolutional neural network architecture, Inception, has set a new state-of-the-art for image classification and detection on the ILSVRC14 challenge.

    Go to source

  • Densely Connected Convolutional Networks

    2017 · 44,970 citations

    DenseNets introduce a new convolutional network architecture where each layer connects to every other layer in a feed-forward manner, significantly increasing the number of direct connections compared to traditional networks.

    Go to source

  • Fully convolutional networks for semantic segmentation

    2015 · 37,045 citations

    This paper introduces fully convolutional networks (FCNs) that can perform semantic segmentation on images of any size, achieving state-of-the-art results with efficient inference.

    Go to source

  • Rich Feature Hierarchies for Accurate Object Detection and Semantic Segmentation

    2014 · 31,870 citations

    A new algorithm, R-CNN, significantly boosts object detection accuracy by over 30% on the PASCAL VOC 2012 dataset, achieving a 53.3% mAP.

    Go to source

  • Rethinking the Inception Architecture for Computer Vision

    2016 · 31,095 citations

    A new network design, Inception, efficiently scales up convolutional networks by factorizing convolutions and using aggressive regularization, achieving state-of-the-art results on the ILSVRC 2012 classification challenge.

    Go to source

  • Swin Transformer: Hierarchical Vision Transformer using Shifted Windows

    2021 IEEE/CVF International Conference on Computer Vision (ICCV) · 2021 · 30,585 citations

    The Swin Transformer introduces a hierarchical architecture with shifted windows, enabling efficient self-attention computation for general-purpose computer vision tasks.

    Go to source

  • Mask R-CNN

    2017 · 28,989 citations

    Mask R-CNN is a new framework that efficiently detects objects and generates high-quality segmentation masks for each instance in an image, extending Faster R-CNN with a mask prediction branch.

    Go to source

  • Feature Pyramid Networks for Object Detection

    2017 · 28,919 citations

    Feature Pyramid Networks (FPNs) efficiently construct a rich multi-scale feature representation from deep convolutional networks with marginal computational cost.

    Go to source

  • Squeeze-and-Excitation Networks

    2018 · 28,761 citations

    A novel 'Squeeze-and-Excitation' (SE) block adaptively recalibrates channel-wise feature responses by modeling interdependencies between channels, significantly improving deep learning model performance.

    Go to source

  • Fast R-CNN

    2015 · 28,041 citations

    Fast R-CNN is a novel object detection method that significantly improves training and testing speeds while increasing accuracy by efficiently classifying object proposals using deep convolutional networks.

    Go to source

  • Focal Loss for Dense Object Detection

    2017 · 26,050 citations

    A novel 'Focal Loss' function effectively addresses the extreme class imbalance in training dense object detectors, enabling them to surpass the accuracy of slower, two-stage detectors while maintaining speed.

    Go to source

  • MobileNetV2: Inverted Residuals and Linear Bottlenecks

    2018 · 25,715 citations

    MobileNetV2 introduces a novel mobile architecture featuring inverted residuals and linear bottlenecks, significantly improving state-of-the-art performance on mobile tasks like image classification, object detection, and semantic segmentation across various model sizes.

    Go to source

  • Learning Multiple Layers of Features from Tiny Images

    2024 · FWCI 670.28 · 25,502 citations

    Researchers developed a multi-layer generative model that learns meaningful visual features from millions of tiny web images, mimicking aspects of the human visual cortex.

    Go to source

  • DeepLab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs

    IEEE Transactions on Pattern Analysis and Machine Intelligence · 2017 · Q1 · SJR 4.00 · FWCI 632.10 · 22,241 citations

    DeepLab introduces 'atrous convolution' to control feature resolution and enlarge filter receptive fields in deep networks, enabling better semantic image segmentation.

    Go to source

  • Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks

    2017 · 22,010 citations

    This paper introduces a novel method for image-to-image translation that does not require paired training data, a common limitation in existing techniques.

    Go to source

  • Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization

    2017 · 21,873 citations

    Grad-CAM generates visual explanations for Convolutional Neural Network (CNN) decisions by highlighting important image regions, making complex models more transparent.

    Go to source

Recently added

  • Abstract Functional Language Logic: A Competitive Mixture of Experts Architecture for Paradox-Free Reasoning and Adaptive Intelligence

    Zenodo (CERN European Organization for Nuclear Research) · 2025 · FWCI 52.50 · 592 citations

    This paper introduces a Competitive Mixture of Experts (CMoE) architecture based on Functional Language Logic (FLL) that replaces probabilistic LLM reasoning with mathematical functional approximators for paradox-free, efficient deduction.

    Go to source

  • VM-UNet: Vision Mamba UNet for Medical Image Segmentation

    ACM Transactions on Multimedia Computing Communications and Applications · 2025 · Q1 · FWCI 229.92 · 493 citations

    VM-UNet, a novel U-shaped architecture for medical image segmentation, leverages State Space Models (SSMs) like Mamba to achieve linear computational complexity and effective long-range modeling, outperforming previous CNN and Transformer models.

    Go to source

  • V0.1IC

    Zenodo (CERN European Organization for Nuclear Research) · 2025 · 499 citations

    A new deep learning framework, Hamming Cube, jointly learns compact binary codes and continuous embeddings while preserving Hamming distance structure.

    Go to source

  • Learning Multiple Layers of Features from Tiny Images

    2024 · FWCI 670.28 · 25,502 citations

    Researchers developed a multi-layer generative model that learns meaningful visual features from millions of tiny web images, mimicking aspects of the human visual cortex.

    Go to source

  • A Multi-Modal Distributed Real-Time IoT System for Urban Traffic Control (Invited Paper)

    Leibniz-Zentrum für Informatik (Schloss Dagstuhl) · 2024 · FWCI 549.48 · 14,339 citations · Open access

    A novel distributed IoT system achieves 77% vehicle detection accuracy and 99.4% emergency vehicle detection accuracy using a two-stage detector and acoustic siren detection on edge devices.

    Go to source

Add this field to your daily feed

Pick your interests and new work in your area arrives every day, summarised. Full summaries live in the app.

Open the app

Other subfields in the same field

All fields