18
I. Athanasiadis et al.
perception of the surrounding environment. The Deep Learning (DL) era brought
great advancements in numerous of Computer vision domains, mainly because
DL-based methods are capable of grasping complex relations and handling huge
amount of data. Specifically, deep Convolutional Neural Networks (CNNs) have
been utilised for the task of object detection. These approaches fall into two
categories, namely the two-stage and the one-stage methods. Modern two-stage
object detection methods such as Faster R-CNN [9] and Mask R-CNN [3] make
use of a trainable network, called Regional Proposal Network (RPN), to propose
regions which potentially enclose ground truth objects in. On the other hand,
in the one-stage methods, the regions are generated and classified in a single
forward pass. The YOLO [8] and the SSD [7] algorithms are the most representative one-stage object detection approaches. In [11] the High Possible Regions
Proposal Network is introduced in which a feature map with empowered edge
information is formed and passed as an additional feature map in the RPN for
more accurate region proposing. Object detection has evolved significantly the
latest years, nevertheless, the accuracy in detecting small objects is still limited.
To address such limitations, in [10] the use of context information is adapted
focusing on detecting small objects. A Generative adversarial networks (GAN)
based approach is presented in [4], which generates super resolved representations of small objects, similar to the bigger ones, where the object detection
algorithms perform better at detecting. The motivation behind this work is to
effectively boost the performance of current two-stage object detection methods in detecting small objects by implementing a more sophisticated anchoring
technique. A suitable baseline method is chosen which we gradually enhance by
combining cascade architecture and a proposed anchoring mechanism targeted
at small-sized objects. Finally in this work, we showcase the ability to amplify
DL-based object detection approaches by complementing them with additional
handcrafted features, targeted explicitly at compensating for the poor visual
representation of the small objects.
2 Methods
State of the art two-stage approaches consist of two discrete modules responsible for region proposing and classifying respectively. At the first stage, a set of
candidate regions of predefined shape and size, called anchors, is uniformly generated across the image. Thereafter, each anchor is validated on its probability
of containing a ground truth object by the RPN. The most confident, in terms
of objectness, anchors constitute the proposed regions which are then passed to
the second stage for further classification.
Baseline: As a baseline approach, Mask R-CNN [3] was chosen for both its
state of the art performance and its efficiency in cases where heavy overlapping
between the relevant objects occurs. By utilising the Feature Pyramid Network
(FPN) approach of [5], Mask R-CNN becomes more appealing to the detection
of small objects.
I. Athanasiadis et al.
perception of the surrounding environment. The Deep Learning (DL) era brought
great advancements in numerous of Computer vision domains, mainly because
DL-based methods are capable of grasping complex relations and handling huge
amount of data. Specifically, deep Convolutional Neural Networks (CNNs) have
been utilised for the task of object detection. These approaches fall into two
categories, namely the two-stage and the one-stage methods. Modern two-stage
object detection methods such as Faster R-CNN [9] and Mask R-CNN [3] make
use of a trainable network, called Regional Proposal Network (RPN), to propose
regions which potentially enclose ground truth objects in. On the other hand,
in the one-stage methods, the regions are generated and classified in a single
forward pass. The YOLO [8] and the SSD [7] algorithms are the most representative one-stage object detection approaches. In [11] the High Possible Regions
Proposal Network is introduced in which a feature map with empowered edge
information is formed and passed as an additional feature map in the RPN for
more accurate region proposing. Object detection has evolved significantly the
latest years, nevertheless, the accuracy in detecting small objects is still limited.
To address such limitations, in [10] the use of context information is adapted
focusing on detecting small objects. A Generative adversarial networks (GAN)
based approach is presented in [4], which generates super resolved representations of small objects, similar to the bigger ones, where the object detection
algorithms perform better at detecting. The motivation behind this work is to
effectively boost the performance of current two-stage object detection methods in detecting small objects by implementing a more sophisticated anchoring
technique. A suitable baseline method is chosen which we gradually enhance by
combining cascade architecture and a proposed anchoring mechanism targeted
at small-sized objects. Finally in this work, we showcase the ability to amplify
DL-based object detection approaches by complementing them with additional
handcrafted features, targeted explicitly at compensating for the poor visual
representation of the small objects.
2 Methods
State of the art two-stage approaches consist of two discrete modules responsible for region proposing and classifying respectively. At the first stage, a set of
candidate regions of predefined shape and size, called anchors, is uniformly generated across the image. Thereafter, each anchor is validated on its probability
of containing a ground truth object by the RPN. The most confident, in terms
of objectness, anchors constitute the proposed regions which are then passed to
the second stage for further classification.
Baseline: As a baseline approach, Mask R-CNN [3] was chosen for both its
state of the art performance and its efficiency in cases where heavy overlapping
between the relevant objects occurs. By utilising the Feature Pyramid Network
(FPN) approach of [5], Mask R-CNN becomes more appealing to the detection
of small objects.
