Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection

Li, Hongyang; Liu, Yu; Ouyang, Wanli; Wang, Xiaogang

File Download

There are no files associated with this item.

Links for fulltext

(May Require Subscription)

Publisher Website: 10.1007/s11263-018-1101-7
Scopus: eid_2-s2.0-85048763445
WOS: WOS:000459493200001
Find via

Supplementary

Citations:
- Scopus: 0
- Web of Science: 0
Appears in Collections:
- HKU Musketeers Foundation Institute of Data Science: Journal/Magazine Articles

Article: Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection

Title	Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection
Authors	Li, Hongyang Liu, Yu Ouyang, Wanli Wang, Xiaogang
Keywords	Computer vision Deep learning Object detection Region proposals
Issue Date	2019
Citation	International Journal of Computer Vision, 2019, v. 127, n. 3, p. 225-238 How to Cite? DOI: http://dx.doi.org/10.1007/s11263-018-1101-7
Abstract	In this paper, we propose a zoom-out-and-in network for generating object proposals. A key observation is that it is difficult to classify anchors of different sizes with the same set of features. Anchors of different sizes should be placed accordingly based on different depth within a network: smaller boxes on high-resolution layers with a smaller stride while larger boxes on low-resolution counterparts with a larger stride. Inspired by the conv/deconv structure, we fully leverage the low-level local details and high-level regional semantics from two feature map streams, which are complimentary to each other, to identify the objectness in an image. A map attention decision (MAD) unit is further proposed to aggressively search for neuron activations among two streams and attend the most contributive ones on the feature learning of the final loss. The unit serves as a decision-maker to adaptively activate maps along certain channels with the solely purpose of optimizing the overall training loss. One advantage of MAD is that the learned weights enforced on each feature channel is predicted on-the-fly based on the input context, which is more suitable than the fixed enforcement of a convolutional kernel. Experimental results on three datasets demonstrate the effectiveness of our proposed algorithm over other state-of-the-arts, in terms of average recall for region proposal and average precision for object detection.
Persistent Identifier	http://hdl.handle.net/10722/351382
ISSN	0920-5691 2023 Impact Factor: 11.6 2023 SCImago Journal Rankings: 6.668
ISI Accession Number ID	WOS:000459493200001

DC Field	Value	Language
dc.contributor.author	Li, Hongyang	-
dc.contributor.author	Liu, Yu	-
dc.contributor.author	Ouyang, Wanli	-
dc.contributor.author	Wang, Xiaogang	-
dc.date.accessioned	2024-11-20T03:55:57Z	-
dc.date.available	2024-11-20T03:55:57Z	-
dc.date.issued	2019	-
dc.identifier.citation	International Journal of Computer Vision, 2019, v. 127, n. 3, p. 225-238	-
dc.identifier.issn	0920-5691	-
dc.identifier.uri	http://hdl.handle.net/10722/351382	-
dc.description.abstract	In this paper, we propose a zoom-out-and-in network for generating object proposals. A key observation is that it is difficult to classify anchors of different sizes with the same set of features. Anchors of different sizes should be placed accordingly based on different depth within a network: smaller boxes on high-resolution layers with a smaller stride while larger boxes on low-resolution counterparts with a larger stride. Inspired by the conv/deconv structure, we fully leverage the low-level local details and high-level regional semantics from two feature map streams, which are complimentary to each other, to identify the objectness in an image. A map attention decision (MAD) unit is further proposed to aggressively search for neuron activations among two streams and attend the most contributive ones on the feature learning of the final loss. The unit serves as a decision-maker to adaptively activate maps along certain channels with the solely purpose of optimizing the overall training loss. One advantage of MAD is that the learned weights enforced on each feature channel is predicted on-the-fly based on the input context, which is more suitable than the fixed enforcement of a convolutional kernel. Experimental results on three datasets demonstrate the effectiveness of our proposed algorithm over other state-of-the-arts, in terms of average recall for region proposal and average precision for object detection.	-
dc.language	eng	-
dc.relation.ispartof	International Journal of Computer Vision	-
dc.subject	Computer vision	-
dc.subject	Deep learning	-
dc.subject	Object detection	-
dc.subject	Region proposals	-
dc.title	Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection	-
dc.type	Article	-
dc.description.nature	link_to_subscribed_fulltext	-
dc.identifier.doi	10.1007/s11263-018-1101-7	-
dc.identifier.scopus	eid_2-s2.0-85048763445	-
dc.identifier.volume	127	-
dc.identifier.issue	3	-
dc.identifier.spage	225	-
dc.identifier.epage	238	-
dc.identifier.eissn	1573-1405	-
dc.identifier.isi	WOS:000459493200001	-

File Download

Links for fulltext

(May Require Subscription)

Supplementary

Article: Zoom Out-and-In Network with Map Attention Decision for Region Proposal and Object Detection

Export via OAI-PMH Interface in XML Formats

OR

Export to Other Non-XML Formats