File Download
There are no files associated with this item.
Links for fulltext
(May Require Subscription)
- Publisher Website: 10.1109/TPAMI.2008.129
- Scopus: eid_2-s2.0-54749131961
- PMID: 18787246
- WOS: WOS:000259110000011
- Find via
Supplementary
- Citations:
- Appears in Collections:
Article: Video event recognition using kernel methods with multilevel temporal alignment
Title | Video event recognition using kernel methods with multilevel temporal alignment |
---|---|
Authors | |
Keywords | Concept ontology Earth mover's distance Event recognition News video Temporally aligned pyramid matching Video indexing |
Issue Date | 2008 |
Citation | IEEE Transactions on Pattern Analysis and Machine Intelligence, 2008, v. 30, n. 11, p. 1985-1997 How to Cite? |
Abstract | In this work, we systematically study the problem of event recognition in unconstrained news video sequences. We adopt the discriminative kernel-based method for which video clip similarity plays an important role. First, we represent a video clip as a bag of orderless descriptors extracted from all of the constituent frames and apply the Earth Mover's Distance (EMD) to integrate similarities among frames from two clips. Observing that a video clip is usually comprised of multiple subclips corresponding to event evolution over time, we further build a multi-level temporal pyramid. At each pyramid level, we integrate the information from different subclips with Integer-valueconstrained EMD to explicitly align the subclips. By fusing the information from the different pyramid levels, we develop Temporally Aligned Pyramid Matching (TAPM) for measuring video similarity. We conduct comprehensive experiments on the Trecvid 2005 corpus, which contains more than 6,800 clips. Our experiments demonstrate that 1) the TAPM multi-level method clearly outperforms single-level EMD, and 2) single-level EMD outperforms keyframe and multi-frame based detection methods by a large margin. In addition, we conduct in-depth investigation of various aspects of the proposed techniques, such as weight selection in single-level EMD, sensitivity to temporal clustering, the effect of temporal alignment, and possible approaches for speedup. © 2008 IEEE. |
Persistent Identifier | http://hdl.handle.net/10722/321355 |
ISSN | 2023 Impact Factor: 20.8 2023 SCImago Journal Rankings: 6.158 |
ISI Accession Number ID |
DC Field | Value | Language |
---|---|---|
dc.contributor.author | Xu, Dong | - |
dc.contributor.author | Chang, Shih Fu | - |
dc.date.accessioned | 2022-11-03T02:18:21Z | - |
dc.date.available | 2022-11-03T02:18:21Z | - |
dc.date.issued | 2008 | - |
dc.identifier.citation | IEEE Transactions on Pattern Analysis and Machine Intelligence, 2008, v. 30, n. 11, p. 1985-1997 | - |
dc.identifier.issn | 0162-8828 | - |
dc.identifier.uri | http://hdl.handle.net/10722/321355 | - |
dc.description.abstract | In this work, we systematically study the problem of event recognition in unconstrained news video sequences. We adopt the discriminative kernel-based method for which video clip similarity plays an important role. First, we represent a video clip as a bag of orderless descriptors extracted from all of the constituent frames and apply the Earth Mover's Distance (EMD) to integrate similarities among frames from two clips. Observing that a video clip is usually comprised of multiple subclips corresponding to event evolution over time, we further build a multi-level temporal pyramid. At each pyramid level, we integrate the information from different subclips with Integer-valueconstrained EMD to explicitly align the subclips. By fusing the information from the different pyramid levels, we develop Temporally Aligned Pyramid Matching (TAPM) for measuring video similarity. We conduct comprehensive experiments on the Trecvid 2005 corpus, which contains more than 6,800 clips. Our experiments demonstrate that 1) the TAPM multi-level method clearly outperforms single-level EMD, and 2) single-level EMD outperforms keyframe and multi-frame based detection methods by a large margin. In addition, we conduct in-depth investigation of various aspects of the proposed techniques, such as weight selection in single-level EMD, sensitivity to temporal clustering, the effect of temporal alignment, and possible approaches for speedup. © 2008 IEEE. | - |
dc.language | eng | - |
dc.relation.ispartof | IEEE Transactions on Pattern Analysis and Machine Intelligence | - |
dc.subject | Concept ontology | - |
dc.subject | Earth mover's distance | - |
dc.subject | Event recognition | - |
dc.subject | News video | - |
dc.subject | Temporally aligned pyramid matching | - |
dc.subject | Video indexing | - |
dc.title | Video event recognition using kernel methods with multilevel temporal alignment | - |
dc.type | Article | - |
dc.description.nature | link_to_subscribed_fulltext | - |
dc.identifier.doi | 10.1109/TPAMI.2008.129 | - |
dc.identifier.pmid | 18787246 | - |
dc.identifier.scopus | eid_2-s2.0-54749131961 | - |
dc.identifier.volume | 30 | - |
dc.identifier.issue | 11 | - |
dc.identifier.spage | 1985 | - |
dc.identifier.epage | 1997 | - |
dc.identifier.isi | WOS:000259110000011 | - |