Temporal Action Detection: From Full Supervision to Zero-Shot Future

Deep learning-based action detection in untrimmed videos: A survey

2022-01-01
Elahe Vahdani, Yingli Tian
Summary
Problem
Method
Results
Takeaways
Abstract

This survey provides a comprehensive overview of deep learning-based temporal action detection (TAD) in untrimmed videos. It covers diverse supervision levels—from fully-supervised to weakly, semi, self, and unsupervised settings—highlighting SOTA methods like VSGN and RTD-Net that achieve high mAP on THUMOS14 and ActivityNet.

Executive Summary

TL;DR: This survey systematically dissects the evolution of Temporal Action Detection (TAD), the task of identifying "what" happened and "when" it started/ended in long, messy videos. It bridges the gap between traditional heavy-supervision models and modern, annotation-efficient learning paradigms like self-supervised and zero-shot detection.

Background Positioning: This work serves as a foundational "map" of the TAD landscape. It moves beyond simple action recognition to address the spatial and temporal nuances of untrimmed video analysis, positioning itself as a critical reference for researchers in sports analytics, surveillance, and robotics.

The Core Conflict: Why TAD is Hard

Most AI researchers are familiar with Action Recognition—classifying a 5-second clip of someone "jumping." However, real-world video is untrimmed.

  1. The "Needle in a Haystack" Problem: Actions of interest often cover only a small fraction (e.g., 30%) of the video. The rest is "background" noise.
  2. Flexible Duration: A "long jump" might last 5 seconds, while "cooking a meal" lasts 20 minutes. Rigid windows fail here.
  3. Annotation Fatigue: Manually marking every start and end frame in a 10-hour surveillance feed is a nightmare.

Methodology: The Technical Arsenal

1. The Architectural Split: Anchor-based vs. Anchor-free

  • Anchor-based (Top-down): Similar to early object detection (Faster R-CNN), these methods place predefined temporal "boxes" across the video.
  • Anchor-free (Bottom-up): These methods predict "boundary scores" (start/end probabilities) for every frame. Insight: This allows for precise, sample-level localization that isn't restricted by predefined lengths.

2. Modeling the "Long View"

Since individual video snippets lack context, the survey highlights three SOTA modeling techniques:

  • Graph Convolutional Networks (GCNs): Treating video segments as nodes to capture complex, non-linear relations between sub-actions.
  • Transformers: Utilizing self-attention to relate a frame at the beginning of a video to one at the very end.
  • Feature Pyramids: Using U-shaped architectures to detect both "micro-actions" (flicking a switch) and "macro-activities" (cleaning a room).

TAD Task Overview Figure 1: Comparison of temporal localization (when) vs. spatio-temporal localization (where and when).

The Shift to "Limited Supervision"

A major contribution of this survey is the detailed breakdown of how we train models with labels they've barely seen:

  • Weakly-Supervised: Training using only a video-level tag (e.g., "this video contains a soccer goal") without saying where the goal is.
  • Attention Mechanisms: Models learn to "attend" only to discriminative frames. Class-agnostic attention is specifically praised for its ability to filter out background noise regardless of the specific action type.

Spatio-temporal Activity Detection Figure 2: Spatio-temporal detection tracks the actor in both time and 2D space.

Experimental Performance Analysis

The survey compares state-of-the-art (SOTA) results across the THUMOS14 and ActivityNet benchmarks.

  • SOTA Leaders: Methods like AFSD and VSGN dominate the fully-supervised leaderboard by balancing precise boundary refinement with multi-scale feature aggregation.
  • The Efficiency Gap: While fully supervised models reach ~56.9% mAP on THUMOS14, weakly supervised models are catching up, with recent entries like D2-Net reaching ~36.0%.

Critical Analysis & The Road Ahead

Future Trends:

  1. Zero-Shot (ZSTAD): Detecting actions the model has never seen during training by using semantic word embeddings.
  2. Domain Transfer: Leveraging the abundance of "trimmed" YouTube clips to improve "untrimmed" detection.
  3. Real-world Deployment: Moving from "offline" batch processing to "online" detection for autonomous vehicles that must react in milliseconds.

Conclusion:

The field is moving away from "brute-force" annotation. The future of video understanding lies in structured self-supervision—models that understand the physics and logic of human movement without needing a human to label every frame.

Find Similar Papers

Try Our Examples

  • Search for recent papers that utilize Zero-Shot Temporal Activity Detection (ZSTAD) to identify unseen action classes in untrimmed videos.
  • Which paper first proposed the Boundary-Matching Network (BMN) for temporal proposal generation, and how have subsequent works like BSN++ improved its efficiency?
  • Looking for studies that apply Temporal Action Detection techniques specifically to anomaly detection within the context of autonomous vehicle sensor streams.
Contents
Temporal Action Detection: From Full Supervision to Zero-Shot Future
1. Executive Summary
2. The Core Conflict: Why TAD is Hard
3. Methodology: The Technical Arsenal
3.1. 1. The Architectural Split: Anchor-based vs. Anchor-free
3.2. 2. Modeling the "Long View"
4. The Shift to "Limited Supervision"
5. Experimental Performance Analysis
6. Critical Analysis & The Road Ahead
6.1. Future Trends:
6.2. Conclusion: