BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20240626T180034Z
LOCATION:3003\, 3rd Floor
DTSTART;TZID=America/Los_Angeles:20240625T103000
DTEND;TZID=America/Los_Angeles:20240625T104500
UID:dac_DAC 2024_sess120_RESEARCH151@linklings.com
SUMMARY:DEFA: Efficient Deformable Attention Acceleration via Pruning-Assi
 sted Grid-Sampling and Multi-Scale Parallel Processing
DESCRIPTION:Research Manuscript\n\nYansong Xu, Dongxu Lyu, Zhenyu Li, Yuzh
 ou Chen, Zilong Wang, Gang Wang, Zhican Wang, Haomin Li, and Guanghui He (
 Shanghai Jiao Tong University)\n\nMulti-scale deformable attention (MSDefo
 rmAttn) has emerged as a key mechanism in various vision tasks, demonstrat
 ing explicit superiority attributed to multi-scale grid-sampling. However,
  this newly introduced operator incurs irregular data access and enormous 
 memory requirement, leading to severe PE under-utilization. Meanwhile, exi
 sting approaches for attention acceleration cannot be directly applied to 
 MSDeformAttn due to lack of support for this distinct procedure. Therefore
 , we propose a dedicated algorithm-architecture co-design dubbed DEFA, the
  first-of-its-kind method for MSDeformAttn acceleration. At the algorithm 
 level, DEFA adopts frequency-weighted pruning and probability-aware prunin
 g for feature maps and sampling points respectively, alleviating the memor
 y footprint by over 80%. At the architecture level, it explores the multi-
 scale parallelism to boost the throughput significantly and further reduce
 s the memory access via fine-grained layer fusion and feature map reusing.
  Extensively evaluated on representative benchmarks, DEFA achieves 10.1-31
 .9× speedup and 20.3-37.7× energy efficiency boost compared to powerful GP
 U platforms. It also rivals the related accelerators by 2.2-3.7× energy ef
 ficiency improvement while providing pioneering support of MSDeformAttn.\n
 \nTopic: AI, Design\n\nKeyword: AI/ML Architecture Design\n\nSession Chair
 s: Andrey Ayupov (Google) and Subhendu ROY (Cadence Design Systems, Inc.)
END:VEVENT
END:VCALENDAR
