BEGIN:VCALENDAR
VERSION:2.0
PRODID:Linklings LLC
BEGIN:VTIMEZONE
TZID:America/Los_Angeles
X-LIC-LOCATION:America/Los_Angeles
BEGIN:DAYLIGHT
TZOFFSETFROM:-0800
TZOFFSETTO:-0700
TZNAME:PDT
DTSTART:19700308T020000
RRULE:FREQ=YEARLY;BYMONTH=3;BYDAY=2SU
END:DAYLIGHT
BEGIN:STANDARD
TZOFFSETFROM:-0700
TZOFFSETTO:-0800
TZNAME:PST
DTSTART:19701101T020000
RRULE:FREQ=YEARLY;BYMONTH=11;BYDAY=1SU
END:STANDARD
END:VTIMEZONE
BEGIN:VEVENT
DTSTAMP:20240626T180034Z
LOCATION:Level 2 Lobby
DTSTART;TZID=America/Los_Angeles:20240626T180000
DTEND;TZID=America/Los_Angeles:20240626T190000
UID:dac_DAC 2024_sess237_RESEARCH1434@linklings.com
SUMMARY:AdaP-CIM: Compute-in-Memory Based Neural Network Accelerator using
  Adaptive Posit and Speculative Alignment
DESCRIPTION:Work-in-Progress Poster\n\nJingyu He, Fengbin Tu, Tim Cheng, a
 nd Chi Ying Tsui (Hong Kong University of Science and Technology (HKUST))\
 n\nLarge neural networks, especially transformer-based models, present two
  critical challenges that exacerbate the memory wall issue in AI accelerat
 or designs. First, the increased dynamic range of the weights requires hig
 her precision quantization formats, leading to higher memory capacity requ
 irements. Second, the exponential growth in model parameters incurs more d
 ata movement, leading to increased latency and power consumption. In this 
 study, we propose two novel approaches to address these problems. First, b
 ased on Posit, we introduce a new format called adaptive Posit (AdaP), whi
 ch dynamically extends the dynamic range of its representation at run time
  with minimal hardware overhead. AdaP, utilizing two exponent encoding sch
 emes, accommodates the data distribution with lower quantization error com
 pared to regular Posit. Second, we propose to use compute-in-memory (CIM) 
 architecture to implement AdaP multiply-and-accumulate (MAC) computation t
 o reduce weight data movement. Traditional CIM proposed for floating-point
 -alike MAC computation uses a comparator tree (CT) to compute the maximum 
 exponent, enabling the CIM to focus on integer MAC. However, the CT-based 
 design has poor scalability as the number of inputs increases. To address 
 this, we propose a speculative input alignment design that significantly r
 educes the delay, area, and power consumption for the max exponent computa
 tion. Software evaluations show that 8-bit AdaP incurs a negligible 0.25% 
 F1 score reduction on the XLM language identification benchmark compared t
 o the full-precision baseline. Hardware synthesis and simulation results f
 urther illustrate that our approach achieves 55% energy efficiency and 2.4
 x area efficiency improvement compared to the state-of-the-art posit proce
 ssing element.\n\nTopic: AI, Autonomous Systems, Cloud, Design, EDA, Embed
 ded Systems, IP, Security
END:VEVENT
END:VCALENDAR
