Alibaba Group Holding Limited
DATA LAYOUT OPTIMIZATION ON PROCESSING IN MEMORY ARCHITECTURE FOR EXECUTING NEURAL NETWORK MODEL
Last updated:
Abstract:
The present disclosure relates to a method for scheduling a computation graph on a processing in memory (PIM) enabled device comprising a memory block assembly. The method comprises allocating a first node of the computation graph on a first memory block of a first array of memory blocks in the memory block assembly and allocating a second node of the computation graph on a second memory block of a second array of memory blocks in the memory block assembly, wherein output data of the first node is used for executing the second node. The memory block assembly can be configured to support data transfer from the first memory block to the second memory block via an internal data coupling in the memory block assembly.
Utility
17 Jan 2020
22 Jul 2021