Intel Corporation
Abstraction layers for scalable distributed machine learning

Last updated:

Abstract:

One embodiment provides for a method of transmitting data between multiple compute nodes of a distributed compute system, the method comprising creating a global view of communication operations to be performed between the multiple compute nodes of the distributed compute system, the global view created using information specific to a machine learning model associated with the distributed compute system; using the global view to determine a communication cost of the communication operations; and automatically determining a number of network endpoints for use in transmitting the data between the multiple compute nodes of the distributed compute system.

Status:
Grant
Type:

Utility

Filling date:

10 Apr 2017

Issue date:

17 Aug 2021