Patent · US11029870B2 · B2 · US
Technologies for dividing work across accelerator devices
- (11) Publication number
- US11029870B2
- (21) Application number
- 15/721,829
- (22) Filing date
- 2017-09-30
- (30) Priority date
- 2016-11-29
- (43) Publication date
- 2021-06-08
- (45) Date of grant
- 2021-06-08
- (51) IPC
- G06F 12/02; G06F 9/38; G06T 1/20; G06T 1/60; H04L 47/20; G06F 11/07; G06F 11/30; G06F 11/34; G06F 12/06; G06F 13/16; G06F 15/80; G06F 16/174; G06F 21/57; G06F 21/62; G06F 21/73; G06F 21/76; G06F 3/06; G06F 7/06; G06F 8/65; G06F 8/654; G06F 8/656; G06F 8/658; G06F 9/4401; G06F 9/48; G06F 9/50; G06F 9/54; G06T 9/00; H01R 13/453; H01R 13/631; H03K 19/173; H03M 7/30; H03M 7/40; H03M 7/42; H04L 12/28; H04L 12/46; H04L 9/08; H05K 7/14
- (52) CPC
- G06F Electric digital data processing: 9/5005, 11/0709, 11/0751, 11/079, 11/1453, 11/3006, 11/3034, 11/3055, 11/3079, 11/3409, 12/023, 12/0284, 12/0692, 13/1652, 13/4022, 13/4027, 15/161, 15/80, 16/1744, 21/44, 21/57, 21/6218, 21/70, 21/73, 21/76, 2212/401, 2212/402, 2221/2107, 3/0604, 3/0608, 3/0611, 3/0613, 3/0617, 3/0641, 3/0647, 3/065, 3/0653, 3/067, 7/06, 8/65, 8/654, 8/656, 8/658, 9/3851, 9/3891, 9/4401, 9/4843, 9/4881, 9/5038, 9/5044, 9/505, 9/5083, 9/544
- G06T Image data processing or generation, in general: 1/20, 1/60, 9/005
- H01R Electrically-conductive connections; structural associations of a plurality of mutually-insulated electrical connecting elements; coupling devices; current collectors: 13/453, 13/4536, 13/4538, 13/631
- H03K Pulse technique: 19/1731
- H03M Coding; decoding; code conversion in general: 7/3084, 7/40, 7/42, 7/60, 7/6011, 7/6017, 7/6029
- H04L Transmission of digital information, e.g. telegraphic communication: 12/2881, 12/4633, 41/044, 41/046, 41/0816, 41/0853, 41/0895, 41/0896, 41/12, 41/142, 41/40, 43/04, 43/06, 43/08, 43/0894, 47/20, 47/2441, 47/78, 47/83, 49/104, 61/2007, 63/1425, 67/10, 67/1014, 67/327, 67/36, 9/0822
- H05K Printed circuits; casings or constructional details of electric apparatus; manufacture of assemblages of electrical components: 7/1452, 7/1487, 7/1492
- (73) Assignee
- Intel Corp
- (72) Inventors
- Susanne M. Balle; Francesc Guim Bernat; Slawomir PUTYRSKI; Joe Grecco; Henry Mitchel; Evan Custodio; Rahul Khanna; Sujoy Sen
- (54) Title
- Technologies for dividing work across accelerator devices
- (57) Abstract
Technologies for dividing work across one or more accelerator devices include a compute device. The compute device is to determine a configuration of each of multiple accelerator devices of the compute device, receive a job to be accelerated from a requester device remote from the compute device, and divide the job into multiple tasks for a parallelization of the multiple tasks among the one or more accelerator devices, as a function of a job analysis of the job and the configuration of each accelerator device. The compute engine is further to schedule the tasks to the one or more accelerator devices based on the job analysis and execute the tasks on the one or more accelerator devices for the parallelization of the multiple tasks to obtain an output of the job.
- Full text
- View on Google Patents
Claims (28)
- A compute device comprising: one or more accelerator devices; and a compute engine to: determine a configuration of at least one accelerator device of the compute device, wherein the configuration is indicative of parallel execution capabilities of the at least one accelerator device, and wherein to determine the configuration includes to determine for the at least one accelerator device whether the at least one accelerator device is capable of accessing a shared data set and whether the at least one accelerator device is capable of accessing a shared memory; receive, from a requester device, a job to be accelerated; divide the job into multiple tasks for a parallel execution of the multiple tasks among the one or more accelerator devices as a function of a job analysis of the job and the configuration of the at least one accelerator device indicative of whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; schedule the multiple tasks to the one or more accelerator devices based on the job analysis; execute the multiple tasks on the one or more accelerator devices for parallel execution of the multiple tasks based on whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; and combine task outputs from the one or more accelerator devices that executed the multiple tasks to obtain an output of the job.
- The compute device of claim 1, wherein the compute engine includes a micro-orchestrator logic unit.
- The compute device of claim 2, wherein to determine the configuration of the at least one accelerator device comprises to determine, by the micro-orchestrator logic unit of the compute device, the configuration of the at least one accelerator device.
- The compute device of claim 1, wherein the compute engine is further to determine an availability of one or more of the accelerator devices.
- The compute device of claim 4, wherein to determine the availability of one or more of the accelerator devices comprises to determine one or more available kernels of each accelerator device.
- The compute device of claim 1, wherein to schedule the multiple tasks to the one or more accelerator devices comprises to schedule parallel execution of the tasks.
- The compute device of claim 1, wherein to schedule the multiple tasks to the one or more accelerator devices comprises to confirm that the one or more accelerator devices are available to simultaneously execute multiple tasks that share data.
- The compute device of claim 1, wherein to execute the multiple tasks comprises to concurrently execute two or more of the tasks on two or more of the accelerator devices of the compute device with a high speed serial interface (HSSI).
- The compute device of claim 1, wherein the compute engine is further to: determine whether an authorization is required from an orchestrator to execute the tasks on the compute device; transmit, in response to a determination that the authorization is required, the job analysis to the orchestrator; receive an authorization from the orchestrator; and receive the tasks to be accelerated.
- The compute device of claim 9, wherein the compute device comprises a first compute device, and wherein to execute the multiple tasks comprises to: determine one or more accelerator devices of a second compute device that is remote to the first compute device to concurrently execute the multiple tasks that share data; and execute one or more of the tasks on the one or more accelerator devices of the first compute device as one or more other tasks of the job are concurrently executed with one or more accelerator devices of the second compute device using a shared virtual memory.
- The compute device of claim 10, wherein to determine the one or more accelerator devices of the second compute device comprises to receive information regarding one or more accelerator devices of the second compute device from an orchestrator server.
- The compute device of claim 9, wherein to execute the tasks comprises to execute the tasks simultaneously on the one or more accelerator devices of the compute device with a high speed serial interface (HSSI).
- One or more machine-readable storage media comprising a plurality of instructions stored thereon that, when executed by at least one compute device cause the at least one compute device to: determine a configuration of at least one accelerator device of the compute device, wherein the configuration is indicative of parallel execution capabilities of at least one accelerator device, and wherein to determine the configuration includes to determine, for at least one accelerator device, whether the at least one accelerator device is capable of accessing a shared data set and whether the at least one accelerator device is capable of accessing a shared memory; receive, from a requester device, a job to be accelerated; divide the job into multiple tasks for a parallel execution of the multiple tasks among the one or more accelerator devices as a function of a job analysis of the job and the configuration of at least one accelerator device indicative of whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; schedule the multiple tasks to the one or more accelerator devices based on the job analysis; cause execution of the multiple tasks on the one or more accelerator devices for the parallel execution of the multiple tasks based on whether the at least one accelerator device is capable of accessing the shared data set and whether the one or more accelerator device is capable of accessing the shared memory; and combine task outputs from the one or more accelerator devices that executed the multiple tasks to obtain an output of the job.
- The one or more machine-readable storage media of claim 13, wherein to determine the configuration of at least one accelerator device comprises to determine, by a micro-orchestrator logic unit of the compute device, the configuration of the at least one accelerator device.
- The one or more machine-readable storage media of claim 13, further comprising a plurality of instructions stored thereon that, in response to being executed, cause the compute device to determine an availability of one or more of the accelerator devices, and the at least one of the accelerator devices is not a general purpose processor.
- The one or more machine-readable storage media of claim 15, wherein to determine the availability of one or more of the accelerator devices comprises to determine one or more available kernels of the at least one accelerator device.
- The one or more machine-readable storage media of claim 13, wherein to schedule the multiple tasks to the one or more accelerator devices comprises to schedule parallel execution of the tasks.
- The one or more machine-readable storage media of claim 13, wherein to schedule the multiple tasks to the one or more accelerator devices comprises to confirm that the one or more accelerator devices are available to simultaneously execute multiple tasks that share data.
- The one or more machine-readable storage media of claim 13, wherein execution of the multiple tasks comprises concurrently execute two or more of the tasks on two or more of the accelerator devices of the compute device with a high speed serial interface (HSSI).
- The one or more machine-readable storage media of claim 13, further comprising a plurality of instructions stored thereon that, in response to being executed, cause the compute device to: determine whether an authorization is required from an orchestrator to execute the tasks on the compute device; transmit, in response to a determination that the authorization is required, the job analysis to the orchestrator; receive an authorization from the orchestrator; and receive the tasks to be accelerated.
- The one or more machine-readable storage media of claim 20, wherein the compute device comprises a first compute device, and wherein to execute the tasks comprises to: determine one or more accelerator devices of a second compute device that is remote to the first compute device to concurrently execute the tasks that share data; and execute one or more of the tasks on the one or more accelerator devices of the first compute device as one or more other tasks of the job are concurrently executed with one or more accelerator devices of the second compute device using a shared virtual memory.
- The one or more machine-readable storage media of claim 21, wherein to determine one or more accelerator devices of the second compute device comprises to receive information regarding one or more accelerator devices of the second compute device from a server.
- The one or more machine-readable storage media of claim 20, wherein to execute the tasks comprises to execute the tasks simultaneously on the one or more accelerator devices of the compute device with a high speed serial interface (HSSI).
- A compute device comprising: circuitry for determining a configuration of at least one of one or more accelerator devices of the compute device, wherein the configuration is indicative of parallel execution capabilities of at least one accelerator device, and wherein determining the configuration includes determining for the at least one accelerator device whether the accelerator device is capable of accessing a shared data set and whether the at least one accelerator device is capable of accessing a shared memory; circuitry for receiving, from a requester device remote from the compute device, a job to be accelerated; means for dividing the job into multiple tasks for a parallelization of the multiple tasks among the one or more accelerator devices, as a function of a job analysis of the job and the configuration of the at least one accelerator device indicative of whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; means for scheduling the multiple tasks to the one or more accelerator devices based on the job analysis; circuitry for executing the multiple tasks on the one or more accelerator devices for the parallelization of the multiple tasks based on whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; and means for combining task outputs from the at least one accelerator devices that executed the multiple tasks to obtain an output of the job.
- A method comprising: determining, by a compute device, a configuration of at least one of one or more accelerator devices of the compute device, wherein the configuration is indicative of parallel execution capabilities of at least one accelerator device, and wherein determining the configuration includes determining for at least one accelerator device whether the accelerator device is capable of accessing a shared data set and whether the accelerator device is capable of accessing a shared memory; receiving, from a requester device and by the compute device, a job to be accelerated; dividing, by the compute device, the job into multiple tasks for a parallelization of the multiple tasks among the one or more accelerator devices, as a function of a job analysis of the job and the configuration of at least one accelerator device indicative of whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; scheduling, by the compute device, the multiple tasks to the one or more accelerator devices based on the job analysis; causing execution, by the compute device, of the multiple tasks on the one or more accelerator devices for the parallelization of the multiple tasks based on whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; and combining, by the compute device, task outputs from the one or more accelerator devices that executed the tasks to obtain an output of the job.
- The method of claim 25, wherein determining the configuration of at least one accelerator device comprises determining, by a micro-orchestrator logic unit of the compute device, the configuration of at least one accelerator device.
- The method of claim 25, further comprising: determining, by the compute device, whether an authorization is required from an orchestrator to execute the tasks on the compute device; transmitting, by the compute device and in response to a determination that the authorization is required, the job analysis to the orchestrator; receiving, by the compute device, an authorization from the orchestrator; and receiving, by the compute device, the tasks to be accelerated.
- The method of claim 27, wherein the compute device comprises a first compute device, and wherein executing the tasks comprises: determining, by the first compute device, one or more accelerator devices of a second compute device that is remote to the first compute device to concurrently execute the multiple tasks that share data; and executing one or more of the tasks on the one or more accelerator devices of the first compute device as one or more other tasks of the job are concurrently executed with one or more accelerator devices of the second compute device using a shared virtual memory.
Description
Demand for accelerator devices has continued to increase because the accelerator devices are becoming more important as they may be used in various technological areas, such as machine learning and genomics. Typical architectures for accelerator devices, such as field programmable gate arrays (FPGAs), cryptography accelerators, graphics accelerators, and/or compression accelerators (referred to herein as “accelerator devices,” “accelerators,” or “accelerator resources”) capable of accelerating the execution of a set of operations in a workload (e.g., processes, applications, services, etc.) may allow static assignment of specified amounts of shared resources of the accelerator device (e.g., high bandwidth memory, data storage, etc.) among different portions of the logic (e.g., circuitry) of the accelerator device. Typically, the workload is allocated with the required processor(s), memory, and accelerator device(s) for the duration of the workload. The workload may use its allocated accelerator device at any point of time; however, in many cases, the accelerator devices will remain idle leading to wastage of resources.
The concepts described herein are illustrated by way of example and not by way of limitation in the accompanying figures. For simplicity and clarity of illustration, elements illustrated in the figures are not necessarily drawn to scale. Where considered appropriate, reference labels have been repeated among the figures to indicate corresponding or analogous elements.
Citations (16)
- US20100191823A1
- US20120054770A1
- US20130179485A1
- US20130232495A1
- US9026765B1
- US20150007182A1
- US20170046179A1
- US20170116004A1
- US20170317945A1
- US10034407B2
- US10045098B2
- US10085358B2
- US20180077235A1
- US20180150298A1
- US20180150330A1
- US20190065281A1
Record as JSON
{
"publication_number": "US11029870B2",
"country": "US",
"kind": "B2",
"title": "Technologies for dividing work across accelerator devices",
"abstract": "Technologies for dividing work across one or more accelerator devices include a compute device. The compute device is to determine a configuration of each of multiple accelerator devices of the compute device, receive a job to be accelerated from a requester device remote from the compute device, and divide the job into multiple tasks for a parallelization of the multiple tasks among the one or more accelerator devices, as a function of a job analysis of the job and the configuration of each accelerator device. The compute engine is further to schedule the tasks to the one or more accelerator devices based on the job analysis and execute the tasks on the one or more accelerator devices for the parallelization of the multiple tasks to obtain an output of the job.",
"claims": [
"1. A compute device comprising: one or more accelerator devices; and a compute engine to: determine a configuration of at least one accelerator device of the compute device, wherein the configuration is indicative of parallel execution capabilities of the at least one accelerator device, and wherein to determine the configuration includes to determine for the at least one accelerator device whether the at least one accelerator device is capable of accessing a shared data set and whether the at least one accelerator device is capable of accessing a shared memory; receive, from a requester device, a job to be accelerated; divide the job into multiple tasks for a parallel execution of the multiple tasks among the one or more accelerator devices as a function of a job analysis of the job and the configuration of the at least one accelerator device indicative of whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; schedule the multiple tasks to the one or more accelerator devices based on the job analysis; execute the multiple tasks on the one or more accelerator devices for parallel execution of the multiple tasks based on whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; and combine task outputs from the one or more accelerator devices that executed the multiple tasks to obtain an output of the job.",
"2. The compute device of claim 1, wherein the compute engine includes a micro-orchestrator logic unit.",
"3. The compute device of claim 2, wherein to determine the configuration of the at least one accelerator device comprises to determine, by the micro-orchestrator logic unit of the compute device, the configuration of the at least one accelerator device.",
"4. The compute device of claim 1, wherein the compute engine is further to determine an availability of one or more of the accelerator devices.",
"5. The compute device of claim 4, wherein to determine the availability of one or more of the accelerator devices comprises to determine one or more available kernels of each accelerator device.",
"6. The compute device of claim 1, wherein to schedule the multiple tasks to the one or more accelerator devices comprises to schedule parallel execution of the tasks.",
"7. The compute device of claim 1, wherein to schedule the multiple tasks to the one or more accelerator devices comprises to confirm that the one or more accelerator devices are available to simultaneously execute multiple tasks that share data.",
"8. The compute device of claim 1, wherein to execute the multiple tasks comprises to concurrently execute two or more of the tasks on two or more of the accelerator devices of the compute device with a high speed serial interface (HSSI).",
"9. The compute device of claim 1, wherein the compute engine is further to: determine whether an authorization is required from an orchestrator to execute the tasks on the compute device; transmit, in response to a determination that the authorization is required, the job analysis to the orchestrator; receive an authorization from the orchestrator; and receive the tasks to be accelerated.",
"10. The compute device of claim 9, wherein the compute device comprises a first compute device, and wherein to execute the multiple tasks comprises to: determine one or more accelerator devices of a second compute device that is remote to the first compute device to concurrently execute the multiple tasks that share data; and execute one or more of the tasks on the one or more accelerator devices of the first compute device as one or more other tasks of the job are concurrently executed with one or more accelerator devices of the second compute device using a shared virtual memory.",
"11. The compute device of claim 10, wherein to determine the one or more accelerator devices of the second compute device comprises to receive information regarding one or more accelerator devices of the second compute device from an orchestrator server.",
"12. The compute device of claim 9, wherein to execute the tasks comprises to execute the tasks simultaneously on the one or more accelerator devices of the compute device with a high speed serial interface (HSSI).",
"13. One or more machine-readable storage media comprising a plurality of instructions stored thereon that, when executed by at least one compute device cause the at least one compute device to: determine a configuration of at least one accelerator device of the compute device, wherein the configuration is indicative of parallel execution capabilities of at least one accelerator device, and wherein to determine the configuration includes to determine, for at least one accelerator device, whether the at least one accelerator device is capable of accessing a shared data set and whether the at least one accelerator device is capable of accessing a shared memory; receive, from a requester device, a job to be accelerated; divide the job into multiple tasks for a parallel execution of the multiple tasks among the one or more accelerator devices as a function of a job analysis of the job and the configuration of at least one accelerator device indicative of whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; schedule the multiple tasks to the one or more accelerator devices based on the job analysis; cause execution of the multiple tasks on the one or more accelerator devices for the parallel execution of the multiple tasks based on whether the at least one accelerator device is capable of accessing the shared data set and whether the one or more accelerator device is capable of accessing the shared memory; and combine task outputs from the one or more accelerator devices that executed the multiple tasks to obtain an output of the job.",
"14. The one or more machine-readable storage media of claim 13, wherein to determine the configuration of at least one accelerator device comprises to determine, by a micro-orchestrator logic unit of the compute device, the configuration of the at least one accelerator device.",
"15. The one or more machine-readable storage media of claim 13, further comprising a plurality of instructions stored thereon that, in response to being executed, cause the compute device to determine an availability of one or more of the accelerator devices, and the at least one of the accelerator devices is not a general purpose processor.",
"16. The one or more machine-readable storage media of claim 15, wherein to determine the availability of one or more of the accelerator devices comprises to determine one or more available kernels of the at least one accelerator device.",
"17. The one or more machine-readable storage media of claim 13, wherein to schedule the multiple tasks to the one or more accelerator devices comprises to schedule parallel execution of the tasks.",
"18. The one or more machine-readable storage media of claim 13, wherein to schedule the multiple tasks to the one or more accelerator devices comprises to confirm that the one or more accelerator devices are available to simultaneously execute multiple tasks that share data.",
"19. The one or more machine-readable storage media of claim 13, wherein execution of the multiple tasks comprises concurrently execute two or more of the tasks on two or more of the accelerator devices of the compute device with a high speed serial interface (HSSI).",
"20. The one or more machine-readable storage media of claim 13, further comprising a plurality of instructions stored thereon that, in response to being executed, cause the compute device to: determine whether an authorization is required from an orchestrator to execute the tasks on the compute device; transmit, in response to a determination that the authorization is required, the job analysis to the orchestrator; receive an authorization from the orchestrator; and receive the tasks to be accelerated.",
"21. The one or more machine-readable storage media of claim 20, wherein the compute device comprises a first compute device, and wherein to execute the tasks comprises to: determine one or more accelerator devices of a second compute device that is remote to the first compute device to concurrently execute the tasks that share data; and execute one or more of the tasks on the one or more accelerator devices of the first compute device as one or more other tasks of the job are concurrently executed with one or more accelerator devices of the second compute device using a shared virtual memory.",
"22. The one or more machine-readable storage media of claim 21, wherein to determine one or more accelerator devices of the second compute device comprises to receive information regarding one or more accelerator devices of the second compute device from a server.",
"23. The one or more machine-readable storage media of claim 20, wherein to execute the tasks comprises to execute the tasks simultaneously on the one or more accelerator devices of the compute device with a high speed serial interface (HSSI).",
"24. A compute device comprising: circuitry for determining a configuration of at least one of one or more accelerator devices of the compute device, wherein the configuration is indicative of parallel execution capabilities of at least one accelerator device, and wherein determining the configuration includes determining for the at least one accelerator device whether the accelerator device is capable of accessing a shared data set and whether the at least one accelerator device is capable of accessing a shared memory; circuitry for receiving, from a requester device remote from the compute device, a job to be accelerated; means for dividing the job into multiple tasks for a parallelization of the multiple tasks among the one or more accelerator devices, as a function of a job analysis of the job and the configuration of the at least one accelerator device indicative of whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; means for scheduling the multiple tasks to the one or more accelerator devices based on the job analysis; circuitry for executing the multiple tasks on the one or more accelerator devices for the parallelization of the multiple tasks based on whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; and means for combining task outputs from the at least one accelerator devices that executed the multiple tasks to obtain an output of the job.",
"25. A method comprising: determining, by a compute device, a configuration of at least one of one or more accelerator devices of the compute device, wherein the configuration is indicative of parallel execution capabilities of at least one accelerator device, and wherein determining the configuration includes determining for at least one accelerator device whether the accelerator device is capable of accessing a shared data set and whether the accelerator device is capable of accessing a shared memory; receiving, from a requester device and by the compute device, a job to be accelerated; dividing, by the compute device, the job into multiple tasks for a parallelization of the multiple tasks among the one or more accelerator devices, as a function of a job analysis of the job and the configuration of at least one accelerator device indicative of whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; scheduling, by the compute device, the multiple tasks to the one or more accelerator devices based on the job analysis; causing execution, by the compute device, of the multiple tasks on the one or more accelerator devices for the parallelization of the multiple tasks based on whether the at least one accelerator device is capable of accessing the shared data set and whether the at least one accelerator device is capable of accessing the shared memory; and combining, by the compute device, task outputs from the one or more accelerator devices that executed the tasks to obtain an output of the job.",
"26. The method of claim 25, wherein determining the configuration of at least one accelerator device comprises determining, by a micro-orchestrator logic unit of the compute device, the configuration of at least one accelerator device.",
"27. The method of claim 25, further comprising: determining, by the compute device, whether an authorization is required from an orchestrator to execute the tasks on the compute device; transmitting, by the compute device and in response to a determination that the authorization is required, the job analysis to the orchestrator; receiving, by the compute device, an authorization from the orchestrator; and receiving, by the compute device, the tasks to be accelerated.",
"28. The method of claim 27, wherein the compute device comprises a first compute device, and wherein executing the tasks comprises: determining, by the first compute device, one or more accelerator devices of a second compute device that is remote to the first compute device to concurrently execute the multiple tasks that share data; and executing one or more of the tasks on the one or more accelerator devices of the first compute device as one or more other tasks of the job are concurrently executed with one or more accelerator devices of the second compute device using a shared virtual memory."
],
"description_excerpt": "Demand for accelerator devices has continued to increase because the accelerator devices are becoming more important as they may be used in various technological areas, such as machine learning and genomics. Typical architectures for accelerator devices, such as field programmable gate arrays (FPGAs), cryptography accelerators, graphics accelerators, and/or compression accelerators (referred to herein as “accelerator devices,” “accelerators,” or “accelerator resources”) capable of accelerating the execution of a set of operations in a workload (e.g., processes, applications, services, etc.) may allow static assignment of specified amounts of shared resources of the accelerator device (e.g., high bandwidth memory, data storage, etc.) among different portions of the logic (e.g., circuitry) of the accelerator device. Typically, the workload is allocated with the required processor(s), memory, and accelerator device(s) for the duration of the workload. The workload may use its allocated accelerator device at any point of time; however, in many cases, the accelerator devices will remain idle leading to wastage of resources.\n\nThe concepts described herein are illustrated by way of example and not by way of limitation in the accompanying figures. For simplicity and clarity of illustration, elements illustrated in the figures are not necessarily drawn to scale. Where considered appropriate, reference labels have been repeated among the figures to indicate corresponding or analogous elements.",
"cpc": [
"G06F 9/5005",
"G06F 11/0709",
"G06F 11/0751",
"G06F 11/079",
"G06F 11/1453",
"G06F 11/3006",
"G06F 11/3034",
"G06F 11/3055",
"G06F 11/3079",
"G06F 11/3409",
"G06F 12/023",
"G06F 12/0284",
"G06F 12/0692",
"G06F 13/1652",
"G06F 13/4022",
"G06F 13/4027",
"G06F 15/161",
"G06F 15/80",
"G06F 16/1744",
"G06F 21/44",
"G06F 21/57",
"G06F 21/6218",
"G06F 21/70",
"G06F 21/73",
"G06F 21/76",
"G06F 2212/401",
"G06F 2212/402",
"G06F 2221/2107",
"G06F 3/0604",
"G06F 3/0608",
"G06F 3/0611",
"G06F 3/0613",
"G06F 3/0617",
"G06F 3/0641",
"G06F 3/0647",
"G06F 3/065",
"G06F 3/0653",
"G06F 3/067",
"G06F 7/06",
"G06F 8/65",
"G06F 8/654",
"G06F 8/656",
"G06F 8/658",
"G06F 9/3851",
"G06F 9/3891",
"G06F 9/4401",
"G06F 9/4843",
"G06F 9/4881",
"G06F 9/5038",
"G06F 9/5044",
"G06F 9/505",
"G06F 9/5083",
"G06F 9/544",
"G06T 1/20",
"G06T 1/60",
"G06T 9/005",
"H01R 13/453",
"H01R 13/4536",
"H01R 13/4538",
"H01R 13/631",
"H03K 19/1731",
"H03M 7/3084",
"H03M 7/40",
"H03M 7/42",
"H03M 7/60",
"H03M 7/6011",
"H03M 7/6017",
"H03M 7/6029",
"H04L 12/2881",
"H04L 12/4633",
"H04L 41/044",
"H04L 41/046",
"H04L 41/0816",
"H04L 41/0853",
"H04L 41/0895",
"H04L 41/0896",
"H04L 41/12",
"H04L 41/142",
"H04L 41/40",
"H04L 43/04",
"H04L 43/06",
"H04L 43/08",
"H04L 43/0894",
"H04L 47/20",
"H04L 47/2441",
"H04L 47/78",
"H04L 47/83",
"H04L 49/104",
"H04L 61/2007",
"H04L 63/1425",
"H04L 67/10",
"H04L 67/1014",
"H04L 67/327",
"H04L 67/36",
"H04L 9/0822",
"H05K 7/1452",
"H05K 7/1487",
"H05K 7/1492"
],
"ipc": [
"G06F 12/02",
"G06F 9/38",
"G06T 1/20",
"G06T 1/60",
"H04L 47/20",
"G06F 11/07",
"G06F 11/30",
"G06F 11/34",
"G06F 12/06",
"G06F 13/16",
"G06F 15/80",
"G06F 16/174",
"G06F 21/57",
"G06F 21/62",
"G06F 21/73",
"G06F 21/76",
"G06F 3/06",
"G06F 7/06",
"G06F 8/65",
"G06F 8/654",
"G06F 8/656",
"G06F 8/658",
"G06F 9/4401",
"G06F 9/48",
"G06F 9/50",
"G06F 9/54",
"G06T 9/00",
"H01R 13/453",
"H01R 13/631",
"H03K 19/173",
"H03M 7/30",
"H03M 7/40",
"H03M 7/42",
"H04L 12/28",
"H04L 12/46",
"H04L 9/08",
"H05K 7/14"
],
"assignees": [
"Intel Corp"
],
"inventors": [
"Susanne M. Balle",
"Francesc Guim Bernat",
"Slawomir PUTYRSKI",
"Joe Grecco",
"Henry Mitchel",
"Evan Custodio",
"Rahul Khanna",
"Sujoy Sen"
],
"filing_date": "2017-09-30",
"publication_date": "2021-06-08",
"grant_date": "2021-06-08",
"priority_date": "2016-11-29",
"application_number": "US-201715721829-A",
"family_id": "62190163",
"cited_by_count": 10,
"citations": [
"US20100191823A1",
"US20120054770A1",
"US20130179485A1",
"US20130232495A1",
"US9026765B1",
"US20150007182A1",
"US20170046179A1",
"US20170116004A1",
"US20170317945A1",
"US10034407B2",
"US10045098B2",
"US10085358B2",
"US20180077235A1",
"US20180150298A1",
"US20180150330A1",
"US20190065281A1"
]
}
Record 1,604 of 8,000 in Patents full text (MLC-0201). Request the full dataset.