MLchartDataset catalogue

Patent · US10404547B2 · B2 · US

Workload optimization, scheduling, and placement for rack-scale architecture computing systems

(11) Publication number
US10404547B2
(21) Application number
15/114,687
(22) Filing date
2015-02-24
(30) Priority date
2014-02-27
(43) Publication date
2019-09-03
(45) Date of grant
2019-09-03
(51) IPC
H04L 12/24; G06F 15/173; H04L 12/26
(52) CPC
  • H04L Transmission of digital information, e.g. telegraphic communication: 41/147, 41/0883, 41/149, 41/22, 41/40, 41/5009, 41/5054, 43/0817, 43/20
(73) Assignee
Intel Corp
(72) Inventors
Katalin K. Bartfai-Walcott; Michael Christopher Woods; Giovani ESTRADA; John Kennedy; Joseph Butler; Slawomir PUTYRSKI; Alexander Leckey; Victor M. BAYON-MOLINO; Connor UPTON; Thijs Metsch
(54) Title
Workload optimization, scheduling, and placement for rack-scale architecture computing systems
(57) Abstract

Technologies for datacenter management include one or more computing racks each including a rack controller. The rack controller may receive system, performance, or health metrics for the components of the computing rack. The rack controller generates regression models to predict component lifespan and may predict logical machine lifespans based on the lifespan of the included hardware components. The rack controller may generate notifications or schedule maintenance sessions based on remaining component or logical machine lifespans. The rack controller may compose logical machines using components having similar remaining lifespans. In some embodiments the rack controller may validate a service level agreement prior to executing an application based on the probability of component failure. A management interface may generate an interactive visualization of the system state and optimize the datacenter schedule based on optimization rules derived from human input in response to the visualization. Other embodiments are described and claimed.

Full text
View on Google Patents

Claims (19)

  1. A rack controller of a computing rack, the rack controller comprising: a processor; and a memory storing a plurality of instructions, which, when executed on the processor, causes the rack controller to: receive a metric associated with a hardware component of the computing rack, the hardware component managed by the rack controller, wherein the metric comprises a system metric, a performance metric, or a health metric; determine a regression model for the hardware component based on the metric associated with the hardware component; determine a mean-time-to-failure (MTTF) value for the hardware component, wherein to determine the MTTF value comprises to (i) determine a predicted metric associated with the hardware component based on the regression model for the hardware component, (ii) determine a service level metric of a service level agreement associated with the hardware component, and (iii) compare the predicted metric with the service level metric to obtain a distance to a point in time in which the predicted metric and the service level metric intersect; and compose a logical machine including the hardware component and a plurality of second hardware components of the computing rack, each of the plurality of second hardware components managed by the rack controller, wherein the plurality of second hardware components is of compute, storage, network, and memory resources each associated with a MTTF value that is associated with a same scheduled maintenance session as the MTTF value for the hardware component.
  2. The rack controller of claim 1, wherein to receive the metric comprises to receive the metric from a metric component of the hardware component.
  3. The rack controller of claim 1, wherein the hardware component comprises a compute resource, a memory resource, a storage resource, or a network resource.
  4. The rack controller of claim 1, wherein to determine the regression model comprises to determine a linear regression model.
  5. The rack controller of claim 1, wherein to determine the regression model comprises to determine a non-linear regression model.
  6. The rack controller of claim 1, wherein to determine the MTTF value for the hardware component comprises to: determine a predicted metric associated with the hardware component based on the regression model; and compare the predicted metric to a corresponding one of the one or more predefined threshold metrics.
  7. The rack controller of claim 1, wherein the plurality of instructions further causes the rack controller to notify a user of the MTTF value for the hardware component.
  8. The rack controller of claim 1, wherein the plurality of instructions further causes the rack controller to determine a future time for a maintenance session associated with the logical machine based on a MTTF value for the logical machine.
  9. The rack controller of claim 8, wherein the plurality of instructions further causes the rack controller to receive a performance indicator associated with a computing application assigned to the logical machine and wherein to determine the future time further comprises to determine the future time based on the performance indicator.
  10. A method for datacenter management, the method comprising: receiving, by a rack controller of a computing rack, a metric associated with a hardware component of the computing rack, wherein the metric comprises a system metric, a performance metric, or a health metric, and wherein the hardware component is managed by the rack controller; determining, by the rack controller, a regression model for the hardware component based on the metric associated with the hardware component; determining, by the rack controller, a mean-time-to-failure (MTTF) value for the hardware component, wherein determining the MTTF value comprises (i) determining a predicted metric associated with the hardware component based on the regression model for the hardware component, (ii) determining a service level metric of a service level agreement associated with the hardware component, and (iii) comparing the predicted metric with the service level metric to obtain a distance to a point in time in which the predicted metric and the service level metric intersect; and composing, by the rack controller, a logical machine including the hardware component and a plurality of second hardware components of the computing rack, each of the plurality of second hardware components managed by the rack controller, wherein the plurality of second hardware components is of compute, storage, network, and memory resources each associated with a MTTF value that is associated with a same scheduled maintenance session as the MTTF value for the hardware component.
  11. The method of claim 10, wherein receiving the metric comprises receiving the metric from a metric component of the hardware component.
  12. The method of claim 10, wherein determining the MTTF value for the hardware component comprises: determining a predicted metric associated with the hardware component based on the regression model; and comparing the predicted metric to a corresponding one of the one or more predefined threshold metrics.
  13. The method of claim 10, further comprising determining, by the rack controller, a future time for a maintenance session associated with the logical machine based on a MTTF value for the logical machine.
  14. The method of claim 13, further comprising: receiving, by the rack controller, a performance indicator associated with a computing application assigned to the logical machine; wherein determining the future time further comprises determining the future time based on the performance indicator.
  15. One or more non-transitory computer-readable storage media comprising a plurality of instructions that in response to being executed cause a rack controller of a computing rack to: receive a metric associated with a hardware component of the computing rack, the hardware component managed by the rack controller, wherein the metric comprises a system metric, a performance metric, or a health metric; determine a regression model for the hardware component based on the metric associated with the hardware component; determine a mean-time-to-failure (MTTF) value for the hardware component, wherein to determine the MTTF value comprises to (i) determine a predicted metric associated with the hardware component based on the regression model for the hardware component, (ii) determine a service level metric of a service level agreement associated with the hardware component, and (iii) compare the predicted metric with the service level metric to obtain a distance to a point in time in which the predicted metric and the service level metric intersect; and compose a logical machine including the hardware component and a plurality of second hardware components of the computing rack, each of the plurality of second hardware components managed by the rack controller, wherein the plurality of second hardware components is associated with a MTTF value that is associated with a same scheduled maintenance session as the MTTF value for the hardware component.
  16. The one or more non-transitory computer-readable storage media of claim 15, wherein to receive the metric comprises to receive the metric from a metric component of the hardware component.
  17. The one or more non-transitory computer-readable storage media of claim 15, wherein to determine the MTTF value for the hardware component comprises to: determine a predicted metric associated with the hardware component based on the regression model; and compare the predicted metric to a predefined corresponding one of the one or more predefined threshold metrics.
  18. The one or more non-transitory computer-readable storage media of claim 15, wherein the plurality of instructions further causes the rack controller to determine a future time for a maintenance session associated with the logical machine based on a MTTF value for the logical machine.
  19. The one or more non-transitory computer-readable storage media of claim 18, wherein the plurality of instructions further causes the rack controller to: receive a performance indicator associated with a computing application assigned to the logical machine; wherein to determine the future time further comprises to determine the future time based on the performance indicator.

Description

“Cloud” computing is a term often used to refer to the provisioning of computing resources as a service, usually by a number of computer servers that are networked together at a location remote from the location from which the services are requested. A cloud datacenter typically refers to the physical arrangement of servers that make up a cloud or a particular portion of a cloud. For example, servers can be physically arranged in the datacenter into rooms, groups, rows, and racks. A datacenter may have one or more “zones,” which may include one or more rooms of servers. Each room may have one or more rows of servers, and each row may include one or more racks. Each rack may include one or more individual server nodes. Servers in zones, rooms, racks, and/or rows may be arranged into virtual groups based on physical infrastructure requirements of the datacenter facility, which may include power, energy, thermal, heat, and/or other requirements.

As the popularity of cloud computing grows, customers are increasingly requiring cloud service providers to include service-level agreements (SLAs) within the terms of their contracts. Such SLAs require cloud service providers to agree to provide the customer with at least a certain level of service, which can be measured by one or more metrics (e.g., system uptime, throughput, etc.). SLA goals (including service delivery objective (SDO) and service level objective (SLO) goals), efficiency targets, compliance objectives, energy targets including facilities, and other environmental and contextual constraints may also all be considered.

Citations (15)

  • US20030154112A1
  • US20040153835A1
  • US20070028147A1
  • JP2005339528A
  • US20090172168A1
  • US20080177613A1
  • JP2008176674A
  • US20090112531A1
  • US20110320591A1
  • US20110119525A1
  • US20110179176A1
  • JP2012203750A
  • US20130138419A1
  • CN103513983A
  • US20150067415A1
Record as JSON
{
  "publication_number": "US10404547B2",
  "country": "US",
  "kind": "B2",
  "title": "Workload optimization, scheduling, and placement for rack-scale architecture computing systems",
  "abstract": "Technologies for datacenter management include one or more computing racks each including a rack controller. The rack controller may receive system, performance, or health metrics for the components of the computing rack. The rack controller generates regression models to predict component lifespan and may predict logical machine lifespans based on the lifespan of the included hardware components. The rack controller may generate notifications or schedule maintenance sessions based on remaining component or logical machine lifespans. The rack controller may compose logical machines using components having similar remaining lifespans. In some embodiments the rack controller may validate a service level agreement prior to executing an application based on the probability of component failure. A management interface may generate an interactive visualization of the system state and optimize the datacenter schedule based on optimization rules derived from human input in response to the visualization. Other embodiments are described and claimed.",
  "claims": [
    "1. A rack controller of a computing rack, the rack controller comprising: a processor; and a memory storing a plurality of instructions, which, when executed on the processor, causes the rack controller to: receive a metric associated with a hardware component of the computing rack, the hardware component managed by the rack controller, wherein the metric comprises a system metric, a performance metric, or a health metric; determine a regression model for the hardware component based on the metric associated with the hardware component; determine a mean-time-to-failure (MTTF) value for the hardware component, wherein to determine the MTTF value comprises to (i) determine a predicted metric associated with the hardware component based on the regression model for the hardware component, (ii) determine a service level metric of a service level agreement associated with the hardware component, and (iii) compare the predicted metric with the service level metric to obtain a distance to a point in time in which the predicted metric and the service level metric intersect; and compose a logical machine including the hardware component and a plurality of second hardware components of the computing rack, each of the plurality of second hardware components managed by the rack controller, wherein the plurality of second hardware components is of compute, storage, network, and memory resources each associated with a MTTF value that is associated with a same scheduled maintenance session as the MTTF value for the hardware component.",
    "2. The rack controller of claim 1, wherein to receive the metric comprises to receive the metric from a metric component of the hardware component.",
    "3. The rack controller of claim 1, wherein the hardware component comprises a compute resource, a memory resource, a storage resource, or a network resource.",
    "4. The rack controller of claim 1, wherein to determine the regression model comprises to determine a linear regression model.",
    "5. The rack controller of claim 1, wherein to determine the regression model comprises to determine a non-linear regression model.",
    "6. The rack controller of claim 1, wherein to determine the MTTF value for the hardware component comprises to: determine a predicted metric associated with the hardware component based on the regression model; and compare the predicted metric to a corresponding one of the one or more predefined threshold metrics.",
    "7. The rack controller of claim 1, wherein the plurality of instructions further causes the rack controller to notify a user of the MTTF value for the hardware component.",
    "8. The rack controller of claim 1, wherein the plurality of instructions further causes the rack controller to determine a future time for a maintenance session associated with the logical machine based on a MTTF value for the logical machine.",
    "9. The rack controller of claim 8, wherein the plurality of instructions further causes the rack controller to receive a performance indicator associated with a computing application assigned to the logical machine and wherein to determine the future time further comprises to determine the future time based on the performance indicator.",
    "10. A method for datacenter management, the method comprising: receiving, by a rack controller of a computing rack, a metric associated with a hardware component of the computing rack, wherein the metric comprises a system metric, a performance metric, or a health metric, and wherein the hardware component is managed by the rack controller; determining, by the rack controller, a regression model for the hardware component based on the metric associated with the hardware component; determining, by the rack controller, a mean-time-to-failure (MTTF) value for the hardware component, wherein determining the MTTF value comprises (i) determining a predicted metric associated with the hardware component based on the regression model for the hardware component, (ii) determining a service level metric of a service level agreement associated with the hardware component, and (iii) comparing the predicted metric with the service level metric to obtain a distance to a point in time in which the predicted metric and the service level metric intersect; and composing, by the rack controller, a logical machine including the hardware component and a plurality of second hardware components of the computing rack, each of the plurality of second hardware components managed by the rack controller, wherein the plurality of second hardware components is of compute, storage, network, and memory resources each associated with a MTTF value that is associated with a same scheduled maintenance session as the MTTF value for the hardware component.",
    "11. The method of claim 10, wherein receiving the metric comprises receiving the metric from a metric component of the hardware component.",
    "12. The method of claim 10, wherein determining the MTTF value for the hardware component comprises: determining a predicted metric associated with the hardware component based on the regression model; and comparing the predicted metric to a corresponding one of the one or more predefined threshold metrics.",
    "13. The method of claim 10, further comprising determining, by the rack controller, a future time for a maintenance session associated with the logical machine based on a MTTF value for the logical machine.",
    "14. The method of claim 13, further comprising: receiving, by the rack controller, a performance indicator associated with a computing application assigned to the logical machine; wherein determining the future time further comprises determining the future time based on the performance indicator.",
    "15. One or more non-transitory computer-readable storage media comprising a plurality of instructions that in response to being executed cause a rack controller of a computing rack to: receive a metric associated with a hardware component of the computing rack, the hardware component managed by the rack controller, wherein the metric comprises a system metric, a performance metric, or a health metric; determine a regression model for the hardware component based on the metric associated with the hardware component; determine a mean-time-to-failure (MTTF) value for the hardware component, wherein to determine the MTTF value comprises to (i) determine a predicted metric associated with the hardware component based on the regression model for the hardware component, (ii) determine a service level metric of a service level agreement associated with the hardware component, and (iii) compare the predicted metric with the service level metric to obtain a distance to a point in time in which the predicted metric and the service level metric intersect; and compose a logical machine including the hardware component and a plurality of second hardware components of the computing rack, each of the plurality of second hardware components managed by the rack controller, wherein the plurality of second hardware components is associated with a MTTF value that is associated with a same scheduled maintenance session as the MTTF value for the hardware component.",
    "16. The one or more non-transitory computer-readable storage media of claim 15, wherein to receive the metric comprises to receive the metric from a metric component of the hardware component.",
    "17. The one or more non-transitory computer-readable storage media of claim 15, wherein to determine the MTTF value for the hardware component comprises to: determine a predicted metric associated with the hardware component based on the regression model; and compare the predicted metric to a predefined corresponding one of the one or more predefined threshold metrics.",
    "18. The one or more non-transitory computer-readable storage media of claim 15, wherein the plurality of instructions further causes the rack controller to determine a future time for a maintenance session associated with the logical machine based on a MTTF value for the logical machine.",
    "19. The one or more non-transitory computer-readable storage media of claim 18, wherein the plurality of instructions further causes the rack controller to: receive a performance indicator associated with a computing application assigned to the logical machine; wherein to determine the future time further comprises to determine the future time based on the performance indicator."
  ],
  "description_excerpt": "“Cloud” computing is a term often used to refer to the provisioning of computing resources as a service, usually by a number of computer servers that are networked together at a location remote from the location from which the services are requested. A cloud datacenter typically refers to the physical arrangement of servers that make up a cloud or a particular portion of a cloud. For example, servers can be physically arranged in the datacenter into rooms, groups, rows, and racks. A datacenter may have one or more “zones,” which may include one or more rooms of servers. Each room may have one or more rows of servers, and each row may include one or more racks. Each rack may include one or more individual server nodes. Servers in zones, rooms, racks, and/or rows may be arranged into virtual groups based on physical infrastructure requirements of the datacenter facility, which may include power, energy, thermal, heat, and/or other requirements.\n\nAs the popularity of cloud computing grows, customers are increasingly requiring cloud service providers to include service-level agreements (SLAs) within the terms of their contracts. Such SLAs require cloud service providers to agree to provide the customer with at least a certain level of service, which can be measured by one or more metrics (e.g., system uptime, throughput, etc.). SLA goals (including service delivery objective (SDO) and service level objective (SLO) goals), efficiency targets, compliance objectives, energy targets including facilities, and other environmental and contextual constraints may also all be considered.",
  "cpc": [
    "H04L 41/147",
    "H04L 41/0883",
    "H04L 41/149",
    "H04L 41/22",
    "H04L 41/40",
    "H04L 41/5009",
    "H04L 41/5054",
    "H04L 43/0817",
    "H04L 43/20"
  ],
  "ipc": [
    "H04L 12/24",
    "G06F 15/173",
    "H04L 12/26"
  ],
  "assignees": [
    "Intel Corp"
  ],
  "inventors": [
    "Katalin K. Bartfai-Walcott",
    "Michael Christopher Woods",
    "Giovani ESTRADA",
    "John Kennedy",
    "Joseph Butler",
    "Slawomir PUTYRSKI",
    "Alexander Leckey",
    "Victor M. BAYON-MOLINO",
    "Connor UPTON",
    "Thijs Metsch"
  ],
  "filing_date": "2015-02-24",
  "publication_date": "2019-09-03",
  "grant_date": "2019-09-03",
  "priority_date": "2014-02-27",
  "application_number": "US-201515114687-A",
  "family_id": "54009540",
  "cited_by_count": 11,
  "citations": [
    "US20030154112A1",
    "US20040153835A1",
    "US20070028147A1",
    "JP2005339528A",
    "US20090172168A1",
    "US20080177613A1",
    "JP2008176674A",
    "US20090112531A1",
    "US20110320591A1",
    "US20110119525A1",
    "US20110179176A1",
    "JP2012203750A",
    "US20130138419A1",
    "CN103513983A",
    "US20150067415A1"
  ]
}

Record 2,610 of 8,000 in Patents full text (MLC-0201). Request the full dataset.