MLchartDataset catalogue

Patent · US9699049B2 · B2 · US

Predictive model for anomaly detection and feedback-based scheduling

(11) Publication number
US9699049B2
(21) Application number
14/586,381
(22) Filing date
2014-12-30
(30) Priority date
2014-09-23
(43) Publication date
2017-07-04
(45) Date of grant
2017-07-04
(51) IPC
G06N 5/02; G06N 20/00; H04L 41/149
(52) CPC
  • H04L Transmission of digital information, e.g. telegraphic communication: 43/08, 41/0896, 41/122, 41/142, 41/147, 41/149, 41/16
  • G06N Computing arrangements based on specific computational models: 20/00, 99/005
(73) Assignee
eBay Inc
(72) Inventors
Chaitali GUPTA; Mayank Bansal; Tzu-Cheng Chuang; Ranjan Sinha; Sami Ben-Romdhane
(54) Title
Predictive model for anomaly detection and feedback-based scheduling
(57) Abstract

In an example embodiment, clusters of nodes in a network are monitored. Then the monitored data may be stored in an open time-series database. Data from the open time-series database is collected and labeled it as training data. Then a model is built through machine learning using the training data. Additional data is retrieved from the open time-series database. The additional data is left as unlabeled. Anomalies in the unlabeled data are computed using the model, producing prediction outcomes and metrics. Finally, the prediction outcomes and the network.

Full text
View on Google Patents

Claims (20)

  1. A system comprising: an open time-series database; a scheduler; a monitoring agent executable by one or more processors and configured to monitor clusters of nodes in a network and store monitored data in the open time-series database; an offline training module comprising: a data collection and preprocessing module configured to collect first data from the open time-series database and to label the first data as training data; and a machine learning model building module configured to build a model through machine learning using the training data; a real-time testing module comprising: a data collection and preprocessing module configured to collect second data from the open time-series database and to leave the second data as unlabeled; and a predictive model engine configured to compute anomalies in the unlabeled data using the model built by the machine learning model and to output prediction outcomes and metrics to the scheduler; and the scheduler configured to use the prediction outcomes and metrics to move or reduce workloads from problematic clusters of nodes in the network.
  2. The system of claim 1, wherein the scheduler comprises an extension to a YARN scheduler.
  3. The system of claim 2, wherein the extension comprises: a scheduler feedback language parser configured to parse feedback information written in a scheduler feedback language.
  4. The system of claim 2, wherein the extension comprises: a feedback agent configured to interact with the predictive model engine to receive feedback information.
  5. The system of claim 2, wherein the extension comprises: a feedback policy module configured to take scheduling rules and generate an execution plan based on the scheduling rules and feedback from a feedback agent.
  6. The system of claim 2, wherein the extension comprises: an action executor configured to execute a scheduling created based on feedback from a feedback agent.
  7. The system of claim 1, wherein the scheduler is contained in a resource manager.
  8. A method comprising: monitoring clusters of nodes in a network; storing monitored data in an open time-series database; collecting data from the open time-series database and labeling it as training data; building a model through machine learning using the training data; collecting additional data from the open time-series database; leaving the additional data as unlabeled; compute anomalies in the unlabeled data using the model, producing prediction outcomes and metrics; and using the prediction outcomes and metrics to move or reduce workloads from problematic clusters of nodes in the network.
  9. The method of claim 8, wherein the computing anomalies includes building a model using a trading data set using Multivariate Gaussian Distribution.
  10. The method of claim 8, wherein the computing anomalies includes applying a Matthews Correlation coefficient as a threshold to reduce false positives.
  11. The method of claim 8, wherein the computing anomalies includes applying a half total error rate as a threshold to reduce false positives.
  12. The method of claim 8, wherein the computing anomalies includes defining a function to calculate an anomaly score of data nodes.
  13. The method of claim 8, wherein the using the prediction outcomes includes: detecting that a data node is anomalous; in response to the detection that the data node is anomalous, locating one or more features contributing to the anomaly.
  14. The method of claim 13, wherein the locating includes deducing one or more features contributing to the anomaly using a single-variate Gaussian Distribution Function.
  15. A non-transitory machine-readable storage medium embodying instructions which, when executed by a machine, cause the machine to execute operations comprising: monitoring clusters of nodes in a network; storing monitored data in an open time-series database; collecting data from the open time-series database and labeling it as training data; building a model through machine learning using the training data; collecting additional data from the open time-series database; leaving the additional data as unlabeled; computing anomalies in the unlabeled data using the model, producing prediction outcomes and metrics; and using the prediction outcomes and metrics to move or reduce workloads from problematic clusters of nodes in the network.
  16. The non-transitory machine-readable storage medium of claim 15, wherein the computing anomalies includes building a model using a trading data set using Multivariate Gaussian Distribution.
  17. The non-transitory machine-readable storage medium of claim 15, wherein the computing anomalies includes applying a Matthews Correlation coefficiant as a threshold to reduce false positives.
  18. The non-transitory machine-readable storage medium of claim 15, wherein the computing anomalies includes applying a half total error rate as a threshold to reduce false positives.
  19. The non-transitory machine-readable storage medium of claim 15, wherein the computing anomalies includes defining a function to calculate an anomaly score of data nodes.
  20. The non-transitory machine-readable storage medium of claim 15, wherein the using the prediction outcomes includes: detecting that a data node is anomalous; in response to the detection that the data node is anomalous, locating one or more features contributing to the anomaly.

Description

This application is a Non-Provisional of and claims the benefit of priority under 35 U.S.C. §119(e) from U.S. Provisional Application Ser. No. 62/054,248, filed on Sep. 23, 2014 which is hereby incorporated by reference herein in its entirety.

A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the Patent and Trademark Office patent files or records, but otherwise reserves all copyright rights whatsoever. The following notice applies to the software and data as described below and in the drawings that form a part of this document: Copyright eBay, Inc. 2013, All Rights Reserved.

The present disclosure relates generally to detecting anomalies in computer server clusters.

In recent years, Hadoop has become the most popular distributed systems platform of choice in both industry and academia, used to distribute computing tasks among a number of different servers. A Hadoop cluster is a special type of computational cluster designed specifically for storing and analyzing large amounts of unstructured data in a distributed computing environment. Hadoop relies on Hadoop Distributed File System (HDFS) to store peta-bytes of data and runs massively parallel MapReduce programs to access them. MapReduce is inspired by functional programming model, and runs “map” and “reduce” tasks on multiple server machines in a cluster. A map function, provided by a user, divides input data into multiple chunks, produces intermediate results.

Citations (9)

  • US20010051862A1
  • US8018860B1
  • US20170019307A1
  • US20120197609A1
  • US20120020216A1
  • US8995249B1
  • US20160359683A1
  • US20160028599A1
  • US20160072726A1
Record as JSON
{
  "publication_number": "US9699049B2",
  "country": "US",
  "kind": "B2",
  "title": "Predictive model for anomaly detection and feedback-based scheduling",
  "abstract": "In an example embodiment, clusters of nodes in a network are monitored. Then the monitored data may be stored in an open time-series database. Data from the open time-series database is collected and labeled it as training data. Then a model is built through machine learning using the training data. Additional data is retrieved from the open time-series database. The additional data is left as unlabeled. Anomalies in the unlabeled data are computed using the model, producing prediction outcomes and metrics. Finally, the prediction outcomes and the network.",
  "claims": [
    "1. A system comprising: an open time-series database; a scheduler; a monitoring agent executable by one or more processors and configured to monitor clusters of nodes in a network and store monitored data in the open time-series database; an offline training module comprising: a data collection and preprocessing module configured to collect first data from the open time-series database and to label the first data as training data; and a machine learning model building module configured to build a model through machine learning using the training data; a real-time testing module comprising: a data collection and preprocessing module configured to collect second data from the open time-series database and to leave the second data as unlabeled; and a predictive model engine configured to compute anomalies in the unlabeled data using the model built by the machine learning model and to output prediction outcomes and metrics to the scheduler; and the scheduler configured to use the prediction outcomes and metrics to move or reduce workloads from problematic clusters of nodes in the network.",
    "2. The system of claim 1, wherein the scheduler comprises an extension to a YARN scheduler.",
    "3. The system of claim 2, wherein the extension comprises: a scheduler feedback language parser configured to parse feedback information written in a scheduler feedback language.",
    "4. The system of claim 2, wherein the extension comprises: a feedback agent configured to interact with the predictive model engine to receive feedback information.",
    "5. The system of claim 2, wherein the extension comprises: a feedback policy module configured to take scheduling rules and generate an execution plan based on the scheduling rules and feedback from a feedback agent.",
    "6. The system of claim 2, wherein the extension comprises: an action executor configured to execute a scheduling created based on feedback from a feedback agent.",
    "7. The system of claim 1, wherein the scheduler is contained in a resource manager.",
    "8. A method comprising: monitoring clusters of nodes in a network; storing monitored data in an open time-series database; collecting data from the open time-series database and labeling it as training data; building a model through machine learning using the training data; collecting additional data from the open time-series database; leaving the additional data as unlabeled; compute anomalies in the unlabeled data using the model, producing prediction outcomes and metrics; and using the prediction outcomes and metrics to move or reduce workloads from problematic clusters of nodes in the network.",
    "9. The method of claim 8, wherein the computing anomalies includes building a model using a trading data set using Multivariate Gaussian Distribution.",
    "10. The method of claim 8, wherein the computing anomalies includes applying a Matthews Correlation coefficient as a threshold to reduce false positives.",
    "11. The method of claim 8, wherein the computing anomalies includes applying a half total error rate as a threshold to reduce false positives.",
    "12. The method of claim 8, wherein the computing anomalies includes defining a function to calculate an anomaly score of data nodes.",
    "13. The method of claim 8, wherein the using the prediction outcomes includes: detecting that a data node is anomalous; in response to the detection that the data node is anomalous, locating one or more features contributing to the anomaly.",
    "14. The method of claim 13, wherein the locating includes deducing one or more features contributing to the anomaly using a single-variate Gaussian Distribution Function.",
    "15. A non-transitory machine-readable storage medium embodying instructions which, when executed by a machine, cause the machine to execute operations comprising: monitoring clusters of nodes in a network; storing monitored data in an open time-series database; collecting data from the open time-series database and labeling it as training data; building a model through machine learning using the training data; collecting additional data from the open time-series database; leaving the additional data as unlabeled; computing anomalies in the unlabeled data using the model, producing prediction outcomes and metrics; and using the prediction outcomes and metrics to move or reduce workloads from problematic clusters of nodes in the network.",
    "16. The non-transitory machine-readable storage medium of claim 15, wherein the computing anomalies includes building a model using a trading data set using Multivariate Gaussian Distribution.",
    "17. The non-transitory machine-readable storage medium of claim 15, wherein the computing anomalies includes applying a Matthews Correlation coefficiant as a threshold to reduce false positives.",
    "18. The non-transitory machine-readable storage medium of claim 15, wherein the computing anomalies includes applying a half total error rate as a threshold to reduce false positives.",
    "19. The non-transitory machine-readable storage medium of claim 15, wherein the computing anomalies includes defining a function to calculate an anomaly score of data nodes.",
    "20. The non-transitory machine-readable storage medium of claim 15, wherein the using the prediction outcomes includes: detecting that a data node is anomalous; in response to the detection that the data node is anomalous, locating one or more features contributing to the anomaly."
  ],
  "description_excerpt": "This application is a Non-Provisional of and claims the benefit of priority under 35 U.S.C. §119(e) from U.S. Provisional Application Ser. No. 62/054,248, filed on Sep. 23, 2014 which is hereby incorporated by reference herein in its entirety.\n\nA portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure, as it appears in the Patent and Trademark Office patent files or records, but otherwise reserves all copyright rights whatsoever. The following notice applies to the software and data as described below and in the drawings that form a part of this document: Copyright eBay, Inc. 2013, All Rights Reserved.\n\nThe present disclosure relates generally to detecting anomalies in computer server clusters.\n\nIn recent years, Hadoop has become the most popular distributed systems platform of choice in both industry and academia, used to distribute computing tasks among a number of different servers. A Hadoop cluster is a special type of computational cluster designed specifically for storing and analyzing large amounts of unstructured data in a distributed computing environment. Hadoop relies on Hadoop Distributed File System (HDFS) to store peta-bytes of data and runs massively parallel MapReduce programs to access them. MapReduce is inspired by functional programming model, and runs “map” and “reduce” tasks on multiple server machines in a cluster. A map function, provided by a user, divides input data into multiple chunks, produces intermediate results.",
  "cpc": [
    "H04L 43/08",
    "G06N 20/00",
    "G06N 99/005",
    "H04L 41/0896",
    "H04L 41/122",
    "H04L 41/142",
    "H04L 41/147",
    "H04L 41/149",
    "H04L 41/16"
  ],
  "ipc": [
    "G06N 5/02",
    "G06N 20/00",
    "H04L 41/149"
  ],
  "assignees": [
    "eBay Inc"
  ],
  "inventors": [
    "Chaitali GUPTA",
    "Mayank Bansal",
    "Tzu-Cheng Chuang",
    "Ranjan Sinha",
    "Sami Ben-Romdhane"
  ],
  "filing_date": "2014-12-30",
  "publication_date": "2017-07-04",
  "grant_date": "2017-07-04",
  "priority_date": "2014-09-23",
  "application_number": "US-201414586381-A",
  "family_id": "55526882",
  "cited_by_count": 53,
  "citations": [
    "US20010051862A1",
    "US8018860B1",
    "US20170019307A1",
    "US20120197609A1",
    "US20120020216A1",
    "US8995249B1",
    "US20160359683A1",
    "US20160028599A1",
    "US20160072726A1"
  ]
}

Record 4,029 of 8,000 in Patents full text (MLC-0201). Request the full dataset.