MLchartDataset catalogue

Patent · US10289925B2 · B2 · US

Object classification in image data using machine learning models

(11) Publication number
US10289925B2
(21) Application number
15/363,835
(22) Filing date
2016-11-29
(30) Priority date
2016-11-29
(43) Publication date
2019-05-14
(45) Date of grant
2019-05-14
(51) IPC
G06K 9/46; G06K 9/62; G06N 20/00; G06K 9/00
(52) CPC
  • G06N Computing arrangements based on specific computational models: 20/00, 99/005
  • G06F Electric digital data processing: 18/241, 18/2431
  • G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/00201, 9/4604, 9/6268
  • G06V Image or video recognition or understanding: 10/44, 10/56, 20/64
(73) Assignee
SAP SE
(72) Inventors
Waqas Ahmad Farooqi; Jonas Lipps; Eckehard Schmidt; Thomas Fricke; Nemrude Verzano
(54) Title
Object classification in image data using machine learning models
(57) Abstract

Combined color and depth data for a field of view is received. Thereafter, using at least one bounding polygon algorithm, at least one proposed bounding polygon is defined for the field of view. It can then be determined, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object. The image data within each bounding polygon that is determined to encapsulate an object can then be provided to a first object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon. Further, the image data within each bounding polygon that is determined to encapsulate an object is provided to a second object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon. A final classification for each bounding polygon is then determined based on the output of the first classifier machine learning model and the output of the second classifier machine learning model.

Full text
View on Google Patents

Claims (17)

  1. A method for implementation by one or more data processors forming part of at least one computing system, the method comprising: receiving combined color and depth data for a field of view; defining, using at least one bounding polygon algorithm, at least one proposed bounding polygon for the field of view; determining, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object; providing the image data within each bounding polygon that is determined to encapsulate an object to a first object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon; providing the image data within each bounding polygon that is determined to encapsulate an object to a second object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon; determining a final classification for each bounding polygon based on the output of the first classifier machine learning model and the output of the second classifier machine learning model; and providing data characterizing the final classification for each bounding polygon; wherein at least one of the binary classifier, the first object classifier, or the second object classifier utilizes a machine learning model that is selected amongst a plurality of machine learning models based on a type of object encapsulated within the corresponding bounding polygon.
  2. The method of claim 1, wherein the at least one first classifier machine learning model is a region and measurements-based convolutional neural network.
  3. The method of claim 1, wherein the combined color and depth image data is RGB-D data.
  4. The method of claim 1, wherein the first object classifier uses metadata characterizing each object.
  5. The method of claim 4, wherein the metadata is extracted from the combined color and image data.
  6. The method of claim 1, wherein the at least one machine learning model of the binary classifier is one or more of: a neural network, a convolutional neural network, a logistic regression model, a support vector machine, decision trees, ensemble model, k-nearest neighbors model, linear regression model, naïve Bayes model, a logistic regression model, and/or a perceptron model.
  7. The method of claim 1, wherein the at least one machine learning model of the first object classifier is one or more of: a neural network, a convolutional neural network, a logistic regression model, a support vector machine, decision trees, ensemble model, k-nearest neighbors model, linear regression model, naïve Bayes model, a logistic regression model, and/or a perceptron model.
  8. The method of claim 1, wherein the at least one machine learning model of the second object classifier is one or more of: a neural network, a convolutional neural network, a logistic regression model, a support vector machine, decision trees, ensemble model, k-nearest neighbors model, linear regression model, naïve Bayes model, a logistic regression model, and/or a perceptron model.
  9. The method of claim 1 further comprising: discarding proposed bounding polygons determined, by the binary classifier, to not include an object.
  10. The method of claim 1, wherein the providing data characterizing the final classification for each bounding polygon comprises at least one of: displaying the data characterizing the final classification for each bounding polygon in an electronic visual display, loading the data characterizing the final classification for each bounding polygon into memory, storing the data characterizing the final classification for each bounding polygon in persistence, or transmitting the data characterizing the final classification for each bounding polygon to a remote computing device.
  11. A method for implementation by one or more data processors forming part of at least one computing device, the method comprising: receiving RGB-data for a field of view; defining, using at least one bounding polygon algorithm, at least one bounding polygon for the field of view; determining, using a binary classifier machine learning model trained using a plurality of images of known objects, whether each bounding polygon encapsulates one of the known objects; providing the image data within each bounding polygon that is determined to encapsulate one of the known objects to a select one or more a plurality of classifier machine learning models trained using a plurality of images of known objects, to classify the known objects; and providing data characterizing the classification of the known objects; wherein: the select one or more of the plurality of classifier machine learning models to which the image data is provided are selected based on metadata associated with the RGB-data; the metadata associated with the RGB-data acts as a pre-classifier.
  12. The method of claim 11, wherein the RGB data is RGB-D data.
  13. A system comprising: at least one data processor; and memory storing instructions which, when executed by the at least one data processor, result in operations comprising: receiving combined color and depth data for a field of view; defining, using at least one bounding polygon algorithm, at least one proposed bounding polygon for the field of view; determining, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object; providing the image data within each bounding polygon that is determined to encapsulate an object to a first object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon; providing the image data within each bounding polygon that is determined to encapsulate an object to a second object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon, the first object classifier being a different type than the second object classifier, the at least one machine learning model of the second object classifier comprising a bag-of-word (BoW) model that treats image features as words; determining a final classification for each bounding polygon based on the output of the first classifier machine learning model and the output of the second classifier machine learning model; and providing data characterizing the final classification for each bounding polygon.
  14. The system of claim 13, wherein the at least one first classifier machine learning model is a region and measurements-based convolutional neural network.
  15. The system of claim 13, wherein the combined color and depth image data is RGB-D data.
  16. The system of claim 13, wherein the first object classifier uses metadata characterizing each object, the metadata being extracted from the combined color and image data.
  17. The system of claim 13, wherein the at least one machine learning model of the binary classifier is one or more of: a neural network, a convolutional neural network, a logistic regression model, a support vector machine, decision trees, ensemble model, k-nearest neighbors model, linear regression model, naïve Bayes model, a logistic regression model, and/or a perceptron model.

Description

The subject matter described herein relates to the classification of objects within image data using machine learning models.

Sensors are increasingly being adopted across multiple computing platforms (including standalone sensors for use in gaming platforms, mobile phones, etc.) to provide multi-dimensional image data (e.g., three dimensional data, etc.). Such image data is computationally analyzed to localize objects and, in some cases, to later identify or otherwise characterize such objects. However, both localization and identification of objects within multi-dimensional image data remains imprecise.

In one aspect, combined color and depth data for a field of view is received. Thereafter, using at least one bounding polygon algorithm, at least one proposed bounding polygon is defined for the field of view. It can then be determined, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object. The image data within each bounding polygon that is determined to encapsulate an object can then be provided to a first object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon.

Citations (4)

  • US20170132450A1
  • US20170161545A1
  • US20180137642A1
  • US20180189611A1
Record as JSON
{
  "publication_number": "US10289925B2",
  "country": "US",
  "kind": "B2",
  "title": "Object classification in image data using machine learning models",
  "abstract": "Combined color and depth data for a field of view is received. Thereafter, using at least one bounding polygon algorithm, at least one proposed bounding polygon is defined for the field of view. It can then be determined, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object. The image data within each bounding polygon that is determined to encapsulate an object can then be provided to a first object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon. Further, the image data within each bounding polygon that is determined to encapsulate an object is provided to a second object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon. A final classification for each bounding polygon is then determined based on the output of the first classifier machine learning model and the output of the second classifier machine learning model.",
  "claims": [
    "1. A method for implementation by one or more data processors forming part of at least one computing system, the method comprising: receiving combined color and depth data for a field of view; defining, using at least one bounding polygon algorithm, at least one proposed bounding polygon for the field of view; determining, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object; providing the image data within each bounding polygon that is determined to encapsulate an object to a first object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon; providing the image data within each bounding polygon that is determined to encapsulate an object to a second object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon; determining a final classification for each bounding polygon based on the output of the first classifier machine learning model and the output of the second classifier machine learning model; and providing data characterizing the final classification for each bounding polygon; wherein at least one of the binary classifier, the first object classifier, or the second object classifier utilizes a machine learning model that is selected amongst a plurality of machine learning models based on a type of object encapsulated within the corresponding bounding polygon.",
    "2. The method of claim 1, wherein the at least one first classifier machine learning model is a region and measurements-based convolutional neural network.",
    "3. The method of claim 1, wherein the combined color and depth image data is RGB-D data.",
    "4. The method of claim 1, wherein the first object classifier uses metadata characterizing each object.",
    "5. The method of claim 4, wherein the metadata is extracted from the combined color and image data.",
    "6. The method of claim 1, wherein the at least one machine learning model of the binary classifier is one or more of: a neural network, a convolutional neural network, a logistic regression model, a support vector machine, decision trees, ensemble model, k-nearest neighbors model, linear regression model, naïve Bayes model, a logistic regression model, and/or a perceptron model.",
    "7. The method of claim 1, wherein the at least one machine learning model of the first object classifier is one or more of: a neural network, a convolutional neural network, a logistic regression model, a support vector machine, decision trees, ensemble model, k-nearest neighbors model, linear regression model, naïve Bayes model, a logistic regression model, and/or a perceptron model.",
    "8. The method of claim 1, wherein the at least one machine learning model of the second object classifier is one or more of: a neural network, a convolutional neural network, a logistic regression model, a support vector machine, decision trees, ensemble model, k-nearest neighbors model, linear regression model, naïve Bayes model, a logistic regression model, and/or a perceptron model.",
    "9. The method of claim 1 further comprising: discarding proposed bounding polygons determined, by the binary classifier, to not include an object.",
    "10. The method of claim 1, wherein the providing data characterizing the final classification for each bounding polygon comprises at least one of: displaying the data characterizing the final classification for each bounding polygon in an electronic visual display, loading the data characterizing the final classification for each bounding polygon into memory, storing the data characterizing the final classification for each bounding polygon in persistence, or transmitting the data characterizing the final classification for each bounding polygon to a remote computing device.",
    "11. A method for implementation by one or more data processors forming part of at least one computing device, the method comprising: receiving RGB-data for a field of view; defining, using at least one bounding polygon algorithm, at least one bounding polygon for the field of view; determining, using a binary classifier machine learning model trained using a plurality of images of known objects, whether each bounding polygon encapsulates one of the known objects; providing the image data within each bounding polygon that is determined to encapsulate one of the known objects to a select one or more a plurality of classifier machine learning models trained using a plurality of images of known objects, to classify the known objects; and providing data characterizing the classification of the known objects; wherein: the select one or more of the plurality of classifier machine learning models to which the image data is provided are selected based on metadata associated with the RGB-data; the metadata associated with the RGB-data acts as a pre-classifier.",
    "12. The method of claim 11, wherein the RGB data is RGB-D data.",
    "13. A system comprising: at least one data processor; and memory storing instructions which, when executed by the at least one data processor, result in operations comprising: receiving combined color and depth data for a field of view; defining, using at least one bounding polygon algorithm, at least one proposed bounding polygon for the field of view; determining, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object; providing the image data within each bounding polygon that is determined to encapsulate an object to a first object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon; providing the image data within each bounding polygon that is determined to encapsulate an object to a second object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon, the first object classifier being a different type than the second object classifier, the at least one machine learning model of the second object classifier comprising a bag-of-word (BoW) model that treats image features as words; determining a final classification for each bounding polygon based on the output of the first classifier machine learning model and the output of the second classifier machine learning model; and providing data characterizing the final classification for each bounding polygon.",
    "14. The system of claim 13, wherein the at least one first classifier machine learning model is a region and measurements-based convolutional neural network.",
    "15. The system of claim 13, wherein the combined color and depth image data is RGB-D data.",
    "16. The system of claim 13, wherein the first object classifier uses metadata characterizing each object, the metadata being extracted from the combined color and image data.",
    "17. The system of claim 13, wherein the at least one machine learning model of the binary classifier is one or more of: a neural network, a convolutional neural network, a logistic regression model, a support vector machine, decision trees, ensemble model, k-nearest neighbors model, linear regression model, naïve Bayes model, a logistic regression model, and/or a perceptron model."
  ],
  "description_excerpt": "The subject matter described herein relates to the classification of objects within image data using machine learning models.\n\nSensors are increasingly being adopted across multiple computing platforms (including standalone sensors for use in gaming platforms, mobile phones, etc.) to provide multi-dimensional image data (e.g., three dimensional data, etc.). Such image data is computationally analyzed to localize objects and, in some cases, to later identify or otherwise characterize such objects. However, both localization and identification of objects within multi-dimensional image data remains imprecise.\n\nIn one aspect, combined color and depth data for a field of view is received. Thereafter, using at least one bounding polygon algorithm, at least one proposed bounding polygon is defined for the field of view. It can then be determined, using a binary classifier having at least one machine learning model trained using a plurality of images of known objects, whether each proposed bounding polygon encapsulates an object. The image data within each bounding polygon that is determined to encapsulate an object can then be provided to a first object classifier having at least one machine learning model trained using a plurality of images of known objects, to classify the object encapsulated within the respective bounding polygon.",
  "cpc": [
    "G06N 20/00",
    "G06F 18/241",
    "G06F 18/2431",
    "G06K 9/00201",
    "G06K 9/4604",
    "G06K 9/6268",
    "G06N 99/005",
    "G06V 10/44",
    "G06V 10/56",
    "G06V 20/64"
  ],
  "ipc": [
    "G06K 9/46",
    "G06K 9/62",
    "G06N 20/00",
    "G06K 9/00"
  ],
  "assignees": [
    "SAP SE"
  ],
  "inventors": [
    "Waqas Ahmad Farooqi",
    "Jonas Lipps",
    "Eckehard Schmidt",
    "Thomas Fricke",
    "Nemrude Verzano"
  ],
  "filing_date": "2016-11-29",
  "publication_date": "2019-05-14",
  "grant_date": "2019-05-14",
  "priority_date": "2016-11-29",
  "application_number": "US-201615363835-A",
  "family_id": "60582358",
  "cited_by_count": 13,
  "citations": [
    "US20170132450A1",
    "US20170161545A1",
    "US20180137642A1",
    "US20180189611A1"
  ]
}

Record 2,816 of 8,000 in Patents full text (MLC-0201). Request the full dataset.