MLchartDataset catalogue

Patent · US11256958B1 · B1 · US

Training with simulated images

(11) Publication number
US11256958B1
(21) Application number
16/519,686
(22) Filing date
2019-07-23
(30) Priority date
2018-08-10
(43) Publication date
2022-02-22
(45) Date of grant
2022-02-22
(51) IPC
G06K 9/62; G06N 20/00; G06T 7/73
(52) CPC
  • G06F Electric digital data processing: 18/2155, 18/217
  • G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 2209/27, 9/6259, 9/6262
  • G06N Computing arrangements based on specific computational models: 20/00, 3/045, 3/0464, 3/08, 3/09
  • G06T Image data processing or generation, in general: 2207/20081, 7/74, 7/75
  • G06V Image or video recognition or understanding: 10/774, 2201/10
(73) Assignee
Apple Inc
(72) Inventors
Melanie S. Subbiah; Jamie R. Lesser; Nicholas E. Apostoloff
(54) Title
Training with simulated images
(57) Abstract

A method that includes obtaining real training samples that include real images that depict real objects, obtaining simulated training samples that include simulated images that depict simulated objects, defining a training dataset that includes at least some of the real training samples and at least some of the simulated training samples, and training a machine learning model to detect subject objects in unannotated input images using the training dataset.

Full text
View on Google Patents

Claims (21)

  1. A method comprising: obtaining real images that depict real objects; determining that a detection failure has occurred in which an object detection system has failed to detect one of the real objects in one of the real images; obtaining failure condition parameters that describe observed conditions from the real image that corresponds to the detection failure; obtaining simulated images that depict simulated objects, by generating a simulated scene using a simulator according to the failure condition parameters and rendering the simulated images of the simulated scene using a rendering engine; defining a training dataset that includes real training samples that are based on the real images and simulated training samples that are based on the simulated images; and training a machine learning model to detect subject objects in unannotated input images using the training dataset.
  2. The method of claim 1, wherein obtaining the real training samples includes capturing the real images in a real-world environment using a camera.
  3. The method of claim 2, wherein generating the simulated scene using the simulator uses three-dimensional models that correspond to the simulated objects and to a simulated environment.
  4. The method of claim 3, wherein the real training samples include annotations indicating locations of the real objects in the real images.
  5. The method of claim 4, wherein the simulated training samples include annotations indicating locations of the simulated objects in the simulated images.
  6. The method of claim 1, wherein: generating the simulated scene using the simulator according to the failure condition parameters comprises initializing a simulation environment in the simulator, and adding multiple groups of simulated objects to a simulated environment according to the failure condition parameters, and rendering the simulated images of the simulated scene using the rendering engine comprises performing multiple iterations of an image generation procedure that includes: determining a location and an orientation with respect to the simulated environment for a virtual camera such that at least one group of simulated objects from the multiple groups of simulated objects is located within a field of view of the virtual camera, positioning the virtual camera with respect to the simulated environment according to the location and orientation, and rendering one of the simulated training images using the virtual camera.
  7. The method of claim 1, further comprising: detecting objects using the trained machine learning model by providing sensor inputs to the trained machine learning model; and controlling operation of a physical system based on the detected objects.
  8. A non-transitory computer-readable storage device including program instructions executable by one or more processors that, when executed, cause the one or more processors to perform operations, the operations comprising: obtaining real images that depict real objects; determining that a detection failure has occurred in which an object detection system has failed to detect one of the real objects in one of the real images; obtaining failure condition parameters that describe observed conditions from the real image that corresponds to the detection failure; obtaining simulated images that depict simulated objects, by generating a simulated scene using a simulator according to the failure condition parameters and rendering the simulated images of the simulated scene using a rendering engine; defining a training dataset that includes real training samples that are based on the real images and simulated training samples that are based on the simulated images; and training a machine learning model to detect subject objects in unannotated input images using the training dataset.
  9. The non-transitory computer-readable storage device of claim 8, wherein obtaining the real training samples includes capturing the real images in a real-world environment using a camera.
  10. The non-transitory computer-readable storage device of claim 9, wherein generating the simulated scene using the simulator uses three-dimensional models that correspond to the simulated objects and to a simulated environment.
  11. The non-transitory computer-readable storage device of claim 10, wherein the real training samples include annotations indicating locations of the real objects in the real images.
  12. The non-transitory computer-readable storage device of claim 11, wherein the simulated training samples include annotations indicating locations of the simulated objects in the simulated images.
  13. The non-transitory computer-readable storage device of claim 8, wherein: generating the simulated scene using the simulator according to the failure condition parameters comprises initializing a simulation environment in the simulator, and adding multiple groups of simulated objects to a simulated environment according to the failure condition parameters, and rendering the simulated images of the simulated scene using the rendering engine comprises performing multiple iterations of an image generation procedure that includes: determining a location and an orientation with respect to the simulated environment for a virtual camera such that at least one group of simulated objects from the multiple groups of simulated objects is located within a field of view of the virtual camera, positioning the virtual camera with respect to the simulated environment according to the location and orientation, and rendering one of the simulated training images using the virtual camera.
  14. The non-transitory computer-readable storage device of claim 8, the operations further comprising: detecting objects using the trained machine learning model by providing sensor inputs to the trained machine learning model; and controlling operation of a physical system based on the detected objects.
  15. An apparatus, comprising: a memory; and one or more processors that are configured to execute instructions that are stored in the memory, wherein the instructions, when executed, cause the one or more processors to: obtain real images that depict real objects, determine that a detection failure has occurred in which an object detection system has failed to detect one of the real objects in one of the real images, obtain failure condition parameters that describe observed conditions from the real image that corresponds to the detection failure, obtain simulated images that depict simulated objects, by generating a simulated scene using a simulator according to the failure condition parameters and rendering the simulated images of the simulated scene using a rendering engine, define a training dataset that includes real training samples that are based on the real images and simulated training samples that are based on the simulated images, and train a machine learning model to detect subject objects in unannotated input images using the training dataset.
  16. The apparatus of claim 15, wherein obtaining the real training samples includes capturing the real images in a real-world environment using a camera.
  17. The apparatus of claim 16, wherein generating the simulated scene using the simulator uses three-dimensional models that correspond to the simulated objects and to a simulated environment.
  18. The apparatus of claim 17, wherein the real training samples include annotations indicating locations of the real objects in the real images.
  19. The apparatus of claim 18, wherein the simulated training samples include annotations indicating locations of the simulated objects in the simulated images.
  20. The apparatus of claim 15, wherein the instructions further cause the one or more processors to: generate the simulated scene using the simulator according to the failure condition parameters comprises initializing a simulation environment in the simulator, and adding multiple groups of simulated objects to a simulated environment according to the failure condition parameters, and render the simulated images of the simulated scene using the rendering engine comprises performing multiple iterations of an image generation procedure that includes: determining a location and an orientation with respect to the simulated environment for a virtual camera such that at least one group of simulated objects from the multiple groups of simulated objects is located within a field of view of the virtual camera, positioning the virtual camera with respect to the simulated environment according to the location and orientation, and rendering one of the simulated training images using the virtual camera.
  21. The apparatus of claim 15, the operations further comprising: detecting objects using the trained machine learning model by providing sensor inputs to the trained machine learning model; and controlling operation of a physical system based on the detected objects.

Description

This disclosure relates to training with simulated images, for example, in robotics and machine learning applications.

Training a machine learning model requires a large training dataset that includes training samples that cover all of the types of situations that the machine learning model is intended to interpret. Because of this, collecting adequate data for training can be time consuming.

Systems and methods for training a machine learning model with simulated images are described herein.

One aspect of the disclosure is a method that includes obtaining real training samples that include real images that depict real objects, obtaining simulated training samples that include simulated images that depict simulated objects, defining a training dataset that includes at least some of the real training samples and at least some of the simulated training samples, and training a machine learning model to detect subject objects in unannotated input images using the training dataset.

Obtaining the real training samples may include capturing the real images in a real-world environment using a camera. Obtaining simulated training samples may include generating a simulated scene that includes a simulation model and subject models that correspond to the simulated objects using a simulator and rendering the simulated images of the simulated scene using a rendering engine. The real training samples may include annotations indicating locations of the real objects in the real images. The simulated training samples may include annotations indicating locations of the simulated objects in the simulated images.

Citations (10)

  • US20150019214A1
  • US20170236013A1
  • WO2018002910A1
  • WO2018071392A1
  • WO2018184187A1
  • US20210201078A1
  • US20180345496A1
  • US20190147582A1
  • US10169678B1
  • US20200050965A1
Record as JSON
{
  "publication_number": "US11256958B1",
  "country": "US",
  "kind": "B1",
  "title": "Training with simulated images",
  "abstract": "A method that includes obtaining real training samples that include real images that depict real objects, obtaining simulated training samples that include simulated images that depict simulated objects, defining a training dataset that includes at least some of the real training samples and at least some of the simulated training samples, and training a machine learning model to detect subject objects in unannotated input images using the training dataset.",
  "claims": [
    "1. A method comprising: obtaining real images that depict real objects; determining that a detection failure has occurred in which an object detection system has failed to detect one of the real objects in one of the real images; obtaining failure condition parameters that describe observed conditions from the real image that corresponds to the detection failure; obtaining simulated images that depict simulated objects, by generating a simulated scene using a simulator according to the failure condition parameters and rendering the simulated images of the simulated scene using a rendering engine; defining a training dataset that includes real training samples that are based on the real images and simulated training samples that are based on the simulated images; and training a machine learning model to detect subject objects in unannotated input images using the training dataset.",
    "2. The method of claim 1, wherein obtaining the real training samples includes capturing the real images in a real-world environment using a camera.",
    "3. The method of claim 2, wherein generating the simulated scene using the simulator uses three-dimensional models that correspond to the simulated objects and to a simulated environment.",
    "4. The method of claim 3, wherein the real training samples include annotations indicating locations of the real objects in the real images.",
    "5. The method of claim 4, wherein the simulated training samples include annotations indicating locations of the simulated objects in the simulated images.",
    "6. The method of claim 1, wherein: generating the simulated scene using the simulator according to the failure condition parameters comprises initializing a simulation environment in the simulator, and adding multiple groups of simulated objects to a simulated environment according to the failure condition parameters, and rendering the simulated images of the simulated scene using the rendering engine comprises performing multiple iterations of an image generation procedure that includes: determining a location and an orientation with respect to the simulated environment for a virtual camera such that at least one group of simulated objects from the multiple groups of simulated objects is located within a field of view of the virtual camera, positioning the virtual camera with respect to the simulated environment according to the location and orientation, and rendering one of the simulated training images using the virtual camera.",
    "7. The method of claim 1, further comprising: detecting objects using the trained machine learning model by providing sensor inputs to the trained machine learning model; and controlling operation of a physical system based on the detected objects.",
    "8. A non-transitory computer-readable storage device including program instructions executable by one or more processors that, when executed, cause the one or more processors to perform operations, the operations comprising: obtaining real images that depict real objects; determining that a detection failure has occurred in which an object detection system has failed to detect one of the real objects in one of the real images; obtaining failure condition parameters that describe observed conditions from the real image that corresponds to the detection failure; obtaining simulated images that depict simulated objects, by generating a simulated scene using a simulator according to the failure condition parameters and rendering the simulated images of the simulated scene using a rendering engine; defining a training dataset that includes real training samples that are based on the real images and simulated training samples that are based on the simulated images; and training a machine learning model to detect subject objects in unannotated input images using the training dataset.",
    "9. The non-transitory computer-readable storage device of claim 8, wherein obtaining the real training samples includes capturing the real images in a real-world environment using a camera.",
    "10. The non-transitory computer-readable storage device of claim 9, wherein generating the simulated scene using the simulator uses three-dimensional models that correspond to the simulated objects and to a simulated environment.",
    "11. The non-transitory computer-readable storage device of claim 10, wherein the real training samples include annotations indicating locations of the real objects in the real images.",
    "12. The non-transitory computer-readable storage device of claim 11, wherein the simulated training samples include annotations indicating locations of the simulated objects in the simulated images.",
    "13. The non-transitory computer-readable storage device of claim 8, wherein: generating the simulated scene using the simulator according to the failure condition parameters comprises initializing a simulation environment in the simulator, and adding multiple groups of simulated objects to a simulated environment according to the failure condition parameters, and rendering the simulated images of the simulated scene using the rendering engine comprises performing multiple iterations of an image generation procedure that includes: determining a location and an orientation with respect to the simulated environment for a virtual camera such that at least one group of simulated objects from the multiple groups of simulated objects is located within a field of view of the virtual camera, positioning the virtual camera with respect to the simulated environment according to the location and orientation, and rendering one of the simulated training images using the virtual camera.",
    "14. The non-transitory computer-readable storage device of claim 8, the operations further comprising: detecting objects using the trained machine learning model by providing sensor inputs to the trained machine learning model; and controlling operation of a physical system based on the detected objects.",
    "15. An apparatus, comprising: a memory; and one or more processors that are configured to execute instructions that are stored in the memory, wherein the instructions, when executed, cause the one or more processors to: obtain real images that depict real objects, determine that a detection failure has occurred in which an object detection system has failed to detect one of the real objects in one of the real images, obtain failure condition parameters that describe observed conditions from the real image that corresponds to the detection failure, obtain simulated images that depict simulated objects, by generating a simulated scene using a simulator according to the failure condition parameters and rendering the simulated images of the simulated scene using a rendering engine, define a training dataset that includes real training samples that are based on the real images and simulated training samples that are based on the simulated images, and train a machine learning model to detect subject objects in unannotated input images using the training dataset.",
    "16. The apparatus of claim 15, wherein obtaining the real training samples includes capturing the real images in a real-world environment using a camera.",
    "17. The apparatus of claim 16, wherein generating the simulated scene using the simulator uses three-dimensional models that correspond to the simulated objects and to a simulated environment.",
    "18. The apparatus of claim 17, wherein the real training samples include annotations indicating locations of the real objects in the real images.",
    "19. The apparatus of claim 18, wherein the simulated training samples include annotations indicating locations of the simulated objects in the simulated images.",
    "20. The apparatus of claim 15, wherein the instructions further cause the one or more processors to: generate the simulated scene using the simulator according to the failure condition parameters comprises initializing a simulation environment in the simulator, and adding multiple groups of simulated objects to a simulated environment according to the failure condition parameters, and render the simulated images of the simulated scene using the rendering engine comprises performing multiple iterations of an image generation procedure that includes: determining a location and an orientation with respect to the simulated environment for a virtual camera such that at least one group of simulated objects from the multiple groups of simulated objects is located within a field of view of the virtual camera, positioning the virtual camera with respect to the simulated environment according to the location and orientation, and rendering one of the simulated training images using the virtual camera.",
    "21. The apparatus of claim 15, the operations further comprising: detecting objects using the trained machine learning model by providing sensor inputs to the trained machine learning model; and controlling operation of a physical system based on the detected objects."
  ],
  "description_excerpt": "This disclosure relates to training with simulated images, for example, in robotics and machine learning applications.\n\nTraining a machine learning model requires a large training dataset that includes training samples that cover all of the types of situations that the machine learning model is intended to interpret. Because of this, collecting adequate data for training can be time consuming.\n\nSystems and methods for training a machine learning model with simulated images are described herein.\n\nOne aspect of the disclosure is a method that includes obtaining real training samples that include real images that depict real objects, obtaining simulated training samples that include simulated images that depict simulated objects, defining a training dataset that includes at least some of the real training samples and at least some of the simulated training samples, and training a machine learning model to detect subject objects in unannotated input images using the training dataset.\n\nObtaining the real training samples may include capturing the real images in a real-world environment using a camera. Obtaining simulated training samples may include generating a simulated scene that includes a simulation model and subject models that correspond to the simulated objects using a simulator and rendering the simulated images of the simulated scene using a rendering engine. The real training samples may include annotations indicating locations of the real objects in the real images. The simulated training samples may include annotations indicating locations of the simulated objects in the simulated images.",
  "cpc": [
    "G06F 18/2155",
    "G06F 18/217",
    "G06K 2209/27",
    "G06K 9/6259",
    "G06K 9/6262",
    "G06N 20/00",
    "G06N 3/045",
    "G06N 3/0464",
    "G06N 3/08",
    "G06N 3/09",
    "G06T 2207/20081",
    "G06T 7/74",
    "G06T 7/75",
    "G06V 10/774",
    "G06V 2201/10"
  ],
  "ipc": [
    "G06K 9/62",
    "G06N 20/00",
    "G06T 7/73"
  ],
  "assignees": [
    "Apple Inc"
  ],
  "inventors": [
    "Melanie S. Subbiah",
    "Jamie R. Lesser",
    "Nicholas E. Apostoloff"
  ],
  "filing_date": "2019-07-23",
  "publication_date": "2022-02-22",
  "grant_date": "2022-02-22",
  "priority_date": "2018-08-10",
  "application_number": "US-201916519686-A",
  "family_id": "80321962",
  "cited_by_count": 34,
  "citations": [
    "US20150019214A1",
    "US20170236013A1",
    "WO2018002910A1",
    "WO2018071392A1",
    "WO2018184187A1",
    "US20210201078A1",
    "US20180345496A1",
    "US20190147582A1",
    "US10169678B1",
    "US20200050965A1"
  ]
}

Record 1,267 of 8,000 in Patents full text (MLC-0201). Request the full dataset.