MLchartDataset catalogue

Patent · US11412133B1 · B1 · US

Autonomously motile device with computer vision

(11) Publication number
US11412133B1
(21) Application number
16/913,498
(22) Filing date
2020-06-26
(30) Priority date
2020-06-26
(43) Publication date
2022-08-09
(45) Date of grant
2022-08-09
(51) IPC
B25J 19/02; B25J 9/16; G05D 1/00; G05D 1/02; G06T 7/246; H04N 5/232; H04N 5/262
(52) CPC
  • H04N Pictorial communication, e.g. television: 23/611, 23/695, 5/23219, 5/23299, 5/2628
  • B25J Manipulators; chambers provided with manipulation devices: 19/023, 9/1697
  • G05D Systems for controlling or regulating non-electric variables: 1/0088, 1/0094, 1/0231, 1/0246
  • G06T Image data processing or generation, in general: 2207/10016, 2207/20081, 2207/20084, 2207/30201, 2207/30244, 7/251, 7/337
(73) Assignee
AMAZON TECH INC
(72) Inventors
Morton Tarun Yohann; KUO CHENG-HAO; THOMAS JIM OOMMEN; ZHOU NING
(54) Title
Autonomously motile device with computer vision
(57) Abstract

A device capable of autonomous motion may process image data determined by one or more cameras to determine one or more properties of objects represented in the image data. The device may determine that two or more computer vision components correspond to a particular property. A first computer vision component may process the image data to determine first output data, and the second computer vision component may process the first output data to determine second output data corresponding to the property.

Full text
View on Google Patents

Claims (18)

  1. A computer-implemented method comprising: receiving, at an autonomously motile device, a command instructing the autonomously motile device to follow a user; determining that processing of the command to follow the user depends upon output from an object-detection component, an object-recognition component, and a position-estimation component; determining first image data using at least a first camera of the autonomously motile device; determining, using the first image data and the object-detection component, a face represented in a first portion of the first image data; determining, using the first portion of the first image data and the object-recognition component, an identifier corresponding to the face; determining that the identifier corresponds to the user; determining, using the first portion of the first image data and the position-estimation component, a first position of the autonomously motile device with respect to the user; and causing the autonomously motile device to move to a second position with respect to the user, wherein the second position is closer to the user than the first position.
  2. The computer-implemented method of claim 1, further comprising, prior to causing the autonomously motile device to move to the second position: determining, using the first image data and the object-recognition component, a body represented in a second portion of the first image data; associating the body with the identifier; determining second image data using at least the first camera of the autonomously motile device; determining, using the second image data and the object-detection component, that the face is not represented in the second image data; determining, using the second image data and the object-detection component, that the body is represented in a first portion of the second image data; determining, using the first portion of the second image data and the position-estimation component, a third position of the autonomously motile device with respect to the user; and determining that the third position is farther from the user than the first position, wherein causing the autonomously motile device to move to the second position is based at least in part on the third position being farther from the user than the first position.
  3. A computer-implemented method comprising: determining, at a device, command data corresponding to a property of a first object in an environment; determining that the command data corresponds to a first computer vision component and a second computer vision component, the first computer vision component corresponding to a first type of computer vision, the second computer vision component corresponding to a second type of computer vision different from the first type of computer vision; determining that an input to the second computer vision component corresponds to an output of the first computer vision component; determining first image data including a first representation of the first object; determining, using the first computer vision component and the first image data, first output data; determining that the first output data corresponds to the first object and a second object; processing, using the second computer vision component, a portion of the first output data corresponding to the first object to determine second output data representing the property; and based at least in part on the second output data, causing the device to perform an action.
  4. The computer-implemented method of claim 3, further comprising: determining, using a first camera of the device during a first time period, second image data including a second representation of the first object; determining, using a second camera of the device during the first time period, third image data including a third representation of the first object; determining, based on the second image data and the third image data, position data corresponding to a position of the first object in the environment; and determining, using a third computer vision component and the position data, third output data representing a position of the first object.
  5. The computer-implemented method of claim 3, further comprising: after determining the first output data, storing, in a storage associated with the device, computer vision data comprising the first image data, the first output data, and a time associated with the first image data; and prior to determining the second output data, determining that the storage includes the computer vision data.
  6. The computer-implemented method of claim 3, further comprising: determining that the first computer vision component corresponds to a first processing resource; determining that the second computer vision component corresponds to a second processing resource; determining that a third computer vision component corresponds to the first processing resource; and after determining the first output data and while determining at least a portion of the second output data, determining, using the third computer vision component and the first image data, third output data.
  7. The computer-implemented method of claim 3, further comprising: determining a time corresponding to the first image data; receiving, from a sensor, sensor data representing a second property of the environment; determining that the sensor data corresponds to the time; and determining third output data using a third computer vision component, the first image data, and the sensor data.
  8. The computer-implemented method of claim 3, further comprising: after determining the first image data, determining second image data including a second representation of the first object; and while determining at least a portion of the first output data, determining, using the first computer vision component and the second image data, third output data.
  9. The computer-implemented method of claim 3, wherein the first output data corresponds to a first position of the first representation, and the computer-implemented method further comprises: after determining the first image data, determining second image data including a second representation of the first object; determining, using the first computer vision component and the second image data, third output data corresponding to a second position of the second representation; determining a difference between the first position and the second position; and causing a camera of the device to move in accordance with the difference.
  10. The computer-implemented method of claim 3, further comprising: receiving, from a camera of the device, camera data having a first size; determining, using an image resize component and the camera data, resized camera data having a second size different from the first size; and determining, using an image rectification component and the resized camera data, rectified camera data corresponding to a shape of a lens of the camera.
  11. A device comprising: at least one processor; and at least one memory including instructions that, when executed by the at least one processor, cause the device to: determine, at the device, command data corresponding to a property of a first object in an environment; determine that the command data corresponds to a first computer vision component and a second computer vision component, the first computer vision component corresponding to a first type of computer vision, the second computer vision component corresponding to a second type of computer vision different from the first type of computer vision; determine that an input to the second computer vision component corresponds to an output of the first computer vision component; determine first image data including a first representation of the first object; determine, using the first computer vision component and the first image data, first output data; determine that the first output data corresponds to the first object and a second object; process, using the second computer vision component, a portion of the first output data corresponding to the first object to determine second output data representing the property; and based at least in part on the second output data, cause the device to perform an action.
  12. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: determine, using a first camera of the device during a first time period, second image data including a second representation of the first object; determine, using a second camera of the device during the first time period, third image data including a third representation of the first object; determine, based on the second image data and the third image data, position data corresponding to a position of the first object in the environment; and determine, using a third computer vision component and the position data, third output data representing the position of the first object.
  13. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: after determining the first output data, store, in a storage associated with the device, computer vision data comprising the first image data, the first output data, and a time associated with the first image data; and prior to determining the second output data, determining that the storage includes the computer vision data.
  14. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: determine that the first computer vision component corresponds to a first processing resource; determine that the second computer vision component corresponds to a second processing resource; determine that a third computer vision component corresponds to the first processing resource; and after determining the first output data and while determining at least a portion of the second output data, determine, using the third computer vision component and the first image data, third output data.
  15. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: determine a time corresponding to the first image data; receive, from a sensor, sensor data representing a second property of the environment; determine that the sensor data corresponds to the time; and determine third output data using a third computer vision component, the first image data, and the sensor data.
  16. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: after determining the first image data, determine second image data including a second representation of the first object; and while determining at least a portion of the first output data, determine, using the first computer vision component and the second image data, third output data.
  17. The device of claim 11, wherein the first output data corresponds to a first position of the first representation, and wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: after determining the first image data, determine second image data including a second representation of the first object; determine, using the first computer vision component and the second image data, third output data corresponding to a second position of the second representation; determine a difference between the first position and the second position; and cause a camera of the device to move in accordance with the difference.
  18. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: receive, from a camera of the device, camera data having a first size; determine, using an image resize component and the camera data, resized camera data having a second size different from the first size; and determine, using an image rectification component and the resized camera data, rectified camera data corresponding to a shape of a lens of the camera.

Description

A computing device may be an autonomously motile device and may include at least one camera for capturing images, which may include representations of objects, in an environment of the computing device. Techniques may be used to process image data received from a camera to determine one or more properties of the object. The device may perform further actions based on the determined properties, such as causing movement of the computing device and/or camera of the computing device.

For a more complete understanding of the present disclosure, reference is now made to the following description taken in conjunction with the accompanying drawings. FIG. 1 illustrate as system and method for processing images according to embodiments of the present disclosure. FIGS. 2A, 2B, and 2C illustrate views of an autonomously motile device according to embodiments of the present disclosure. FIG. 3 illustrates a view of an autonomously motile device in an environment according to embodiments of the present disclosure. FIGS. 4A, 4B, and 4C illustrate images captured by an autonomously motile device in an environment according to embodiments of the present disclosure. FIGS. 5A and 5B illustrate components for image processing by an autonomously motile device according to embodiments of the present disclosure. FIGS. 6A, 6B, and 6C illustrate computer vision components according to embodiments of the present disclosure. FIGS. 7A and 7B illustrate processing stages of computer vision components according to embodiments of the present disclosure.

Citations (13)

  • US10121494B1
  • US10427306B1
  • US10471611B2
  • US10803667B1
  • US10955860B2
  • US2002044691A1
  • US2012134579A1
  • US2015125032A1
  • US2015253428A1
  • US2016004909A1
  • US2019291277A1
  • US2021193116A1
  • US9876993B2
Record as JSON
{
  "publication_number": "US11412133B1",
  "country": "US",
  "kind": "B1",
  "title": "Autonomously motile device with computer vision",
  "abstract": "A device capable of autonomous motion may process image data determined by one or more cameras to determine one or more properties of objects represented in the image data. The device may determine that two or more computer vision components correspond to a particular property. A first computer vision component may process the image data to determine first output data, and the second computer vision component may process the first output data to determine second output data corresponding to the property.",
  "claims": [
    "1. A computer-implemented method comprising: receiving, at an autonomously motile device, a command instructing the autonomously motile device to follow a user; determining that processing of the command to follow the user depends upon output from an object-detection component, an object-recognition component, and a position-estimation component; determining first image data using at least a first camera of the autonomously motile device; determining, using the first image data and the object-detection component, a face represented in a first portion of the first image data; determining, using the first portion of the first image data and the object-recognition component, an identifier corresponding to the face; determining that the identifier corresponds to the user; determining, using the first portion of the first image data and the position-estimation component, a first position of the autonomously motile device with respect to the user; and causing the autonomously motile device to move to a second position with respect to the user, wherein the second position is closer to the user than the first position.",
    "2. The computer-implemented method of claim 1, further comprising, prior to causing the autonomously motile device to move to the second position: determining, using the first image data and the object-recognition component, a body represented in a second portion of the first image data; associating the body with the identifier; determining second image data using at least the first camera of the autonomously motile device; determining, using the second image data and the object-detection component, that the face is not represented in the second image data; determining, using the second image data and the object-detection component, that the body is represented in a first portion of the second image data; determining, using the first portion of the second image data and the position-estimation component, a third position of the autonomously motile device with respect to the user; and determining that the third position is farther from the user than the first position, wherein causing the autonomously motile device to move to the second position is based at least in part on the third position being farther from the user than the first position.",
    "3. A computer-implemented method comprising: determining, at a device, command data corresponding to a property of a first object in an environment; determining that the command data corresponds to a first computer vision component and a second computer vision component, the first computer vision component corresponding to a first type of computer vision, the second computer vision component corresponding to a second type of computer vision different from the first type of computer vision; determining that an input to the second computer vision component corresponds to an output of the first computer vision component; determining first image data including a first representation of the first object; determining, using the first computer vision component and the first image data, first output data; determining that the first output data corresponds to the first object and a second object; processing, using the second computer vision component, a portion of the first output data corresponding to the first object to determine second output data representing the property; and based at least in part on the second output data, causing the device to perform an action.",
    "4. The computer-implemented method of claim 3, further comprising: determining, using a first camera of the device during a first time period, second image data including a second representation of the first object; determining, using a second camera of the device during the first time period, third image data including a third representation of the first object; determining, based on the second image data and the third image data, position data corresponding to a position of the first object in the environment; and determining, using a third computer vision component and the position data, third output data representing a position of the first object.",
    "5. The computer-implemented method of claim 3, further comprising: after determining the first output data, storing, in a storage associated with the device, computer vision data comprising the first image data, the first output data, and a time associated with the first image data; and prior to determining the second output data, determining that the storage includes the computer vision data.",
    "6. The computer-implemented method of claim 3, further comprising: determining that the first computer vision component corresponds to a first processing resource; determining that the second computer vision component corresponds to a second processing resource; determining that a third computer vision component corresponds to the first processing resource; and after determining the first output data and while determining at least a portion of the second output data, determining, using the third computer vision component and the first image data, third output data.",
    "7. The computer-implemented method of claim 3, further comprising: determining a time corresponding to the first image data; receiving, from a sensor, sensor data representing a second property of the environment; determining that the sensor data corresponds to the time; and determining third output data using a third computer vision component, the first image data, and the sensor data.",
    "8. The computer-implemented method of claim 3, further comprising: after determining the first image data, determining second image data including a second representation of the first object; and while determining at least a portion of the first output data, determining, using the first computer vision component and the second image data, third output data.",
    "9. The computer-implemented method of claim 3, wherein the first output data corresponds to a first position of the first representation, and the computer-implemented method further comprises: after determining the first image data, determining second image data including a second representation of the first object; determining, using the first computer vision component and the second image data, third output data corresponding to a second position of the second representation; determining a difference between the first position and the second position; and causing a camera of the device to move in accordance with the difference.",
    "10. The computer-implemented method of claim 3, further comprising: receiving, from a camera of the device, camera data having a first size; determining, using an image resize component and the camera data, resized camera data having a second size different from the first size; and determining, using an image rectification component and the resized camera data, rectified camera data corresponding to a shape of a lens of the camera.",
    "11. A device comprising: at least one processor; and at least one memory including instructions that, when executed by the at least one processor, cause the device to: determine, at the device, command data corresponding to a property of a first object in an environment; determine that the command data corresponds to a first computer vision component and a second computer vision component, the first computer vision component corresponding to a first type of computer vision, the second computer vision component corresponding to a second type of computer vision different from the first type of computer vision; determine that an input to the second computer vision component corresponds to an output of the first computer vision component; determine first image data including a first representation of the first object; determine, using the first computer vision component and the first image data, first output data; determine that the first output data corresponds to the first object and a second object; process, using the second computer vision component, a portion of the first output data corresponding to the first object to determine second output data representing the property; and based at least in part on the second output data, cause the device to perform an action.",
    "12. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: determine, using a first camera of the device during a first time period, second image data including a second representation of the first object; determine, using a second camera of the device during the first time period, third image data including a third representation of the first object; determine, based on the second image data and the third image data, position data corresponding to a position of the first object in the environment; and determine, using a third computer vision component and the position data, third output data representing the position of the first object.",
    "13. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: after determining the first output data, store, in a storage associated with the device, computer vision data comprising the first image data, the first output data, and a time associated with the first image data; and prior to determining the second output data, determining that the storage includes the computer vision data.",
    "14. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: determine that the first computer vision component corresponds to a first processing resource; determine that the second computer vision component corresponds to a second processing resource; determine that a third computer vision component corresponds to the first processing resource; and after determining the first output data and while determining at least a portion of the second output data, determine, using the third computer vision component and the first image data, third output data.",
    "15. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: determine a time corresponding to the first image data; receive, from a sensor, sensor data representing a second property of the environment; determine that the sensor data corresponds to the time; and determine third output data using a third computer vision component, the first image data, and the sensor data.",
    "16. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: after determining the first image data, determine second image data including a second representation of the first object; and while determining at least a portion of the first output data, determine, using the first computer vision component and the second image data, third output data.",
    "17. The device of claim 11, wherein the first output data corresponds to a first position of the first representation, and wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: after determining the first image data, determine second image data including a second representation of the first object; determine, using the first computer vision component and the second image data, third output data corresponding to a second position of the second representation; determine a difference between the first position and the second position; and cause a camera of the device to move in accordance with the difference.",
    "18. The device of claim 11, wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the device to: receive, from a camera of the device, camera data having a first size; determine, using an image resize component and the camera data, resized camera data having a second size different from the first size; and determine, using an image rectification component and the resized camera data, rectified camera data corresponding to a shape of a lens of the camera."
  ],
  "description_excerpt": "A computing device may be an autonomously motile device and may include at least one camera for capturing images, which may include representations of objects, in an environment of the computing device. Techniques may be used to process image data received from a camera to determine one or more properties of the object. The device may perform further actions based on the determined properties, such as causing movement of the computing device and/or camera of the computing device.\n\nFor a more complete understanding of the present disclosure, reference is now made to the following description taken in conjunction with the accompanying drawings. FIG. 1 illustrate as system and method for processing images according to embodiments of the present disclosure. FIGS. 2A, 2B, and 2C illustrate views of an autonomously motile device according to embodiments of the present disclosure. FIG. 3 illustrates a view of an autonomously motile device in an environment according to embodiments of the present disclosure. FIGS. 4A, 4B, and 4C illustrate images captured by an autonomously motile device in an environment according to embodiments of the present disclosure. FIGS. 5A and 5B illustrate components for image processing by an autonomously motile device according to embodiments of the present disclosure. FIGS. 6A, 6B, and 6C illustrate computer vision components according to embodiments of the present disclosure. FIGS. 7A and 7B illustrate processing stages of computer vision components according to embodiments of the present disclosure.",
  "cpc": [
    "H04N 23/611",
    "B25J 19/023",
    "B25J 9/1697",
    "G05D 1/0088",
    "G05D 1/0094",
    "G05D 1/0231",
    "G05D 1/0246",
    "G06T 2207/10016",
    "G06T 2207/20081",
    "G06T 2207/20084",
    "G06T 2207/30201",
    "G06T 2207/30244",
    "G06T 7/251",
    "G06T 7/337",
    "H04N 23/695",
    "H04N 5/23219",
    "H04N 5/23299",
    "H04N 5/2628"
  ],
  "ipc": [
    "B25J 19/02",
    "B25J 9/16",
    "G05D 1/00",
    "G05D 1/02",
    "G06T 7/246",
    "H04N 5/232",
    "H04N 5/262"
  ],
  "assignees": [
    "AMAZON TECH INC"
  ],
  "inventors": [
    "Morton Tarun Yohann",
    "KUO CHENG-HAO",
    "THOMAS JIM OOMMEN",
    "ZHOU NING"
  ],
  "filing_date": "2020-06-26",
  "publication_date": "2022-08-09",
  "grant_date": "2022-08-09",
  "priority_date": "2020-06-26",
  "application_number": "US-202016913498-A",
  "family_id": "82706089",
  "citations": [
    "US10121494B1",
    "US10427306B1",
    "US10471611B2",
    "US10803667B1",
    "US10955860B2",
    "US2002044691A1",
    "US2012134579A1",
    "US2015125032A1",
    "US2015253428A1",
    "US2016004909A1",
    "US2019291277A1",
    "US2021193116A1",
    "US9876993B2"
  ]
}

Record 441 of 5,000 in Patents full text (MLC-0201). Request the full dataset.