MLchartDataset catalogue

Patent · US10427306B1 · B1 · US

Multimodal object identification

(11) Publication number
US10427306B1
(21) Application number
15/643,138
(22) Filing date
2017-07-06
(30) Priority date
2017-07-06
(43) Publication date
2019-10-01
(45) Date of grant
2019-10-01
(51) IPC
B25J 19/02; B25J 9/16; G05B 19/048
(52) CPC
  • B25J Manipulators; chambers provided with manipulation devices: 9/1697, 11/0005, 11/008, 13/003, 19/026
  • G05B Control or regulating systems in general; functional elements of such systems; monitoring or testing arrangements for such systems or elements: 19/048, 2219/23021, 2219/35444
  • G06F Electric digital data processing: 3/016, 3/017, 3/167
(73) Assignee
X DEV LLC
(72) Inventors
QUINLAN MICHAEL JOSEPH; COHEN GABRIEL A
(54) Title
Multimodal object identification
(57) Abstract

Methods, systems, and apparatus for receiving a command for controlling a robot, the command referencing an object, receiving sensor data for a portion of an environment of the robot, identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data, accessing map data indicating locations of objects within a space, searching the map data for the object, wherein the search of the map data is restricted to the spatial region, determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region, and in response to determining that the object referenced in the command is present in the spatial region, controlling the robot to perform an action with respect to the object referenced in the command.

Full text
View on Google Patents

Claims (20)

  1. A computer-implemented method comprising: receiving a command for controlling a robot, the command referencing an object; receiving sensor data for a portion of an environment of the robot, the sensor data being captured by a sensor of the robot; identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data; in response to identifying the gesture, accessing map data indicating locations of objects within a space, the map data being generated before receiving the command; searching the map data for the object referenced in the command, wherein the search of the map data is restricted, based on the identified gesture, to the spatial region indicated by the gesture; determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region indicated by the gesture; and in response to determining that the object referenced in the command is present in the spatial region indicated by the gesture, controlling the robot to perform an action with respect to the object referenced in the command.
  2. The computer-implemented method of claim 1, wherein the identified gesture is one of an arm wave, a hand gesture, or a glance.
  3. The computer-implemented method of claim 1, wherein receiving the command for controlling the robot, the command referencing the object comprises: receiving audio data captured by a microphone of the robot that corresponds to the command; identifying, based on performing speech recognition on the audio data, one or more candidate objects that each correspond to a respective candidate transcription of at least a portion of the audio data; accessing an inventory of one or more objects within the space; and identifying the object from among the inventory of the one or more objects within the space based at least on comparing each candidate transcription of at least a portion of the audio data to the one or more objects of the inventory.
  4. The computer-implemented method of claim 1, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: determining a location of the robot within the space; and determining the spatial region based at least on the gesture of the human and the location of the robot within the space.
  5. The computer-implemented method of claim 1, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: determining an orientation of the robot within the space when receiving the sensor data; and determining the spatial region based at least on the gesture of the human and the orientation of the robot within the space.
  6. The computer-implemented method of claim 1, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: detecting one or more predetermined shapes from the sensor data, each of the one or more predetermined shapes corresponding to a gesture of a human; determining one or more locations of the detected one or more predetermined shapes within the sensor data; and determining the spatial region based at least on the one or more predetermined shapes and the one or more locations of the detected one or more predetermined shapes within the sensor data.
  7. The computer-implemented method of claim 1, comprising: receiving a second command for controlling the robot, the command referencing a second object; receiving second sensor data for a portion of the environment of the robot, the second sensor data being captured by the sensor of the robot; identifying, from the second sensor data, a second gesture of a human that indicates a second spatial region located outside of the portion of the environment described by the second sensor data; searching the map data for the second object referenced in the second command, wherein the search of the map data is restricted, based on the identified second gesture, to the second spatial region indicated by the second gesture; determining, based at least on searching the map data for the second object referenced in the second command, that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture; based at least on determining that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture, searching the map data for the second object referenced in the second command, wherein the search of the map is restricted to a third spatial region; determining, based at least on searching the map data for the second object referenced in the second command, that the second object referenced in the second command is present in the third spatial region; and in response to determining that the second object referenced in the second command is present in the third spatial region, controlling the robot to perform a second action with respect to the second object referenced in the second command.
  8. The computer-implemented method of claim 7, wherein the third spatial region is larger than the second spatial region indicated by the second gesture.
  9. The computer-implemented method of claim 1, comprising: receiving a second command for controlling the robot, the command referencing a second object; receiving second sensor data for a portion of the environment of the robot, the second sensor data being captured by the sensor of the robot; identifying, from the second sensor data, a second gesture of a human that indicates a second spatial region located outside of the portion of the environment described by the second sensor data; searching the map data for the second object referenced in the second command, wherein the search of the map data is restricted, based on the identified second gesture, to the second spatial region indicated by the second gesture; determining, based at least on searching the map data for the second object referenced in the second command, that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture; and based at least on determining that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture, controlling the robot to indicate that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture.
  10. The computer-implemented method of claim 1, wherein controlling the robot to perform the action with respect to the object referenced in the command comprises controlling the robot to perform one of retrieving the object referenced in the command, move the object referenced in the command to a predetermined location, move the object referenced in the command to a location indicated by the command, or navigate to the location of the object referenced in the command.
  11. The computer-implemented method of claim 1, comprising: determining a location of the object referenced in the command within the spatial region, wherein the location of the object referenced in the command within the spatial region is represented by a set of coordinates within the space.
  12. The computer-implemented method of claim 1, wherein the sensor data is one of image data, infrared image data, light detection and ranging (LIDAR) data, thermal image data, night vision image data, or motion data.
  13. The computer-implemented method of claim 1, wherein the sensor data for the portion of the environment of the robot is image data for a field of view of a camera of the robot, the image data being captured by the camera of the robot.
  14. A system comprising: one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: receiving a command for controlling a robot, the command referencing an object; receiving sensor data for a portion of an environment of the robot, the sensor data being captured by a sensor of the robot; identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data; in response to identifying the gesture, accessing map data indicating locations of objects within a space, the map data being generated before receiving the command; searching the map data for the object referenced in the command, wherein the search of the map data is restricted, based on the identified gesture, to the spatial region indicated by the gesture; determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region indicated by the gesture; and in response to determining that the object referenced in the command is present in the spatial region indicated by the gesture, controlling the robot to perform an action with respect to the object referenced in the command.
  15. The system of claim 14, wherein the identified gesture is one of an arm wave, a hand gesture, or a glance.
  16. The system of claim 14, wherein receiving the command for controlling the robot, the command referencing the object comprises: receiving audio data captured by a microphone of the robot that corresponds to the command; identifying, based on performing speech recognition on the audio data, one or more candidate objects that each correspond to a respective candidate transcription of at least a portion of the audio data; accessing an inventory of one or more objects within the space; and identifying the object from among the inventory of the one or more objects within the space based at least on comparing each candidate transcription of at least a portion of the audio data to the one or more objects of the inventory.
  17. The system of claim 14, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: determining a location of the robot within the space; and determining the spatial region based at least on the gesture of the human and the location of the robot within the space.
  18. The system of claim 14, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: determining an orientation of the robot within the space when receiving the sensor data; and determining the spatial region based at least on the gesture of the human and the orientation of the robot within the space.
  19. The system of claim 14, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: detecting one or more predetermined shapes from the sensor data, each of the one or more predetermined shapes corresponding to a gesture of a human; determining one or more locations of the detected one or more predetermined shapes within the sensor data; and determining the spatial region based at least on the one or more predetermined shapes and the one or more locations of the detected one or more predetermined shapes within the sensor data.
  20. A non-transitory computer-readable storage device storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising: receiving a command for controlling a robot, the command referencing an object; receiving sensor data for a portion of an environment of the robot, the sensor data being captured by a sensor of the robot; identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data; in response to identifying the gesture, accessing map data indicating locations of objects within a space, the map data being generated before receiving the command; searching the map data for the object referenced in the command, wherein the search of the map data is restricted, based on the identified gesture, to the spatial region indicated by the gesture; determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region indicated by the gesture; and in response to determining that the object referenced in the command is present in the spatial region indicated by the gesture, controlling the robot to perform an action with respect to the object referenced in the command.

Description

This specification relates to object identification, and one particular implementation relates to identifying objects based on human-robot interactions.

Human-computer interfaces that permit users to provide natural language or gestural inputs are becoming exceedingly pervasive. For example, a personal assistant application can receiving human speech and identify a command based on an analysis of that speech. The personal assistant application can perform or trigger operations in response to the identified command. Similarly, computer applications may receive images or video of a user and can detect human gestures from the images or video. The computer can interpret those gestures as commands and may perform or trigger operations responsive to the identified commands. These techniques are also being applied in the field of robotics to enable human-robot interactions. For example, users may be able to provide gestures or speech inputs to a robot to command the robot to perform specific actions. In some examples, a user may refer to a particular object in a command, either by word or by gesture. In response to such a command, the robot may be required to identify a physical object in its environment that corresponds to the particular object referenced in the command. Challenges may arise where multiple objects in the robot's environment correspond to the particular object referenced in the command. In those instances, the robot may be required to disambiguate between the multiple objects, to identify a particular one of the multiple objects that the user likely intended to reference in their command.

Citations (20)

  • CN104851211A
  • US2007192910A1
  • US2011054691A1
  • US2011118877A1
  • US2012173018A1
  • US2012185090A1
  • US2012197439A1
  • US2012316679A1
  • US2013325244A1
  • US2015025681A1
  • US2015032252A1
  • US2015283702A1
  • US2015362988A1
  • US2015378444A1
  • US2016140963A1
  • US2018192845A1
  • US7881936B2
  • US8140188B2
  • US9165566B2
  • US9427863B2
Record as JSON
{
  "publication_number": "US10427306B1",
  "country": "US",
  "kind": "B1",
  "title": "Multimodal object identification",
  "abstract": "Methods, systems, and apparatus for receiving a command for controlling a robot, the command referencing an object, receiving sensor data for a portion of an environment of the robot, identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data, accessing map data indicating locations of objects within a space, searching the map data for the object, wherein the search of the map data is restricted to the spatial region, determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region, and in response to determining that the object referenced in the command is present in the spatial region, controlling the robot to perform an action with respect to the object referenced in the command.",
  "claims": [
    "1. A computer-implemented method comprising: receiving a command for controlling a robot, the command referencing an object; receiving sensor data for a portion of an environment of the robot, the sensor data being captured by a sensor of the robot; identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data; in response to identifying the gesture, accessing map data indicating locations of objects within a space, the map data being generated before receiving the command; searching the map data for the object referenced in the command, wherein the search of the map data is restricted, based on the identified gesture, to the spatial region indicated by the gesture; determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region indicated by the gesture; and in response to determining that the object referenced in the command is present in the spatial region indicated by the gesture, controlling the robot to perform an action with respect to the object referenced in the command.",
    "2. The computer-implemented method of claim 1, wherein the identified gesture is one of an arm wave, a hand gesture, or a glance.",
    "3. The computer-implemented method of claim 1, wherein receiving the command for controlling the robot, the command referencing the object comprises: receiving audio data captured by a microphone of the robot that corresponds to the command; identifying, based on performing speech recognition on the audio data, one or more candidate objects that each correspond to a respective candidate transcription of at least a portion of the audio data; accessing an inventory of one or more objects within the space; and identifying the object from among the inventory of the one or more objects within the space based at least on comparing each candidate transcription of at least a portion of the audio data to the one or more objects of the inventory.",
    "4. The computer-implemented method of claim 1, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: determining a location of the robot within the space; and determining the spatial region based at least on the gesture of the human and the location of the robot within the space.",
    "5. The computer-implemented method of claim 1, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: determining an orientation of the robot within the space when receiving the sensor data; and determining the spatial region based at least on the gesture of the human and the orientation of the robot within the space.",
    "6. The computer-implemented method of claim 1, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: detecting one or more predetermined shapes from the sensor data, each of the one or more predetermined shapes corresponding to a gesture of a human; determining one or more locations of the detected one or more predetermined shapes within the sensor data; and determining the spatial region based at least on the one or more predetermined shapes and the one or more locations of the detected one or more predetermined shapes within the sensor data.",
    "7. The computer-implemented method of claim 1, comprising: receiving a second command for controlling the robot, the command referencing a second object; receiving second sensor data for a portion of the environment of the robot, the second sensor data being captured by the sensor of the robot; identifying, from the second sensor data, a second gesture of a human that indicates a second spatial region located outside of the portion of the environment described by the second sensor data; searching the map data for the second object referenced in the second command, wherein the search of the map data is restricted, based on the identified second gesture, to the second spatial region indicated by the second gesture; determining, based at least on searching the map data for the second object referenced in the second command, that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture; based at least on determining that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture, searching the map data for the second object referenced in the second command, wherein the search of the map is restricted to a third spatial region; determining, based at least on searching the map data for the second object referenced in the second command, that the second object referenced in the second command is present in the third spatial region; and in response to determining that the second object referenced in the second command is present in the third spatial region, controlling the robot to perform a second action with respect to the second object referenced in the second command.",
    "8. The computer-implemented method of claim 7, wherein the third spatial region is larger than the second spatial region indicated by the second gesture.",
    "9. The computer-implemented method of claim 1, comprising: receiving a second command for controlling the robot, the command referencing a second object; receiving second sensor data for a portion of the environment of the robot, the second sensor data being captured by the sensor of the robot; identifying, from the second sensor data, a second gesture of a human that indicates a second spatial region located outside of the portion of the environment described by the second sensor data; searching the map data for the second object referenced in the second command, wherein the search of the map data is restricted, based on the identified second gesture, to the second spatial region indicated by the second gesture; determining, based at least on searching the map data for the second object referenced in the second command, that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture; and based at least on determining that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture, controlling the robot to indicate that the second object referenced in the second command is absent from the second spatial region indicated by the second gesture.",
    "10. The computer-implemented method of claim 1, wherein controlling the robot to perform the action with respect to the object referenced in the command comprises controlling the robot to perform one of retrieving the object referenced in the command, move the object referenced in the command to a predetermined location, move the object referenced in the command to a location indicated by the command, or navigate to the location of the object referenced in the command.",
    "11. The computer-implemented method of claim 1, comprising: determining a location of the object referenced in the command within the spatial region, wherein the location of the object referenced in the command within the spatial region is represented by a set of coordinates within the space.",
    "12. The computer-implemented method of claim 1, wherein the sensor data is one of image data, infrared image data, light detection and ranging (LIDAR) data, thermal image data, night vision image data, or motion data.",
    "13. The computer-implemented method of claim 1, wherein the sensor data for the portion of the environment of the robot is image data for a field of view of a camera of the robot, the image data being captured by the camera of the robot.",
    "14. A system comprising: one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: receiving a command for controlling a robot, the command referencing an object; receiving sensor data for a portion of an environment of the robot, the sensor data being captured by a sensor of the robot; identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data; in response to identifying the gesture, accessing map data indicating locations of objects within a space, the map data being generated before receiving the command; searching the map data for the object referenced in the command, wherein the search of the map data is restricted, based on the identified gesture, to the spatial region indicated by the gesture; determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region indicated by the gesture; and in response to determining that the object referenced in the command is present in the spatial region indicated by the gesture, controlling the robot to perform an action with respect to the object referenced in the command.",
    "15. The system of claim 14, wherein the identified gesture is one of an arm wave, a hand gesture, or a glance.",
    "16. The system of claim 14, wherein receiving the command for controlling the robot, the command referencing the object comprises: receiving audio data captured by a microphone of the robot that corresponds to the command; identifying, based on performing speech recognition on the audio data, one or more candidate objects that each correspond to a respective candidate transcription of at least a portion of the audio data; accessing an inventory of one or more objects within the space; and identifying the object from among the inventory of the one or more objects within the space based at least on comparing each candidate transcription of at least a portion of the audio data to the one or more objects of the inventory.",
    "17. The system of claim 14, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: determining a location of the robot within the space; and determining the spatial region based at least on the gesture of the human and the location of the robot within the space.",
    "18. The system of claim 14, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: determining an orientation of the robot within the space when receiving the sensor data; and determining the spatial region based at least on the gesture of the human and the orientation of the robot within the space.",
    "19. The system of claim 14, wherein identifying the gesture of the human that indicates the spatial region located outside of the portion of the environment described by the sensor data comprises: detecting one or more predetermined shapes from the sensor data, each of the one or more predetermined shapes corresponding to a gesture of a human; determining one or more locations of the detected one or more predetermined shapes within the sensor data; and determining the spatial region based at least on the one or more predetermined shapes and the one or more locations of the detected one or more predetermined shapes within the sensor data.",
    "20. A non-transitory computer-readable storage device storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising: receiving a command for controlling a robot, the command referencing an object; receiving sensor data for a portion of an environment of the robot, the sensor data being captured by a sensor of the robot; identifying, from the sensor data, a gesture of a human that indicates a spatial region located outside of the portion of the environment described by the sensor data; in response to identifying the gesture, accessing map data indicating locations of objects within a space, the map data being generated before receiving the command; searching the map data for the object referenced in the command, wherein the search of the map data is restricted, based on the identified gesture, to the spatial region indicated by the gesture; determining, based at least on searching the map data for the object referenced in the command, that the object referenced in the command is present in the spatial region indicated by the gesture; and in response to determining that the object referenced in the command is present in the spatial region indicated by the gesture, controlling the robot to perform an action with respect to the object referenced in the command."
  ],
  "description_excerpt": "This specification relates to object identification, and one particular implementation relates to identifying objects based on human-robot interactions.\n\nHuman-computer interfaces that permit users to provide natural language or gestural inputs are becoming exceedingly pervasive. For example, a personal assistant application can receiving human speech and identify a command based on an analysis of that speech. The personal assistant application can perform or trigger operations in response to the identified command. Similarly, computer applications may receive images or video of a user and can detect human gestures from the images or video. The computer can interpret those gestures as commands and may perform or trigger operations responsive to the identified commands. These techniques are also being applied in the field of robotics to enable human-robot interactions. For example, users may be able to provide gestures or speech inputs to a robot to command the robot to perform specific actions. In some examples, a user may refer to a particular object in a command, either by word or by gesture. In response to such a command, the robot may be required to identify a physical object in its environment that corresponds to the particular object referenced in the command. Challenges may arise where multiple objects in the robot's environment correspond to the particular object referenced in the command. In those instances, the robot may be required to disambiguate between the multiple objects, to identify a particular one of the multiple objects that the user likely intended to reference in their command.",
  "cpc": [
    "B25J 9/1697",
    "B25J 11/0005",
    "B25J 11/008",
    "B25J 13/003",
    "B25J 19/026",
    "G05B 19/048",
    "G05B 2219/23021",
    "G05B 2219/35444",
    "G06F 3/016",
    "G06F 3/017",
    "G06F 3/167"
  ],
  "ipc": [
    "B25J 19/02",
    "B25J 9/16",
    "G05B 19/048"
  ],
  "assignees": [
    "X DEV LLC"
  ],
  "inventors": [
    "QUINLAN MICHAEL JOSEPH",
    "COHEN GABRIEL A"
  ],
  "filing_date": "2017-07-06",
  "publication_date": "2019-10-01",
  "grant_date": "2019-10-01",
  "priority_date": "2017-07-06",
  "application_number": "US-201715643138-A",
  "family_id": "68063883",
  "citations": [
    "CN104851211A",
    "US2007192910A1",
    "US2011054691A1",
    "US2011118877A1",
    "US2012173018A1",
    "US2012185090A1",
    "US2012197439A1",
    "US2012316679A1",
    "US2013325244A1",
    "US2015025681A1",
    "US2015032252A1",
    "US2015283702A1",
    "US2015362988A1",
    "US2015378444A1",
    "US2016140963A1",
    "US2018192845A1",
    "US7881936B2",
    "US8140188B2",
    "US9165566B2",
    "US9427863B2"
  ]
}

Record 1,019 of 5,000 in Patents full text (MLC-0201). Request the full dataset.