MLchartDataset catalogue

Patent · US10943120B2 · B2 · US

Enhanced pose determination for display device

(11) Publication number
US10943120B2
(21) Application number
16/221,065
(22) Filing date
2018-12-14
(30) Priority date
2017-12-15
(43) Publication date
2021-03-09
(45) Date of grant
2021-03-09
(51) IPC
G02B 27/01; G06F 3/01; G06F 3/0481; G06K 9/00; G09G 5/00
(52) CPC
  • G06F Electric digital data processing: 1/163, 3/012, 3/013, 3/0304, 3/04815
  • G02B Optical elements, systems or apparatus: 2027/0112, 2027/0174, 2027/0178, 27/017, 27/0172
  • G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/00597, 9/00671
  • G06V Image or video recognition or understanding: 20/20, 40/18
(73) Assignee
Magic Leap Inc
(72) Inventors
Martin Georg Zahnert; Joao Antonio Pereira Faro; Miguel Andres Granados Velasquez; Dominik Michael Kasper; Ashwin Swaminathan; Anush Mohan; Prateek Singhal
(54) Title
Enhanced pose determination for display device
(57) Abstract

To determine the head pose of a user, a head-mounted display system having an imaging device can obtain a current image of a real-world environment, with points corresponding to salient points which will be used to determine the head pose. The salient points are patch-based and include: a first salient point being projected onto the current image from a previous image, and with a second salient point included in the current image being extracted from the current image. Each salient point is subsequently matched with real-world points based on descriptor-based map information indicating locations of salient points in the real-world environment. The orientation of the imaging devices is determined based on the matching and based on the relative positions of the salient points in the view captured in the current image. The orientation may be used to extrapolate the head pose of the wearer of the head-mounted display system.

Full text
View on Google Patents

Claims (27)

  1. A system comprising: one or more imaging devices; one or more processors; and one or more computer storage media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: obtaining, via the one or more imaging devices, a current image of a real-world environment, the current image including a plurality of points for determining pose; accessing a first patch associated with a first salient point which is being tracked by the system, the first salient point being included in a prior image of the real-world environment, wherein the first salient point represents a first feature of the real-world environment; projecting the first patch onto the current image, wherein the first salient point is matched with a corresponding one of the plurality of points included in the current image, such that a position at which the first feature is represented in the current image is determined to correspond to the one of the plurality of points; extracting a second salient point from the current image, such that the second salient point is tracked by the system, the second salient point representing a second feature of the real-world environment; providing respective descriptors for the salient points associated with the current image, the salient points comprising the first salient point and the second salient point; matching, based on the descriptors, the salient points associated with the current image with real-world locations specified in a descriptor-based map of the real-world environment, such that real-world locations associated with the first feature and the second feature are identified; and determining, based on the matching, a pose associated with the system, the pose indicating at least an orientation of the one or more imaging devices in the real-world environment.
  2. The system of claim 1, wherein the operations further comprise adjusting a position of the first patch projected onto the current image, wherein the first patch includes a portion of the previous image encompassing the first salient point and an area of the previous image around the first salient point, and wherein adjusting the position comprises: locating a second patch in the current image similar to the first patch, wherein the first salient point is positioned in a similar location within the second patch as the first patch.
  3. The system of claim 2, wherein locating the second patch comprises minimizing a difference between the first patch in the previous image and the second patch in the current image.
  4. The system of claim 2, wherein projecting the first patch onto the current image is based, at least in part, on information from an inertial measurement unit of the system.
  5. The system of claim 1, wherein extracting the second salient point comprises: determining that an image area of the current image has less than a threshold number of salient points projected from the previous image; and extracting one or more additional salient points from the image area, the extracted salient points including the second salient point.
  6. The system of claim 5, wherein the image area comprises an entirety of the current image, or wherein the image comprises a subset of the current image.
  7. The system of claim 5, wherein the image area comprises a subset of the current image, and wherein the system is configured to adjust a size associated with the subset based on one or more of processing constraints or differences between one or more prior determined poses.
  8. The system of claim 1, wherein matching salient points associated with the current image with real-world locations specified in the map of the real-world environment comprises: accessing the descriptor-based map, the descriptor-based map comprising real-world locations of salient points and associated descriptors; and matching descriptors for salient points of the current image with descriptors for salient points at real-world locations.
  9. The system of claim 8, wherein the operations further comprise: projecting salient points provided in the descriptor-based map onto the current image, wherein the projection is based on one or more of an inertial measurement unit, an extended kalman filter, or visual-inertial odometry.
  10. The system of claim 1, wherein the system is configured to generate the descriptor-based map using at least the one or more imaging devices.
  11. The system of claim 1, wherein determining the pose is based on the real-world locations of salient points and the relative positions of the salient points in the view captured in the current image.
  12. The system of claim 1, wherein the operations further comprise: generating patches associated with respective salient points extracted from the current image, such that for a subsequent image to the current image, the patches comprise the salient points available to be projected onto the subsequent image.
  13. The system of claim 1, wherein providing descriptors comprises generating descriptors for each of the salient points.
  14. A head-mounted augmented reality display system comprising: one or more outwardly-facing imaging devices configured to obtain images of a real-world environment; one or more processors, the processors configured to: obtain a current image of the real-world environment; perform frame-to-frame tracking on the current image, such that patch-based salient points included in a previous image are projected onto the current image, each salient point representing a respective feature of the real-world environment; perform map-to-frame tracking on the current image, wherein map-to-frame tracking comprises: obtaining descriptors of the salient points in the current image, the salient points comprising the patch-based salient points, and matching the obtained descriptors of the salient points with descriptors stored in a map database, the stored descriptors corresponding to descriptors of features of the real-world environment, and the map database storing real-world locations associated with the features, such that real-world locations of the salient points are identified; and determine a pose associated with the display device, the pose indicating at least an orientation of the one or more imaging devices in the real-world environment.
  15. A method comprising: obtaining, via one or more imaging devices, a current image of a real-world environment, the current image including a plurality of points for determining pose; accessing a first patch associated with a first salient point, the first salient point being included in a prior image of the real-world environment, wherein the first salient point represents a first feature of the real-world environment; projecting the first patch onto the current image, wherein the first salient point is matched with a corresponding one of the plurality of points included in the current image, such that a position at which the first feature is represented in the current image is determined to correspond to the one of the plurality of points; extracting a second salient point from the current image, the second salient point representing a second feature of the real-world environment; providing respective descriptors for the salient points associated with the current image, the salient points comprising the first salient point and the second salient point; matching, based on the descriptors, the salient points associated with the current image with real-world locations specified in a descriptor-based map of the real-world environment, such that real-world locations associated with the first feature and the second feature are identified; and determining, based on the matching, a pose associated with the system, the pose indicating at least an orientation of the one or more imaging devices in the real-world environment.
  16. The method of claim 15, further comprising adjusting a position of the first patch projected onto the current image, wherein the first patch includes a portion of the previous image encompassing the first salient point and an area of the previous image around the first salient point, and wherein adjusting the position comprises locating a second patch in the current image similar to the first patch, wherein the first salient point is positioned in a similar location within the second patch as the first patch.
  17. The method of claim 16, wherein locating the second patch comprises determining a patch in the current image with a minimum of differences with the first patch.
  18. The method of claim 16, wherein projecting the first patch onto the current image is based, at least in part, on information from an inertial measurement unit of the display device.
  19. The method of claim 15, wherein extracting the second salient point comprises: determining that an image area of the current image has less than a threshold number of salient points projected from the previous image; and extracting one or more additional salient points from the image area, the extracted salient points including the second salient point.
  20. The method of claim 19, wherein the image area comprises an entirety of the current image, or wherein the image comprises a subset of the current image.
  21. The method of claim 19, wherein the image area comprises a subset of the current image, and wherein the processors are configured to adjust a size associated with the subset based on one or more of processing constraints or differences between one or more prior determined poses.
  22. The method of claim 15, wherein matching salient points associated with the current image with real-world locations specified in the map of the real-world environment comprises: accessing the descriptor-based map, the descriptor-based map comprising real-world locations of salient points and associated descriptors; and matching descriptors for salient points of the current image with descriptors for salient points at real-world locations.
  23. The method of claim 22, further comprising: projecting salient points provided in the descriptor-based map onto the current image, wherein the projection is based on one or more of an inertial measurement unit, an extended kalman filter, or visual-inertial odometry.
  24. The method of claim 15, wherein determining the pose is based on the real-world locations of salient points and the relative positions of the salient points in the view captured in the current image.
  25. The method of claim 15, further comprising: generating patches associated with respective salient points extracted from the current image, such that for a subsequent image to the current image, the patches comprise the salient points available to be projected onto the subsequent image.
  26. The method of claim 15, wherein providing descriptors comprises generating descriptors for each of the salient points.
  27. The method of claim 15, further comprising generating the descriptor-based map using at least the one or more imaging devices.

Description

The present disclosure relates to display systems and, more particularly, to augmented reality display systems.

Modern computing and display technologies have facilitated the development of systems for so called “virtual reality” or “augmented reality” experiences, wherein digitally reproduced images or portions thereof are presented to a user in a manner wherein they seem to be, or may be perceived as, real. A virtual reality, or “VR”, scenario typically involves presentation of digital or virtual image information without transparency to other actual real-world visual input; an augmented reality, or “AR”, scenario typically involves presentation of digital or virtual image information as an augmentation to visualization of the actual world around the user. A mixed reality, or “MR”, scenario is a type of AR scenario and typically involves virtual objects that are integrated into, and responsive to, the natural world. For example, in an MR scenario, AR image content may be blocked by or otherwise be perceived as interacting with objects in the real world.

Referring to FIG. 1, an augmented reality scene 10 is depicted wherein a user of an AR technology sees a real-world park- like setting 20 featuring people, trees, buildings in the background, and a concrete platform 30. In addition to these items, the user of the AR technology also perceives that he “sees” “virtual content” such as a robot statue 40 standing upon the real- world platform 30, and a cartoon- like avatar character 50 flying by which seems to be a personification of a bumble bee, even though these elements 40, 50 do not exist in the real world.

Citations (40)

  • US9081426B2
  • US6850221B1
  • US20120127062A1
  • US20160011419A1
  • US9348143B2
  • US20120169887A1
  • US20130125027A1
  • US20130082922A1
  • US9215293B2
  • US20130117377A1
  • US9547174B2
  • US20140177023A1
  • US20140218468A1
  • US9851563B2
  • US20150309263A2
  • US9671566B2
  • US20150016777A1
  • US20130342671A1
  • US20140071539A1
  • US9740006B2
  • US20150268415A1
  • US20140306866A1
  • US20140267420A1
  • US9417452B2
  • US20140368645A1
  • US20150103306A1
  • US9470906B2
  • US9791700B2
  • US20150205126A1
  • US20150178939A1
  • US9874749B2
  • US20150222883A1
  • US20150222884A1
  • US20160026253A1
  • US20150302652A1
  • US20150326570A1
  • US20150346495A1
  • US20150346490A1
  • US9857591B2
  • WO2019118886A1
Record as JSON
{
  "publication_number": "US10943120B2",
  "country": "US",
  "kind": "B2",
  "title": "Enhanced pose determination for display device",
  "abstract": "To determine the head pose of a user, a head-mounted display system having an imaging device can obtain a current image of a real-world environment, with points corresponding to salient points which will be used to determine the head pose. The salient points are patch-based and include: a first salient point being projected onto the current image from a previous image, and with a second salient point included in the current image being extracted from the current image. Each salient point is subsequently matched with real-world points based on descriptor-based map information indicating locations of salient points in the real-world environment. The orientation of the imaging devices is determined based on the matching and based on the relative positions of the salient points in the view captured in the current image. The orientation may be used to extrapolate the head pose of the wearer of the head-mounted display system.",
  "claims": [
    "1. A system comprising: one or more imaging devices; one or more processors; and one or more computer storage media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: obtaining, via the one or more imaging devices, a current image of a real-world environment, the current image including a plurality of points for determining pose; accessing a first patch associated with a first salient point which is being tracked by the system, the first salient point being included in a prior image of the real-world environment, wherein the first salient point represents a first feature of the real-world environment; projecting the first patch onto the current image, wherein the first salient point is matched with a corresponding one of the plurality of points included in the current image, such that a position at which the first feature is represented in the current image is determined to correspond to the one of the plurality of points; extracting a second salient point from the current image, such that the second salient point is tracked by the system, the second salient point representing a second feature of the real-world environment; providing respective descriptors for the salient points associated with the current image, the salient points comprising the first salient point and the second salient point; matching, based on the descriptors, the salient points associated with the current image with real-world locations specified in a descriptor-based map of the real-world environment, such that real-world locations associated with the first feature and the second feature are identified; and determining, based on the matching, a pose associated with the system, the pose indicating at least an orientation of the one or more imaging devices in the real-world environment.",
    "2. The system of claim 1, wherein the operations further comprise adjusting a position of the first patch projected onto the current image, wherein the first patch includes a portion of the previous image encompassing the first salient point and an area of the previous image around the first salient point, and wherein adjusting the position comprises: locating a second patch in the current image similar to the first patch, wherein the first salient point is positioned in a similar location within the second patch as the first patch.",
    "3. The system of claim 2, wherein locating the second patch comprises minimizing a difference between the first patch in the previous image and the second patch in the current image.",
    "4. The system of claim 2, wherein projecting the first patch onto the current image is based, at least in part, on information from an inertial measurement unit of the system.",
    "5. The system of claim 1, wherein extracting the second salient point comprises: determining that an image area of the current image has less than a threshold number of salient points projected from the previous image; and extracting one or more additional salient points from the image area, the extracted salient points including the second salient point.",
    "6. The system of claim 5, wherein the image area comprises an entirety of the current image, or wherein the image comprises a subset of the current image.",
    "7. The system of claim 5, wherein the image area comprises a subset of the current image, and wherein the system is configured to adjust a size associated with the subset based on one or more of processing constraints or differences between one or more prior determined poses.",
    "8. The system of claim 1, wherein matching salient points associated with the current image with real-world locations specified in the map of the real-world environment comprises: accessing the descriptor-based map, the descriptor-based map comprising real-world locations of salient points and associated descriptors; and matching descriptors for salient points of the current image with descriptors for salient points at real-world locations.",
    "9. The system of claim 8, wherein the operations further comprise: projecting salient points provided in the descriptor-based map onto the current image, wherein the projection is based on one or more of an inertial measurement unit, an extended kalman filter, or visual-inertial odometry.",
    "10. The system of claim 1, wherein the system is configured to generate the descriptor-based map using at least the one or more imaging devices.",
    "11. The system of claim 1, wherein determining the pose is based on the real-world locations of salient points and the relative positions of the salient points in the view captured in the current image.",
    "12. The system of claim 1, wherein the operations further comprise: generating patches associated with respective salient points extracted from the current image, such that for a subsequent image to the current image, the patches comprise the salient points available to be projected onto the subsequent image.",
    "13. The system of claim 1, wherein providing descriptors comprises generating descriptors for each of the salient points.",
    "14. A head-mounted augmented reality display system comprising: one or more outwardly-facing imaging devices configured to obtain images of a real-world environment; one or more processors, the processors configured to: obtain a current image of the real-world environment; perform frame-to-frame tracking on the current image, such that patch-based salient points included in a previous image are projected onto the current image, each salient point representing a respective feature of the real-world environment; perform map-to-frame tracking on the current image, wherein map-to-frame tracking comprises: obtaining descriptors of the salient points in the current image, the salient points comprising the patch-based salient points, and matching the obtained descriptors of the salient points with descriptors stored in a map database, the stored descriptors corresponding to descriptors of features of the real-world environment, and the map database storing real-world locations associated with the features, such that real-world locations of the salient points are identified; and determine a pose associated with the display device, the pose indicating at least an orientation of the one or more imaging devices in the real-world environment.",
    "15. A method comprising: obtaining, via one or more imaging devices, a current image of a real-world environment, the current image including a plurality of points for determining pose; accessing a first patch associated with a first salient point, the first salient point being included in a prior image of the real-world environment, wherein the first salient point represents a first feature of the real-world environment; projecting the first patch onto the current image, wherein the first salient point is matched with a corresponding one of the plurality of points included in the current image, such that a position at which the first feature is represented in the current image is determined to correspond to the one of the plurality of points; extracting a second salient point from the current image, the second salient point representing a second feature of the real-world environment; providing respective descriptors for the salient points associated with the current image, the salient points comprising the first salient point and the second salient point; matching, based on the descriptors, the salient points associated with the current image with real-world locations specified in a descriptor-based map of the real-world environment, such that real-world locations associated with the first feature and the second feature are identified; and determining, based on the matching, a pose associated with the system, the pose indicating at least an orientation of the one or more imaging devices in the real-world environment.",
    "16. The method of claim 15, further comprising adjusting a position of the first patch projected onto the current image, wherein the first patch includes a portion of the previous image encompassing the first salient point and an area of the previous image around the first salient point, and wherein adjusting the position comprises locating a second patch in the current image similar to the first patch, wherein the first salient point is positioned in a similar location within the second patch as the first patch.",
    "17. The method of claim 16, wherein locating the second patch comprises determining a patch in the current image with a minimum of differences with the first patch.",
    "18. The method of claim 16, wherein projecting the first patch onto the current image is based, at least in part, on information from an inertial measurement unit of the display device.",
    "19. The method of claim 15, wherein extracting the second salient point comprises: determining that an image area of the current image has less than a threshold number of salient points projected from the previous image; and extracting one or more additional salient points from the image area, the extracted salient points including the second salient point.",
    "20. The method of claim 19, wherein the image area comprises an entirety of the current image, or wherein the image comprises a subset of the current image.",
    "21. The method of claim 19, wherein the image area comprises a subset of the current image, and wherein the processors are configured to adjust a size associated with the subset based on one or more of processing constraints or differences between one or more prior determined poses.",
    "22. The method of claim 15, wherein matching salient points associated with the current image with real-world locations specified in the map of the real-world environment comprises: accessing the descriptor-based map, the descriptor-based map comprising real-world locations of salient points and associated descriptors; and matching descriptors for salient points of the current image with descriptors for salient points at real-world locations.",
    "23. The method of claim 22, further comprising: projecting salient points provided in the descriptor-based map onto the current image, wherein the projection is based on one or more of an inertial measurement unit, an extended kalman filter, or visual-inertial odometry.",
    "24. The method of claim 15, wherein determining the pose is based on the real-world locations of salient points and the relative positions of the salient points in the view captured in the current image.",
    "25. The method of claim 15, further comprising: generating patches associated with respective salient points extracted from the current image, such that for a subsequent image to the current image, the patches comprise the salient points available to be projected onto the subsequent image.",
    "26. The method of claim 15, wherein providing descriptors comprises generating descriptors for each of the salient points.",
    "27. The method of claim 15, further comprising generating the descriptor-based map using at least the one or more imaging devices."
  ],
  "description_excerpt": "The present disclosure relates to display systems and, more particularly, to augmented reality display systems.\n\nModern computing and display technologies have facilitated the development of systems for so called “virtual reality” or “augmented reality” experiences, wherein digitally reproduced images or portions thereof are presented to a user in a manner wherein they seem to be, or may be perceived as, real. A virtual reality, or “VR”, scenario typically involves presentation of digital or virtual image information without transparency to other actual real-world visual input; an augmented reality, or “AR”, scenario typically involves presentation of digital or virtual image information as an augmentation to visualization of the actual world around the user. A mixed reality, or “MR”, scenario is a type of AR scenario and typically involves virtual objects that are integrated into, and responsive to, the natural world. For example, in an MR scenario, AR image content may be blocked by or otherwise be perceived as interacting with objects in the real world.\n\nReferring to FIG. 1, an augmented reality scene 10 is depicted wherein a user of an AR technology sees a real-world park- like setting 20 featuring people, trees, buildings in the background, and a concrete platform 30. In addition to these items, the user of the AR technology also perceives that he “sees” “virtual content” such as a robot statue 40 standing upon the real- world platform 30, and a cartoon- like avatar character 50 flying by which seems to be a personification of a bumble bee, even though these elements 40, 50 do not exist in the real world.",
  "cpc": [
    "G06F 1/163",
    "G02B 2027/0112",
    "G02B 2027/0174",
    "G02B 2027/0178",
    "G02B 27/017",
    "G02B 27/0172",
    "G06F 3/012",
    "G06F 3/013",
    "G06F 3/0304",
    "G06F 3/04815",
    "G06K 9/00597",
    "G06K 9/00671",
    "G06V 20/20",
    "G06V 40/18"
  ],
  "ipc": [
    "G02B 27/01",
    "G06F 3/01",
    "G06F 3/0481",
    "G06K 9/00",
    "G09G 5/00"
  ],
  "assignees": [
    "Magic Leap Inc"
  ],
  "inventors": [
    "Martin Georg Zahnert",
    "Joao Antonio Pereira Faro",
    "Miguel Andres Granados Velasquez",
    "Dominik Michael Kasper",
    "Ashwin Swaminathan",
    "Anush Mohan",
    "Prateek Singhal"
  ],
  "filing_date": "2018-12-14",
  "publication_date": "2021-03-09",
  "grant_date": "2021-03-09",
  "priority_date": "2017-12-15",
  "application_number": "US-201816221065-A",
  "family_id": "66814528",
  "cited_by_count": 5,
  "citations": [
    "US9081426B2",
    "US6850221B1",
    "US20120127062A1",
    "US20160011419A1",
    "US9348143B2",
    "US20120169887A1",
    "US20130125027A1",
    "US20130082922A1",
    "US9215293B2",
    "US20130117377A1",
    "US9547174B2",
    "US20140177023A1",
    "US20140218468A1",
    "US9851563B2",
    "US20150309263A2",
    "US9671566B2",
    "US20150016777A1",
    "US20130342671A1",
    "US20140071539A1",
    "US9740006B2",
    "US20150268415A1",
    "US20140306866A1",
    "US20140267420A1",
    "US9417452B2",
    "US20140368645A1",
    "US20150103306A1",
    "US9470906B2",
    "US9791700B2",
    "US20150205126A1",
    "US20150178939A1",
    "US9874749B2",
    "US20150222883A1",
    "US20150222884A1",
    "US20160026253A1",
    "US20150302652A1",
    "US20150326570A1",
    "US20150346495A1",
    "US20150346490A1",
    "US9857591B2",
    "WO2019118886A1"
  ]
}

Record 1,725 of 8,000 in Patents full text (MLC-0201). Request the full dataset.