MLchartDataset catalogue

Patent · US11019449B2 · B2 · US

Six degrees of freedom and three degrees of freedom backward compatibility

(11) Publication number
US11019449B2
(21) Application number
16/567,700
(22) Filing date
2019-09-11
(30) Priority date
2018-10-06
(43) Publication date
2021-05-25
(45) Date of grant
2021-05-25
(51) IPC
G06F 3/01; G06F 3/0346; G06K 9/00; G06T 19/00; H04S 7/00; H04W 4/02; H04W 4/029
(52) CPC
  • H04S Stereophonic systems: 7/304, 2400/11, 2400/15, 2420/01, 2420/11, 7/303
  • G06F Electric digital data processing: 1/163, 3/011, 3/012, 3/0304, 3/0346
  • G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/00671
  • G06T Image data processing or generation, in general: 19/006
  • G06V Image or video recognition or understanding: 20/20
  • H04R Loudspeakers, microphones, gramophone pick-ups or like acoustic electromechanical transducers; electric hearing AIDS; public address systems: 2460/07, 2499/15
  • H04W Wireless communication networks: 4/027, 4/029
(73) Assignee
Qualcomm Inc
(72) Inventors
Moo Young Kim; Nils Günther Peters; S M Akramus Salehin; Siddhartha Goutham Swaminathan; Dipanjan Sen
(54) Title
Six degrees of freedom and three degrees of freedom backward compatibility
(57) Abstract

A device and method for backward compatibility for virtual reality (VR), mixed reality (MR), augmented reality (AR), computer vision, and graphics systems. The device and method enable rendering audio data with more degrees of freedom on devices that support fewer degrees of freedom. The device includes memory configured to store audio data representative of a soundfield captured at a plurality of capture locations, metadata that enables the audio data to be rendered to support N degrees of freedom, and adaptation metadata that enables the audio data to be rendered to support M degrees of freedom. The device also includes one or more processors coupled to the memory, and configured to adapt, based on the adaptation metadata, the audio data to provide the M degrees of freedom, and generate speaker feeds based on the adapted audio data.

Full text
View on Google Patents

Claims (23)

  1. A device comprising: a memory configured to store audio data representative of a soundfield captured at a plurality of capture locations, metadata that enables the audio data to be rendered to support N degrees of freedom, and adaptation metadata that enables the audio data to be rendered to support M degrees of freedom, wherein N is a first integer number and M is a second integer number that is different than the first integer number; and one or more processors coupled to the memory, and configured to: determine a user location by causing a display device to display a plurality of user locations or a trajectory, and receive, from a user, an input indicative of one of the plurality of locations or indicative of a position on the trajectory; adapt, based on the adaptation metadata, the audio data to provide the M degrees of freedom; and generate speaker feeds based on the adapted audio data.
  2. The device of claim 1, wherein N=6 and M is less than N.
  3. The device of claim 1, further comprising the display device.
  4. The device of claim 1, wherein to determine the user location, the one or more processors are configured to: select based on the position of the trajectory, one of a plurality of locations as the user location.
  5. The device of claim 2, wherein the 6 degrees of freedom comprise yaw, pitch, roll, and translational distance defined in either a two-dimensional spatial coordinate space or a three-dimensional spatial coordinate space.
  6. The device of claim 5, wherein the M degrees of freedom includes three degrees of freedom, the three degrees of freedom comprising yaw, pitch, and roll.
  7. The device of claim 5, wherein M=0.
  8. The device of claim 1, wherein the adaptation metadata includes the user location and a user orientation, and wherein the one or more processors are configured to adapt, based on the user location and the user orientation, the audio data to provide zero degrees of freedom.
  9. The device of claim 1, wherein the adaptation metadata includes the user location, and wherein the one or more processors are configured to: determine, based on the user location, an effects matrix that provides the M degrees of freedom; and apply the effects matrix to the audio data to adapt the soundfield.
  10. The device of claim 9, wherein the one or more processors are further configured to multiply the effects matrix by a rendering matrix to obtain an updated rendering matrix, and wherein the one or more processors are further configured to apply the updated rendering matrix to the audio data to: provide the M degrees of freedom; and generate the speaker feeds.
  11. The device of claim 1, wherein the one or more processors are further configured to store the user location to the memory as the adaptation metadata.
  12. The device of claim 1, wherein the one or more processors are further configured to obtain a bitstream that specifies the adaptation metadata, the adaptation metadata including the user location associated the audio data.
  13. The device of claim 1, wherein the one or more processors are further configured to obtain a rotation indication indicative of a rotational head movement of the user interfacing with the device, and wherein the one or more processors are further configured to adapt, based on the rotation indication and the adaptation metadata, the audio data to provide three degrees of freedom.
  14. The device of claim 1, wherein the one or more processors are coupled to speakers of a wearable device, wherein the one or more processors are configured to apply a binaural renderer to adapted higher order ambisonic audio data to generate the speaker feeds, and wherein the one or more processors are further configured to output the speaker feeds to the speakers.
  15. The device of claim 14, wherein the wearable device comprises a watch, glasses, headphones, an augmented reality (AR) headset, a virtual reality (VR) headset, or an extended reality (XR) headset.
  16. The device of claim 1, wherein the audio data comprises higher order ambisonic coefficients associated with spherical basis functions having an order of one or less, higher order ambisonic coefficients associated with spherical basis functions having a mixed order and suborder, or higher order ambisonic coefficients associated with spherical basis functions having an order greater than one.
  17. The device of claim 1, wherein the audio data comprises one or more audio objects.
  18. The device of claim 1, further comprising: one or more speakers configured to reproduce, based on the speaker feeds, the soundfield.
  19. The device of claim 18, wherein the device is one of a vehicle, unmanned vehicle, a robot, or a handset.
  20. The device of claim 1, wherein the one or more processors comprise processing circuitry.
  21. The device of claim 20, wherein the processing circuitry comprises one or more application specific integrated circuits.
  22. A method comprising: storing audio data representative of a soundfield captured at a plurality of capture locations; storing metadata that enables the audio data to be rendered to support N degrees of freedom; storing adaptation metadata that enables the audio data to be rendered to support M degrees of freedom, wherein N is a first integer number and M is a second integer number that is different than the first integer number; displaying a plurality of user locations or a trajectory on a display device; receiving, from a user, an input indicative of one of the plurality of user locations or indicative of a position on the trajectory; determining a user location based on the input adapting, based on the adaptation metadata, the audio data to provide the M degrees of freedom; and generating speaker feeds based on the adapted audio data.
  23. A device comprising: means for storing audio data representative of a soundfield captured at a plurality of capture locations; means for storing metadata that enables the audio data to be rendered to support N degrees of freedom; means for storing adaptation metadata that enables the audio data to be rendered to support M degrees of freedom, wherein N is a first integer number and M is a second integer number that is different than the first integer number; means for causing a display device to display a plurality of user locations or a trajectory on the display device; means for receiving, from a user, an input indicative of one of the plurality of user locations or indicative of a position on the trajectory; means for determining a user location based on the input; means for adapting, based on the adaptation metadata, the audio data to provide the M degrees of freedom; and means for generating speaker feeds based on the adapted audio data.

Description

This application claims the benefit of U.S. Provisional Application No. 62/742,324 filed Oct. 6, 2018, the entire content of which is hereby incorporated by reference.

This disclosure relates to processing of media data, such as audio data.

In recent years, there is an increasing interest in Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR) technologies. Advances to image processing and computer vision technologies in the wireless space, have led to better rendering and computational resources allocated to improving the visual quality and immersive visual experience of these technologies.

In VR technologies, virtual information may be presented to a user using a head-mounted display such that the user may visually experience an artificial world on a screen in front of their eyes. In AR technologies, the real-world is augmented by visual objects that are super-imposed, or, overlaid on physical objects in the real-world. The augmentation may insert new visual objects or mask visual objects to the real-world environment. In MR technologies, the boundary between what's real or synthetic/virtual and visually experienced by a user is becoming difficult to discern.

This disclosure relates generally to auditory aspects of the user experience of computer-mediated reality systems, including virtual reality (VR), mixed reality (MR), augmented reality (AR), computer vision, and graphics systems. More specifically, the techniques may enable rendering of audio data for VR, MR, AR, etc. that accounts for five or more degrees of freedom on devices or systems that support fewer than five degrees of freedom.

Citations (34)

  • US20030227476A1
  • US20080130904A1
  • US20110249821A1
  • US20130202129A1
  • US20190253691A1
  • US20140168277A1
  • US20150070274A1
  • US20160225377A1
  • US20170180905A1
  • US20170047071A1
  • US20160227340A1
  • US20170078825A1
  • WO2017066300A2
  • US9591427B1
  • US20170295446A1
  • US20180046431A1
  • US20180068664A1
  • US20180091917A1
  • US20190230436A1
  • US20180139565A1
  • US20180217806A1
  • US20180249274A1
  • US20180288558A1
  • US20200094141A1
  • US20200162833A1
  • US20200145778A1
  • US10405126B2
  • US20200211574A1
  • US20190379876A1
  • US20200260206A1
  • WO2019075425A1
  • US20190306651A1
  • US20190313200A1
  • US20190387350A1
Record as JSON
{
  "publication_number": "US11019449B2",
  "country": "US",
  "kind": "B2",
  "title": "Six degrees of freedom and three degrees of freedom backward compatibility",
  "abstract": "A device and method for backward compatibility for virtual reality (VR), mixed reality (MR), augmented reality (AR), computer vision, and graphics systems. The device and method enable rendering audio data with more degrees of freedom on devices that support fewer degrees of freedom. The device includes memory configured to store audio data representative of a soundfield captured at a plurality of capture locations, metadata that enables the audio data to be rendered to support N degrees of freedom, and adaptation metadata that enables the audio data to be rendered to support M degrees of freedom. The device also includes one or more processors coupled to the memory, and configured to adapt, based on the adaptation metadata, the audio data to provide the M degrees of freedom, and generate speaker feeds based on the adapted audio data.",
  "claims": [
    "1. A device comprising: a memory configured to store audio data representative of a soundfield captured at a plurality of capture locations, metadata that enables the audio data to be rendered to support N degrees of freedom, and adaptation metadata that enables the audio data to be rendered to support M degrees of freedom, wherein N is a first integer number and M is a second integer number that is different than the first integer number; and one or more processors coupled to the memory, and configured to: determine a user location by causing a display device to display a plurality of user locations or a trajectory, and receive, from a user, an input indicative of one of the plurality of locations or indicative of a position on the trajectory; adapt, based on the adaptation metadata, the audio data to provide the M degrees of freedom; and generate speaker feeds based on the adapted audio data.",
    "2. The device of claim 1, wherein N=6 and M is less than N.",
    "3. The device of claim 1, further comprising the display device.",
    "4. The device of claim 1, wherein to determine the user location, the one or more processors are configured to: select based on the position of the trajectory, one of a plurality of locations as the user location.",
    "5. The device of claim 2, wherein the 6 degrees of freedom comprise yaw, pitch, roll, and translational distance defined in either a two-dimensional spatial coordinate space or a three-dimensional spatial coordinate space.",
    "6. The device of claim 5, wherein the M degrees of freedom includes three degrees of freedom, the three degrees of freedom comprising yaw, pitch, and roll.",
    "7. The device of claim 5, wherein M=0.",
    "8. The device of claim 1, wherein the adaptation metadata includes the user location and a user orientation, and wherein the one or more processors are configured to adapt, based on the user location and the user orientation, the audio data to provide zero degrees of freedom.",
    "9. The device of claim 1, wherein the adaptation metadata includes the user location, and wherein the one or more processors are configured to: determine, based on the user location, an effects matrix that provides the M degrees of freedom; and apply the effects matrix to the audio data to adapt the soundfield.",
    "10. The device of claim 9, wherein the one or more processors are further configured to multiply the effects matrix by a rendering matrix to obtain an updated rendering matrix, and wherein the one or more processors are further configured to apply the updated rendering matrix to the audio data to: provide the M degrees of freedom; and generate the speaker feeds.",
    "11. The device of claim 1, wherein the one or more processors are further configured to store the user location to the memory as the adaptation metadata.",
    "12. The device of claim 1, wherein the one or more processors are further configured to obtain a bitstream that specifies the adaptation metadata, the adaptation metadata including the user location associated the audio data.",
    "13. The device of claim 1, wherein the one or more processors are further configured to obtain a rotation indication indicative of a rotational head movement of the user interfacing with the device, and wherein the one or more processors are further configured to adapt, based on the rotation indication and the adaptation metadata, the audio data to provide three degrees of freedom.",
    "14. The device of claim 1, wherein the one or more processors are coupled to speakers of a wearable device, wherein the one or more processors are configured to apply a binaural renderer to adapted higher order ambisonic audio data to generate the speaker feeds, and wherein the one or more processors are further configured to output the speaker feeds to the speakers.",
    "15. The device of claim 14, wherein the wearable device comprises a watch, glasses, headphones, an augmented reality (AR) headset, a virtual reality (VR) headset, or an extended reality (XR) headset.",
    "16. The device of claim 1, wherein the audio data comprises higher order ambisonic coefficients associated with spherical basis functions having an order of one or less, higher order ambisonic coefficients associated with spherical basis functions having a mixed order and suborder, or higher order ambisonic coefficients associated with spherical basis functions having an order greater than one.",
    "17. The device of claim 1, wherein the audio data comprises one or more audio objects.",
    "18. The device of claim 1, further comprising: one or more speakers configured to reproduce, based on the speaker feeds, the soundfield.",
    "19. The device of claim 18, wherein the device is one of a vehicle, unmanned vehicle, a robot, or a handset.",
    "20. The device of claim 1, wherein the one or more processors comprise processing circuitry.",
    "21. The device of claim 20, wherein the processing circuitry comprises one or more application specific integrated circuits.",
    "22. A method comprising: storing audio data representative of a soundfield captured at a plurality of capture locations; storing metadata that enables the audio data to be rendered to support N degrees of freedom; storing adaptation metadata that enables the audio data to be rendered to support M degrees of freedom, wherein N is a first integer number and M is a second integer number that is different than the first integer number; displaying a plurality of user locations or a trajectory on a display device; receiving, from a user, an input indicative of one of the plurality of user locations or indicative of a position on the trajectory; determining a user location based on the input adapting, based on the adaptation metadata, the audio data to provide the M degrees of freedom; and generating speaker feeds based on the adapted audio data.",
    "23. A device comprising: means for storing audio data representative of a soundfield captured at a plurality of capture locations; means for storing metadata that enables the audio data to be rendered to support N degrees of freedom; means for storing adaptation metadata that enables the audio data to be rendered to support M degrees of freedom, wherein N is a first integer number and M is a second integer number that is different than the first integer number; means for causing a display device to display a plurality of user locations or a trajectory on the display device; means for receiving, from a user, an input indicative of one of the plurality of user locations or indicative of a position on the trajectory; means for determining a user location based on the input; means for adapting, based on the adaptation metadata, the audio data to provide the M degrees of freedom; and means for generating speaker feeds based on the adapted audio data."
  ],
  "description_excerpt": "This application claims the benefit of U.S. Provisional Application No. 62/742,324 filed Oct. 6, 2018, the entire content of which is hereby incorporated by reference.\n\nThis disclosure relates to processing of media data, such as audio data.\n\nIn recent years, there is an increasing interest in Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR) technologies. Advances to image processing and computer vision technologies in the wireless space, have led to better rendering and computational resources allocated to improving the visual quality and immersive visual experience of these technologies.\n\nIn VR technologies, virtual information may be presented to a user using a head-mounted display such that the user may visually experience an artificial world on a screen in front of their eyes. In AR technologies, the real-world is augmented by visual objects that are super-imposed, or, overlaid on physical objects in the real-world. The augmentation may insert new visual objects or mask visual objects to the real-world environment. In MR technologies, the boundary between what's real or synthetic/virtual and visually experienced by a user is becoming difficult to discern.\n\nThis disclosure relates generally to auditory aspects of the user experience of computer-mediated reality systems, including virtual reality (VR), mixed reality (MR), augmented reality (AR), computer vision, and graphics systems. More specifically, the techniques may enable rendering of audio data for VR, MR, AR, etc. that accounts for five or more degrees of freedom on devices or systems that support fewer than five degrees of freedom.",
  "cpc": [
    "H04S 7/304",
    "G06F 1/163",
    "G06F 3/011",
    "G06F 3/012",
    "G06F 3/0304",
    "G06F 3/0346",
    "G06K 9/00671",
    "G06T 19/006",
    "G06V 20/20",
    "H04R 2460/07",
    "H04R 2499/15",
    "H04S 2400/11",
    "H04S 2400/15",
    "H04S 2420/01",
    "H04S 2420/11",
    "H04S 7/303",
    "H04W 4/027",
    "H04W 4/029"
  ],
  "ipc": [
    "G06F 3/01",
    "G06F 3/0346",
    "G06K 9/00",
    "G06T 19/00",
    "H04S 7/00",
    "H04W 4/02",
    "H04W 4/029"
  ],
  "assignees": [
    "Qualcomm Inc"
  ],
  "inventors": [
    "Moo Young Kim",
    "Nils Günther Peters",
    "S M Akramus Salehin",
    "Siddhartha Goutham Swaminathan",
    "Dipanjan Sen"
  ],
  "filing_date": "2019-09-11",
  "publication_date": "2021-05-25",
  "grant_date": "2021-05-25",
  "priority_date": "2018-10-06",
  "application_number": "US-201916567700-A",
  "family_id": "70051401",
  "cited_by_count": 5,
  "citations": [
    "US20030227476A1",
    "US20080130904A1",
    "US20110249821A1",
    "US20130202129A1",
    "US20190253691A1",
    "US20140168277A1",
    "US20150070274A1",
    "US20160225377A1",
    "US20170180905A1",
    "US20170047071A1",
    "US20160227340A1",
    "US20170078825A1",
    "WO2017066300A2",
    "US9591427B1",
    "US20170295446A1",
    "US20180046431A1",
    "US20180068664A1",
    "US20180091917A1",
    "US20190230436A1",
    "US20180139565A1",
    "US20180217806A1",
    "US20180249274A1",
    "US20180288558A1",
    "US20200094141A1",
    "US20200162833A1",
    "US20200145778A1",
    "US10405126B2",
    "US20200211574A1",
    "US20190379876A1",
    "US20200260206A1",
    "WO2019075425A1",
    "US20190306651A1",
    "US20190313200A1",
    "US20190387350A1"
  ]
}

Record 1,619 of 8,000 in Patents full text (MLC-0201). Request the full dataset.