Patent · US11837249B2 · B2 · US
Visually presenting auditory information
- (11) Publication number
- US11837249B2
- (21) Application number
- 17/396,657
- (22) Filing date
- 2021-08-07
- (30) Priority date
- 2016-07-16
- (43) Publication date
- 2023-12-05
- (45) Date of grant
- 2023-12-05
- (51) IPC
- G06F 40/10; G10L 17/00; G10L 21/0364; H04R 1/40; H04R 25/00; H04R 3/00
- (52) CPC
- G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 21/0364, 15/1822, 15/26, 17/00, 25/66
- G06F Electric digital data processing: 40/10
- H04R Loudspeakers, microphones, gramophone pick-ups or like acoustic electromechanical transducers; electric hearing AIDS; public address systems: 1/406, 2201/023, 25/407, 3/005, 5/0335
- (72) Inventors
- Ron Zass
- (54) Title
- Visually presenting auditory information
- (57) Abstract
Systems, methods and non-transitory computer readable media for processing audio and visually presenting information are provided. Audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus may be obtained. The audio data may be analyzed to obtain textual information. The audio data may be analyzed to associate different portions of the textual information with different speakers. A head mounted display system may be used to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information.
- Full text
- View on Google Patents
Claims (20)
- A non-transitory computer readable medium storing data and computer implementable instructions that when executed by at least one processor cause the at least one processor to perform a method for processing audio and visually presenting information, the method comprising: obtaining audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus; analyzing the audio data to obtain textual information; analyzing the audio data to associate different portions of the textual information with different speakers; using a head mounted display system to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information; analyzing the audio data to identify different voice properties in different parts of the audio data; associating different parts of the textual information with different voice properties; determining information based on particular voice properties; and presenting the information determined based on the particular voice properties along the part of the textual information associated with the particular voice properties.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises: analyzing the audio data to identify a nonverbal sound; and generating textual description of the nonverbal sound.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises: analyzing the audio data to identify a melody; and generating textual description of the melody.
- The non-transitory computer readable medium of claim 1, wherein the head mounted display system is an augmented reality display system.
- The non-transitory computer readable medium of claim 4, wherein the association of the presentation regions with the speakers is configured to overlay the portions of the textual information associated with a speaker over the speaker in the augmented reality display system.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises: analyzing the audio data to identify an item in the audio data; representing the item using a graphical symbol; and displaying the graphical symbol in conjunction with the textual information.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises displaying the different portions of the textual information using different sets of visual display parameters.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises: associating the different parts of the textual information with different textual formats based on the association of the different parts of the textual information with the different voice properties.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises associating the different portions of the textual information with different textual formats based on the association of the different portions of the textual information with the different speakers.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises: determining that a portion of a speech is said in a specific linguistic tone; selecting a visual display parameter for a part of the textual information associated with the portion of the speech based on the specific linguistic tone; and displaying the part of the textual information using the selected visual display parameter.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises visually presenting information associated with a speaker in conjunction with a portion of the textual information associated with the speaker.
- The non-transitory computer readable medium of claim 1, wherein the association of the presentation regions with the speakers is based on spatial orientation of the speakers.
- The non-transitory computer readable medium of claim 1, wherein the association of the presentation regions with the speakers is based on positions of the speakers.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises: associating a part of the textual information with speech produced by a user; and avoiding presenting the part of the textual information associated with the speech produced by the user.
- The non-transitory computer readable medium of claim 1, wherein the method further comprises: determining that a part of the textual information is associated with speech that do not involve the user; and avoiding presenting the part of the textual information associated.
- The non-transitory computer readable medium of claim 15, wherein the speech that do not involve the user is a conversation that do not involve the user.
- The non-transitory computer readable medium of claim 15, wherein the speech that do not involve the user is a speech not directed at the user.
- A system for processing audio and visually presenting information, the method comprising: at least one processing unit configured to: obtain audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus; analyze the audio data to obtain textual information; analyze the audio data to associate different portions of the textual information with different speakers; use a head mounted display system to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information; analyze the audio data to identify different voice properties in different parts of the audio data; associate different parts of the textual information with different voice properties; determine information based on particular voice properties; and present the information determined based on the particular voice properties along the part of the textual information associated with the particular voice properties.
- A method for processing audio and visually presenting information, the method comprising: obtaining audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus; analyzing the audio data to obtain textual information; analyzing the audio data to associate different portions of the textual information with different speakers; using a head mounted display system to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information; analyzing the audio data to identify different voice properties in different parts of the audio data; associating different parts of the textual information with different voice properties; determining information based on particular voice properties; and presenting the information determined based on the particular voice properties along the part of the textual information associated with the particular voice properties.
- A non-transitory computer readable medium storing data and computer implementable instructions that when executed by at least one processor cause the at least one processor to perform a method for processing audio and visually presenting information, the method comprising: obtaining audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus; analyzing the audio data to obtain textual information; analyzing the audio data to associate different portions of the textual information with different speakers; using a head mounted display system to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information; analyzing the audio data to identify a nonverbal sound; and generating textual description of the nonverbal sound.
Description
The disclosed embodiments generally relate to systems and methods for processing audio. More particularly, the disclosed embodiments relate to systems and methods for processing audio to visually present auditory information.
Audio as well as other sensors are now part of numerous devices, from intelligent personal assistant devices to mobile phones, and the availability of audio data and other information produced by these devices is increasing.
Various conditions may cause difficulties in maintaining socially appropriate eye contact, including social phobia, autism, Asperger syndrome, and so forth. Inappropriate eye contact may include avoidance of eye contact, abnormal eye contact pattern, excessive eye contact, starring inappropriately, and so forth.
Tantrums, including temper tantrums, meltdowns and sensory meltdowns, are emotional outbursts characterized by stubbornness, crying, screaming, defiance, ranting, hitting, and tirades. In the general population, tantrums are more common in childhood, and are the result of frustration. In some conditions, including autism and Asperger's syndrome, tantrums may be the response to sensory overloads.
Echolalia is a speech disorder characterized by meaningless repetition of vocalization and speech made by another person. Palilalia is a speech disorder characterized by meaningless repetition of vocalization and speech made by the same person. The repetition may be of syllables, words, utterances, phrases, sentences, and so forth.
Citations (89)
- US5950157A
- US5940798A
- US20030018475A1
- US20020103649A1
- US6707921B2
- US6882971B2
- US20040107098A1
- US20060238877A1
- US20070172805A1
- US20060167691A1
- US10084920B1
- US20080201141A1
- US20100250257A1
- US8825468B2
- US8259954B2
- US20100262419A1
- US20090306981A1
- US20090319265A1
- US20100174533A1
- US20120020503A1
- US8441356B1
- US8719016B1
- US20100280336A1
- US20110087491A1
- US20110228914A1
- US20120128186A1
- US20120053929A1
- US20140247926A1
- US20120116772A1
- US20120215532A1
- US8326338B1
- US20120265537A1
- US20130080168A1
- US9355648B2
- US8183997B1
- US20120128683A1
- US9213705B1
- US20150011842A1
- US20150058013A1
- US9536536B2
- US9746916B2
- US9799336B2
- US20140081634A1
- US20140122077A1
- US9412375B2
- US20140163960A1
- US20140236596A1
- US9282284B2
- US20140379352A1
- US9641942B2
- US20150036856A1
- US20150073309A1
- US20150099946A1
- US9870357B2
- US20180285312A1
- US20150302867A1
- US20150334346A1
- US20150356836A1
- US20150373477A1
- US20160019895A1
- US20170208415A1
- US20160064002A1
- US20160098993A1
- US20160133257A1
- US10241741B2
- US20160350664A1
- US20160379638A1
- US20170041699A1
- US9894320B2
- US20170076000A1
- US10269372B1
- US20170103748A1
- US20170105662A1
- US20170111303A1
- US20170221500A1
- US9749583B1
- US20170364599A1
- US10446166B2
- US20180020285A1
- US20210366505A1
- US20180077483A1
- US10268447B1
- US10692485B1
- US20180240458A1
- US20190279654A1
- US10803852B2
- US10878802B2
- US20200050863A1
- US20200066294A1
Record as JSON
{
"publication_number": "US11837249B2",
"country": "US",
"kind": "B2",
"title": "Visually presenting auditory information",
"abstract": "Systems, methods and non-transitory computer readable media for processing audio and visually presenting information are provided. Audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus may be obtained. The audio data may be analyzed to obtain textual information. The audio data may be analyzed to associate different portions of the textual information with different speakers. A head mounted display system may be used to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information.",
"claims": [
"1. A non-transitory computer readable medium storing data and computer implementable instructions that when executed by at least one processor cause the at least one processor to perform a method for processing audio and visually presenting information, the method comprising: obtaining audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus; analyzing the audio data to obtain textual information; analyzing the audio data to associate different portions of the textual information with different speakers; using a head mounted display system to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information; analyzing the audio data to identify different voice properties in different parts of the audio data; associating different parts of the textual information with different voice properties; determining information based on particular voice properties; and presenting the information determined based on the particular voice properties along the part of the textual information associated with the particular voice properties.",
"2. The non-transitory computer readable medium of claim 1, wherein the method further comprises: analyzing the audio data to identify a nonverbal sound; and generating textual description of the nonverbal sound.",
"3. The non-transitory computer readable medium of claim 1, wherein the method further comprises: analyzing the audio data to identify a melody; and generating textual description of the melody.",
"4. The non-transitory computer readable medium of claim 1, wherein the head mounted display system is an augmented reality display system.",
"5. The non-transitory computer readable medium of claim 4, wherein the association of the presentation regions with the speakers is configured to overlay the portions of the textual information associated with a speaker over the speaker in the augmented reality display system.",
"6. The non-transitory computer readable medium of claim 1, wherein the method further comprises: analyzing the audio data to identify an item in the audio data; representing the item using a graphical symbol; and displaying the graphical symbol in conjunction with the textual information.",
"7. The non-transitory computer readable medium of claim 1, wherein the method further comprises displaying the different portions of the textual information using different sets of visual display parameters.",
"8. The non-transitory computer readable medium of claim 1, wherein the method further comprises: associating the different parts of the textual information with different textual formats based on the association of the different parts of the textual information with the different voice properties.",
"9. The non-transitory computer readable medium of claim 1, wherein the method further comprises associating the different portions of the textual information with different textual formats based on the association of the different portions of the textual information with the different speakers.",
"10. The non-transitory computer readable medium of claim 1, wherein the method further comprises: determining that a portion of a speech is said in a specific linguistic tone; selecting a visual display parameter for a part of the textual information associated with the portion of the speech based on the specific linguistic tone; and displaying the part of the textual information using the selected visual display parameter.",
"11. The non-transitory computer readable medium of claim 1, wherein the method further comprises visually presenting information associated with a speaker in conjunction with a portion of the textual information associated with the speaker.",
"12. The non-transitory computer readable medium of claim 1, wherein the association of the presentation regions with the speakers is based on spatial orientation of the speakers.",
"13. The non-transitory computer readable medium of claim 1, wherein the association of the presentation regions with the speakers is based on positions of the speakers.",
"14. The non-transitory computer readable medium of claim 1, wherein the method further comprises: associating a part of the textual information with speech produced by a user; and avoiding presenting the part of the textual information associated with the speech produced by the user.",
"15. The non-transitory computer readable medium of claim 1, wherein the method further comprises: determining that a part of the textual information is associated with speech that do not involve the user; and avoiding presenting the part of the textual information associated.",
"16. The non-transitory computer readable medium of claim 15, wherein the speech that do not involve the user is a conversation that do not involve the user.",
"17. The non-transitory computer readable medium of claim 15, wherein the speech that do not involve the user is a speech not directed at the user.",
"18. A system for processing audio and visually presenting information, the method comprising: at least one processing unit configured to: obtain audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus; analyze the audio data to obtain textual information; analyze the audio data to associate different portions of the textual information with different speakers; use a head mounted display system to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information; analyze the audio data to identify different voice properties in different parts of the audio data; associate different parts of the textual information with different voice properties; determine information based on particular voice properties; and present the information determined based on the particular voice properties along the part of the textual information associated with the particular voice properties.",
"19. A method for processing audio and visually presenting information, the method comprising: obtaining audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus; analyzing the audio data to obtain textual information; analyzing the audio data to associate different portions of the textual information with different speakers; using a head mounted display system to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information; analyzing the audio data to identify different voice properties in different parts of the audio data; associating different parts of the textual information with different voice properties; determining information based on particular voice properties; and presenting the information determined based on the particular voice properties along the part of the textual information associated with the particular voice properties.",
"20. A non-transitory computer readable medium storing data and computer implementable instructions that when executed by at least one processor cause the at least one processor to perform a method for processing audio and visually presenting information, the method comprising: obtaining audio data captured by one or more audio sensors included in a wearable apparatus from an environment of a wearer of the wearable apparatus; analyzing the audio data to obtain textual information; analyzing the audio data to associate different portions of the textual information with different speakers; using a head mounted display system to present each portion of the textual information in a presentation region associated with the speaker associated with the portion of the textual information; analyzing the audio data to identify a nonverbal sound; and generating textual description of the nonverbal sound."
],
"description_excerpt": "The disclosed embodiments generally relate to systems and methods for processing audio. More particularly, the disclosed embodiments relate to systems and methods for processing audio to visually present auditory information.\n\nAudio as well as other sensors are now part of numerous devices, from intelligent personal assistant devices to mobile phones, and the availability of audio data and other information produced by these devices is increasing.\n\nVarious conditions may cause difficulties in maintaining socially appropriate eye contact, including social phobia, autism, Asperger syndrome, and so forth. Inappropriate eye contact may include avoidance of eye contact, abnormal eye contact pattern, excessive eye contact, starring inappropriately, and so forth.\n\nTantrums, including temper tantrums, meltdowns and sensory meltdowns, are emotional outbursts characterized by stubbornness, crying, screaming, defiance, ranting, hitting, and tirades. In the general population, tantrums are more common in childhood, and are the result of frustration. In some conditions, including autism and Asperger's syndrome, tantrums may be the response to sensory overloads.\n\nEcholalia is a speech disorder characterized by meaningless repetition of vocalization and speech made by another person. Palilalia is a speech disorder characterized by meaningless repetition of vocalization and speech made by the same person. The repetition may be of syllables, words, utterances, phrases, sentences, and so forth.",
"cpc": [
"G10L 21/0364",
"G06F 40/10",
"G10L 15/1822",
"G10L 15/26",
"G10L 17/00",
"G10L 25/66",
"H04R 1/406",
"H04R 2201/023",
"H04R 25/407",
"H04R 3/005",
"H04R 5/0335"
],
"ipc": [
"G06F 40/10",
"G10L 17/00",
"G10L 21/0364",
"H04R 1/40",
"H04R 25/00",
"H04R 3/00"
],
"inventors": [
"Ron Zass"
],
"filing_date": "2021-08-07",
"publication_date": "2023-12-05",
"grant_date": "2023-12-05",
"priority_date": "2016-07-16",
"application_number": "US-202117396657-A",
"family_id": "69586473",
"cited_by_count": 6,
"citations": [
"US5950157A",
"US5940798A",
"US20030018475A1",
"US20020103649A1",
"US6707921B2",
"US6882971B2",
"US20040107098A1",
"US20060238877A1",
"US20070172805A1",
"US20060167691A1",
"US10084920B1",
"US20080201141A1",
"US20100250257A1",
"US8825468B2",
"US8259954B2",
"US20100262419A1",
"US20090306981A1",
"US20090319265A1",
"US20100174533A1",
"US20120020503A1",
"US8441356B1",
"US8719016B1",
"US20100280336A1",
"US20110087491A1",
"US20110228914A1",
"US20120128186A1",
"US20120053929A1",
"US20140247926A1",
"US20120116772A1",
"US20120215532A1",
"US8326338B1",
"US20120265537A1",
"US20130080168A1",
"US9355648B2",
"US8183997B1",
"US20120128683A1",
"US9213705B1",
"US20150011842A1",
"US20150058013A1",
"US9536536B2",
"US9746916B2",
"US9799336B2",
"US20140081634A1",
"US20140122077A1",
"US9412375B2",
"US20140163960A1",
"US20140236596A1",
"US9282284B2",
"US20140379352A1",
"US9641942B2",
"US20150036856A1",
"US20150073309A1",
"US20150099946A1",
"US9870357B2",
"US20180285312A1",
"US20150302867A1",
"US20150334346A1",
"US20150356836A1",
"US20150373477A1",
"US20160019895A1",
"US20170208415A1",
"US20160064002A1",
"US20160098993A1",
"US20160133257A1",
"US10241741B2",
"US20160350664A1",
"US20160379638A1",
"US20170041699A1",
"US9894320B2",
"US20170076000A1",
"US10269372B1",
"US20170103748A1",
"US20170105662A1",
"US20170111303A1",
"US20170221500A1",
"US9749583B1",
"US20170364599A1",
"US10446166B2",
"US20180020285A1",
"US20210366505A1",
"US20180077483A1",
"US10268447B1",
"US10692485B1",
"US20180240458A1",
"US20190279654A1",
"US10803852B2",
"US10878802B2",
"US20200050863A1",
"US20200066294A1"
]
}
Record 595 of 8,000 in Patents full text (MLC-0201). Request the full dataset.