Patent · US11115630B1 · B1 · US
Custom and automated audio prompts for devices
- (11) Publication number
- US11115630B1
- (21) Application number
- 16/359,520
- (22) Filing date
- 2019-03-20
- (30) Priority date
- 2018-03-28
- (43) Publication date
- 2021-09-07
- (45) Date of grant
- 2021-09-07
- (51) IPC
- G06F 3/16; G06K 9/00; H04L 29/08; H04N 7/18
- (52) CPC
- H04N Pictorial communication, e.g. television: 7/186, 7/188
- G06F Electric digital data processing: 3/167
- G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 2009/00738, 9/00671, 9/00711, 9/00771
- G06V Image or video recognition or understanding: 20/20, 20/40, 20/44, 20/52
- H04L Transmission of digital information, e.g. telegraphic communication: 67/125
- (73) Assignee
- Amazon Technologies Inc
- (72) Inventors
- Elliott Lemberger; John Modestine; Kevin Park; Richard Carter Mosher; Trevor Grolle; Kirk David Bacon
- (54) Title
- Custom and automated audio prompts for devices
- (57) Abstract
A network-connected security device is communicatively coupled to an audio/video (A/V) recording and communication device having a camera and a speaker. A method receives video data captured by the camera, and performs an object recognition algorithm upon the received video data to identify an object therein. The method performs a table lookup using the identified object, into a data structure that associates objects with at least one description of a predefined voice message. The method selects a description of a predefined voice message associated with the identified object, and transmits the selected description's predefined voice message to the A/V recording and communication device for output through the speaker.
- Full text
- View on Google Patents
Claims (20)
- A method comprising: receiving audio prompt data from a user device; receiving, from the user device, a request to associate the audio prompt data with an object; storing first identifier data associated with the audio prompt data; storing second identifier data associated with the object; receiving first image data generated by an electronic device; determining that the first image data represents the object; based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and sending the audio prompt data to the electronic device.
- The method as recited in claim 1, further comprising: receiving second image data representing the object, and wherein the determining that the first image data represents the object comprises determining, based at least in part on the second image data, that the first image data represents the object.
- The method as recited in claim 1, wherein the object is a person, and wherein the method further comprises: receiving second image data representing the person, and wherein: the second identifier data represents an identity of the person; and the determining that the first image data represents the person comprises determining, based at least in part on the second image data, the identity of the person represented by the first image data.
- The method as recited in claim 1, wherein the first identifier data represents a description of the audio prompt data, and wherein the method further comprises: based at least in part on the determining that the first image data represents the object, determining that the description is associated with the second identifier data, and wherein the selecting the audio prompt data comprises selecting, based at least in part on the description being associated with the second identifier data, the audio prompt data.
- The method as recited in claim 1, further comprising: determining that the first identifier data is associated with the audio prompt data; determining that the first identifier data is associated with additional audio prompt data; determining a first value associated with the audio prompt data; determining a second value associated with the additional audio prompt data; and determining that the first value is greater than the second value, and wherein the selecting the audio prompt data is further based at least in part on the determining that the first value is greater than the second value.
- The method as recited in claim 1, further comprising: sending a message to the user device, the message including at least an image represented by the first image data and a description of the audio prompt data; and receiving, from the user device, a request to output the audio prompt data, and wherein the sending the audio prompt data to the electronic device is based at least in part on the receiving the request to output the audio prompt data.
- The method as recited in claim 1, further comprising: receiving audio data generated by the electronic device; and identifying user speech represented by the audio data, and wherein the selecting the audio prompt data is further based at least in part on the identifying the user speech.
- The method as recited in claim 1, wherein the receiving the request to associate the audio prompt data with the object comprises receiving, from the user device, at least: the first identifier data associated with the audio prompt data; and the second identifier associated with the object.
- The method as recited in claim 1, wherein the storing the first identifier data associated with the audio prompt data comprises storing the first identifier data that represents a description associated with the audio prompt data, the description including one or more words that identify the object.
- An electronic device comprising: a camera; one or more speakers; one or more processors; and one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the electronic device to perform operations comprising: receiving, from a system, audio prompt data; receiving, from the system, a request to associate the audio prompt data with an object; storing first identifier data associated with the audio prompt data; storing second identifier data associated with the object; generating first image data using the camera; determining that the first image data represents the object; based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and outputting, using the one or more speakers, sound represented by the audio prompt data.
- The electronic device as recited in claim 10, the one or more computer-readable media storing further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising: receiving second image data representing the object, and wherein the determining that the first image data represents the object comprises determining, based at least in part on the second image data, that the first image data represents the object.
- The electronic device as recited in claim 10, wherein the second identifier data represents in identity of the object, the object being a person, and wherein the one or more computer-readable media store further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising: receiving second image data representing the person and wherein: the determining that the first image data represents the person comprises determining, based at least in part on the second image data, the identity of the person represented by the first image data; and the selecting the audio prompt data comprises selecting, based at least in part on the identity, the audio prompt data.
- The electronic device as recited in claim 10, wherein the receiving the request to associate the audio prompt data with the object comprises receiving, from the system, at least: the first identifier data associated with the audio prompt data; and the second identifier data associated with the object.
- The electronic device as recited in claim 10, wherein: the first identifier data represents a description associated with the audio prompt data, the description including one or more words that identify the object; the one or more computer-readable media store further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising determining, based at least in part on the first image data representing the object, that the description includes the one or more words that identify the object; and the selecting the audio prompt data is based at least in part on the determining that the description includes the one or more words that identify the object.
- A method comprising: receiving, from a system, audio prompt data; receiving, from the system, a request to associate the audio prompt data with an object; storing identifier data associated with the audio prompt data, the identifier data representing a description that includes one or more words that identify the object; generating first image data using a camera; determining that the first image data represents the object; based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and outputting sound represented by the audio prompt data.
- The method as recited in claim 15, wherein the receiving the request to associate the audio prompt data with the object comprises at least receiving, from the system, the identifier data associated with the audio prompt data.
- The method as recited in claim 15, further comprising: receiving second image data representing the object, wherein the determining that the first image data represents the object comprises at least: analyzing the first image data using at least the second image data; and determining that the first image data represents the object.
- The method as recited in claim 15, wherein the determining that the first image data represents the object comprises at least determining that the first image data represents one or more characteristics; and determining that the one or more characteristics are associated with the object.
- The method as recited in claim 15, wherein the selecting the audio prompt data comprises at least: determining that the identifier data represents the description; determining that the description includes the one or more words that identify the object; and selecting the audio prompt data based at least in part on the description including the one or more words that identify the object.
- The method as recited in claim 15, wherein the outputting of sound represented by the audio prompt data comprises outputting the sound that includes user speech, the audio prompt data representing the user speech.
Description
Home security is a concern for many homeowners and renters. Those seeking to protect or monitor their homes often wish to have video and audio communications with visitors, for example, those visiting an external door or entryway. Audio/Video (A/V) recording and communication devices, such as doorbells, provide this functionality, and can also aid in crime detection and prevention. For example, audio and/or video captured by an A/V recording and communication device can be uploaded to the cloud and recorded on a remote server. Subsequent review of the A/V footage can aid law enforcement in capturing perpetrators of home burglaries and other crimes. Further, the presence of one or more A/V recording and communication devices on the exterior of a home, such as a doorbell unit at the entrance to the home, acts as a powerful deterrent against would-be burglars.
The various embodiments of the present custom and automated audio prompts for network-connected security devices now will be discussed in detail with an emphasis on highlighting the advantageous features. These embodiments depict the novel and non-obvious custom and automated audio prompts for network-connected security devices shown in the accompanying drawings, which are for illustrative purposes only. These drawings include the following figures, in which like numerals indicate like parts:
FIG. 1 is a functional block diagram illustrating a system for streaming and storing A/V content captured by an audio/video (A/V) recording and communication device according to various aspects of the present disclosure;
Citations (128)
- US4764953A
- US6072402A
- US5760848A
- US5428388A
- GB2286283A
- US6456322B1
- EP0944883A1
- WO1998039894A1
- US6192257B1
- US6429893B1
- US6271752B1
- US6633231B1
- US20020147982A1
- US6476858B1
- WO2001013638A1
- GB2354394A
- JP2001103463A
- GB2357387A
- US20020094111A1
- WO2001093220A1
- US6970183B1
- JP2002033839A
- JP2002125059A
- WO2002085019A1
- JP2002344640A
- JP2002342863A
- JP2002354137A
- JP2002368890A
- US20030043047A1
- WO2003028375A1
- US6658091B1
- US7683929B2
- JP2003283696A
- WO2003096696A1
- US7065196B2
- US20040095254A1
- JP2004128835A
- US20070103541A1
- US8139098B2
- US7193644B2
- US20050285934A1
- US8144183B2
- US8154581B2
- US6753774B2
- US20040086093A1
- US20040085205A1
- US20040135686A1
- US20040085450A1
- CN2585521Y
- US7062291B2
- US7738917B2
- GB2400958A
- EP1480462A1
- US7450638B2
- US20050111660A1
- US7085361B2
- US20060010199A1
- JP2005341040A
- US7109860B2
- WO2006038760A1
- US7304572B2
- US20060022816A1
- US7683924B2
- JP2006147650A
- US20060139449A1
- WO2006067782A1
- US20060156361A1
- CN2792061Y
- US7643056B2
- JP2006262342A
- US20070008081A1
- US7382249B2
- WO2007125143A1
- US8619136B2
- JP2009008925A
- US20100225455A1
- US20130057695A1
- US20150035987A1
- US20140267716A1
- US9058738B1
- US9065987B2
- US8872915B1
- US8937659B1
- US8941736B1
- US8947530B1
- US8823795B1
- US8953040B1
- US9013575B2
- US9049352B2
- US9053622B2
- US9736284B2
- US9060104B2
- US9060103B2
- US8780201B1
- US9196133B2
- US9094584B2
- US9113051B1
- US9113052B1
- US9118819B1
- US9142214B2
- US9160987B1
- US9165444B2
- US9342936B2
- US9247219B2
- US8842180B1
- US9237318B2
- US9179107B1
- US9179108B1
- US9172922B1
- US9743049B2
- US9230424B1
- US9179109B1
- US9799183B2
- US9197867B1
- US9172921B1
- US9508239B1
- US9786133B2
- US20150163463A1
- US20160134932A1
- US9253455B1
- US9769435B2
- US9172920B1
- US20170162225A1
- US20170359423A1
- US20170358186A1
- US20180176512A1
- US20180232201A1
- US20190087646A1
Record as JSON
{
"publication_number": "US11115630B1",
"country": "US",
"kind": "B1",
"title": "Custom and automated audio prompts for devices",
"abstract": "A network-connected security device is communicatively coupled to an audio/video (A/V) recording and communication device having a camera and a speaker. A method receives video data captured by the camera, and performs an object recognition algorithm upon the received video data to identify an object therein. The method performs a table lookup using the identified object, into a data structure that associates objects with at least one description of a predefined voice message. The method selects a description of a predefined voice message associated with the identified object, and transmits the selected description's predefined voice message to the A/V recording and communication device for output through the speaker.",
"claims": [
"1. A method comprising: receiving audio prompt data from a user device; receiving, from the user device, a request to associate the audio prompt data with an object; storing first identifier data associated with the audio prompt data; storing second identifier data associated with the object; receiving first image data generated by an electronic device; determining that the first image data represents the object; based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and sending the audio prompt data to the electronic device.",
"2. The method as recited in claim 1, further comprising: receiving second image data representing the object, and wherein the determining that the first image data represents the object comprises determining, based at least in part on the second image data, that the first image data represents the object.",
"3. The method as recited in claim 1, wherein the object is a person, and wherein the method further comprises: receiving second image data representing the person, and wherein: the second identifier data represents an identity of the person; and the determining that the first image data represents the person comprises determining, based at least in part on the second image data, the identity of the person represented by the first image data.",
"4. The method as recited in claim 1, wherein the first identifier data represents a description of the audio prompt data, and wherein the method further comprises: based at least in part on the determining that the first image data represents the object, determining that the description is associated with the second identifier data, and wherein the selecting the audio prompt data comprises selecting, based at least in part on the description being associated with the second identifier data, the audio prompt data.",
"5. The method as recited in claim 1, further comprising: determining that the first identifier data is associated with the audio prompt data; determining that the first identifier data is associated with additional audio prompt data; determining a first value associated with the audio prompt data; determining a second value associated with the additional audio prompt data; and determining that the first value is greater than the second value, and wherein the selecting the audio prompt data is further based at least in part on the determining that the first value is greater than the second value.",
"6. The method as recited in claim 1, further comprising: sending a message to the user device, the message including at least an image represented by the first image data and a description of the audio prompt data; and receiving, from the user device, a request to output the audio prompt data, and wherein the sending the audio prompt data to the electronic device is based at least in part on the receiving the request to output the audio prompt data.",
"7. The method as recited in claim 1, further comprising: receiving audio data generated by the electronic device; and identifying user speech represented by the audio data, and wherein the selecting the audio prompt data is further based at least in part on the identifying the user speech.",
"8. The method as recited in claim 1, wherein the receiving the request to associate the audio prompt data with the object comprises receiving, from the user device, at least: the first identifier data associated with the audio prompt data; and the second identifier associated with the object.",
"9. The method as recited in claim 1, wherein the storing the first identifier data associated with the audio prompt data comprises storing the first identifier data that represents a description associated with the audio prompt data, the description including one or more words that identify the object.",
"10. An electronic device comprising: a camera; one or more speakers; one or more processors; and one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the electronic device to perform operations comprising: receiving, from a system, audio prompt data; receiving, from the system, a request to associate the audio prompt data with an object; storing first identifier data associated with the audio prompt data; storing second identifier data associated with the object; generating first image data using the camera; determining that the first image data represents the object; based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and outputting, using the one or more speakers, sound represented by the audio prompt data.",
"11. The electronic device as recited in claim 10, the one or more computer-readable media storing further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising: receiving second image data representing the object, and wherein the determining that the first image data represents the object comprises determining, based at least in part on the second image data, that the first image data represents the object.",
"12. The electronic device as recited in claim 10, wherein the second identifier data represents in identity of the object, the object being a person, and wherein the one or more computer-readable media store further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising: receiving second image data representing the person and wherein: the determining that the first image data represents the person comprises determining, based at least in part on the second image data, the identity of the person represented by the first image data; and the selecting the audio prompt data comprises selecting, based at least in part on the identity, the audio prompt data.",
"13. The electronic device as recited in claim 10, wherein the receiving the request to associate the audio prompt data with the object comprises receiving, from the system, at least: the first identifier data associated with the audio prompt data; and the second identifier data associated with the object.",
"14. The electronic device as recited in claim 10, wherein: the first identifier data represents a description associated with the audio prompt data, the description including one or more words that identify the object; the one or more computer-readable media store further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising determining, based at least in part on the first image data representing the object, that the description includes the one or more words that identify the object; and the selecting the audio prompt data is based at least in part on the determining that the description includes the one or more words that identify the object.",
"15. A method comprising: receiving, from a system, audio prompt data; receiving, from the system, a request to associate the audio prompt data with an object; storing identifier data associated with the audio prompt data, the identifier data representing a description that includes one or more words that identify the object; generating first image data using a camera; determining that the first image data represents the object; based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and outputting sound represented by the audio prompt data.",
"16. The method as recited in claim 15, wherein the receiving the request to associate the audio prompt data with the object comprises at least receiving, from the system, the identifier data associated with the audio prompt data.",
"17. The method as recited in claim 15, further comprising: receiving second image data representing the object, wherein the determining that the first image data represents the object comprises at least: analyzing the first image data using at least the second image data; and determining that the first image data represents the object.",
"18. The method as recited in claim 15, wherein the determining that the first image data represents the object comprises at least determining that the first image data represents one or more characteristics; and determining that the one or more characteristics are associated with the object.",
"19. The method as recited in claim 15, wherein the selecting the audio prompt data comprises at least: determining that the identifier data represents the description; determining that the description includes the one or more words that identify the object; and selecting the audio prompt data based at least in part on the description including the one or more words that identify the object.",
"20. The method as recited in claim 15, wherein the outputting of sound represented by the audio prompt data comprises outputting the sound that includes user speech, the audio prompt data representing the user speech."
],
"description_excerpt": "Home security is a concern for many homeowners and renters. Those seeking to protect or monitor their homes often wish to have video and audio communications with visitors, for example, those visiting an external door or entryway. Audio/Video (A/V) recording and communication devices, such as doorbells, provide this functionality, and can also aid in crime detection and prevention. For example, audio and/or video captured by an A/V recording and communication device can be uploaded to the cloud and recorded on a remote server. Subsequent review of the A/V footage can aid law enforcement in capturing perpetrators of home burglaries and other crimes. Further, the presence of one or more A/V recording and communication devices on the exterior of a home, such as a doorbell unit at the entrance to the home, acts as a powerful deterrent against would-be burglars.\n\nThe various embodiments of the present custom and automated audio prompts for network-connected security devices now will be discussed in detail with an emphasis on highlighting the advantageous features. These embodiments depict the novel and non-obvious custom and automated audio prompts for network-connected security devices shown in the accompanying drawings, which are for illustrative purposes only. These drawings include the following figures, in which like numerals indicate like parts:\n\nFIG. 1 is a functional block diagram illustrating a system for streaming and storing A/V content captured by an audio/video (A/V) recording and communication device according to various aspects of the present disclosure;",
"cpc": [
"H04N 7/186",
"G06F 3/167",
"G06K 2009/00738",
"G06K 9/00671",
"G06K 9/00711",
"G06K 9/00771",
"G06V 20/20",
"G06V 20/40",
"G06V 20/44",
"G06V 20/52",
"H04L 67/125",
"H04N 7/188"
],
"ipc": [
"G06F 3/16",
"G06K 9/00",
"H04L 29/08",
"H04N 7/18"
],
"assignees": [
"Amazon Technologies Inc"
],
"inventors": [
"Elliott Lemberger",
"John Modestine",
"Kevin Park",
"Richard Carter Mosher",
"Trevor Grolle",
"Kirk David Bacon"
],
"filing_date": "2019-03-20",
"publication_date": "2021-09-07",
"grant_date": "2021-09-07",
"priority_date": "2018-03-28",
"application_number": "US-201916359520-A",
"family_id": "77559196",
"cited_by_count": 13,
"citations": [
"US4764953A",
"US6072402A",
"US5760848A",
"US5428388A",
"GB2286283A",
"US6456322B1",
"EP0944883A1",
"WO1998039894A1",
"US6192257B1",
"US6429893B1",
"US6271752B1",
"US6633231B1",
"US20020147982A1",
"US6476858B1",
"WO2001013638A1",
"GB2354394A",
"JP2001103463A",
"GB2357387A",
"US20020094111A1",
"WO2001093220A1",
"US6970183B1",
"JP2002033839A",
"JP2002125059A",
"WO2002085019A1",
"JP2002344640A",
"JP2002342863A",
"JP2002354137A",
"JP2002368890A",
"US20030043047A1",
"WO2003028375A1",
"US6658091B1",
"US7683929B2",
"JP2003283696A",
"WO2003096696A1",
"US7065196B2",
"US20040095254A1",
"JP2004128835A",
"US20070103541A1",
"US8139098B2",
"US7193644B2",
"US20050285934A1",
"US8144183B2",
"US8154581B2",
"US6753774B2",
"US20040086093A1",
"US20040085205A1",
"US20040135686A1",
"US20040085450A1",
"CN2585521Y",
"US7062291B2",
"US7738917B2",
"GB2400958A",
"EP1480462A1",
"US7450638B2",
"US20050111660A1",
"US7085361B2",
"US20060010199A1",
"JP2005341040A",
"US7109860B2",
"WO2006038760A1",
"US7304572B2",
"US20060022816A1",
"US7683924B2",
"JP2006147650A",
"US20060139449A1",
"WO2006067782A1",
"US20060156361A1",
"CN2792061Y",
"US7643056B2",
"JP2006262342A",
"US20070008081A1",
"US7382249B2",
"WO2007125143A1",
"US8619136B2",
"JP2009008925A",
"US20100225455A1",
"US20130057695A1",
"US20150035987A1",
"US20140267716A1",
"US9058738B1",
"US9065987B2",
"US8872915B1",
"US8937659B1",
"US8941736B1",
"US8947530B1",
"US8823795B1",
"US8953040B1",
"US9013575B2",
"US9049352B2",
"US9053622B2",
"US9736284B2",
"US9060104B2",
"US9060103B2",
"US8780201B1",
"US9196133B2",
"US9094584B2",
"US9113051B1",
"US9113052B1",
"US9118819B1",
"US9142214B2",
"US9160987B1",
"US9165444B2",
"US9342936B2",
"US9247219B2",
"US8842180B1",
"US9237318B2",
"US9179107B1",
"US9179108B1",
"US9172922B1",
"US9743049B2",
"US9230424B1",
"US9179109B1",
"US9799183B2",
"US9197867B1",
"US9172921B1",
"US9508239B1",
"US9786133B2",
"US20150163463A1",
"US20160134932A1",
"US9253455B1",
"US9769435B2",
"US9172920B1",
"US20170162225A1",
"US20170359423A1",
"US20170358186A1",
"US20180176512A1",
"US20180232201A1",
"US20190087646A1"
]
}
Record 1,491 of 8,000 in Patents full text (MLC-0201). Request the full dataset.