Patent · US9526127B1 · B1 · US
Affecting the behavior of a user device based on a user's gaze
- (11) Publication number
- US9526127B1
- (21) Application number
- 13/676,517
- (22) Filing date
- 2012-11-14
- (30) Priority date
- 2011-11-18
- (43) Publication date
- 2016-12-20
- (45) Date of grant
- 2016-12-20
- (51) IPC
- H04W 88/02
- (52) CPC
- H04W Wireless communication networks: 88/02
- A63F Card, board, or roulette games; indoor games using small moving playing bodies; video games; games not otherwise provided for: 2300/1093
- G06F Electric digital data processing: 3/167
- G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 15/22, 15/25, 2015/223, 25/78
- H04N Pictorial communication, e.g. television: 21/44218
- (73) Assignee
- Google LLC
- (72) Inventors
- Gabriel Taubman; William J. Byrne
- (54) Title
- Affecting the behavior of a user device based on a user's gaze
- (57) Abstract
A device may determine whether a user is facing a display screen associated with the device; and present feedback to the user. When presenting the feedback, the device may present visual information, that is based on the feedback, when determining that the user is facing the display screen associated with the device, and present audio information, that is based on the feedback, when determining that the user is not facing the display screen associated with the device. At least a portion of the audio information might not be presented when the visual information is presented when the user is facing the display screen associated with the device.
- Full text
- View on Google Patents
Claims (20)
- A computer-implemented method comprising: receiving, by a user device, audio data corresponding to an utterance of a user; determining that a first portion of a transcription of the utterance includes a keyword that is associated with a voice command; after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command, determining, by the user device, that an image of the user does not include one or more features that are characteristic of the user facing a display of the user device; and in response to determining that the image of the user does not include one or more features that are characteristic of the user facing the display of the user device, determining, by the user device, to prevent a second portion of the transcription of the utterance from being processed as a voice command.
- The method of claim 1, comprising: discarding the second portion of the transcription without inputting the second portion of the transcription to a dialog engine based on determining to prevent the second portion of the transcription of the utterance from being processed as a voice command.
- The method of claim 1, comprising: generating the image of the user by a camera on the user device after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command.
- The method of claim 1, wherein the voice command is a command for the user to take an action other than taking a picture.
- The method of claim 1, wherein determining the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining a direction of a gaze of the user; and classifying the gaze of the user as a gaze that is not directed toward the display of the user device.
- The method of claim 1, wherein determining that the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining that one or more particular facial features are not visible in the image.
- The method of claim 1, wherein determining that the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining an orientation of two or more particular facial features of the user with respect to each other.
- The method of claim 1, wherein determining that the image includes one or more features that are characteristic of the user facing a display of the user device comprises: comparing the image to a different image in which the user is labeled as facing the display of the mobile device.
- A system comprising: one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: receiving, by a user device, audio data corresponding to an utterance of a user; determining that a first portion of a transcription of the utterance includes a keyword that is associated with a voice command; after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command, determining, by the user device, that an image of the user does not include one or more features that are characteristic of the user facing a display of the user device; and in response to determining that the image of the user does not include one or more features that are characteristic of the user facing the display of the user device, determining, by the user device, to prevent a second portion of the transcription of the utterance from being processed as a voice command.
- The system of claim 9, wherein the operations further comprise: discarding the second portion of the transcription without inputting the second portion of the transcription to a dialog engine based on determining to prevent the second portion of the transcription of the utterance from being processed as a voice command.
- The system of claim 9, wherein the operations further comprise: generating the image of the user by a camera on the user device after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command.
- The system of claim 9, wherein the voice command is a command for the user to take an action other than taking a picture.
- The system of claim 9, wherein determining the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining a direction of a gaze of the user; and classifying the gaze of the user as a gaze that is not directed toward the display of the user device.
- The system of claim 9, wherein determining that the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining that one or more particular facial features are not visible in the image.
- The system of claim 9, wherein determining that the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining an orientation of two or more particular facial features of the user with respect to each other.
- The system of claim 9, wherein determining that the image includes one or more features that are characteristic of the user facing a display of the user device comprises: comparing the image to a different image in which the user is labeled as facing the display of the mobile device.
- A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising: receiving, by a user device, audio data corresponding to an utterance of a user; determining that a first portion of a transcription of the utterance includes a keyword that is associated with a voice command; after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command, determining, by the user device, that an image of the user does not include one or more features that are characteristic of the user facing a display of the user device; and in response to determining that the image of the user does not include one or more features that are characteristic of the user facing the display of the user device, determining, by the user device, to prevent a second portion of the transcription of the utterance from being processed as a voice command.
- The medium of claim 17, wherein the operations further comprise: discarding the second portion of the transcription without inputting the second portion of the transcription to a dialog engine based on determining to prevent the second portion of the transcription of the utterance from being processed as a voice command.
- The medium of claim 17, wherein the operations further comprise: generating the image of the user by a camera on the user device after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command.
- The medium of claim 17, wherein the voice command is a command for the user to take an action other than taking a picture.
Description
User devices, such as mobile telephones, can provide users with information. For example, when a mobile telephone rings, the mobile telephone may provide a visual notification (e.g., a screen of the mobile telephone may display call information), an audible notification (e.g., a ring tone), and/or a sensory notification (e.g., the mobile telephone may vibrate).
According to one general implementation, one or more notifications of a user device may be identified as redundant, based on estimating whether the gaze of the user of the user device is directed towards a display of the user device. For example, if a user's gaze is deemed to be directed towards the user device, the user may not need to hear a ring tone and/or feel a vibration, since the user can already see that a call has been received. Accordingly, the ring tone and/or vibration may be identified as redundant, and may not be output. In another example, the user's gaze may be directed away from the display of the user device and may not be able to see a visual notification. Accordingly, the visual notification may be identified as redundant, and may not be output.
Citations (14)
- US6754373B1
- US20020105575A1
- US20060111916A1
- US20090315827A1
- US20080146289A1
- US20130013316A1
- US20100205667A1
- US20120007713A1
- US8451314B1
- US20110216093A1
- US8635066B2
- US20120157114A1
- US20120259638A1
- US20120300061A1
Record as JSON
{
"publication_number": "US9526127B1",
"country": "US",
"kind": "B1",
"title": "Affecting the behavior of a user device based on a user's gaze",
"abstract": "A device may determine whether a user is facing a display screen associated with the device; and present feedback to the user. When presenting the feedback, the device may present visual information, that is based on the feedback, when determining that the user is facing the display screen associated with the device, and present audio information, that is based on the feedback, when determining that the user is not facing the display screen associated with the device. At least a portion of the audio information might not be presented when the visual information is presented when the user is facing the display screen associated with the device.",
"claims": [
"1. A computer-implemented method comprising: receiving, by a user device, audio data corresponding to an utterance of a user; determining that a first portion of a transcription of the utterance includes a keyword that is associated with a voice command; after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command, determining, by the user device, that an image of the user does not include one or more features that are characteristic of the user facing a display of the user device; and in response to determining that the image of the user does not include one or more features that are characteristic of the user facing the display of the user device, determining, by the user device, to prevent a second portion of the transcription of the utterance from being processed as a voice command.",
"2. The method of claim 1, comprising: discarding the second portion of the transcription without inputting the second portion of the transcription to a dialog engine based on determining to prevent the second portion of the transcription of the utterance from being processed as a voice command.",
"3. The method of claim 1, comprising: generating the image of the user by a camera on the user device after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command.",
"4. The method of claim 1, wherein the voice command is a command for the user to take an action other than taking a picture.",
"5. The method of claim 1, wherein determining the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining a direction of a gaze of the user; and classifying the gaze of the user as a gaze that is not directed toward the display of the user device.",
"6. The method of claim 1, wherein determining that the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining that one or more particular facial features are not visible in the image.",
"7. The method of claim 1, wherein determining that the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining an orientation of two or more particular facial features of the user with respect to each other.",
"8. The method of claim 1, wherein determining that the image includes one or more features that are characteristic of the user facing a display of the user device comprises: comparing the image to a different image in which the user is labeled as facing the display of the mobile device.",
"9. A system comprising: one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising: receiving, by a user device, audio data corresponding to an utterance of a user; determining that a first portion of a transcription of the utterance includes a keyword that is associated with a voice command; after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command, determining, by the user device, that an image of the user does not include one or more features that are characteristic of the user facing a display of the user device; and in response to determining that the image of the user does not include one or more features that are characteristic of the user facing the display of the user device, determining, by the user device, to prevent a second portion of the transcription of the utterance from being processed as a voice command.",
"10. The system of claim 9, wherein the operations further comprise: discarding the second portion of the transcription without inputting the second portion of the transcription to a dialog engine based on determining to prevent the second portion of the transcription of the utterance from being processed as a voice command.",
"11. The system of claim 9, wherein the operations further comprise: generating the image of the user by a camera on the user device after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command.",
"12. The system of claim 9, wherein the voice command is a command for the user to take an action other than taking a picture.",
"13. The system of claim 9, wherein determining the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining a direction of a gaze of the user; and classifying the gaze of the user as a gaze that is not directed toward the display of the user device.",
"14. The system of claim 9, wherein determining that the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining that one or more particular facial features are not visible in the image.",
"15. The system of claim 9, wherein determining that the image does not include one or more features that are characteristic of the user facing a display of the user device comprises: determining an orientation of two or more particular facial features of the user with respect to each other.",
"16. The system of claim 9, wherein determining that the image includes one or more features that are characteristic of the user facing a display of the user device comprises: comparing the image to a different image in which the user is labeled as facing the display of the mobile device.",
"17. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising: receiving, by a user device, audio data corresponding to an utterance of a user; determining that a first portion of a transcription of the utterance includes a keyword that is associated with a voice command; after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command, determining, by the user device, that an image of the user does not include one or more features that are characteristic of the user facing a display of the user device; and in response to determining that the image of the user does not include one or more features that are characteristic of the user facing the display of the user device, determining, by the user device, to prevent a second portion of the transcription of the utterance from being processed as a voice command.",
"18. The medium of claim 17, wherein the operations further comprise: discarding the second portion of the transcription without inputting the second portion of the transcription to a dialog engine based on determining to prevent the second portion of the transcription of the utterance from being processed as a voice command.",
"19. The medium of claim 17, wherein the operations further comprise: generating the image of the user by a camera on the user device after determining that the first portion of the transcription of the utterance includes the keyword that is associated with the voice command.",
"20. The medium of claim 17, wherein the voice command is a command for the user to take an action other than taking a picture."
],
"description_excerpt": "User devices, such as mobile telephones, can provide users with information. For example, when a mobile telephone rings, the mobile telephone may provide a visual notification (e.g., a screen of the mobile telephone may display call information), an audible notification (e.g., a ring tone), and/or a sensory notification (e.g., the mobile telephone may vibrate).\n\nAccording to one general implementation, one or more notifications of a user device may be identified as redundant, based on estimating whether the gaze of the user of the user device is directed towards a display of the user device. For example, if a user's gaze is deemed to be directed towards the user device, the user may not need to hear a ring tone and/or feel a vibration, since the user can already see that a call has been received. Accordingly, the ring tone and/or vibration may be identified as redundant, and may not be output. In another example, the user's gaze may be directed away from the display of the user device and may not be able to see a visual notification. Accordingly, the visual notification may be identified as redundant, and may not be output.",
"cpc": [
"H04W 88/02",
"A63F 2300/1093",
"G06F 3/167",
"G10L 15/22",
"G10L 15/25",
"G10L 2015/223",
"G10L 25/78",
"H04N 21/44218"
],
"ipc": [
"H04W 88/02"
],
"assignees": [
"Google LLC"
],
"inventors": [
"Gabriel Taubman",
"William J. Byrne"
],
"filing_date": "2012-11-14",
"publication_date": "2016-12-20",
"grant_date": "2016-12-20",
"priority_date": "2011-11-18",
"application_number": "US-201213676517-A",
"family_id": "57538764",
"cited_by_count": 132,
"citations": [
"US6754373B1",
"US20020105575A1",
"US20060111916A1",
"US20090315827A1",
"US20080146289A1",
"US20130013316A1",
"US20100205667A1",
"US20120007713A1",
"US8451314B1",
"US20110216093A1",
"US8635066B2",
"US20120157114A1",
"US20120259638A1",
"US20120300061A1"
]
}
Record 4,408 of 8,000 in Patents full text (MLC-0201). Request the full dataset.