MLchartDataset catalogue

Patent · US10657951B2 · B2 · US

Controlling synthesized speech output from a voice-controlled device

(11) Publication number
US10657951B2
(21) Application number
15/854,142
(22) Filing date
2017-12-26
(30) Priority date
2017-12-26
(43) Publication date
2020-05-19
(45) Date of grant
2020-05-19
(51) IPC
B25J 11/00; B25J 13/00; B25J 9/00; G06K 9/00; G10L 13/00; G10L 13/02; G10L 13/027; G10L 15/08; G10L 15/18; G10L 15/22; G10L 15/26; G10L 21/0208; G10L 25/84
(52) CPC
  • G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 13/02, 13/00, 15/18, 15/22, 15/26, 2015/088, 2015/228, 2021/02087, 25/84
  • B25J Manipulators; chambers provided with manipulation devices: 11/0005, 13/003, 9/0003
  • G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/00228, 9/00288
  • G06V Image or video recognition or understanding: 40/161, 40/172
(73) Assignee
International Business Machines Corp
(72) Inventors
Shang Qing Guo; Jonathan Lenchner
(54) Title
Controlling synthesized speech output from a voice-controlled device
(57) Abstract

A system, a computer program product, and method for controlling synthesized speech output on a voice-controlled device. One or more sensors are used to detect whether one person or more than one person is within a first settable distance from the voice-controlled device. Next a determination is made whether the audio input received is recognized as speech. In response to only one person being detected within the settable distance, begin outputting synthesized speech based on the audio input without waiting for an attention word to be recognized and otherwise wait for additional criteria before outputting synthesized speech based on the speech input. The additional criteria includes determining that more than one person is detected and recognizing that the attention word is received before outputting synthesized speech based on the audio input. Another criteria includes more than one person being detected and determining that the person is looking at the voice-controlled device before outputting synthesized speech based on the audio input.

Full text
View on Google Patents

Claims (20)

  1. A computer-implemented method on a voice-controlled device for controlling synthesized speech output, the method comprising: detecting, with at least one sensor, whether one person or more than one person is within a first settable distance from the voice-controlled device; determining whether audio input being received is speech; and in response to only one person being detected within the first settable distance and the audio input being speech, initiating an output of synthesized speech based on the audio input without waiting for an attention word to be recognized and otherwise waiting for additional criteria before outputting synthesized speech.
  2. The computer-implemented method of claim 1, wherein the additional criteria includes determining that more than one person is detected and recognizing that the attention word is received before outputting synthesized speech.
  3. The computer-implemented method of claim 1, wherein in response to more than one person being detected, further comprising: determining that one person of the more than one person being detected is within a second settable distance from the voice-controlled device and determining that the one person is looking at the voice-controlled device before outputting synthesized speech.
  4. The computer-implemented method of claim 3, wherein the second settable distance is closer to the voice-controlled device as compared with the first settable distance.
  5. The computer-implemented method of claim 3, wherein the voice-controlled device is a robot with eyes or a face and in response to the one person looking at the eyes or the face of the robot then outputting synthesized speech.
  6. The computer-implemented method of claim 3, wherein the at least one sensor is a visible light camera sensor, an infra-red camera sensor, laser pulse-based radar sensor, an acoustical sensor, or a combination thereof.
  7. The computer-implemented method of claim 1, further comprising: capturing audio input while outputting synthesized speech; in response to the captured audio being recognized as speech, pausing the outputting of synthesized speech; or in response to the captured being not recognized as speech and a volume of the audio input being above a settable background noise threshold, pausing the outputting of synthesized speech.
  8. The computer-implemented method of claim 7, further comprising: filtering out the outputting of synthesized speech based on a synthesized speech input from the audio input.
  9. The computer-implemented method of claim 7, further comprising: in response to the captured audio being not recognized as speech and a volume of the audio input below or equal to a settable background noise threshold, and the pausing of the output of synthesized speech being within a settable pause timeframe, resuming the output of synthesized speech.
  10. A voice-controlled device for controlling synthesized speech output, the computer system comprising: a processor device; and a memory operably coupled to the processor device and storing computer-executable instructions causing: detecting, with at least one sensor, whether one person or more than one person is within a first settable distance from the voice-controlled device; determining whether audio input is being received is speech; and in response to only one person being detected within the first settable distance and the audio input being speech, initiating an output of synthesized speech without waiting for an attention word to be recognized and otherwise waiting for additional criteria before outputting synthesized speech.
  11. The voice-controlled device of claim 10, wherein the additional criteria includes determining that more than one person is detected and recognizing that the attention word is received before outputting synthesized speech.
  12. The voice-controlled device of claim 10, wherein in response to more than one person being detected, further comprising: determining that one person of the more than one person being detected is within a second settable distance from the voice-controlled device and determining that the one person is looking at the voice-controlled device before outputting synthesized speech.
  13. The voice-controlled device of claim 12, wherein the second settable distance is closer to the voice-controlled device as compared with the first settable distance.
  14. The voice-controlled device of claim 12, wherein the voice-controlled device is a robot with eyes or a face and in response to the one person looking at the eyes or the face of the robot then outputting synthesized speech.
  15. The voice-controlled device of claim 12, wherein the at least one sensor is a visible light camera sensor, an infra-red camera sensor, laser pulse-based radar sensor, an acoustical sensor, or a combination thereof.
  16. A computer program product for controlling synthesized speech output, the computer program product comprising: a non-transitory computer readable storage medium readable by a processing device and storing program instructions for execution by the processing device, said program instructions comprising: detecting, with at least one sensor, whether one person or more than one person is within a first settable distance from a voice-controlled device; determining whether audio input is being received is speech; and in response to only one person being detected within the first settable distance and the audio input being speech, initiating an output of synthesized speech without waiting for an attention word to be recognized and otherwise waiting for additional criteria before outputting synthesized speech.
  17. The computer program product of claim 16, wherein the additional criteria includes determining that more than one person is detected and recognizing that the attention word is received before outputting synthesized speech.
  18. The computer program product of claim 16, wherein in response to more than one person being detected, further comprising: determining that one person of the more than one person being detected is within a second settable distance from the voice-controlled device and determining that the one person is looking at the voice-controlled device before outputting synthesized speech.
  19. The computer program product of claim 18, wherein the second settable distance is closer to the voice-controlled device as compared with the first settable distance.
  20. The computer program product of claim 18, wherein the voice-controlled device is a robot with eyes or a face and in response to the one person looking at the eyes or the face of the robot then outputting synthesized speech.

Description

The present invention generally relates to voice-controlled devices and more specifically relates to controlling synthesized speech output therefrom.

Speech recognition systems are rapidly increasing in significance in many areas of data and communications technology. In recent years, speech recognition has advanced to the point where it is used by millions of people across various applications. Speech recognition applications now include interactive voice response systems, voice dialing, data entry, dictation mode systems including medical transcription, automotive applications, etc. There are also “command and control” applications that utilize speech recognition for controlling tasks such as adjusting the climate control in a vehicle or requesting a smart phone to play a particular song.

Voice controlled devices, such as robots, smart speakers, intelligent personal assistants are typically placed in an environment where there are humans. Other voice controlled devices may also be present. These voice controlled devices face several obstacles about when to speak or respond. For example the voice controlled devices may hear an utterance and not know for sure whether the utterance is directed to it or not. Likewise the voice controlled devices may be in the presence of one or more people and not know whether or not it is appropriate to start up a conversation. Further the voice controlled devices may start speaking and while speaking it may discern a human or another voice controlled device speaking.

Citations (11)

  • US6363301B1
  • US6967455B2
  • US20070192910A1
  • US20140095007A1
  • US20130094656A1
  • WO2013168363A1
  • US20150120060A1
  • EP2933071A1
  • WO2016174594A1
  • US20160350553A1
  • US10027662B1
Record as JSON
{
  "publication_number": "US10657951B2",
  "country": "US",
  "kind": "B2",
  "title": "Controlling synthesized speech output from a voice-controlled device",
  "abstract": "A system, a computer program product, and method for controlling synthesized speech output on a voice-controlled device. One or more sensors are used to detect whether one person or more than one person is within a first settable distance from the voice-controlled device. Next a determination is made whether the audio input received is recognized as speech. In response to only one person being detected within the settable distance, begin outputting synthesized speech based on the audio input without waiting for an attention word to be recognized and otherwise wait for additional criteria before outputting synthesized speech based on the speech input. The additional criteria includes determining that more than one person is detected and recognizing that the attention word is received before outputting synthesized speech based on the audio input. Another criteria includes more than one person being detected and determining that the person is looking at the voice-controlled device before outputting synthesized speech based on the audio input.",
  "claims": [
    "1. A computer-implemented method on a voice-controlled device for controlling synthesized speech output, the method comprising: detecting, with at least one sensor, whether one person or more than one person is within a first settable distance from the voice-controlled device; determining whether audio input being received is speech; and in response to only one person being detected within the first settable distance and the audio input being speech, initiating an output of synthesized speech based on the audio input without waiting for an attention word to be recognized and otherwise waiting for additional criteria before outputting synthesized speech.",
    "2. The computer-implemented method of claim 1, wherein the additional criteria includes determining that more than one person is detected and recognizing that the attention word is received before outputting synthesized speech.",
    "3. The computer-implemented method of claim 1, wherein in response to more than one person being detected, further comprising: determining that one person of the more than one person being detected is within a second settable distance from the voice-controlled device and determining that the one person is looking at the voice-controlled device before outputting synthesized speech.",
    "4. The computer-implemented method of claim 3, wherein the second settable distance is closer to the voice-controlled device as compared with the first settable distance.",
    "5. The computer-implemented method of claim 3, wherein the voice-controlled device is a robot with eyes or a face and in response to the one person looking at the eyes or the face of the robot then outputting synthesized speech.",
    "6. The computer-implemented method of claim 3, wherein the at least one sensor is a visible light camera sensor, an infra-red camera sensor, laser pulse-based radar sensor, an acoustical sensor, or a combination thereof.",
    "7. The computer-implemented method of claim 1, further comprising: capturing audio input while outputting synthesized speech; in response to the captured audio being recognized as speech, pausing the outputting of synthesized speech; or in response to the captured being not recognized as speech and a volume of the audio input being above a settable background noise threshold, pausing the outputting of synthesized speech.",
    "8. The computer-implemented method of claim 7, further comprising: filtering out the outputting of synthesized speech based on a synthesized speech input from the audio input.",
    "9. The computer-implemented method of claim 7, further comprising: in response to the captured audio being not recognized as speech and a volume of the audio input below or equal to a settable background noise threshold, and the pausing of the output of synthesized speech being within a settable pause timeframe, resuming the output of synthesized speech.",
    "10. A voice-controlled device for controlling synthesized speech output, the computer system comprising: a processor device; and a memory operably coupled to the processor device and storing computer-executable instructions causing: detecting, with at least one sensor, whether one person or more than one person is within a first settable distance from the voice-controlled device; determining whether audio input is being received is speech; and in response to only one person being detected within the first settable distance and the audio input being speech, initiating an output of synthesized speech without waiting for an attention word to be recognized and otherwise waiting for additional criteria before outputting synthesized speech.",
    "11. The voice-controlled device of claim 10, wherein the additional criteria includes determining that more than one person is detected and recognizing that the attention word is received before outputting synthesized speech.",
    "12. The voice-controlled device of claim 10, wherein in response to more than one person being detected, further comprising: determining that one person of the more than one person being detected is within a second settable distance from the voice-controlled device and determining that the one person is looking at the voice-controlled device before outputting synthesized speech.",
    "13. The voice-controlled device of claim 12, wherein the second settable distance is closer to the voice-controlled device as compared with the first settable distance.",
    "14. The voice-controlled device of claim 12, wherein the voice-controlled device is a robot with eyes or a face and in response to the one person looking at the eyes or the face of the robot then outputting synthesized speech.",
    "15. The voice-controlled device of claim 12, wherein the at least one sensor is a visible light camera sensor, an infra-red camera sensor, laser pulse-based radar sensor, an acoustical sensor, or a combination thereof.",
    "16. A computer program product for controlling synthesized speech output, the computer program product comprising: a non-transitory computer readable storage medium readable by a processing device and storing program instructions for execution by the processing device, said program instructions comprising: detecting, with at least one sensor, whether one person or more than one person is within a first settable distance from a voice-controlled device; determining whether audio input is being received is speech; and in response to only one person being detected within the first settable distance and the audio input being speech, initiating an output of synthesized speech without waiting for an attention word to be recognized and otherwise waiting for additional criteria before outputting synthesized speech.",
    "17. The computer program product of claim 16, wherein the additional criteria includes determining that more than one person is detected and recognizing that the attention word is received before outputting synthesized speech.",
    "18. The computer program product of claim 16, wherein in response to more than one person being detected, further comprising: determining that one person of the more than one person being detected is within a second settable distance from the voice-controlled device and determining that the one person is looking at the voice-controlled device before outputting synthesized speech.",
    "19. The computer program product of claim 18, wherein the second settable distance is closer to the voice-controlled device as compared with the first settable distance.",
    "20. The computer program product of claim 18, wherein the voice-controlled device is a robot with eyes or a face and in response to the one person looking at the eyes or the face of the robot then outputting synthesized speech."
  ],
  "description_excerpt": "The present invention generally relates to voice-controlled devices and more specifically relates to controlling synthesized speech output therefrom.\n\nSpeech recognition systems are rapidly increasing in significance in many areas of data and communications technology. In recent years, speech recognition has advanced to the point where it is used by millions of people across various applications. Speech recognition applications now include interactive voice response systems, voice dialing, data entry, dictation mode systems including medical transcription, automotive applications, etc. There are also “command and control” applications that utilize speech recognition for controlling tasks such as adjusting the climate control in a vehicle or requesting a smart phone to play a particular song.\n\nVoice controlled devices, such as robots, smart speakers, intelligent personal assistants are typically placed in an environment where there are humans. Other voice controlled devices may also be present. These voice controlled devices face several obstacles about when to speak or respond. For example the voice controlled devices may hear an utterance and not know for sure whether the utterance is directed to it or not. Likewise the voice controlled devices may be in the presence of one or more people and not know whether or not it is appropriate to start up a conversation. Further the voice controlled devices may start speaking and while speaking it may discern a human or another voice controlled device speaking.",
  "cpc": [
    "G10L 13/02",
    "B25J 11/0005",
    "B25J 13/003",
    "B25J 9/0003",
    "G06K 9/00228",
    "G06K 9/00288",
    "G06V 40/161",
    "G06V 40/172",
    "G10L 13/00",
    "G10L 15/18",
    "G10L 15/22",
    "G10L 15/26",
    "G10L 2015/088",
    "G10L 2015/228",
    "G10L 2021/02087",
    "G10L 25/84"
  ],
  "ipc": [
    "B25J 11/00",
    "B25J 13/00",
    "B25J 9/00",
    "G06K 9/00",
    "G10L 13/00",
    "G10L 13/02",
    "G10L 13/027",
    "G10L 15/08",
    "G10L 15/18",
    "G10L 15/22",
    "G10L 15/26",
    "G10L 21/0208",
    "G10L 25/84"
  ],
  "assignees": [
    "International Business Machines Corp"
  ],
  "inventors": [
    "Shang Qing Guo",
    "Jonathan Lenchner"
  ],
  "filing_date": "2017-12-26",
  "publication_date": "2020-05-19",
  "grant_date": "2020-05-19",
  "priority_date": "2017-12-26",
  "application_number": "US-201715854142-A",
  "family_id": "66950580",
  "cited_by_count": 2,
  "citations": [
    "US6363301B1",
    "US6967455B2",
    "US20070192910A1",
    "US20140095007A1",
    "US20130094656A1",
    "WO2013168363A1",
    "US20150120060A1",
    "EP2933071A1",
    "WO2016174594A1",
    "US20160350553A1",
    "US10027662B1"
  ]
}

Record 2,158 of 8,000 in Patents full text (MLC-0201). Request the full dataset.