MLchartDataset catalogue

Patent · US10977452B2 · B2 · US

Multi-lingual virtual personal assistant

(11) Publication number
US10977452B2
(21) Application number
16/509,428
(22) Filing date
2019-07-11
(30) Priority date
2015-12-22
(43) Publication date
2021-04-13
(45) Date of grant
2021-04-13
(51) IPC
G06F 17/27; G06F 17/28; G06F 16/9032; G10L 15/07; G10L 15/18; G10L 15/22; G06F 40/30; G06F 40/58
(52) CPC
  • G06F Electric digital data processing: 40/58, 16/90332, 40/30
  • G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 15/07, 15/1815, 15/1822, 15/22, 2015/228, 25/63
(73) Assignee
SRI International Inc
(72) Inventors
Wen Wang; Dimitra Vergyri; Girish Acharya
(54) Title
Multi-lingual virtual personal assistant
(57) Abstract

Provided are systems, computer-implemented methods, and computer-program products for a multi-lingual device, capable of receiving verbal input in multiple languages, and further capable of providing conversational responses in multiple languages. In various implementations, the multi-lingual device includes an automatic speech recognition engine capable of receiving verbal input in a first natural language and providing a textual representation of the input and a confidence value for the recognition. The multi-lingual device can also include a machine translation engine, capable of translating textual input from the first natural language into a second natural language. The machine translation engine can output a confidence value for the translation. The multi-lingual device can further include natural language processing, capable of translating from the second natural language to a computer-based language. Input in the computer-based language can be processed, and the multi-lingual device can take an action based on the result of the processing.

Full text
View on Google Patents

Claims (17)

  1. A method, comprising: identifying a first portion of an audio input as comprising a first input language and a second portion of the audio input as comprising a second input language; wherein at least one of the first input language and the second input language is different than a processing language used to process the audio input; wherein each of the first input language, the second input language, and the processing language is a natural language; translating at least one of the first and second portions of the audio input into the processing language to produce at least one translated portion of the audio input; using at least one model trained to recognize machine translation errors, performing semantic analysis on the at least one translated portion of the audio input to determine and output semantic information corresponding to the audio input; using the semantic information, formulating computer language output to cause a device to perform an action; in response to at least one of the first input language and the second input language corresponding to the processing language used to process the audio input, tagging a portion of the audio input that is in the processing language and skipping the translating for the tagged portion of the audio input; wherein the method is performed by one or more computing devices.
  2. The method of claim 1, further comprising, in response to both the first input language and the second input language being different than the processing language used to process the audio input, translating both the first and second portions of the audio input into the processing language to produce first and second translated portions of the audio input; and using both the first and second translated portions of the audio input to determine and output the semantic information and formulate the computer language output.
  3. The method of claim 1, further comprising producing system-generated output of the action in the processing language.
  4. The method of claim 1, further comprising translating system-generated output of the action into at least one of the first input language and the second input language.
  5. The method of claim 1, further comprising performing the translating using at least two different machine translation engines and combining outputs of the at least two different machine translation engines to produce the at least one translated portion of the audio input.
  6. The method of claim 5, further comprising assigning confidence values to the outputs of the at least two different machine translation engines and using the confidence values to determine the at least one translated portion of the audio input.
  7. A system, comprising: one or more processors capable of executing instructions; and one or more non-transitory computer-readable media coupled to the processor and including instructions that, if executed by the one or more processors, cause the system to be capable of performing operations comprising: identifying a first portion of an audio input as comprising a first input language and a second portion of the audio input as comprising a second input language; at least one of the first input language and the second input language being different than a processing language used to process the audio input; each of the first input language, the second input language, and the processing language being a natural language; translating at least one of the first and second portions of the audio input into the processing language to produce at least one translated portion of the audio input; using at least one model trained to recognize machine translation errors, performing semantic analysis on the at least one translated portion of the audio input to determine and output semantic information corresponding to the audio input; using the semantic information, formulating computer language output to cause a device to perform an action; in response to at least one of the first input language and the second input language corresponding to the processing language used to process the audio input, tagging a portion of the audio input that is in the processing language and skipping the translating for the tagged portion of the audio input.
  8. The system of claim 7, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising, in response to both the first input language and the second input language being different than the processing language used to process the audio input, translating both the first and second portions of the audio input into the processing language to produce first and second translated portions of the audio input; and using both the first and second translated portions of the audio input to determine and output the semantic information and formulate the computer language output.
  9. The system of claim 7, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising producing system-generated output of the action in the processing language.
  10. The system of claim 7, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising translating system-generated output of the action into at least one of the first input language and the second input language.
  11. The system of claim 7, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising performing the translating using at least two different machine translation engines and combining outputs of the at least two different machine translation engines to produce the at least one translated portion of the audio input.
  12. The system of claim 11, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising assigning confidence values to the outputs of the at least two different machine translation engines and using the confidence values to determine the at least one translated portion of the audio input.
  13. One or more non-transitory machine-readable storage media comprising instructions that, if executed by one or more processors, cause the one or more processors to be capable of performing operations comprising: identifying a first portion of an audio input as comprising a first input language and a second portion of the audio input as comprising a second input language; at least one of the first input language and the second input language being different than a processing language used to process the audio input; each of the first input language, the second input language, and the processing language being a natural language; translating at least one of the first and second portions of the audio input into the processing language to produce at least one translated portion of the audio input; using at least one model trained to recognize machine translation errors, performing semantic analysis on the at least one translated portion of the audio input to determine and output semantic information corresponding to the audio input; using the semantic information, formulating computer language output to cause a device to perform an action; in response to at least one of the first input language and the second input language corresponding to the processing language used to process the audio input, tagging a portion of the audio input that is in the processing language and skipping the translating for the tagged portion of the audio input.
  14. The one or more non-transitory machine-readable storage media of claim 13, wherein the instructions, if executed by the one or more processors, cause the one or more processors to be capable of performing operations further comprising, in response to both the first input language and the second input language being different than the processing language used to process the audio input, translating both the first and second portions of the audio input into the processing language to produce first and second translated portions of the audio input; and using both the first and second translated portions of the audio input to determine and output the semantic information and formulate the computer language output.
  15. The one or more non-transitory machine-readable storage media of claim 13, wherein the instructions, if executed by the one or more processors, cause the one or more processors to be capable of performing operations further comprising producing system-generated output of the action in the processing language.
  16. The one or more non-transitory machine-readable storage media of claim 13, wherein the instructions, if executed by the one or more processors, cause the one or more processors to be capable of performing operations further comprising translating system-generated output of the action into at least one of the first input language and the second input language.
  17. The one or more non-transitory machine-readable storage media of claim 13, wherein the instructions, if executed by the one or more processors, cause the one or more processors to be capable of performing operations further comprising performing the translating using at least two different machine translation engines; combining outputs of the at least two different machine translation engines to produce the at least one translated portion of the audio input; and assigning confidence values to the outputs of the at least two different machine translation engines and using the confidence values to determine the at least one translated portion of the audio input.

Description

Provided are systems, methods, such as computer-implemented methods, and computer-program products for a multi-lingual device, capable of receiving verbal input in multiple languages, and further capable of providing conversational responses in multiple languages.

In various systems, methods, and/or computer-program products, a multi-lingual device can be configured to receive verbal input. The verbal input can provided in a first language, which is a natural language spoken by humans. The multi-lingual device can further be configured to determine original text from the verbal input. The text can be determined using an automatic speech recognition engine of the multi-lingual device. The original text can be output in the first language. The multi-lingual device can further be configured to determine a confidence value for the original text. The confidence value for the original text can use a statistical association between the original text and the verbal input. The automatic speech recognition engine can output the original text according to the confidence value for the original text. The multi-lingual device can further be configured to determine translated text corresponding to the original text. The translated text can be determined using a machine translation engine of the multi-lingual device. The machine translation engine can translate the original text to a second language, which is also a natural language. The multi-lingual device can further be configured to determine a confidence value for the translated text.

Citations (9)

  • US20080071518A1
  • US8706471B2
  • US20150051897A1
  • US20090281789A1
  • US20100185434A1
  • US20130211818A1
  • US20130262076A1
  • US20140303957A1
  • US20140365200A1
Record as JSON
{
  "publication_number": "US10977452B2",
  "country": "US",
  "kind": "B2",
  "title": "Multi-lingual virtual personal assistant",
  "abstract": "Provided are systems, computer-implemented methods, and computer-program products for a multi-lingual device, capable of receiving verbal input in multiple languages, and further capable of providing conversational responses in multiple languages. In various implementations, the multi-lingual device includes an automatic speech recognition engine capable of receiving verbal input in a first natural language and providing a textual representation of the input and a confidence value for the recognition. The multi-lingual device can also include a machine translation engine, capable of translating textual input from the first natural language into a second natural language. The machine translation engine can output a confidence value for the translation. The multi-lingual device can further include natural language processing, capable of translating from the second natural language to a computer-based language. Input in the computer-based language can be processed, and the multi-lingual device can take an action based on the result of the processing.",
  "claims": [
    "1. A method, comprising: identifying a first portion of an audio input as comprising a first input language and a second portion of the audio input as comprising a second input language; wherein at least one of the first input language and the second input language is different than a processing language used to process the audio input; wherein each of the first input language, the second input language, and the processing language is a natural language; translating at least one of the first and second portions of the audio input into the processing language to produce at least one translated portion of the audio input; using at least one model trained to recognize machine translation errors, performing semantic analysis on the at least one translated portion of the audio input to determine and output semantic information corresponding to the audio input; using the semantic information, formulating computer language output to cause a device to perform an action; in response to at least one of the first input language and the second input language corresponding to the processing language used to process the audio input, tagging a portion of the audio input that is in the processing language and skipping the translating for the tagged portion of the audio input; wherein the method is performed by one or more computing devices.",
    "2. The method of claim 1, further comprising, in response to both the first input language and the second input language being different than the processing language used to process the audio input, translating both the first and second portions of the audio input into the processing language to produce first and second translated portions of the audio input; and using both the first and second translated portions of the audio input to determine and output the semantic information and formulate the computer language output.",
    "3. The method of claim 1, further comprising producing system-generated output of the action in the processing language.",
    "4. The method of claim 1, further comprising translating system-generated output of the action into at least one of the first input language and the second input language.",
    "5. The method of claim 1, further comprising performing the translating using at least two different machine translation engines and combining outputs of the at least two different machine translation engines to produce the at least one translated portion of the audio input.",
    "6. The method of claim 5, further comprising assigning confidence values to the outputs of the at least two different machine translation engines and using the confidence values to determine the at least one translated portion of the audio input.",
    "7. A system, comprising: one or more processors capable of executing instructions; and one or more non-transitory computer-readable media coupled to the processor and including instructions that, if executed by the one or more processors, cause the system to be capable of performing operations comprising: identifying a first portion of an audio input as comprising a first input language and a second portion of the audio input as comprising a second input language; at least one of the first input language and the second input language being different than a processing language used to process the audio input; each of the first input language, the second input language, and the processing language being a natural language; translating at least one of the first and second portions of the audio input into the processing language to produce at least one translated portion of the audio input; using at least one model trained to recognize machine translation errors, performing semantic analysis on the at least one translated portion of the audio input to determine and output semantic information corresponding to the audio input; using the semantic information, formulating computer language output to cause a device to perform an action; in response to at least one of the first input language and the second input language corresponding to the processing language used to process the audio input, tagging a portion of the audio input that is in the processing language and skipping the translating for the tagged portion of the audio input.",
    "8. The system of claim 7, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising, in response to both the first input language and the second input language being different than the processing language used to process the audio input, translating both the first and second portions of the audio input into the processing language to produce first and second translated portions of the audio input; and using both the first and second translated portions of the audio input to determine and output the semantic information and formulate the computer language output.",
    "9. The system of claim 7, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising producing system-generated output of the action in the processing language.",
    "10. The system of claim 7, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising translating system-generated output of the action into at least one of the first input language and the second input language.",
    "11. The system of claim 7, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising performing the translating using at least two different machine translation engines and combining outputs of the at least two different machine translation engines to produce the at least one translated portion of the audio input.",
    "12. The system of claim 11, wherein the instructions, if executed by the one or more processors, cause the system to be capable of performing operations further comprising assigning confidence values to the outputs of the at least two different machine translation engines and using the confidence values to determine the at least one translated portion of the audio input.",
    "13. One or more non-transitory machine-readable storage media comprising instructions that, if executed by one or more processors, cause the one or more processors to be capable of performing operations comprising: identifying a first portion of an audio input as comprising a first input language and a second portion of the audio input as comprising a second input language; at least one of the first input language and the second input language being different than a processing language used to process the audio input; each of the first input language, the second input language, and the processing language being a natural language; translating at least one of the first and second portions of the audio input into the processing language to produce at least one translated portion of the audio input; using at least one model trained to recognize machine translation errors, performing semantic analysis on the at least one translated portion of the audio input to determine and output semantic information corresponding to the audio input; using the semantic information, formulating computer language output to cause a device to perform an action; in response to at least one of the first input language and the second input language corresponding to the processing language used to process the audio input, tagging a portion of the audio input that is in the processing language and skipping the translating for the tagged portion of the audio input.",
    "14. The one or more non-transitory machine-readable storage media of claim 13, wherein the instructions, if executed by the one or more processors, cause the one or more processors to be capable of performing operations further comprising, in response to both the first input language and the second input language being different than the processing language used to process the audio input, translating both the first and second portions of the audio input into the processing language to produce first and second translated portions of the audio input; and using both the first and second translated portions of the audio input to determine and output the semantic information and formulate the computer language output.",
    "15. The one or more non-transitory machine-readable storage media of claim 13, wherein the instructions, if executed by the one or more processors, cause the one or more processors to be capable of performing operations further comprising producing system-generated output of the action in the processing language.",
    "16. The one or more non-transitory machine-readable storage media of claim 13, wherein the instructions, if executed by the one or more processors, cause the one or more processors to be capable of performing operations further comprising translating system-generated output of the action into at least one of the first input language and the second input language.",
    "17. The one or more non-transitory machine-readable storage media of claim 13, wherein the instructions, if executed by the one or more processors, cause the one or more processors to be capable of performing operations further comprising performing the translating using at least two different machine translation engines; combining outputs of the at least two different machine translation engines to produce the at least one translated portion of the audio input; and assigning confidence values to the outputs of the at least two different machine translation engines and using the confidence values to determine the at least one translated portion of the audio input."
  ],
  "description_excerpt": "Provided are systems, methods, such as computer-implemented methods, and computer-program products for a multi-lingual device, capable of receiving verbal input in multiple languages, and further capable of providing conversational responses in multiple languages.\n\nIn various systems, methods, and/or computer-program products, a multi-lingual device can be configured to receive verbal input. The verbal input can provided in a first language, which is a natural language spoken by humans. The multi-lingual device can further be configured to determine original text from the verbal input. The text can be determined using an automatic speech recognition engine of the multi-lingual device. The original text can be output in the first language. The multi-lingual device can further be configured to determine a confidence value for the original text. The confidence value for the original text can use a statistical association between the original text and the verbal input. The automatic speech recognition engine can output the original text according to the confidence value for the original text. The multi-lingual device can further be configured to determine translated text corresponding to the original text. The translated text can be determined using a machine translation engine of the multi-lingual device. The machine translation engine can translate the original text to a second language, which is also a natural language. The multi-lingual device can further be configured to determine a confidence value for the translated text.",
  "cpc": [
    "G06F 40/58",
    "G06F 16/90332",
    "G06F 40/30",
    "G10L 15/07",
    "G10L 15/1815",
    "G10L 15/1822",
    "G10L 15/22",
    "G10L 2015/228",
    "G10L 25/63"
  ],
  "ipc": [
    "G06F 17/27",
    "G06F 17/28",
    "G06F 16/9032",
    "G10L 15/07",
    "G10L 15/18",
    "G10L 15/22",
    "G06F 40/30",
    "G06F 40/58"
  ],
  "assignees": [
    "SRI International Inc"
  ],
  "inventors": [
    "Wen Wang",
    "Dimitra Vergyri",
    "Girish Acharya"
  ],
  "filing_date": "2019-07-11",
  "publication_date": "2021-04-13",
  "grant_date": "2021-04-13",
  "priority_date": "2015-12-22",
  "application_number": "US-201916509428-A",
  "family_id": "59091205",
  "cited_by_count": 23,
  "citations": [
    "US20080071518A1",
    "US8706471B2",
    "US20150051897A1",
    "US20090281789A1",
    "US20100185434A1",
    "US20130211818A1",
    "US20130262076A1",
    "US20140303957A1",
    "US20140365200A1"
  ]
}

Record 1,675 of 8,000 in Patents full text (MLC-0201). Request the full dataset.