Patent · US10691898B2 · B2 · US
Synchronization method for visual information and auditory information and information processing device
- (11) Publication number
- US10691898B2
- (21) Application number
- 15/771,460
- (22) Filing date
- 2015-10-29
- (30) Priority date
- 2015-10-29
- (43) Publication date
- 2020-06-23
- (45) Date of grant
- 2020-06-23
- (51) IPC
- B25J 9/16; G06F 40/194; G06F 40/40; G06F 40/45; G06F 40/58; G10L 13/00; G10L 13/04; G10L 15/22; G10L 15/26; G10L 21/055; H04N 21/233; H04N 21/234; H04N 21/2343; H04N 21/439; H04N 21/44; H04N 21/4402
- (52) CPC
- G06F Electric digital data processing: 40/40, 40/194, 40/45, 40/58
- B25J Manipulators; chambers provided with manipulation devices: 9/1692, 9/1694, 9/1697
- G05B Control or regulating systems in general; functional elements of such systems; monitoring or testing arrangements for such systems or elements: 2219/40531
- G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 13/00, 13/043, 15/22, 15/26, 21/055
- H04N Pictorial communication, e.g. television: 21/233, 21/23418, 21/234336, 21/4394, 21/44008, 21/440236
- (73) Assignee
- Hitachi Ltd
- (72) Inventors
- Qinghua Sun; Takeshi Homma; Takashi Sumiyoshi; Masahito Togami
- (54) Title
- Synchronization method for visual information and auditory information and information processing device
- (57) Abstract
Disclosed is a method for synchronizing visual information and auditory information characterized by extracting visual information included in video, recognizing auditory information in a first language that is included in a speech in the first language, associating the visual information with the auditory information in the first language, translating the auditory information in the first language to auditory information in a second language, and editing at least one of the visual information with the auditory information in the second language so as to associate the visual information and the auditory information in the second language with each other.
- Full text
- View on Google Patents
Claims (14)
- A method for synchronizing visual information and first auditory information, comprising: extracting the visual information included in an image, the visual information including a first gesture and a second gesture occurring after the first gesture; recognizing first auditory information in a first language that is included in a speech in the first language; associating the visual information with the first auditory information in the first language; translating the first auditory information in the first language into second auditory information of a second language; and editing the visual information and the second auditory information in the second language so as to associate the visual information with the second auditory information in the second language, wherein the editing of the visual information includes editing the visual information so that the second gesture occurs before the first gesture.
- The method for synchronizing visual information and first auditory information according to claim 1, wherein editing to associate the visual information with the second auditory information in the second language includes editing to evaluate a time lag between the visual information and the second auditory information in the second language to reduce the particular time lag.
- The method for synchronizing visual information and first auditory information according to claim 1, wherein an editing method that edits at least one of the visual information and the second auditory information in the second language selects one or more of the most appropriate editing methods from a plurality of editing methods.
- The method for synchronizing visual information and first auditory information according to claim 3, wherein a method for selecting the most appropriate methods uses results of evaluating the reduction in the naturalness of the visual information and of the second auditory information in the second language due to each of the editing methods.
- The method for synchronizing visual information and first auditory information according to claim 4, wherein the evaluation of reduction in the naturalness of the visual information evaluates the naturalness by using at least one of the factors of continuity of the image, naturalness of the image, and continuity of robot motion corresponding to the image.
- The method for synchronizing visual information and first auditory information according to claim 4, wherein the evaluation of reduction in the naturalness of the second auditory information in the second language evaluates the naturalness by using at least one of the factors of continuity of a speech in the second language that includes the second auditory information in the second language, naturalness of the second language speech, consistency in the meaning of the speech in the second language and the speech in the first language, and ease in understanding the meaning of the speech in the second language.
- The method for synchronizing visual information and first auditory information according to claim 3, wherein the method for selecting the most appropriate methods selects an editing method with less reduction in the naturalness of the visual information and of the second auditory information in the second language due to editing of the visual information and editing of the second auditory information in the second language.
- The method for synchronizing visual information and first auditory information according to claim 3, wherein the editing method for editing the visual information changes the timing of the visual information by using at least one of the methods of temporarily stopping reproduction of the image, editing using CG of the image, changing the rate of robot motion corresponding to the image, and changing the order of robot motions corresponding to the image.
- The method for synchronizing visual information and first auditory information according to claim 3, wherein the editing method for editing the second auditory information in the second language changes the timing of the second auditory information by using at least one of the methods of temporarily stopping reproduction of the speech in the second language that includes the second auditory information in the second language, changing the reproduction order of the speech in the second language, changing the order of spoken words of the speech in the second language, and changing the speech content of the speech in the second language.
- An information processing device, comprising: a memory coupled to a processor, the memory storing instructions that when executed by the processor, configure the processor to: input input image data including first visual information as well as input speech data in a first language that includes first auditory information, the first visual information including a first gesture and a second gesture occurring after the first gesture, output output visual data including second visual information corresponding to the first visual information, as well as output speech data in a second language that includes second auditory information corresponding to the first auditory information, detect the first visual information from the input image data, recognize the first auditory information from the input speech data, associate the first visual information with the first auditory information, convert the input speech data in the first language into the output speech data in the second language, and control the output timing of the second visual information and the second auditory information so as to reduce the output timing difference between the second visual information and the second auditory information, both of which correspond to the first visual information and the first auditory information that are associated with each other, wherein the optimization unit includes an image editing unit for editing output image data which is the output visual data so that the second gesture occurs before the first gesture.
- The information processing device according to claim 10, wherein the processor is further configured to execute at least one of edit robot motion command data which is the output visual data, edit text data to generate the output speech data, and edit the output speech data.
- The information processing device according to claim 11, wherein the processor is further configured to: compare the difference in output timing between the second visual information and the second auditory information with a threshold, both of which correspond to the first visual information and the first auditory information that are associated with each other, and upon determining that the difference in output timing is greater than the threshold, edit the output visual data, edit the text data, and edit the output speech.
- The information processing device according to claim 11, wherein the first visual information is the same as the second visual information, and wherein the processor is further configured to: output image data by performing editing on the input image data to change the temporal relationship in such a way that the first visual information included in the input image data is replaced with the second visual information.
- The information processing device according to claim 11, wherein the processor is further configured to output speech data by changing text data for generating the output speech data by changing a candidate of a translation result in the translation unit, or by changing the order of words in the text data.
Description
The present invention relates to a technique suitable for application to devices that perform automatic speech translation, and the like. More specifically, the present invention relates to a technique for automatically generating auditory information (translated speech) in a second language as well as visual information (edited video, reproduction of motion by a robot, and the like) for the listener, from input first language auditory information (speech) and input visual information (motion of the speaker, and the like).
Against the background of recent significant progress in techniques such as speech recognition, machine translation, and speech synthesis, speech translation systems, which are a combination of these techniques, have been put into practical use. In such systems, an input in a first language is converted into a text in the first language by speech recognition technique. Further, the text in the first language is translated into a text in a second language by machine translation, and then is converted into a speech in the second language by a speech synthesis module corresponding to the second language. The practical application of this technique will eliminate the language barrier, thus allowing people to freely communicate with foreigners.
At the same time, in addition to auditory information from the ears, visual information from the eyes such as facial expression and gesture can greatly contribute to the transmission of meaning. For example, a gesture such as “pointing” can greatly contribute to the understanding of meaning.
Citations (18)
- JPH06253197A
- JPH08220985A
- US20030085901A1
- US20070165022A1
- JP2001224002A
- JP2002123282A
- JP2004230479A
- US20060136226A1
- JP2008306691A
- US20110138433A1
- US20120105719A1
- US20120276504A1
- US20140160134A1
- US8874429B1
- US20140372100A1
- US20150181306A1
- US20150199978A1
- US20160042766A1
Record as JSON
{
"publication_number": "US10691898B2",
"country": "US",
"kind": "B2",
"title": "Synchronization method for visual information and auditory information and information processing device",
"abstract": "Disclosed is a method for synchronizing visual information and auditory information characterized by extracting visual information included in video, recognizing auditory information in a first language that is included in a speech in the first language, associating the visual information with the auditory information in the first language, translating the auditory information in the first language to auditory information in a second language, and editing at least one of the visual information with the auditory information in the second language so as to associate the visual information and the auditory information in the second language with each other.",
"claims": [
"1. A method for synchronizing visual information and first auditory information, comprising: extracting the visual information included in an image, the visual information including a first gesture and a second gesture occurring after the first gesture; recognizing first auditory information in a first language that is included in a speech in the first language; associating the visual information with the first auditory information in the first language; translating the first auditory information in the first language into second auditory information of a second language; and editing the visual information and the second auditory information in the second language so as to associate the visual information with the second auditory information in the second language, wherein the editing of the visual information includes editing the visual information so that the second gesture occurs before the first gesture.",
"2. The method for synchronizing visual information and first auditory information according to claim 1, wherein editing to associate the visual information with the second auditory information in the second language includes editing to evaluate a time lag between the visual information and the second auditory information in the second language to reduce the particular time lag.",
"3. The method for synchronizing visual information and first auditory information according to claim 1, wherein an editing method that edits at least one of the visual information and the second auditory information in the second language selects one or more of the most appropriate editing methods from a plurality of editing methods.",
"4. The method for synchronizing visual information and first auditory information according to claim 3, wherein a method for selecting the most appropriate methods uses results of evaluating the reduction in the naturalness of the visual information and of the second auditory information in the second language due to each of the editing methods.",
"5. The method for synchronizing visual information and first auditory information according to claim 4, wherein the evaluation of reduction in the naturalness of the visual information evaluates the naturalness by using at least one of the factors of continuity of the image, naturalness of the image, and continuity of robot motion corresponding to the image.",
"6. The method for synchronizing visual information and first auditory information according to claim 4, wherein the evaluation of reduction in the naturalness of the second auditory information in the second language evaluates the naturalness by using at least one of the factors of continuity of a speech in the second language that includes the second auditory information in the second language, naturalness of the second language speech, consistency in the meaning of the speech in the second language and the speech in the first language, and ease in understanding the meaning of the speech in the second language.",
"7. The method for synchronizing visual information and first auditory information according to claim 3, wherein the method for selecting the most appropriate methods selects an editing method with less reduction in the naturalness of the visual information and of the second auditory information in the second language due to editing of the visual information and editing of the second auditory information in the second language.",
"8. The method for synchronizing visual information and first auditory information according to claim 3, wherein the editing method for editing the visual information changes the timing of the visual information by using at least one of the methods of temporarily stopping reproduction of the image, editing using CG of the image, changing the rate of robot motion corresponding to the image, and changing the order of robot motions corresponding to the image.",
"9. The method for synchronizing visual information and first auditory information according to claim 3, wherein the editing method for editing the second auditory information in the second language changes the timing of the second auditory information by using at least one of the methods of temporarily stopping reproduction of the speech in the second language that includes the second auditory information in the second language, changing the reproduction order of the speech in the second language, changing the order of spoken words of the speech in the second language, and changing the speech content of the speech in the second language.",
"10. An information processing device, comprising: a memory coupled to a processor, the memory storing instructions that when executed by the processor, configure the processor to: input input image data including first visual information as well as input speech data in a first language that includes first auditory information, the first visual information including a first gesture and a second gesture occurring after the first gesture, output output visual data including second visual information corresponding to the first visual information, as well as output speech data in a second language that includes second auditory information corresponding to the first auditory information, detect the first visual information from the input image data, recognize the first auditory information from the input speech data, associate the first visual information with the first auditory information, convert the input speech data in the first language into the output speech data in the second language, and control the output timing of the second visual information and the second auditory information so as to reduce the output timing difference between the second visual information and the second auditory information, both of which correspond to the first visual information and the first auditory information that are associated with each other, wherein the optimization unit includes an image editing unit for editing output image data which is the output visual data so that the second gesture occurs before the first gesture.",
"11. The information processing device according to claim 10, wherein the processor is further configured to execute at least one of edit robot motion command data which is the output visual data, edit text data to generate the output speech data, and edit the output speech data.",
"12. The information processing device according to claim 11, wherein the processor is further configured to: compare the difference in output timing between the second visual information and the second auditory information with a threshold, both of which correspond to the first visual information and the first auditory information that are associated with each other, and upon determining that the difference in output timing is greater than the threshold, edit the output visual data, edit the text data, and edit the output speech.",
"13. The information processing device according to claim 11, wherein the first visual information is the same as the second visual information, and wherein the processor is further configured to: output image data by performing editing on the input image data to change the temporal relationship in such a way that the first visual information included in the input image data is replaced with the second visual information.",
"14. The information processing device according to claim 11, wherein the processor is further configured to output speech data by changing text data for generating the output speech data by changing a candidate of a translation result in the translation unit, or by changing the order of words in the text data."
],
"description_excerpt": "The present invention relates to a technique suitable for application to devices that perform automatic speech translation, and the like. More specifically, the present invention relates to a technique for automatically generating auditory information (translated speech) in a second language as well as visual information (edited video, reproduction of motion by a robot, and the like) for the listener, from input first language auditory information (speech) and input visual information (motion of the speaker, and the like).\n\nAgainst the background of recent significant progress in techniques such as speech recognition, machine translation, and speech synthesis, speech translation systems, which are a combination of these techniques, have been put into practical use. In such systems, an input in a first language is converted into a text in the first language by speech recognition technique. Further, the text in the first language is translated into a text in a second language by machine translation, and then is converted into a speech in the second language by a speech synthesis module corresponding to the second language. The practical application of this technique will eliminate the language barrier, thus allowing people to freely communicate with foreigners.\n\nAt the same time, in addition to auditory information from the ears, visual information from the eyes such as facial expression and gesture can greatly contribute to the transmission of meaning. For example, a gesture such as “pointing” can greatly contribute to the understanding of meaning.",
"cpc": [
"G06F 40/40",
"B25J 9/1692",
"B25J 9/1694",
"B25J 9/1697",
"G05B 2219/40531",
"G06F 40/194",
"G06F 40/45",
"G06F 40/58",
"G10L 13/00",
"G10L 13/043",
"G10L 15/22",
"G10L 15/26",
"G10L 21/055",
"H04N 21/233",
"H04N 21/23418",
"H04N 21/234336",
"H04N 21/4394",
"H04N 21/44008",
"H04N 21/440236"
],
"ipc": [
"B25J 9/16",
"G06F 40/194",
"G06F 40/40",
"G06F 40/45",
"G06F 40/58",
"G10L 13/00",
"G10L 13/04",
"G10L 15/22",
"G10L 15/26",
"G10L 21/055",
"H04N 21/233",
"H04N 21/234",
"H04N 21/2343",
"H04N 21/439",
"H04N 21/44",
"H04N 21/4402"
],
"assignees": [
"Hitachi Ltd"
],
"inventors": [
"Qinghua Sun",
"Takeshi Homma",
"Takashi Sumiyoshi",
"Masahito Togami"
],
"filing_date": "2015-10-29",
"publication_date": "2020-06-23",
"grant_date": "2020-06-23",
"priority_date": "2015-10-29",
"application_number": "US-201515771460-A",
"family_id": "58630010",
"cited_by_count": 2,
"citations": [
"JPH06253197A",
"JPH08220985A",
"US20030085901A1",
"US20070165022A1",
"JP2001224002A",
"JP2002123282A",
"JP2004230479A",
"US20060136226A1",
"JP2008306691A",
"US20110138433A1",
"US20120105719A1",
"US20120276504A1",
"US20140160134A1",
"US8874429B1",
"US20140372100A1",
"US20150181306A1",
"US20150199978A1",
"US20160042766A1"
]
}
Record 2,105 of 8,000 in Patents full text (MLC-0201). Request the full dataset.