MLchartDataset catalogue

Patent · US10667045B1 · B1 · US

Robot and auto data processing method thereof

(11) Publication number
US10667045B1
(21) Application number
16/447,986
(22) Filing date
2019-06-21
(30) Priority date
2018-12-28
(43) Publication date
2020-05-26
(45) Date of grant
2020-05-26
(51) IPC
H04R 1/02; H04R 1/40
(52) CPC
  • H04R Loudspeakers, microphones, gramophone pick-ups or like acoustic electromechanical transducers; electric hearing AIDS; public address systems: 1/406, 1/028, 2201/401, 2410/01, 2420/01, 2430/20
  • B25J Manipulators; chambers provided with manipulation devices: 11/0005
  • G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 2021/02082, 2021/02166, 21/0208, 21/0216
(73) Assignee
UBTECH ROBOTICS CORP LTD
(72) Inventors
XIONG YOUJUN; Xing Fanglin
(54) Title
Robot and auto data processing method thereof
(57) Abstract

The present disclosure provides a robot and an audio data processing method thereof. The robot includes a body part, a main control module, and a sound pickup module electrically coupled to the main control module. The sound pickup module includes N microphones distributed around the body part to collect audio data. The main control module is configured to obtain the audio data of a sound source from a part of the N microphones collecting the audio data of the sound source without blocked by the body part, and perform a sound source localization and a sound pickup based on the obtained audio data. The 360-degree wake-up and sound source localization of the robot and the beam-forming of directional beams are realized. In addition, the sound pickup is realized without forming microphone holes on the head of the robot, hence the aesthetics of the robot will not be affected.

Full text
View on Google Patents

Claims (14)

  1. A robot, comprising: at least one body part; a main control module comprising a data buffer pool; and a sound pickup module electrically coupled to the main control module, wherein the sound pickup module comprises N microphones distributed around the body part to collect audio data, where N≥3 and N is an integer, and wherein when collecting the audio data, a part of the N microphones is capable of receiving a direct sound from a sound source, but the rest part of the N microphones is incapable of receiving the direct sound but reflect sounds of the sound source; wherein, the main control module is configured to obtain first audio data of the sound source collected by the N microphones, perform a sound source localization based on the first audio data, obtain second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound, and perform a sound pickup and a voice recognition based on the second audio data; wherein, the main control module is further configured to store X channels of reference audio data and N channels of audio data to the data buffer pool, obtain a first group of the audio data from the data buffer pool as the first audio data of the sound source collected by the N microphones, to use a first predetermined algorithm to locate a sound source, and obtain a second group of the audio data from the data buffer pool as the second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound, to use a second predetermined algorithm to perform a beam-forming and an audio noise reduction; wherein the N channels of audio data is six channels of audio data, and the X channels of reference audio data is two channels of reference audio data; wherein, audio data obtained by a first microphone in microphones arrays is taken as first audio data, audio data obtained by a second microphone in the microphones arrays is taken as second audio data, audio data obtained by a third microphone in the microphones arrays is taken as third audio data, audio data obtained by a fourth microphone in the microphones arrays is taken as fourth audio data, audio data obtained by a fifth microphone in the microphones arrays is taken as fifth audio data, audio data obtained by a sixth microphone in the microphones arrays is taken as sixth audio data, first channel reference audio data in the two channels of the reference audio data is taken as seventh audio data, and second channel reference audio data in the two channels of the reference audio data is taken as eighth audio data; wherein, the first group of the audio data comprises the first audio data, the second audio data, the third audio data, the fourth audio data, the fifth audio data, the sixth audio data, the seventh audio data, and the eighth audio data; and wherein, the second group of the audio data comprises the first audio data, the second audio data, the third audio data, the sixth audio data, the seventh audio data, and the eighth audio data.
  2. The robot of claim 1, wherein the sound pickup module further comprises: a MIC small board electrically coupled to each of the microphone array and the main control module, wherein the MIC small board is coupled to perform an analog-to-digital conversion on the N channels of audio data collected by the microphones, encode the converted audio data, and transmit the encoded audio data to the main control module.
  3. The robot of claim 2, wherein the MIC small board comprises: an analog-to-digital converter electrically coupled to the microphone arrays and the main control module, wherein the analog-to-digital converter performs the analog-to-digital conversion on the N channels of audio data.
  4. The robot of claim 2, further comprising: a power amplifier electrically coupled to the main control module; wherein the main control module is configured to generate the X channels of reference audio data based on audio data obtained from the power amplifier to transmit to the MIC small board, and the MIC small board is further configured to perform an analog-to-digital conversion on the X channels of reference audio data, encode the converted X channels of reference audio data, and transmit the encoded X channels of reference audio data to the main control module.
  5. The robot of claim 4, wherein, the main control module is further configured to obtain the audio data played by the power amplifier and generate the X channels of reference audio data based on the audio data played by the power amplifier.
  6. The robot of claim 5, wherein, number of the X channels of reference audio data is same as number of channels of the audio data played by the power amplifier, the main control module is electrically coupled to the MIC small board directly through data lines, and amount of the data lines corresponds to amount of the X channels of the reference audio data.
  7. The robot of claim 5, wherein, the MIC small board is further configured to perform a data fusion, fuse received reference audio data with the N channels of audio data, and transmit fused audio data to the main control module.
  8. The robot of claim 1, wherein the body part is a neck, the microphone array comprises six microphones, the six microphones are disposed around the neck and are distributed on a circumference centered on any point on a longitudinal axis of the body part.
  9. The robot of claim 1, wherein the body part comprises at least one of neck and a trunk.
  10. The robot of claim 1, wherein the main control module is further configured to determine an angle difference between a sound source position and a current position through the sound source localization, control the robot to turn according to the angle difference, wake up the robot, and perform the sound pickup and the voice recognition based on the second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound.
  11. The robot of claim 1, wherein the at least one body part includes a head, a neck, and a trunk, and the N microphones are distributed around each of the at least one body part in a non-even manner, or, the N microphones are distributed around the head, the trunk, or two or more of the at least one body part.
  12. The robot of claim 1, wherein the main control module is a development board, and the data buffer pool is configured in a software layer of the development board.
  13. A computer-implemented audio data processing method based on a robot comprising: at least one body part; a main control module; and a sound pickup module electrically coupled to the main control module, wherein the sound pickup module comprises N microphones distributed around the body part to collect audio data, where N≥3 and N is an integer, and wherein when collecting the audio data, a part of the N microphones is capable of receiving a direct sound from a sound source, but the rest part of the N microphones is incapable of receiving the direct sound but reflect sounds of the sound source; wherein, the main control module is configured to obtain first audio data of the sound source collected by the N microphones, perform a sound source localization based on the first audio data, obtain second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound, and perform a sound pickup and a voice recognition based on the second audio data; the method comprising executing on a processor of the robot the steps of: collecting audio data through the N microphones of the sound pickup module; transmitting the N channels of audio data collected by the N microphones to the main control module; storing, by the main control module, the N channels of audio data to a data buffer pool; and performing, by the main control module, the sound source localization and the sound pickup based on the audio data; wherein the step of storing, by the main control module, the N channels of audio data to the data buffer pool and the step of performing, by the main control module, the sound source localization and the sound pickup based on the audio data further comprise: storing X channels of reference audio data and the N channels of audio data to the data buffer pool; obtaining a first group of the audio data from the data buffer pool as the first audio data of the sound source collected by the N microphones, to use a first predetermined algorithm to locate a sound source; and obtaining a second group of the audio data from the data buffer pool as the second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound, to use a second predetermined algorithm to perform a beam-forming and an audio noise reduction; wherein the N channels of audio data is six channels of audio data, and the X channels of reference audio data is two channels of reference audio data; wherein, audio data obtained by a first microphone in microphones arrays is taken as first audio data, audio data obtained by a second microphone in the microphones arrays is taken as second audio data, audio data obtained by a third microphone in the microphones arrays is taken as third audio data, audio data obtained by a fourth microphone in the microphones arrays is taken as fourth audio data, audio data obtained by a fifth microphone in the microphones arrays is taken as fifth audio data, audio data obtained by a sixth microphone in the microphones arrays is taken as sixth audio data, first channel reference audio data in the two channels of reference audio data is taken as seventh audio data, and second channel reference audio data in the two channels of reference audio data is taken as eighth audio data; wherein, the first group of the audio data comprises the first audio data, the second audio data, the third audio data, the fourth audio data, the fifth audio data, the sixth audio data, the seventh audio data, and the eighth audio data; and wherein, the second group of the audio data comprises the first audio data, the second audio data, the third audio data, the sixth audio data, the seventh audio data, and the eighth audio data.
  14. The method of claim 13, wherein, the main control module determines an angle difference between a sound source position and a current position through the sound source localization, control the robot to turn according to the angle difference, wake up the robot, and perform the sound pickup and the voice recognition based on the second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound.

Description

1. Technical Field The present disclosure relates to intelligent robot technology, and particularly to a robot and an audio data processing method thereof. 2. Description of Related Art When designing a robot, if the position of a microphone array is not arranged correctly, the voice interaction will be affected. Because the most basic requirement and prerequisite for the beam-forming of the microphone array is that sounds should directly reach each microphone in the microphone array. Therefore, if an annular microphone array is disposed at the neck of the robot, the neck of the robot will hide the microphones behind the neck, which causes the sounds to be reflected by the neck and can not directly reach the microphone behind the neck of the robot, thus affecting the effect of sound pickup. In order to resolve the above-mentioned problems, it is generally to place an annular microphone array on the head of the robot, or to use an annular microphone array and a linear microphone array at the same time, where the annular microphone array is disposed at the neck of the robot for realizing the 360-degree wake-up and 360-degree sound source localization of the robot, and the linear microphone array is disposed on the head of the robot for beam-forming so as to perform sound pickup.

However, disposing the annular microphone array on the head of the robot will cause a limit to the height of the robot.

Citations (11)

  • JP2007221300A
  • JP2007241157A
  • JP2008278399A
  • JP2011069901A
  • US2005058300A1
  • US2010150364A1
  • US2017243577A1
  • US2018374494A1
  • US2019104360A1
  • US2019250245A1
  • US2019364375A1
Record as JSON
{
  "publication_number": "US10667045B1",
  "country": "US",
  "kind": "B1",
  "title": "Robot and auto data processing method thereof",
  "abstract": "The present disclosure provides a robot and an audio data processing method thereof. The robot includes a body part, a main control module, and a sound pickup module electrically coupled to the main control module. The sound pickup module includes N microphones distributed around the body part to collect audio data. The main control module is configured to obtain the audio data of a sound source from a part of the N microphones collecting the audio data of the sound source without blocked by the body part, and perform a sound source localization and a sound pickup based on the obtained audio data. The 360-degree wake-up and sound source localization of the robot and the beam-forming of directional beams are realized. In addition, the sound pickup is realized without forming microphone holes on the head of the robot, hence the aesthetics of the robot will not be affected.",
  "claims": [
    "1. A robot, comprising: at least one body part; a main control module comprising a data buffer pool; and a sound pickup module electrically coupled to the main control module, wherein the sound pickup module comprises N microphones distributed around the body part to collect audio data, where N≥3 and N is an integer, and wherein when collecting the audio data, a part of the N microphones is capable of receiving a direct sound from a sound source, but the rest part of the N microphones is incapable of receiving the direct sound but reflect sounds of the sound source; wherein, the main control module is configured to obtain first audio data of the sound source collected by the N microphones, perform a sound source localization based on the first audio data, obtain second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound, and perform a sound pickup and a voice recognition based on the second audio data; wherein, the main control module is further configured to store X channels of reference audio data and N channels of audio data to the data buffer pool, obtain a first group of the audio data from the data buffer pool as the first audio data of the sound source collected by the N microphones, to use a first predetermined algorithm to locate a sound source, and obtain a second group of the audio data from the data buffer pool as the second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound, to use a second predetermined algorithm to perform a beam-forming and an audio noise reduction; wherein the N channels of audio data is six channels of audio data, and the X channels of reference audio data is two channels of reference audio data; wherein, audio data obtained by a first microphone in microphones arrays is taken as first audio data, audio data obtained by a second microphone in the microphones arrays is taken as second audio data, audio data obtained by a third microphone in the microphones arrays is taken as third audio data, audio data obtained by a fourth microphone in the microphones arrays is taken as fourth audio data, audio data obtained by a fifth microphone in the microphones arrays is taken as fifth audio data, audio data obtained by a sixth microphone in the microphones arrays is taken as sixth audio data, first channel reference audio data in the two channels of the reference audio data is taken as seventh audio data, and second channel reference audio data in the two channels of the reference audio data is taken as eighth audio data; wherein, the first group of the audio data comprises the first audio data, the second audio data, the third audio data, the fourth audio data, the fifth audio data, the sixth audio data, the seventh audio data, and the eighth audio data; and wherein, the second group of the audio data comprises the first audio data, the second audio data, the third audio data, the sixth audio data, the seventh audio data, and the eighth audio data.",
    "2. The robot of claim 1, wherein the sound pickup module further comprises: a MIC small board electrically coupled to each of the microphone array and the main control module, wherein the MIC small board is coupled to perform an analog-to-digital conversion on the N channels of audio data collected by the microphones, encode the converted audio data, and transmit the encoded audio data to the main control module.",
    "3. The robot of claim 2, wherein the MIC small board comprises: an analog-to-digital converter electrically coupled to the microphone arrays and the main control module, wherein the analog-to-digital converter performs the analog-to-digital conversion on the N channels of audio data.",
    "4. The robot of claim 2, further comprising: a power amplifier electrically coupled to the main control module; wherein the main control module is configured to generate the X channels of reference audio data based on audio data obtained from the power amplifier to transmit to the MIC small board, and the MIC small board is further configured to perform an analog-to-digital conversion on the X channels of reference audio data, encode the converted X channels of reference audio data, and transmit the encoded X channels of reference audio data to the main control module.",
    "5. The robot of claim 4, wherein, the main control module is further configured to obtain the audio data played by the power amplifier and generate the X channels of reference audio data based on the audio data played by the power amplifier.",
    "6. The robot of claim 5, wherein, number of the X channels of reference audio data is same as number of channels of the audio data played by the power amplifier, the main control module is electrically coupled to the MIC small board directly through data lines, and amount of the data lines corresponds to amount of the X channels of the reference audio data.",
    "7. The robot of claim 5, wherein, the MIC small board is further configured to perform a data fusion, fuse received reference audio data with the N channels of audio data, and transmit fused audio data to the main control module.",
    "8. The robot of claim 1, wherein the body part is a neck, the microphone array comprises six microphones, the six microphones are disposed around the neck and are distributed on a circumference centered on any point on a longitudinal axis of the body part.",
    "9. The robot of claim 1, wherein the body part comprises at least one of neck and a trunk.",
    "10. The robot of claim 1, wherein the main control module is further configured to determine an angle difference between a sound source position and a current position through the sound source localization, control the robot to turn according to the angle difference, wake up the robot, and perform the sound pickup and the voice recognition based on the second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound.",
    "11. The robot of claim 1, wherein the at least one body part includes a head, a neck, and a trunk, and the N microphones are distributed around each of the at least one body part in a non-even manner, or, the N microphones are distributed around the head, the trunk, or two or more of the at least one body part.",
    "12. The robot of claim 1, wherein the main control module is a development board, and the data buffer pool is configured in a software layer of the development board.",
    "13. A computer-implemented audio data processing method based on a robot comprising: at least one body part; a main control module; and a sound pickup module electrically coupled to the main control module, wherein the sound pickup module comprises N microphones distributed around the body part to collect audio data, where N≥3 and N is an integer, and wherein when collecting the audio data, a part of the N microphones is capable of receiving a direct sound from a sound source, but the rest part of the N microphones is incapable of receiving the direct sound but reflect sounds of the sound source; wherein, the main control module is configured to obtain first audio data of the sound source collected by the N microphones, perform a sound source localization based on the first audio data, obtain second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound, and perform a sound pickup and a voice recognition based on the second audio data; the method comprising executing on a processor of the robot the steps of: collecting audio data through the N microphones of the sound pickup module; transmitting the N channels of audio data collected by the N microphones to the main control module; storing, by the main control module, the N channels of audio data to a data buffer pool; and performing, by the main control module, the sound source localization and the sound pickup based on the audio data; wherein the step of storing, by the main control module, the N channels of audio data to the data buffer pool and the step of performing, by the main control module, the sound source localization and the sound pickup based on the audio data further comprise: storing X channels of reference audio data and the N channels of audio data to the data buffer pool; obtaining a first group of the audio data from the data buffer pool as the first audio data of the sound source collected by the N microphones, to use a first predetermined algorithm to locate a sound source; and obtaining a second group of the audio data from the data buffer pool as the second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound, to use a second predetermined algorithm to perform a beam-forming and an audio noise reduction; wherein the N channels of audio data is six channels of audio data, and the X channels of reference audio data is two channels of reference audio data; wherein, audio data obtained by a first microphone in microphones arrays is taken as first audio data, audio data obtained by a second microphone in the microphones arrays is taken as second audio data, audio data obtained by a third microphone in the microphones arrays is taken as third audio data, audio data obtained by a fourth microphone in the microphones arrays is taken as fourth audio data, audio data obtained by a fifth microphone in the microphones arrays is taken as fifth audio data, audio data obtained by a sixth microphone in the microphones arrays is taken as sixth audio data, first channel reference audio data in the two channels of reference audio data is taken as seventh audio data, and second channel reference audio data in the two channels of reference audio data is taken as eighth audio data; wherein, the first group of the audio data comprises the first audio data, the second audio data, the third audio data, the fourth audio data, the fifth audio data, the sixth audio data, the seventh audio data, and the eighth audio data; and wherein, the second group of the audio data comprises the first audio data, the second audio data, the third audio data, the sixth audio data, the seventh audio data, and the eighth audio data.",
    "14. The method of claim 13, wherein, the main control module determines an angle difference between a sound source position and a current position through the sound source localization, control the robot to turn according to the angle difference, wake up the robot, and perform the sound pickup and the voice recognition based on the second audio data of the sound source collected by the part of the N microphones which is capable of receiving the direct sound."
  ],
  "description_excerpt": "1. Technical Field The present disclosure relates to intelligent robot technology, and particularly to a robot and an audio data processing method thereof. 2. Description of Related Art When designing a robot, if the position of a microphone array is not arranged correctly, the voice interaction will be affected. Because the most basic requirement and prerequisite for the beam-forming of the microphone array is that sounds should directly reach each microphone in the microphone array. Therefore, if an annular microphone array is disposed at the neck of the robot, the neck of the robot will hide the microphones behind the neck, which causes the sounds to be reflected by the neck and can not directly reach the microphone behind the neck of the robot, thus affecting the effect of sound pickup. In order to resolve the above-mentioned problems, it is generally to place an annular microphone array on the head of the robot, or to use an annular microphone array and a linear microphone array at the same time, where the annular microphone array is disposed at the neck of the robot for realizing the 360-degree wake-up and 360-degree sound source localization of the robot, and the linear microphone array is disposed on the head of the robot for beam-forming so as to perform sound pickup.\n\nHowever, disposing the annular microphone array on the head of the robot will cause a limit to the height of the robot.",
  "cpc": [
    "H04R 1/406",
    "B25J 11/0005",
    "G10L 2021/02082",
    "G10L 2021/02166",
    "G10L 21/0208",
    "G10L 21/0216",
    "H04R 1/028",
    "H04R 2201/401",
    "H04R 2410/01",
    "H04R 2420/01",
    "H04R 2430/20"
  ],
  "ipc": [
    "H04R 1/02",
    "H04R 1/40"
  ],
  "assignees": [
    "UBTECH ROBOTICS CORP LTD"
  ],
  "inventors": [
    "XIONG YOUJUN",
    "Xing Fanglin"
  ],
  "filing_date": "2019-06-21",
  "publication_date": "2020-05-26",
  "grant_date": "2020-05-26",
  "priority_date": "2018-12-28",
  "application_number": "US-201916447986-A",
  "family_id": "70549763",
  "citations": [
    "JP2007221300A",
    "JP2007241157A",
    "JP2008278399A",
    "JP2011069901A",
    "US2005058300A1",
    "US2010150364A1",
    "US2017243577A1",
    "US2018374494A1",
    "US2019104360A1",
    "US2019250245A1",
    "US2019364375A1"
  ]
}

Record 860 of 5,000 in Patents full text (MLC-0201). Request the full dataset.