MLchartDataset catalogue

Patent · US10702991B2 · B2 · US

Apparatus, robot, method and recording medium having program recorded thereon

(11) Publication number
US10702991B2
(21) Application number
15/899,372
(22) Filing date
2018-02-20
(30) Priority date
2017-03-08
(43) Publication date
2020-07-07
(45) Date of grant
2020-07-07
(51) IPC
B25J 11/00; G06K 9/00; G06N 3/00; G10L 15/08; G10L 25/63
(52) CPC
  • B25J Manipulators; chambers provided with manipulation devices: 11/0015, 11/001
  • G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/00718
  • G06N Computing arrangements based on specific computational models: 3/008
  • G06V Image or video recognition or understanding: 20/41
  • G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 15/08, 2015/088, 25/63, 25/84
(73) Assignee
Panasonic Intellectual Property Management Co Ltd
(72) Inventors
Michiko Sasagawa; Ryouta Miyazaki
(54) Title
Apparatus, robot, method and recording medium having program recorded thereon
(57) Abstract

An apparatus, robot, method and recording medium is provided, wherein when it is determined that a speech of an adult includes a warning word, whether the adult is angry or scolding is determined based on a physical feature value of the speech of the adult. When it is determined that the adult is angry, at least any of the following processes is performed: (a) a process of causing a loudspeaker to output a first sound, (b) a process of causing an apparatus to perform a first operation, and (c) a process of causing a display to perform a first display.

Full text
View on Google Patents

Claims (8)

  1. An apparatus for processing a voice, the apparatus comprising: a microphone which acquires a sound around the apparatus; a processor and a memory that stores a set of executable instructions, wherein when the processor executes the instructions, the apparatus operates as: an extractor which extracts a voice from the acquired sound and determines whether the voice includes a speech of an adult; a voice recognizer which recognizes the speech of the adult when it is determined that the voice includes the speech of the adult and determines whether the speech of the adult includes a phrase included in a dictionary stored in the memory, wherein the voice recognizer further determines whether the speech of the adult includes a phrase corresponding to a name of a child, based on data indicating the name of the child stored in the memory, a speech analyzer configured to determine, after the voice recognizer determines that the speech of the adult includes the phrase included in the dictionary, whether the adult is angry and whether the adult is scolding, based on a physical feature value of the speech of the adult, wherein when it is determined that the speech of the adult includes a phrase corresponding to the name of the child, the speech analyzer further determines that the child is a target person at whom the adult is scolding or angry, a controller which causes the apparatus to perform a first process to alert the adult of the adult being angry when it is determined that the adult is angry, a loudspeaker, wherein the first process includes either of (i) a process of causing the loudspeaker to output a first sound and (ii) a process of causing the apparatus to perform a first operation, a display, wherein the first process further includes either of (i) a process of causing the display to perform a first display and (ii) the process of causing the apparatus to perform the first operation, and a camera which acquires video data around the apparatus, wherein the first process further includes either of (i) a process of causing the camera to take an image of the adult and (ii) the process of causing the apparatus to perform the first operation, wherein when the processor executes the instructions, the apparatus is further configured to operate as: a comparator which determines whether a person included in the video data is the child, based on video data corresponding to the child stored in the memory, and a video analyzer which determines, based on the video data, whether an orientation of the child has been changed in a second period after the speech of the adult is recognized when it is determined that the adult is scolding the child and the person included in the video data is the child and determines, based on the video data, whether the child is continuously holding an object by hand in the second period when it is determined that the orientation of the child has not been changed, wherein: in the second period, when it is determined that the orientation of the child has not been changed or when it is determined that the child is continuously holding the object by hand, the controller causes the apparatus to perform a second process, and the second process includes any of (i) a process of causing the loudspeaker to output a second sound, (ii) a process of causing the apparatus to perform a second operation, (iii) a process of causing the apparatus to perform the second operation and (iv) a process of causing the display to perform a second display.
  2. The apparatus according to claim 1, wherein the second sound includes a predetermined alarm sound.
  3. The apparatus according to claim 1, wherein the second sound includes predetermined music.
  4. The apparatus according to claim 1, wherein the second sound includes a voice prompting the child to stop an action the child is currently doing.
  5. The apparatus according to claim 1, wherein the second sound includes a voice asking the child what the child is doing now.
  6. The apparatus according to claim 1, wherein the second operation includes an operation of causing the display to be oriented toward the child.
  7. The apparatus according to claim 1, wherein the second operation is an operation of causing the apparatus to be oriented toward the child.
  8. The apparatus according to claim 1, wherein the second display includes a display symbolically representing eyes and a mouth on the apparatus, and the display corresponds to a predetermined facial expression on the apparatus.

Description

The present disclosure relates to a voice processing apparatus, robot, method, and recording medium having a program recorded thereon.

In recent years, technologies for user emotion recognition by processing voice emitted from a user have been actively conducted. Examples of a conventional emotion recognition method include a method using language information about a voice emitted from a speaker, a method using prosodic characteristics of sound of the voice, and a method of performing a facial expression analysis from a face image.

Japanese Patent No. 4015424 discloses an example of technology of emotion recognition based on the language information about a voice emitted from a user. Specifically, Japanese Patent No. 4015424 discloses a technique as follows. When a user is asked a question, “Do you have fun in playing football?”, and makes a reply, “Playing football is very boring”, “football” is extracted as a keyword, and a phase including the keyword includes words indicating a negative feeling, “very boring”. Thus, an inference is made that the user is not interested in football, and a question about topics other than football is made.

Also, Japanese Unexamined Patent Application Publication No. 2006-123136 discloses an example of technology of determining an emotion from inputted voice and face image of a user and outputting a response in accordance with the determined emotion. Specifically, Japanese Unexamined Patent Application Publication No.

Citations (21)

  • JP2003202892A
  • JP2006123136A
  • US20110118870A1
  • JP2009131928A
  • US20130016815A1
  • US8903176B2
  • US20140095151A1
  • US20140093849A1
  • US9846843B2
  • US20150310878A1
  • US20160019915A1
  • US20180285641A1
  • US20160372110A1
  • US20160379107A1
  • US20180331839A1
  • US20170244942A1
  • US20170310820A1
  • US20180166076A1
  • US20180240454A1
  • US20180257236A1
  • US10020076B1
Record as JSON
{
  "publication_number": "US10702991B2",
  "country": "US",
  "kind": "B2",
  "title": "Apparatus, robot, method and recording medium having program recorded thereon",
  "abstract": "An apparatus, robot, method and recording medium is provided, wherein when it is determined that a speech of an adult includes a warning word, whether the adult is angry or scolding is determined based on a physical feature value of the speech of the adult. When it is determined that the adult is angry, at least any of the following processes is performed: (a) a process of causing a loudspeaker to output a first sound, (b) a process of causing an apparatus to perform a first operation, and (c) a process of causing a display to perform a first display.",
  "claims": [
    "1. An apparatus for processing a voice, the apparatus comprising: a microphone which acquires a sound around the apparatus; a processor and a memory that stores a set of executable instructions, wherein when the processor executes the instructions, the apparatus operates as: an extractor which extracts a voice from the acquired sound and determines whether the voice includes a speech of an adult; a voice recognizer which recognizes the speech of the adult when it is determined that the voice includes the speech of the adult and determines whether the speech of the adult includes a phrase included in a dictionary stored in the memory, wherein the voice recognizer further determines whether the speech of the adult includes a phrase corresponding to a name of a child, based on data indicating the name of the child stored in the memory, a speech analyzer configured to determine, after the voice recognizer determines that the speech of the adult includes the phrase included in the dictionary, whether the adult is angry and whether the adult is scolding, based on a physical feature value of the speech of the adult, wherein when it is determined that the speech of the adult includes a phrase corresponding to the name of the child, the speech analyzer further determines that the child is a target person at whom the adult is scolding or angry, a controller which causes the apparatus to perform a first process to alert the adult of the adult being angry when it is determined that the adult is angry, a loudspeaker, wherein the first process includes either of (i) a process of causing the loudspeaker to output a first sound and (ii) a process of causing the apparatus to perform a first operation, a display, wherein the first process further includes either of (i) a process of causing the display to perform a first display and (ii) the process of causing the apparatus to perform the first operation, and a camera which acquires video data around the apparatus, wherein the first process further includes either of (i) a process of causing the camera to take an image of the adult and (ii) the process of causing the apparatus to perform the first operation, wherein when the processor executes the instructions, the apparatus is further configured to operate as: a comparator which determines whether a person included in the video data is the child, based on video data corresponding to the child stored in the memory, and a video analyzer which determines, based on the video data, whether an orientation of the child has been changed in a second period after the speech of the adult is recognized when it is determined that the adult is scolding the child and the person included in the video data is the child and determines, based on the video data, whether the child is continuously holding an object by hand in the second period when it is determined that the orientation of the child has not been changed, wherein: in the second period, when it is determined that the orientation of the child has not been changed or when it is determined that the child is continuously holding the object by hand, the controller causes the apparatus to perform a second process, and the second process includes any of (i) a process of causing the loudspeaker to output a second sound, (ii) a process of causing the apparatus to perform a second operation, (iii) a process of causing the apparatus to perform the second operation and (iv) a process of causing the display to perform a second display.",
    "2. The apparatus according to claim 1, wherein the second sound includes a predetermined alarm sound.",
    "3. The apparatus according to claim 1, wherein the second sound includes predetermined music.",
    "4. The apparatus according to claim 1, wherein the second sound includes a voice prompting the child to stop an action the child is currently doing.",
    "5. The apparatus according to claim 1, wherein the second sound includes a voice asking the child what the child is doing now.",
    "6. The apparatus according to claim 1, wherein the second operation includes an operation of causing the display to be oriented toward the child.",
    "7. The apparatus according to claim 1, wherein the second operation is an operation of causing the apparatus to be oriented toward the child.",
    "8. The apparatus according to claim 1, wherein the second display includes a display symbolically representing eyes and a mouth on the apparatus, and the display corresponds to a predetermined facial expression on the apparatus."
  ],
  "description_excerpt": "The present disclosure relates to a voice processing apparatus, robot, method, and recording medium having a program recorded thereon.\n\nIn recent years, technologies for user emotion recognition by processing voice emitted from a user have been actively conducted. Examples of a conventional emotion recognition method include a method using language information about a voice emitted from a speaker, a method using prosodic characteristics of sound of the voice, and a method of performing a facial expression analysis from a face image.\n\nJapanese Patent No. 4015424 discloses an example of technology of emotion recognition based on the language information about a voice emitted from a user. Specifically, Japanese Patent No. 4015424 discloses a technique as follows. When a user is asked a question, “Do you have fun in playing football?”, and makes a reply, “Playing football is very boring”, “football” is extracted as a keyword, and a phase including the keyword includes words indicating a negative feeling, “very boring”. Thus, an inference is made that the user is not interested in football, and a question about topics other than football is made.\n\nAlso, Japanese Unexamined Patent Application Publication No. 2006-123136 discloses an example of technology of determining an emotion from inputted voice and face image of a user and outputting a response in accordance with the determined emotion. Specifically, Japanese Unexamined Patent Application Publication No.",
  "cpc": [
    "B25J 11/0015",
    "B25J 11/001",
    "G06K 9/00718",
    "G06N 3/008",
    "G06V 20/41",
    "G10L 15/08",
    "G10L 2015/088",
    "G10L 25/63",
    "G10L 25/84"
  ],
  "ipc": [
    "B25J 11/00",
    "G06K 9/00",
    "G06N 3/00",
    "G10L 15/08",
    "G10L 25/63"
  ],
  "assignees": [
    "Panasonic Intellectual Property Management Co Ltd"
  ],
  "inventors": [
    "Michiko Sasagawa",
    "Ryouta Miyazaki"
  ],
  "filing_date": "2018-02-20",
  "publication_date": "2020-07-07",
  "grant_date": "2020-07-07",
  "priority_date": "2017-03-08",
  "application_number": "US-201815899372-A",
  "family_id": "61526555",
  "cited_by_count": 0,
  "citations": [
    "JP2003202892A",
    "JP2006123136A",
    "US20110118870A1",
    "JP2009131928A",
    "US20130016815A1",
    "US8903176B2",
    "US20140095151A1",
    "US20140093849A1",
    "US9846843B2",
    "US20150310878A1",
    "US20160019915A1",
    "US20180285641A1",
    "US20160372110A1",
    "US20160379107A1",
    "US20180331839A1",
    "US20170244942A1",
    "US20170310820A1",
    "US20180166076A1",
    "US20180240454A1",
    "US20180257236A1",
    "US10020076B1"
  ]
}

Record 2,087 of 8,000 in Patents full text (MLC-0201). Request the full dataset.