MLchartDataset catalogue

Patent · US10152968B1 · B1 · US

Systems and methods for speech-based monitoring and/or control of automation devices

(11) Publication number
US10152968B1
(21) Application number
15/194,500
(22) Filing date
2016-06-27
(30) Priority date
2015-06-26
(43) Publication date
2018-12-11
(45) Date of grant
2018-12-11
(51) IPC
G06F 3/16; G10L 15/00; G10L 15/183; G10L 15/22
(52) CPC
  • G10L Speech analysis techniques or speech synthesis; speech recognition; speech or voice processing techniques; speech or audio coding or decoding: 15/22, 15/00, 15/16, 15/183, 2015/223, 2015/228
  • G06F Electric digital data processing: 3/167
(73) Assignee
Iconics Inc
(72) Inventors
Russell L. Agrusa; Vojtech Kresl; Christopher N. Elsbree; Marco Tagliaferri; Lukas Volf
(54) Title
Systems and methods for speech-based monitoring and/or control of automation devices
(57) Abstract

Systems and methods for speech-based monitoring and/or control of automation devices are described. A speech-based method for monitoring and/or control of automation devices may include steps of determining a type of automation device to which first speech relates based, at least in part, on a location associated with the first speech; selecting a topic-specific speech recognition model adapted to recognize speech related to the determined type of automation device; using the topic-specific speech recognition model to recognize second speech provided at the location, wherein recognizing the second speech comprises identifying a query or command relating to the type of automation device and represented by the second speech; and issuing the query or command represented by the second speech to an automation device of the determined type.

Full text
View on Google Patents

Claims (22)

  1. A computer-implemented method comprising: inferring, based on data indicating locations of a plurality of automation devices and on a location associated with a user, that first speech of the user is directed to a particular automation device of a particular type, wherein the particular automation device is installed and operating at one of the indicated locations, and wherein the location of the particular automation device is proximate to the location associated with the user while the user is uttering the first speech; selecting a topic-specific speech recognition model adapted to recognize speech related to the determined type of automation device; using the topic-specific speech recognition model to recognize second speech provided at the location, wherein recognizing the second speech comprises identifying a query or command relating to the type of automation device and represented by the second speech; and issuing the query or command represented by the second speech to the particular automation device of the particular type, thereby prompting the particular automation device to perform an act responsive to the command or query, wherein the type of automation device is a manufacturing process control type, an industrial process control type, an energy-production process control type, a water treatment process control type, an environmental regulation process control type, and/or a utility process control type.
  2. The method of claim 1, wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining first location data representing the location associated with the user; providing the first location data to a search engine operable to search for an automation device disposed at a location proximate to the location represented by the first location data; and determining, based on results provided by the search engine, that the particular automation device of the particular type is disposed proximate to the location represented by the first location data.
  3. The method of claim 1, wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining first location data representing the location associated with the user; obtaining second location data representing one or more locations of one or more respective automation devices, wherein the one or more automation devices include a first automation device located nearest to the location represented by the first location data; and based on the first location data and the second location data, identifying the first automation device as the particular automation device to which the first speech is directed.
  4. The method of claim 1, wherein the location associated with the user is the location of an acoustic sensor operable to sense the first speech, and wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining identification data identifying the acoustic sensor; providing the identification data to a search engine operable to search for an automation device disposed proximate to the acoustic sensor identified by the identification data; and determining, based on results provided by the search engine, that the particular automation device of the particular type is disposed proximate to the acoustic sensor.
  5. The method of claim 1, wherein the location associated with the user comprises a location of a device including an acoustic sensor operable to detect the first speech, a location from which the first speech originates, or a location of the user by whom the first speech is provided.
  6. The method of claim 5, wherein the acoustic sensor comprises a microphone.
  7. The method of claim 1, wherein inferring the particular automation device to which the first speech is directed is performed before the first speech is uttered.
  8. The method of claim 1, wherein the particular automation device to which the first speech is directed is further inferred based, at least in part, on content of the first speech.
  9. The method of claim 1, wherein the topic-specific speech recognition model includes an acoustic model and a language model, and wherein the language model is adapted to model speech related to the particular type of the particular automation device.
  10. A system comprising: at least one memory for storing computer-executable instructions; and at least one processing unit for executing the instructions, wherein execution of the instructions causes the at least one processing unit to perform operations comprising: inferring, based on data indicating locations of a plurality of automation devices and on a location associated with a user, that first speech of the user is directed to a particular automation device of a particular type, wherein the particular automation device is installed and operating at one of the indicated locations, and wherein the location of the particular automation device is proximate to the location associated with the user while the user is uttering the first speech; selecting a topic-specific speech recognition model adapted to recognize speech related to the determined type of automation device, using the topic-specific speech recognition model to recognize second speech provided at the location, wherein recognizing the second speech comprises identifying a query or command relating to the type of automation device and represented by the second speech, and issuing the query or command represented by the second speech to the particular automation device of the particular type, thereby prompting the particular automation device to perform an act responsive to the command or query, wherein the type of automation device is a manufacturing process control type, an industrial process control type, an energy-production process control type, a water treatment process control type, an environmental regulation process control type, and/or a utility process control type.
  11. The system of claim 10, wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining first location data representing the location associated with the user; providing the first location data to a search engine operable to search for an automation device disposed at a location proximate to the location represented by the first location data; and determining, based on results provided by the search engine, that the particular automation device of the particular type is disposed proximate to the location represented by the first location data.
  12. The system of claim 10, wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining first location data representing the location associated with the user; obtaining second location data representing one or more locations of one or more respective automation devices, wherein the one or more automation devices include a first automation device located nearest to the location represented by the first location data; and based on the first location data and the second location data, identifying the first automation device as the particular automation device to which the first speech is directed.
  13. The system of claim 10, wherein the location associated with the user is the location of an acoustic sensor operable to sense the first speech, and wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining identification data identifying the acoustic sensor; providing the identification data to a search engine operable to search for an automation device disposed proximate to the acoustic sensor identified by the identification data; and determining, based on results provided by the search engine, that the particular automation device of the particular type is disposed proximate to the acoustic sensor.
  14. The system of claim 10, further comprising an acoustic sensor operable to detect speech, wherein the system is a mobile device, and wherein the location associated with the user is a location of the mobile device.
  15. The system of claim 14, wherein the acoustic sensor comprises a microphone.
  16. The system of claim 10, wherein inferring the particular automation device to which the first speech is directed is performed before the first speech is uttered.
  17. The system of claim 10, wherein the particular automation device to which the first speech is directed is further inferred based, at least in part, on content of the first speech.
  18. The method of claim 1, wherein the topic-specific speech recognition model includes an acoustic model and a language model, and wherein the language model is adapted to model speech related to the particular type of the particular automation device.
  19. The method of claim 1, wherein issuing the query or command to the particular automation device comprises sending the query or command to a process controller associated with the particular automation device, and wherein prompting the particular automation device to perform an act responsive to the command or query comprises prompting the process controller to control the particular automation device to perform an act responsive to the command or query.
  20. The method of claim 1, wherein the plurality of automation devices includes two or more proximate automation devices installed and operating at locations proximate to the location associated with the user, and wherein inferring that first speech of the user is directed to the particular automation device includes: inferring, independent of content of the second speech, that the first speech is directed to an unknown one of the two or more proximate automation devices; and based on the content of the first speech, selecting, from the two or more proximate automation devices, the particular automation device to which the first speech is directed.
  21. The method of claim 1, further comprising: prior to issuing the query or command to the particular automation device, prompting a user to confirm the query or command and an identity of the particular automation device, wherein prompting the user includes presenting a confirmation prompt specifying the query or command and the particular automation device to which the query or command will be issued.
  22. The method of claim 1, wherein the first speech comprises the second speech.

Description

The present disclosure relates generally to speech recognition and, more particularly, to systems and methods for speech-based monitoring and/or control of automation devices.

Speech recognition systems generally use speech recognition models to recognize speech. For automatic dictation or speech-to-text processing applications, the recognized speech can then be converted into corresponding text. Alternatively or in addition, for natural language processing (NLP) applications, the recognized speech can be interpreted and action can be taken based thereon.

A speech recognition model generally includes an acoustic model and a language model. Acoustic models generally model the relationships between audio signals (e.g., electrical signals representing sounds) and phonemes or other linguistic units of speech. An acoustic model may be created by using audio recordings of speech and corresponding transcriptions of the speech to train a predictive model (e.g., a statistical model) to identify the linguistic units represented by audio signals. Different acoustic models may be specifically trained for use by a particular user (or group of users), or for use in a particular environment.

Citations (12)

  • US5317673A
  • US5581659A
  • US6151574A
  • US20040260538A1
  • US20050021337A1
  • US20100318412A1
  • US20120123904A1
  • US20140257803A1
  • US9153231B1
  • US20150039299A1
  • US20150127327A1
  • US9324321B2
Record as JSON
{
  "publication_number": "US10152968B1",
  "country": "US",
  "kind": "B1",
  "title": "Systems and methods for speech-based monitoring and/or control of automation devices",
  "abstract": "Systems and methods for speech-based monitoring and/or control of automation devices are described. A speech-based method for monitoring and/or control of automation devices may include steps of determining a type of automation device to which first speech relates based, at least in part, on a location associated with the first speech; selecting a topic-specific speech recognition model adapted to recognize speech related to the determined type of automation device; using the topic-specific speech recognition model to recognize second speech provided at the location, wherein recognizing the second speech comprises identifying a query or command relating to the type of automation device and represented by the second speech; and issuing the query or command represented by the second speech to an automation device of the determined type.",
  "claims": [
    "1. A computer-implemented method comprising: inferring, based on data indicating locations of a plurality of automation devices and on a location associated with a user, that first speech of the user is directed to a particular automation device of a particular type, wherein the particular automation device is installed and operating at one of the indicated locations, and wherein the location of the particular automation device is proximate to the location associated with the user while the user is uttering the first speech; selecting a topic-specific speech recognition model adapted to recognize speech related to the determined type of automation device; using the topic-specific speech recognition model to recognize second speech provided at the location, wherein recognizing the second speech comprises identifying a query or command relating to the type of automation device and represented by the second speech; and issuing the query or command represented by the second speech to the particular automation device of the particular type, thereby prompting the particular automation device to perform an act responsive to the command or query, wherein the type of automation device is a manufacturing process control type, an industrial process control type, an energy-production process control type, a water treatment process control type, an environmental regulation process control type, and/or a utility process control type.",
    "2. The method of claim 1, wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining first location data representing the location associated with the user; providing the first location data to a search engine operable to search for an automation device disposed at a location proximate to the location represented by the first location data; and determining, based on results provided by the search engine, that the particular automation device of the particular type is disposed proximate to the location represented by the first location data.",
    "3. The method of claim 1, wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining first location data representing the location associated with the user; obtaining second location data representing one or more locations of one or more respective automation devices, wherein the one or more automation devices include a first automation device located nearest to the location represented by the first location data; and based on the first location data and the second location data, identifying the first automation device as the particular automation device to which the first speech is directed.",
    "4. The method of claim 1, wherein the location associated with the user is the location of an acoustic sensor operable to sense the first speech, and wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining identification data identifying the acoustic sensor; providing the identification data to a search engine operable to search for an automation device disposed proximate to the acoustic sensor identified by the identification data; and determining, based on results provided by the search engine, that the particular automation device of the particular type is disposed proximate to the acoustic sensor.",
    "5. The method of claim 1, wherein the location associated with the user comprises a location of a device including an acoustic sensor operable to detect the first speech, a location from which the first speech originates, or a location of the user by whom the first speech is provided.",
    "6. The method of claim 5, wherein the acoustic sensor comprises a microphone.",
    "7. The method of claim 1, wherein inferring the particular automation device to which the first speech is directed is performed before the first speech is uttered.",
    "8. The method of claim 1, wherein the particular automation device to which the first speech is directed is further inferred based, at least in part, on content of the first speech.",
    "9. The method of claim 1, wherein the topic-specific speech recognition model includes an acoustic model and a language model, and wherein the language model is adapted to model speech related to the particular type of the particular automation device.",
    "10. A system comprising: at least one memory for storing computer-executable instructions; and at least one processing unit for executing the instructions, wherein execution of the instructions causes the at least one processing unit to perform operations comprising: inferring, based on data indicating locations of a plurality of automation devices and on a location associated with a user, that first speech of the user is directed to a particular automation device of a particular type, wherein the particular automation device is installed and operating at one of the indicated locations, and wherein the location of the particular automation device is proximate to the location associated with the user while the user is uttering the first speech; selecting a topic-specific speech recognition model adapted to recognize speech related to the determined type of automation device, using the topic-specific speech recognition model to recognize second speech provided at the location, wherein recognizing the second speech comprises identifying a query or command relating to the type of automation device and represented by the second speech, and issuing the query or command represented by the second speech to the particular automation device of the particular type, thereby prompting the particular automation device to perform an act responsive to the command or query, wherein the type of automation device is a manufacturing process control type, an industrial process control type, an energy-production process control type, a water treatment process control type, an environmental regulation process control type, and/or a utility process control type.",
    "11. The system of claim 10, wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining first location data representing the location associated with the user; providing the first location data to a search engine operable to search for an automation device disposed at a location proximate to the location represented by the first location data; and determining, based on results provided by the search engine, that the particular automation device of the particular type is disposed proximate to the location represented by the first location data.",
    "12. The system of claim 10, wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining first location data representing the location associated with the user; obtaining second location data representing one or more locations of one or more respective automation devices, wherein the one or more automation devices include a first automation device located nearest to the location represented by the first location data; and based on the first location data and the second location data, identifying the first automation device as the particular automation device to which the first speech is directed.",
    "13. The system of claim 10, wherein the location associated with the user is the location of an acoustic sensor operable to sense the first speech, and wherein inferring that the first speech of the user is directed to the particular automation device of the particular type comprises: obtaining identification data identifying the acoustic sensor; providing the identification data to a search engine operable to search for an automation device disposed proximate to the acoustic sensor identified by the identification data; and determining, based on results provided by the search engine, that the particular automation device of the particular type is disposed proximate to the acoustic sensor.",
    "14. The system of claim 10, further comprising an acoustic sensor operable to detect speech, wherein the system is a mobile device, and wherein the location associated with the user is a location of the mobile device.",
    "15. The system of claim 14, wherein the acoustic sensor comprises a microphone.",
    "16. The system of claim 10, wherein inferring the particular automation device to which the first speech is directed is performed before the first speech is uttered.",
    "17. The system of claim 10, wherein the particular automation device to which the first speech is directed is further inferred based, at least in part, on content of the first speech.",
    "18. The method of claim 1, wherein the topic-specific speech recognition model includes an acoustic model and a language model, and wherein the language model is adapted to model speech related to the particular type of the particular automation device.",
    "19. The method of claim 1, wherein issuing the query or command to the particular automation device comprises sending the query or command to a process controller associated with the particular automation device, and wherein prompting the particular automation device to perform an act responsive to the command or query comprises prompting the process controller to control the particular automation device to perform an act responsive to the command or query.",
    "20. The method of claim 1, wherein the plurality of automation devices includes two or more proximate automation devices installed and operating at locations proximate to the location associated with the user, and wherein inferring that first speech of the user is directed to the particular automation device includes: inferring, independent of content of the second speech, that the first speech is directed to an unknown one of the two or more proximate automation devices; and based on the content of the first speech, selecting, from the two or more proximate automation devices, the particular automation device to which the first speech is directed.",
    "21. The method of claim 1, further comprising: prior to issuing the query or command to the particular automation device, prompting a user to confirm the query or command and an identity of the particular automation device, wherein prompting the user includes presenting a confirmation prompt specifying the query or command and the particular automation device to which the query or command will be issued.",
    "22. The method of claim 1, wherein the first speech comprises the second speech."
  ],
  "description_excerpt": "The present disclosure relates generally to speech recognition and, more particularly, to systems and methods for speech-based monitoring and/or control of automation devices.\n\nSpeech recognition systems generally use speech recognition models to recognize speech. For automatic dictation or speech-to-text processing applications, the recognized speech can then be converted into corresponding text. Alternatively or in addition, for natural language processing (NLP) applications, the recognized speech can be interpreted and action can be taken based thereon.\n\nA speech recognition model generally includes an acoustic model and a language model. Acoustic models generally model the relationships between audio signals (e.g., electrical signals representing sounds) and phonemes or other linguistic units of speech. An acoustic model may be created by using audio recordings of speech and corresponding transcriptions of the speech to train a predictive model (e.g., a statistical model) to identify the linguistic units represented by audio signals. Different acoustic models may be specifically trained for use by a particular user (or group of users), or for use in a particular environment.",
  "cpc": [
    "G10L 15/22",
    "G06F 3/167",
    "G10L 15/00",
    "G10L 15/16",
    "G10L 15/183",
    "G10L 2015/223",
    "G10L 2015/228"
  ],
  "ipc": [
    "G06F 3/16",
    "G10L 15/00",
    "G10L 15/183",
    "G10L 15/22"
  ],
  "assignees": [
    "Iconics Inc"
  ],
  "inventors": [
    "Russell L. Agrusa",
    "Vojtech Kresl",
    "Christopher N. Elsbree",
    "Marco Tagliaferri",
    "Lukas Volf"
  ],
  "filing_date": "2016-06-27",
  "publication_date": "2018-12-11",
  "grant_date": "2018-12-11",
  "priority_date": "2015-06-26",
  "application_number": "US-201615194500-A",
  "family_id": "64535889",
  "cited_by_count": 28,
  "citations": [
    "US5317673A",
    "US5581659A",
    "US6151574A",
    "US20040260538A1",
    "US20050021337A1",
    "US20100318412A1",
    "US20120123904A1",
    "US20140257803A1",
    "US9153231B1",
    "US20150039299A1",
    "US20150127327A1",
    "US9324321B2"
  ]
}

Record 3,133 of 8,000 in Patents full text (MLC-0201). Request the full dataset.