MLchartDataset catalogue

Patent · US2023315856A1 · A1 · US

Methods and apparatus for augmenting training data using large language models

(11) Publication number
US2023315856A1
(21) Application number
17/710,127
(22) Filing date
2022-03-31
(30) Priority date
2022-03-31
(43) Publication date
2023-10-05
(51) IPC
G06F 21/57; G06F 40/205; G06N 3/08; G06F 40/30
(52) CPC
  • G06F Electric digital data processing: 21/57, 16/243, 16/24522, 2221/034, 40/205, 40/274, 40/289, 40/30, 40/35
  • G06N Computing arrangements based on specific computational models: 3/045, 3/08
(73) Assignee
Sophos Ltd
(72) Inventors
Younghoo Lee; Miklós Sándor BÉKY; Joshua Daniel Saxe
(54) Title
Methods and apparatus for augmenting training data using large language models
(57) Abstract

In some embodiments, a processor receives natural language data for performing an identified cybersecurity task. The processor can provide the natural language data to a first machine learning (ML) model. The first ML model can automatically infer a template query based on the natural language data. The processor can receive user input indicating a finalized query and to provide the finalized query as input to a system configured to perform the identified computational task. The processor can provide the finalized query as a reference phrase to a second ML model, the second ML model configured to generate a set of natural language phrases similar to the reference phrase. The processor can generate supplemental training data using the set of natural language phrases similar to the reference phrase to augment training data used to improve performance of the first ML model and/or the second ML model.

Full text
View on Google Patents

Claims (1)

  1. A method, comprising: receiving, via an interface, a natural language request for performing an identified task in a management system, the management system being associated with a set of system commands, the set of system commands being configured to perform one or more computational tasks associated with the management system, the identified task being from the one or more computational tasks, and the management system operating within an identified context; extracting, based on the identified context, a set of features from the natural language request; providing the set of features to a first machine learning (ML) model to infer a template command based on the natural language request, the first ML model trained using first training data, the first training data including a set of natural language phrases associated with the context, the first ML model trained to receive features based on the set of natural language phrases as input, the template system command being associated with a set of system commands; receiving, as output from the first ML model, the template command associated with the natural language request; displaying, via the interface, the template command in an editable form to be edited or approved by a user; receiving, via the interface, a final command, based on the template command, the final command approved by the user; providing the final command as a reference input to a second ML model, the second ML model configured to generate a set of natural language phrases semantically related to the reference input; receiving, from the second ML model, the set of natural language phrases semantically related to the reference input; generating a second training data based on the set of natural language phrases semantically related to the reference input; augmenting the first training data by adding the second training data to the first training data to generate third training data; and providing the final command to the management system to implement the identified task. 2. The method of claim 1, wherein the providing the set of features to the first ML model to infer the template command based on the natural language request includes: parsing the natural language request into a set of portions to predict, using the first ML model, the template command based on the natural language request, the template command including the set of portions, each portion from the set of portions being associated with a parameter from a set of parameters. 3. The method of claim 2, wherein the displaying the template command in the editable form to be edited or approved by the user includes: displaying, via the interface, the set of portions of the template command, the interface configured to receive changes, provided by the user, to a portion from the set of portions of the template query to form the final command. 4. The method of claim 1, wherein receiving the natural language includes: receiving an incomplete portion of the natural language request; automatically inferring using a third ML model, based on the incomplete portion of the natural language request, a set of potential options, each option from the set of potential options providing a remaining portion of the natural language request; and displaying, via the interface, the set of potential options to the user as a selectable list, the interface configured to allow the user to select at least one option from the set of potential options for providing a remaining portion of the natural language request, receiving a selection of one option from the set of options to generate a complete natural language request to perform the identified task; generating a fourth training data based on the incomplete portion of the natural language request, and the one option from the set of options; and augmenting the first training by adding the fourth training data to the first training data to generate fifth training data. 5. The method of claim 2, wherein receiving the final command includes receiving changes to each portion from the set of portions of the template command to form the final command, the method further comprising: generating a fourth training data including the changes to each portion from the set of portions of the template command, the fourth training data configured to improve performance of at least one of the first ML model at inferring the template command or the second ML model at generating a set of natural language phrases semantically related to the reference input. 6. (canceled) 7. An apparatus, comprising: a memory including processor-executable instructions; and one or more hardware processors in communication with the memory that, having executed the processor-executable instructions, are configured to: receive first training data, the first training data being based on an identified context associated with cybersecurity; train, using the first training data, a machine learning (ML) model to receive a set of natural language phrases as input and to infer, based on each natural language phrase from the set of natural language phrases, a template system command associated with a set of system commands, the set of system commands being configured to perform one or more computational tasks associated with a cybersecurity management system operating within the identified context; receive, via an interface, a sample natural language phrase associated with a user request for performing an identified computational task, the identified computational task being associated with implementing measures for malware detection and mitigation; provide the sample natural language phrase as input to the ML model such that the ML model infers a template command based on the sample natural language phrase; receive, via the interface, edits to the template command provided by a user; generate second training data based on the sample natural language phrase and the edits to the template command provided by the user; augment the first training data by adding the second training data to the first training data to generate third training data; generate, based on the edits to the template command a finalized command associated with the identified computational task; provide the finalized command to the management system such that the management system performs the identified computational task; and modify a setting in the management system based on the performance of the identified computational task. 8. The apparatus of claim 7, wherein the ML model is configured to generate the template command by parsing the sample natural language phrase into a set of portions, each portion from the set of portions being associated with a parameter from a set of parameters, the one or more processors further configured to: cause the template command, including the set of parameters, to be displayed, via the interface, the template command, including the set of parameters, being editable by the user. 9. The apparatus of claim 8, wherein at least one of the set of portions or the set of parameters is determined based on a syntax associated with the natural language data. 10. The apparatus of claim 8, wherein the one or more processors are configured to cause the display of the set of parameters such that each portion from the set of portions is displayed as one option from a plurality of options associated with a parameter from the set of parameters, the processor configured to: provide control tools, via the interface, the control tools configured to receive, from the user, a selection of at least one option from the plurality of options associated with each parameter from the set of parameters, the selection forming the edits to the template command. 11. (canceled) 12. The apparatus of claim 7, wherein the ML model is configured to receive a partially complete portion of the sample natural language phrase associated with a user request for performing the identified computational task and automatically infer, based on the partially complete portion, at least a part of a remaining portion of the sample natural language phrase. 13. The apparatus of claim 7, wherein the one or more processors are further configured to: provide at least one of the sample natural language phrase, the template command, the edits to the template commands, or the finalized command, to a second ML model as a reference input, the second ML model being trained on natural language data associated with the identified context; invoke the second ML model to generate a set of natural language phrases semantically related to the reference input; generate a fourth training data based on the set of natural language phrases semantically related to the reference input; and augment the at least one of the first training data or the third training data by adding the fourth training data to the at least one of the first training data or the third training data to generate fifth training data. 14. The apparatus of claim 7, wherein the identified computational task includes at least one of blocking a communication or a host, applying a patch to a set of hosts, rebooting a machine, or executing a rule at an identified endpoint. 15. The apparatus of claim 7, wherein the ML model is configured to generate the template command by parsing the sample natural language phrase associated with the user request into a set of portions, each portion from the set of portions being associated with a parameter from a set of parameters, the set of parameters including at least one of an action parameter, an object parameter, or a descriptor parameter associated with the object parameter. 16. A non-transitory processor-readable medium storing code representing instructions to be executed by one or more processors, the instructions comprising code to cause the one or more processors to: receive, via an interface, a first portion of a natural language request, the first portion of the natural language request indicating an identified context associated with a task to be performed via a management system; provide, the first portion of the natural language request to a first machine learning (ML) model, the first ML model configured to predict from the portion of the natural language request and the identified context, a set of options, each option from the set of options indicating a second portion of the natural language request to generate a complete natural language request, the first ML model having been trained using first training data associated with the identified context; display, via the interface, the set of options as a list from which a user may select an option from the set of options; receive, via the interface, a selection of the option from the set of options to generate a complete natural language request to perform an identified task; generate second training data based on the first portion of the natural language request and the option from the set of options; provide the complete natural language request as an input to a second ML model; generate, using the second ML model, a template command based on the complete natural language request; augment the first training by adding the second training data to the first training data to generate third training data, the third training data configured to improve performance of at least one of the first ML model at predicting the option from the set of options to generate the complete natural language request, or the second ML model at generating the template command based on the complete natural language request; provide the template command to the management system to perform the task via the management system; and receive an indication confirming a completion of the task via the management system. 17. The non-transitory processor-readable medium of claim 16, wherein the second ML model is the same as the first ML model. 18. The non-transitory processor-readable medium of claim 16, wherein the second ML model is different than the first ML model. 19. The non-transitory processor-readable medium of claim 16, further comprising code to cause the one or more processors to: receive, from the second ML model, the template command including a set of portions forming the template command, each portion from the set of portions being associated with a parameter from a set of parameters, the second ML model configured to predict the set of parameters based on the identified context associated with the complete natural language request; and display, via the interface, the template query including the set of portions forming the template query and the set of parameters. 20. The non-transitory processor-readable medium of claim 16, wherein at least one of the first ML model or the second ML model is a Generative Pre-trained Transformer 3 (GPT-3) model. 21. The non-transitory processor-readable medium of claim 16, wherein the second ML model is trained on parsing natural language data obtained from a plurality of language corpuses, the code to cause the one or more processors to generate the template query further comprising code to cause the one or more processors to: train the second ML model, using natural language data associated with potential cybersecurity measures, to receive a natural language phrase associated with a task to implement the cybersecurity measures, the code to cause the processor to provide the complete natural language request to the second ML model to generate the template command includes code to cause the processor to provide the complete natural language request to the second ML model to generate, based on the natural language phrase, the template command based on a set of parameters including an action parameter associated with the cybersecurity measures. 22. The non-transitory processor-readable medium of claim 21, wherein the action parameter associated with the cybersecurity measures includes at least one of blocking a communication or a host, allowing a communication, filtering a set of data, sorting a set of data, displaying a set of data, applying a patch, rebooting an endpoint, or executing a rule.

Description

The embodiments described herein relate to methods and apparatus for natural language-based querying or manipulating cybersecurity management systems that are used to monitor hardware, software and/or communications for virus and/or malware detection to ensure data integrity and/or to prevent or detect potential attacks.

The embodiments described herein relate to methods and apparatus for receiving natural language tasks and converting them to complex commands and/or queries for manipulating cybersecurity management systems using Machine Learning (ML) models. The embodiments described herein relate to methods and apparatus for training the ML models using training data, and generating and/or augmenting the training data to improve performance of the ML models at inferring complex commands from natural language phrases.

Some known malicious artifacts can be embedded and distributed in several forms (e.g., text files, audio files, video files, data files, executable files, uniform resource locators (URLs) providing the address of a resource on the Internet, etc.) that are seemingly harmless in appearance but hard to detect and can be prone to cause severe damage or compromise of sensitive hardware, data, information, and/or the like. Management systems, for example, cybersecurity systems, can be configured to manage network or digital security of data and/or computational resources (e.g., servers, endpoints, etc.), often associated with organizations.

Citations (5)

  • US20130179460A1
  • US20190034410A1
  • US20190205322A1
  • US11183175B2
  • US20230106226A1
Record as JSON
{
  "publication_number": "US2023315856A1",
  "country": "US",
  "kind": "A1",
  "title": "Methods and apparatus for augmenting training data using large language models",
  "abstract": "In some embodiments, a processor receives natural language data for performing an identified cybersecurity task. The processor can provide the natural language data to a first machine learning (ML) model. The first ML model can automatically infer a template query based on the natural language data. The processor can receive user input indicating a finalized query and to provide the finalized query as input to a system configured to perform the identified computational task. The processor can provide the finalized query as a reference phrase to a second ML model, the second ML model configured to generate a set of natural language phrases similar to the reference phrase. The processor can generate supplemental training data using the set of natural language phrases similar to the reference phrase to augment training data used to improve performance of the first ML model and/or the second ML model.",
  "claims": [
    "1. A method, comprising: receiving, via an interface, a natural language request for performing an identified task in a management system, the management system being associated with a set of system commands, the set of system commands being configured to perform one or more computational tasks associated with the management system, the identified task being from the one or more computational tasks, and the management system operating within an identified context; extracting, based on the identified context, a set of features from the natural language request; providing the set of features to a first machine learning (ML) model to infer a template command based on the natural language request, the first ML model trained using first training data, the first training data including a set of natural language phrases associated with the context, the first ML model trained to receive features based on the set of natural language phrases as input, the template system command being associated with a set of system commands; receiving, as output from the first ML model, the template command associated with the natural language request; displaying, via the interface, the template command in an editable form to be edited or approved by a user; receiving, via the interface, a final command, based on the template command, the final command approved by the user; providing the final command as a reference input to a second ML model, the second ML model configured to generate a set of natural language phrases semantically related to the reference input; receiving, from the second ML model, the set of natural language phrases semantically related to the reference input; generating a second training data based on the set of natural language phrases semantically related to the reference input; augmenting the first training data by adding the second training data to the first training data to generate third training data; and providing the final command to the management system to implement the identified task. 2. The method of claim 1, wherein the providing the set of features to the first ML model to infer the template command based on the natural language request includes: parsing the natural language request into a set of portions to predict, using the first ML model, the template command based on the natural language request, the template command including the set of portions, each portion from the set of portions being associated with a parameter from a set of parameters. 3. The method of claim 2, wherein the displaying the template command in the editable form to be edited or approved by the user includes: displaying, via the interface, the set of portions of the template command, the interface configured to receive changes, provided by the user, to a portion from the set of portions of the template query to form the final command. 4. The method of claim 1, wherein receiving the natural language includes: receiving an incomplete portion of the natural language request; automatically inferring using a third ML model, based on the incomplete portion of the natural language request, a set of potential options, each option from the set of potential options providing a remaining portion of the natural language request; and displaying, via the interface, the set of potential options to the user as a selectable list, the interface configured to allow the user to select at least one option from the set of potential options for providing a remaining portion of the natural language request, receiving a selection of one option from the set of options to generate a complete natural language request to perform the identified task; generating a fourth training data based on the incomplete portion of the natural language request, and the one option from the set of options; and augmenting the first training by adding the fourth training data to the first training data to generate fifth training data. 5. The method of claim 2, wherein receiving the final command includes receiving changes to each portion from the set of portions of the template command to form the final command, the method further comprising: generating a fourth training data including the changes to each portion from the set of portions of the template command, the fourth training data configured to improve performance of at least one of the first ML model at inferring the template command or the second ML model at generating a set of natural language phrases semantically related to the reference input. 6. (canceled) 7. An apparatus, comprising: a memory including processor-executable instructions; and one or more hardware processors in communication with the memory that, having executed the processor-executable instructions, are configured to: receive first training data, the first training data being based on an identified context associated with cybersecurity; train, using the first training data, a machine learning (ML) model to receive a set of natural language phrases as input and to infer, based on each natural language phrase from the set of natural language phrases, a template system command associated with a set of system commands, the set of system commands being configured to perform one or more computational tasks associated with a cybersecurity management system operating within the identified context; receive, via an interface, a sample natural language phrase associated with a user request for performing an identified computational task, the identified computational task being associated with implementing measures for malware detection and mitigation; provide the sample natural language phrase as input to the ML model such that the ML model infers a template command based on the sample natural language phrase; receive, via the interface, edits to the template command provided by a user; generate second training data based on the sample natural language phrase and the edits to the template command provided by the user; augment the first training data by adding the second training data to the first training data to generate third training data; generate, based on the edits to the template command a finalized command associated with the identified computational task; provide the finalized command to the management system such that the management system performs the identified computational task; and modify a setting in the management system based on the performance of the identified computational task. 8. The apparatus of claim 7, wherein the ML model is configured to generate the template command by parsing the sample natural language phrase into a set of portions, each portion from the set of portions being associated with a parameter from a set of parameters, the one or more processors further configured to: cause the template command, including the set of parameters, to be displayed, via the interface, the template command, including the set of parameters, being editable by the user. 9. The apparatus of claim 8, wherein at least one of the set of portions or the set of parameters is determined based on a syntax associated with the natural language data. 10. The apparatus of claim 8, wherein the one or more processors are configured to cause the display of the set of parameters such that each portion from the set of portions is displayed as one option from a plurality of options associated with a parameter from the set of parameters, the processor configured to: provide control tools, via the interface, the control tools configured to receive, from the user, a selection of at least one option from the plurality of options associated with each parameter from the set of parameters, the selection forming the edits to the template command. 11. (canceled) 12. The apparatus of claim 7, wherein the ML model is configured to receive a partially complete portion of the sample natural language phrase associated with a user request for performing the identified computational task and automatically infer, based on the partially complete portion, at least a part of a remaining portion of the sample natural language phrase. 13. The apparatus of claim 7, wherein the one or more processors are further configured to: provide at least one of the sample natural language phrase, the template command, the edits to the template commands, or the finalized command, to a second ML model as a reference input, the second ML model being trained on natural language data associated with the identified context; invoke the second ML model to generate a set of natural language phrases semantically related to the reference input; generate a fourth training data based on the set of natural language phrases semantically related to the reference input; and augment the at least one of the first training data or the third training data by adding the fourth training data to the at least one of the first training data or the third training data to generate fifth training data. 14. The apparatus of claim 7, wherein the identified computational task includes at least one of blocking a communication or a host, applying a patch to a set of hosts, rebooting a machine, or executing a rule at an identified endpoint. 15. The apparatus of claim 7, wherein the ML model is configured to generate the template command by parsing the sample natural language phrase associated with the user request into a set of portions, each portion from the set of portions being associated with a parameter from a set of parameters, the set of parameters including at least one of an action parameter, an object parameter, or a descriptor parameter associated with the object parameter. 16. A non-transitory processor-readable medium storing code representing instructions to be executed by one or more processors, the instructions comprising code to cause the one or more processors to: receive, via an interface, a first portion of a natural language request, the first portion of the natural language request indicating an identified context associated with a task to be performed via a management system; provide, the first portion of the natural language request to a first machine learning (ML) model, the first ML model configured to predict from the portion of the natural language request and the identified context, a set of options, each option from the set of options indicating a second portion of the natural language request to generate a complete natural language request, the first ML model having been trained using first training data associated with the identified context; display, via the interface, the set of options as a list from which a user may select an option from the set of options; receive, via the interface, a selection of the option from the set of options to generate a complete natural language request to perform an identified task; generate second training data based on the first portion of the natural language request and the option from the set of options; provide the complete natural language request as an input to a second ML model; generate, using the second ML model, a template command based on the complete natural language request; augment the first training by adding the second training data to the first training data to generate third training data, the third training data configured to improve performance of at least one of the first ML model at predicting the option from the set of options to generate the complete natural language request, or the second ML model at generating the template command based on the complete natural language request; provide the template command to the management system to perform the task via the management system; and receive an indication confirming a completion of the task via the management system. 17. The non-transitory processor-readable medium of claim 16, wherein the second ML model is the same as the first ML model. 18. The non-transitory processor-readable medium of claim 16, wherein the second ML model is different than the first ML model. 19. The non-transitory processor-readable medium of claim 16, further comprising code to cause the one or more processors to: receive, from the second ML model, the template command including a set of portions forming the template command, each portion from the set of portions being associated with a parameter from a set of parameters, the second ML model configured to predict the set of parameters based on the identified context associated with the complete natural language request; and display, via the interface, the template query including the set of portions forming the template query and the set of parameters. 20. The non-transitory processor-readable medium of claim 16, wherein at least one of the first ML model or the second ML model is a Generative Pre-trained Transformer 3 (GPT-3) model. 21. The non-transitory processor-readable medium of claim 16, wherein the second ML model is trained on parsing natural language data obtained from a plurality of language corpuses, the code to cause the one or more processors to generate the template query further comprising code to cause the one or more processors to: train the second ML model, using natural language data associated with potential cybersecurity measures, to receive a natural language phrase associated with a task to implement the cybersecurity measures, the code to cause the processor to provide the complete natural language request to the second ML model to generate the template command includes code to cause the processor to provide the complete natural language request to the second ML model to generate, based on the natural language phrase, the template command based on a set of parameters including an action parameter associated with the cybersecurity measures. 22. The non-transitory processor-readable medium of claim 21, wherein the action parameter associated with the cybersecurity measures includes at least one of blocking a communication or a host, allowing a communication, filtering a set of data, sorting a set of data, displaying a set of data, applying a patch, rebooting an endpoint, or executing a rule."
  ],
  "description_excerpt": "The embodiments described herein relate to methods and apparatus for natural language-based querying or manipulating cybersecurity management systems that are used to monitor hardware, software and/or communications for virus and/or malware detection to ensure data integrity and/or to prevent or detect potential attacks.\n\nThe embodiments described herein relate to methods and apparatus for receiving natural language tasks and converting them to complex commands and/or queries for manipulating cybersecurity management systems using Machine Learning (ML) models. The embodiments described herein relate to methods and apparatus for training the ML models using training data, and generating and/or augmenting the training data to improve performance of the ML models at inferring complex commands from natural language phrases.\n\nSome known malicious artifacts can be embedded and distributed in several forms (e.g., text files, audio files, video files, data files, executable files, uniform resource locators (URLs) providing the address of a resource on the Internet, etc.) that are seemingly harmless in appearance but hard to detect and can be prone to cause severe damage or compromise of sensitive hardware, data, information, and/or the like. Management systems, for example, cybersecurity systems, can be configured to manage network or digital security of data and/or computational resources (e.g., servers, endpoints, etc.), often associated with organizations.",
  "cpc": [
    "G06F 21/57",
    "G06F 16/243",
    "G06F 16/24522",
    "G06F 2221/034",
    "G06F 40/205",
    "G06F 40/274",
    "G06F 40/289",
    "G06F 40/30",
    "G06F 40/35",
    "G06N 3/045",
    "G06N 3/08"
  ],
  "ipc": [
    "G06F 21/57",
    "G06F 40/205",
    "G06N 3/08",
    "G06F 40/30"
  ],
  "assignees": [
    "Sophos Ltd"
  ],
  "inventors": [
    "Younghoo Lee",
    "Miklós Sándor BÉKY",
    "Joshua Daniel Saxe"
  ],
  "filing_date": "2022-03-31",
  "publication_date": "2023-10-05",
  "priority_date": "2022-03-31",
  "application_number": "US-202217710127-A",
  "family_id": "86052694",
  "cited_by_count": 50,
  "citations": [
    "US20130179460A1",
    "US20190034410A1",
    "US20190205322A1",
    "US11183175B2",
    "US20230106226A1"
  ]
}

Record 657 of 8,000 in Patents full text (MLC-0201). Request the full dataset.