MLchartDataset catalogue

Patent · US2026106886A1 · A1 · US

Cve labeling for exploits using proof-of-concept and llm

(11) Publication number
US2026106886A1
(21) Application number
19/422,162
(22) Filing date
2025-12-16
(43) Publication date
2026-04-16
(52) CPC
  • H04L Transmission of digital information, e.g. telegraphic communication: 63/1433
(54) Title
Cve labeling for exploits using proof-of-concept and llm
(57) Abstract

The disclosed system determines a CVE identifier based on match confidence against entries in a database of indicators of vulnerability exploits that are mapped to corresponding CVE identifiers. The system builds and maintains the database by generating these exploit indicators from various cybersecurity data having associated vulnerability identifiers. The system extracts elements from the cybersecurity data to construct different exploit indicators and then stores them in the database mapped to corresponding CVE identifiers. Depending upon the cybersecurity data from which elements are extracted, different types of indicators of an exploit may be generated for a same vulnerability.

Full text
View on Google Patents

Claims (1)

  1. A method comprising; building a database of exploit representations mapped to vulnerability identifiers, wherein building the database of exploit representations comprises, generating uniform resource identifier (URI)-based exploit representations from URIs in published vulnerability descriptions that have vulnerability identifiers and from URIs in exploit proof of concepts (PoCs) that have vulnerability identifiers; generating header-based exploit representations from header field names in the exploit PoCs; and generating body-based exploit representations from keywords detected in malicious packet payload samples; and for a malicious packet that is not associated with a vulnerability identifier, determining a vulnerability identifier in the database for labeling the malicious packet. 2. The method of claim 1, wherein generating URI-based exploit representations from URIs in published vulnerability descriptions that have vulnerability identifiers comprises: prompting a language model to extract from a URI in a published vulnerability description a hostname component, a path component, and one or more parameters in a query component of the URI; aggregating the extracted components or an extracted component with an extracted query parameter to form the URI-based exploit representation; and mapping the vulnerability identifier of the published vulnerability description to the aggregation. 3. The method of claim 1, wherein generating URI-based exploit representations from URIs in exploit PoCs that have vulnerability identifiers comprises parsing a URI in an exploit PoC to extract a hostname component, a path component, and one or more parameters in a query component of the URI and aggregating the extracted components or an extracted component and an extracted query parameter to form the URI-based exploit representation and to map the vulnerability identifier of the exploit PoC to the aggregation. 4. The method of claim 1, wherein generating header-based exploit representations from header field names in the exploit PoCs comprises parsing a header of an exploit PoC to extract a plurality of header field names and to aggregate the plurality of extracted header field names to form a header-based exploit representation and mapping the vulnerability identifier of the exploit PoC to the header-based exploit representation. 5. The method of claim 1, wherein querying the database with the first and second malicious packet representations comprises querying the database for exploit representations that at least partially match either of the first and the second malicious packet representations. 6. The method of claim 5, further comprising scoring similarity of each of the exploit representations returned based on the query, wherein determining the most similar of the exploit representations comprises determining similarity based on the similarity scores. 7. The method of claim 6, wherein the exploit representations returned based on the query at least comprise a first URI-based exploit representation and a first header-based exploit representation, wherein scoring similarity comprises weighting the first URI-based exploit representation more than the first header-based exploit representation. 8. The method of claim 1, wherein determining a vulnerability identifier in the database for labeling the malicious packet comprises: generating a first representation of the malicious packet from a URI if in the malicious packet, a second representation with header field names in the malicious packet, and a third representation with keywords from a payload of the malicious packet; querying the database with the generated representations of the malicious packet; determining a most similar of exploit representations returned based on the query; and indicating the vulnerability identifier mapped to the most similar exploit representation for labeling the malicious packet. 9. A non-transitory, machine-readable medium having program code stored thereon, the program code comprising instructions to: build a database of exploit representations from published vulnerability descriptions that have vulnerability identifiers and from exploit proof of concepts (PoCs) that have vulnerability identifiers, wherein the instructions to build the database of exploit representations comprise instructions to, for each of the published vulnerability descriptions and exploit PoCs, generate a uniform resource identifier (URI)-based exploit representation from a set of one or more URI components in the published vulnerability description or the exploit PoC and assign the corresponding vulnerability identifier to the URI-based exploit representation; for each of the exploit PoCs, generate a header-based exploit representation from a set of header field names in the exploit PoC and assign the corresponding vulnerability identifier to the header-based exploit representation; generate a body-based exploit representation from keywords in a body or payload of the exploit PoC and assign the corresponding vulnerability identifier to the body-based exploit representation; and in response to receipt of a malicious packet, determine a vulnerability identifier indicated in the database for labeling the malicious packet. 10. The non-transitory, machine-readable medium of claim 9, wherein the instructions to search the database for a most similar of the exploit representations comprise instructions to search the database for exploit representations that at least partially match one of the malicious packet representations and then determine the most similar of a plurality of exploit representations from the database that at least partially match based on extent of matching. 11. The non-transitory, machine-readable medium of claim 10, wherein the program code further comprises instructions to score similarity of each of the plurality of exploit representations with respect to the corresponding one of the malicious packet representations, wherein the instructions to determine the most similar of the plurality of exploit representations is based on the similarity scores. 12. The non-transitory, machine-readable medium of claim 11, wherein the plurality of exploit representations comprises at least two of a first URI-based exploit representation, a first body-based exploit representation, and a first header-based exploit representation, wherein the instructions to score similarity comprise instructions to weight the first URI-based exploit representation and/or the first body-based exploit representation more than the first header-based exploit representation. 13. The non-transitory, machine-readable medium of claim 9, wherein the instructions to generate a URI-based exploit representation from a URI in a published vulnerability description comprise instructions to prompt a language model to extract one or more components of the URI from the published vulnerability description and form the URI-based exploit representation from the extracted one or more URI components. 14. The non-transitory, machine-readable medium of claim 9, wherein the instructions to generate a URI-based exploit representation from a URI in a published vulnerability description comprise instructions to prompt a language model to extract a hostname component and a path component from a URI in the published vulnerability description. 15. The non-transitory, machine-readable medium of claim 14, wherein the instructions to prompt the language model comprise instructions to prompt the language model to extract one or more parameters from a query component of a URI in the published vulnerability description, wherein the URI-based exploit representation is based on an extracted query parameter as well as the extracted hostname and path components. 16. The non-transitory, machine-readable medium of claim 9, wherein the instructions to generate a header-based exploit representation of an exploit PoC comprise instructions to extract a set of header field names from the exploit PoC and aggregate the set of header field names. 17. The non-transitory, machine-readable medium of claim 9, wherein the instructions to determine a vulnerability identifier indicated in the database for labelling the malicious packet comprise instructions to: generate a first representation of the malicious packet with URI components in the malicious packet, a second representation with header field names in the malicious packet, and a third representation with a set of keywords detected in a payload of the malicious packet; search the database for a most similar of the exploit representations with respect to the malicious packet representations; and indicate the vulnerability identifier associated with the most similar exploit representation for labeling the malicious packet. 18. A method comprising: building a database of exploit representations mapped to vulnerability identifiers, wherein building the database of exploit representations comprises, generating exploit markers from uniform resource identifiers (URIs), header field names, and keywords found in at least one of published vulnerability descriptions that have vulnerability identifiers and exploit proof of concepts (PoCs) that have vulnerability identifiers; and creating mappings between the exploit markers and the vulnerability identifiers; and based on indication of a malicious packet that is not associated with a vulnerability identifier, determining a vulnerability identifier in the database for labeling the malicious packet. 19. The method of claim 18, wherein generating exploit markers comprises: generating URI-based exploit markers from URIs in published vulnerability descriptions and from URIs in exploit PoCs; generating header-based exploit markers from header field names in the exploit PoCs; and generating body-based exploit markers from keywords detected in the exploit PoCs and/or malicious packet payload samples. 20. The method of claim 19, wherein generating URI-based exploit markers from URIs in published vulnerability descriptions comprises prompting a language model to extract from a URI in a published vulnerability description a hostname component, a path component, and one or more parameters in a query component of the URI and aggregating the extracted components or an extracted component with an extracted query parameter to form the URI-based exploit marker, wherein generating URI-based exploit markers from URIs in exploit PoCs comprises parsing a URI in an exploit PoC to extract a hostname component, a path component, and one or more parameters in a query component of the URI and aggregating the extracted components or an extracted component with an extracted query parameter to form the URI-based exploit marker, wherein generating header-based exploit markers from header field names in the exploit PoCs comprises parsing a header of an exploit PoC to extract a plurality of header field names and to aggregate the plurality of extracted field names to form a header-based exploit representation and map the vulnerability identifier of the exploit PoC to the header-based exploit representation. 21. The method of claim 18, wherein determining a vulnerability identifier in the database for labeling the malicious packet comprises: generating one or more representations of the malicious packet depending on which of a URI, request header field names, and a request body are included within the malicious packet; querying the database with each representation of the malicious packet; scoring match confidence for each exploit marker returned based on the query; and indicating the vulnerability identifier mapped to the exploit marker with the highest match confidence for labeling the malicious packet.
Record as JSON
{
  "publication_number": "US2026106886A1",
  "country": "US",
  "kind": "A1",
  "title": "Cve labeling for exploits using proof-of-concept and llm",
  "abstract": "The disclosed system determines a CVE identifier based on match confidence against entries in a database of indicators of vulnerability exploits that are mapped to corresponding CVE identifiers. The system builds and maintains the database by generating these exploit indicators from various cybersecurity data having associated vulnerability identifiers. The system extracts elements from the cybersecurity data to construct different exploit indicators and then stores them in the database mapped to corresponding CVE identifiers. Depending upon the cybersecurity data from which elements are extracted, different types of indicators of an exploit may be generated for a same vulnerability.",
  "claims": [
    "1. A method comprising; building a database of exploit representations mapped to vulnerability identifiers, wherein building the database of exploit representations comprises, generating uniform resource identifier (URI)-based exploit representations from URIs in published vulnerability descriptions that have vulnerability identifiers and from URIs in exploit proof of concepts (PoCs) that have vulnerability identifiers; generating header-based exploit representations from header field names in the exploit PoCs; and generating body-based exploit representations from keywords detected in malicious packet payload samples; and for a malicious packet that is not associated with a vulnerability identifier, determining a vulnerability identifier in the database for labeling the malicious packet. 2. The method of claim 1, wherein generating URI-based exploit representations from URIs in published vulnerability descriptions that have vulnerability identifiers comprises: prompting a language model to extract from a URI in a published vulnerability description a hostname component, a path component, and one or more parameters in a query component of the URI; aggregating the extracted components or an extracted component with an extracted query parameter to form the URI-based exploit representation; and mapping the vulnerability identifier of the published vulnerability description to the aggregation. 3. The method of claim 1, wherein generating URI-based exploit representations from URIs in exploit PoCs that have vulnerability identifiers comprises parsing a URI in an exploit PoC to extract a hostname component, a path component, and one or more parameters in a query component of the URI and aggregating the extracted components or an extracted component and an extracted query parameter to form the URI-based exploit representation and to map the vulnerability identifier of the exploit PoC to the aggregation. 4. The method of claim 1, wherein generating header-based exploit representations from header field names in the exploit PoCs comprises parsing a header of an exploit PoC to extract a plurality of header field names and to aggregate the plurality of extracted header field names to form a header-based exploit representation and mapping the vulnerability identifier of the exploit PoC to the header-based exploit representation. 5. The method of claim 1, wherein querying the database with the first and second malicious packet representations comprises querying the database for exploit representations that at least partially match either of the first and the second malicious packet representations. 6. The method of claim 5, further comprising scoring similarity of each of the exploit representations returned based on the query, wherein determining the most similar of the exploit representations comprises determining similarity based on the similarity scores. 7. The method of claim 6, wherein the exploit representations returned based on the query at least comprise a first URI-based exploit representation and a first header-based exploit representation, wherein scoring similarity comprises weighting the first URI-based exploit representation more than the first header-based exploit representation. 8. The method of claim 1, wherein determining a vulnerability identifier in the database for labeling the malicious packet comprises: generating a first representation of the malicious packet from a URI if in the malicious packet, a second representation with header field names in the malicious packet, and a third representation with keywords from a payload of the malicious packet; querying the database with the generated representations of the malicious packet; determining a most similar of exploit representations returned based on the query; and indicating the vulnerability identifier mapped to the most similar exploit representation for labeling the malicious packet. 9. A non-transitory, machine-readable medium having program code stored thereon, the program code comprising instructions to: build a database of exploit representations from published vulnerability descriptions that have vulnerability identifiers and from exploit proof of concepts (PoCs) that have vulnerability identifiers, wherein the instructions to build the database of exploit representations comprise instructions to, for each of the published vulnerability descriptions and exploit PoCs, generate a uniform resource identifier (URI)-based exploit representation from a set of one or more URI components in the published vulnerability description or the exploit PoC and assign the corresponding vulnerability identifier to the URI-based exploit representation; for each of the exploit PoCs, generate a header-based exploit representation from a set of header field names in the exploit PoC and assign the corresponding vulnerability identifier to the header-based exploit representation; generate a body-based exploit representation from keywords in a body or payload of the exploit PoC and assign the corresponding vulnerability identifier to the body-based exploit representation; and in response to receipt of a malicious packet, determine a vulnerability identifier indicated in the database for labeling the malicious packet. 10. The non-transitory, machine-readable medium of claim 9, wherein the instructions to search the database for a most similar of the exploit representations comprise instructions to search the database for exploit representations that at least partially match one of the malicious packet representations and then determine the most similar of a plurality of exploit representations from the database that at least partially match based on extent of matching. 11. The non-transitory, machine-readable medium of claim 10, wherein the program code further comprises instructions to score similarity of each of the plurality of exploit representations with respect to the corresponding one of the malicious packet representations, wherein the instructions to determine the most similar of the plurality of exploit representations is based on the similarity scores. 12. The non-transitory, machine-readable medium of claim 11, wherein the plurality of exploit representations comprises at least two of a first URI-based exploit representation, a first body-based exploit representation, and a first header-based exploit representation, wherein the instructions to score similarity comprise instructions to weight the first URI-based exploit representation and/or the first body-based exploit representation more than the first header-based exploit representation. 13. The non-transitory, machine-readable medium of claim 9, wherein the instructions to generate a URI-based exploit representation from a URI in a published vulnerability description comprise instructions to prompt a language model to extract one or more components of the URI from the published vulnerability description and form the URI-based exploit representation from the extracted one or more URI components. 14. The non-transitory, machine-readable medium of claim 9, wherein the instructions to generate a URI-based exploit representation from a URI in a published vulnerability description comprise instructions to prompt a language model to extract a hostname component and a path component from a URI in the published vulnerability description. 15. The non-transitory, machine-readable medium of claim 14, wherein the instructions to prompt the language model comprise instructions to prompt the language model to extract one or more parameters from a query component of a URI in the published vulnerability description, wherein the URI-based exploit representation is based on an extracted query parameter as well as the extracted hostname and path components. 16. The non-transitory, machine-readable medium of claim 9, wherein the instructions to generate a header-based exploit representation of an exploit PoC comprise instructions to extract a set of header field names from the exploit PoC and aggregate the set of header field names. 17. The non-transitory, machine-readable medium of claim 9, wherein the instructions to determine a vulnerability identifier indicated in the database for labelling the malicious packet comprise instructions to: generate a first representation of the malicious packet with URI components in the malicious packet, a second representation with header field names in the malicious packet, and a third representation with a set of keywords detected in a payload of the malicious packet; search the database for a most similar of the exploit representations with respect to the malicious packet representations; and indicate the vulnerability identifier associated with the most similar exploit representation for labeling the malicious packet. 18. A method comprising: building a database of exploit representations mapped to vulnerability identifiers, wherein building the database of exploit representations comprises, generating exploit markers from uniform resource identifiers (URIs), header field names, and keywords found in at least one of published vulnerability descriptions that have vulnerability identifiers and exploit proof of concepts (PoCs) that have vulnerability identifiers; and creating mappings between the exploit markers and the vulnerability identifiers; and based on indication of a malicious packet that is not associated with a vulnerability identifier, determining a vulnerability identifier in the database for labeling the malicious packet. 19. The method of claim 18, wherein generating exploit markers comprises: generating URI-based exploit markers from URIs in published vulnerability descriptions and from URIs in exploit PoCs; generating header-based exploit markers from header field names in the exploit PoCs; and generating body-based exploit markers from keywords detected in the exploit PoCs and/or malicious packet payload samples. 20. The method of claim 19, wherein generating URI-based exploit markers from URIs in published vulnerability descriptions comprises prompting a language model to extract from a URI in a published vulnerability description a hostname component, a path component, and one or more parameters in a query component of the URI and aggregating the extracted components or an extracted component with an extracted query parameter to form the URI-based exploit marker, wherein generating URI-based exploit markers from URIs in exploit PoCs comprises parsing a URI in an exploit PoC to extract a hostname component, a path component, and one or more parameters in a query component of the URI and aggregating the extracted components or an extracted component with an extracted query parameter to form the URI-based exploit marker, wherein generating header-based exploit markers from header field names in the exploit PoCs comprises parsing a header of an exploit PoC to extract a plurality of header field names and to aggregate the plurality of extracted field names to form a header-based exploit representation and map the vulnerability identifier of the exploit PoC to the header-based exploit representation. 21. The method of claim 18, wherein determining a vulnerability identifier in the database for labeling the malicious packet comprises: generating one or more representations of the malicious packet depending on which of a URI, request header field names, and a request body are included within the malicious packet; querying the database with each representation of the malicious packet; scoring match confidence for each exploit marker returned based on the query; and indicating the vulnerability identifier mapped to the exploit marker with the highest match confidence for labeling the malicious packet."
  ],
  "cpc": [
    "H04L 63/1433"
  ],
  "filing_date": "2025-12-16",
  "publication_date": "2026-04-16",
  "application_number": "US-202519422162-A"
}

Record 13 of 5,000 in Patents full text (MLC-0201). Request the full dataset.