MLchartDataset catalogue

Patent · US11625422B2 · B2 · US

Context based surface form generation for cognitive system dictionaries

(11) Publication number
US11625422B2
(21) Application number
16/700,218
(22) Filing date
2019-12-02
(30) Priority date
2019-12-02
(43) Publication date
2023-04-11
(45) Date of grant
2023-04-11
(51) IPC
G06F 16/31; G06F 16/33; G06F 16/332; G06F 16/335; G06F 40/169; G16H 15/00
(52) CPC
  • G06F Electric digital data processing: 16/3329, 16/328, 16/3344, 16/335, 40/169, 40/216, 40/237, 40/284, 40/295, 40/30
  • G16H Healthcare informatics, i.e. information and communication technology [ICT] specially adapted for the handling or processing of medical or healthcare data: 15/00, 20/70, 50/20, 70/60
(73) Assignee
Merative US LP
(72) Inventors
Robert C. Sizemore; Jennifer L. La Rocca; Sterling R. Smith; Mario J. Lorenzo; Kristin E. McNeil; David B. Werts
(54) Title
Context based surface form generation for cognitive system dictionaries
(57) Abstract

A mechanism is provided in a data processing system to implement an annotator for annotating content using context-based surface forms. The mechanism receives a dictionary data structure of surface forms comprising a plurality of regular expressions and input content. The mechanism compares a given span of text in the input content to each regular expression in the dictionary data structure. Responsive to the given span of text matching a given regular expression, an annotator annotates the span of text with a content indicator corresponding to a content category associated with the dictionary data structure. The mechanism performs a natural language processing operation on the input content based on results of the annotation.

Full text
View on Google Patents

Claims (20)

  1. A method, in a data processing system comprising at least one processor and at least one memory, wherein the at least one memory comprises instructions that are executed by the at least one processor to configure the at least one processor to implement an annotator for annotating content using context-based surface forms, the method comprising: generating a plurality of predefined content indicator data structures, wherein each predefined content indicator data structure comprises a dictionary of context-based surface forms of content indicators corresponding to a corresponding medical condition in a plurality of different medical conditions, wherein each of the context-based surface forms comprise one or more regular expressions and corresponding context that specifies a criterion which, when satisfied, causes an annotator to annotate a corresponding portion of text in which the one or more regular expressions are present; generating a user specific content indicator dictionary (USCID) data structure for a patient at least by selecting a subset of the predefined content indicator data structures that correspond to one or more medical conditions, in the plurality of different medical conditions, that are associated with the patient; receiving input content; comparing a given span of text in the input content to each regular expression in each context-based surface form of the subset of predefined content indicator data structures in the user specific content indicator dictionary data structure for the patient; responsive to the given span of text matching at least one regular expression and at least one corresponding context, annotating the span of text with a content indicator corresponding to a content category associated with the user specific content indicator dictionary data structure; and suppressing or promoting content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition.
  2. The method of claim 1, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents prescribed targets.
  3. The method of claim 2, wherein the metacharacter syntax comprises wildcards.
  4. The method of claim 1, wherein the annotator comprises a regular expression processor that translates a regular expression in an implemented syntax into an internal representation that can be executed and matched against a string representing the span of text being searched.
  5. A computer program product comprising a computer readable storage medium having a computer readable program stored therein, wherein the computer readable program, when executed on a data processing system, causes the data processing system to implement an annotator for annotating content using context-based surface forms, wherein the computer readable program causes the data processing system to: generate a plurality of predefined content indicator data structures, wherein each predefined content indicator data structure comprises a dictionary of context-based surface forms of content indicators corresponding to a corresponding medical condition in a plurality of different medical conditions, wherein each of the context-based surface forms comprise one or more regular expressions and corresponding context that specifies a criterion which, when satisfied, causes an annotator to annotate a corresponding portion of text in which the one or more regular expressions are present; generate a user specific content indicator dictionary (USCID) data structure for a patient at least by selecting a subset of the predefined content indicator data structures that correspond to one or more medical conditions, in the plurality of different medical conditions, that are associated with the patient; receive input content; compare a given span of text in the input content to each regular expression in each context-based surface form of the subset of predefined content indicator data structures in the user specific content indicator dictionary data structure for the patient; responsive to the given span of text matching at least one regular expression and at least one corresponding context, annotate the span of text with a content indicator corresponding to a content category associated with the user specific content indicator dictionary data structure; and suppress or promote content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition.
  6. The computer program product of claim 5, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents prescribed targets.
  7. The computer program product of claim 6, wherein the metacharacter syntax comprises wildcards.
  8. The computer program product of claim 5, wherein the annotator comprises a regular expression processor that translates a regular expression in an implemented syntax into an internal representation that can be executed and matched against a string representing the span of text being searched.
  9. An apparatus comprising: at least one processor; and at least one memory coupled to the at least one processor, wherein the at least one memory comprises instructions which, when executed by the at least one processor, cause the at least one processor to implement an annotator for annotating content using context-based surface forms, wherein the instructions cause the processor to: generate a plurality of predefined content indicator data structures, wherein each predefined content indicator data structure comprises a dictionary of context-based surface forms of content indicators corresponding to a corresponding medical condition in a plurality of different medical conditions, wherein each of the context-based surface forms comprise one or more regular expressions and corresponding context that specifies a criterion which, when satisfied, causes an annotator to annotate a corresponding portion of text in which the one or more regular expressions are present; generate a user specific content indicator dictionary (USCID) data structure for a patient at least by selecting a subset of the predefined content indicator data structures that correspond to one or more medical conditions, in the plurality of different medical conditions, that are associated with the patient; receive input content; compare a given span of text in the input content to each regular expression in each context-based surface form of the subset of predefined content indicator data structures in the user specific content indicator dictionary data structure for the patient; responsive to the given span of text matching at least one regular expression and at least one corresponding context, annotate the span of text with a content indicator corresponding to a content category associated with the user specific content indicator dictionary data structure; and suppress or promote content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition.
  10. The apparatus of claim 9, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents prescribed targets.
  11. The apparatus of claim 9, wherein the annotator comprises a regular expression processor that translates a regular expression in an implemented syntax into an internal representation that can be executed and matched against a string representing the span of text being searched.
  12. The method of claim 1, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents context-based surface forms.
  13. The method of claim 1, wherein generating the USCID data structure comprises: receiving, by a surface form editor, user input specifying one or more surface forms; generating, by the surface form editor, one or more regular expressions for the one or more surface forms, wherein each regular expression is associated with an annotation type; and storing, by the surface form editor, the one or more regular expressions in a file that is specified during runtime so the surface forms can be processed.
  14. The computer program product of claim 5, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents context-based surface forms.
  15. The computer program product of claim 5, wherein generating the USCID data structure comprises: receiving, by a surface form editor, user input specifying one or more surface forms; generating, by the surface form editor, one or more regular expressions for the one or more surface forms, wherein each regular expression is associated with an annotation type; and storing, by the surface form editor, the one or more regular expressions in a file that is specified during runtime so the surface forms can be processed.
  16. The apparatus of claim 9, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents context-based surface forms.
  17. The apparatus of claim 9, wherein generating the USCID data structure comprises: receiving, by a surface form editor, user input specifying one or more surface forms; generating, by the surface form editor, one or more regular expressions for the one or more surface forms, wherein each regular expression is associated with an annotation type; and storing, by the surface form editor, the one or more regular expressions in a file that is specified during runtime so the surface forms can be processed.
  18. The method of claim 1, wherein suppressing or promoting content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition comprises: in response to the annotated content corresponding to a negative content indicator data structure of the USCID data structure, replacing the annotated content with content corresponding to a positive content indicator data structure of the USCID data structure.
  19. The computer program product of claim 5, wherein suppressing or promoting content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition comprises: in response to the annotated content corresponding to a negative content indicator data structure of the USCID data structure, replacing the annotated content with content corresponding to a positive content indicator data structure of the USCID data structure.
  20. The apparatus of claim 9, wherein suppressing or promoting content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition comprises: in response to the annotated content corresponding to a negative content indicator data structure of the USCID data structure, replacing the annotated content with content corresponding to a positive content indicator data structure of the USCID data structure.

Description

The present application relates generally to an improved data processing apparatus and method and more specifically to mechanisms for context based surface form generation for cognitive system dictionaries.

In modern computing environments such as the Internet, computing tools are present to specifically track and analyze a user's interaction with content so that advertisers can target those individuals with advertisements designed to entice the individual to purchase a product or service, view specific content, or the like. For example, these tools may analyze various factors such as search histories, click-throughs, content viewing, electronic product purchases, electronic shopping cart contents, social networking posts, etc. to determine what types of items/services and/or content a user may be interested in and then present to them corresponding advertisements designed to entice them into performing an action to reward the advertiser, e.g., make a purchase, view content, etc.

In response, other computing tools have been developed to avoid such content, such as pop-up advertisement blockers, adult content blockers, and the like. These computing tools are keyed to broad categories of content, e.g., adult content, or to particular mechanisms for presenting the content, e.g., pop-ups, banner ads, etc.

Citations (45)

  • US5559693A
  • US6310629B1
  • US6236959B1
  • US7076438B1
  • US20070050150A1
  • US9495957B2
  • US8024173B1
  • US20140279746A1
  • US20100131507A1
  • US20110125734A1
  • US8515193B1
  • US20130226843A1
  • US20140058738A1
  • US20190342602A1
  • US10121076B2
  • US20150100308A1
  • US20190141411A1
  • US10121256B2
  • US20160085742A1
  • US9754021B2
  • US20160203267A1
  • US20160371587A1
  • US9740956B2
  • US9616568B1
  • US20170091164A1
  • US20170124037A1
  • WO2017201676A1
  • US20170358256A1
  • US20180039699A1
  • US20180121618A1
  • US20180143970A1
  • US20180260680A1
  • US10268688B2
  • US20180336318A1
  • US20190022863A1
  • US20190080055A1
  • US20190108912A1
  • US20190156953A1
  • US20190156821A1
  • US20190250895A1
  • CN108647591A
  • US10769503B1
  • CN108765450A
  • CN109063723A
  • US11288319B1
Record as JSON
{
  "publication_number": "US11625422B2",
  "country": "US",
  "kind": "B2",
  "title": "Context based surface form generation for cognitive system dictionaries",
  "abstract": "A mechanism is provided in a data processing system to implement an annotator for annotating content using context-based surface forms. The mechanism receives a dictionary data structure of surface forms comprising a plurality of regular expressions and input content. The mechanism compares a given span of text in the input content to each regular expression in the dictionary data structure. Responsive to the given span of text matching a given regular expression, an annotator annotates the span of text with a content indicator corresponding to a content category associated with the dictionary data structure. The mechanism performs a natural language processing operation on the input content based on results of the annotation.",
  "claims": [
    "1. A method, in a data processing system comprising at least one processor and at least one memory, wherein the at least one memory comprises instructions that are executed by the at least one processor to configure the at least one processor to implement an annotator for annotating content using context-based surface forms, the method comprising: generating a plurality of predefined content indicator data structures, wherein each predefined content indicator data structure comprises a dictionary of context-based surface forms of content indicators corresponding to a corresponding medical condition in a plurality of different medical conditions, wherein each of the context-based surface forms comprise one or more regular expressions and corresponding context that specifies a criterion which, when satisfied, causes an annotator to annotate a corresponding portion of text in which the one or more regular expressions are present; generating a user specific content indicator dictionary (USCID) data structure for a patient at least by selecting a subset of the predefined content indicator data structures that correspond to one or more medical conditions, in the plurality of different medical conditions, that are associated with the patient; receiving input content; comparing a given span of text in the input content to each regular expression in each context-based surface form of the subset of predefined content indicator data structures in the user specific content indicator dictionary data structure for the patient; responsive to the given span of text matching at least one regular expression and at least one corresponding context, annotating the span of text with a content indicator corresponding to a content category associated with the user specific content indicator dictionary data structure; and suppressing or promoting content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition.",
    "2. The method of claim 1, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents prescribed targets.",
    "3. The method of claim 2, wherein the metacharacter syntax comprises wildcards.",
    "4. The method of claim 1, wherein the annotator comprises a regular expression processor that translates a regular expression in an implemented syntax into an internal representation that can be executed and matched against a string representing the span of text being searched.",
    "5. A computer program product comprising a computer readable storage medium having a computer readable program stored therein, wherein the computer readable program, when executed on a data processing system, causes the data processing system to implement an annotator for annotating content using context-based surface forms, wherein the computer readable program causes the data processing system to: generate a plurality of predefined content indicator data structures, wherein each predefined content indicator data structure comprises a dictionary of context-based surface forms of content indicators corresponding to a corresponding medical condition in a plurality of different medical conditions, wherein each of the context-based surface forms comprise one or more regular expressions and corresponding context that specifies a criterion which, when satisfied, causes an annotator to annotate a corresponding portion of text in which the one or more regular expressions are present; generate a user specific content indicator dictionary (USCID) data structure for a patient at least by selecting a subset of the predefined content indicator data structures that correspond to one or more medical conditions, in the plurality of different medical conditions, that are associated with the patient; receive input content; compare a given span of text in the input content to each regular expression in each context-based surface form of the subset of predefined content indicator data structures in the user specific content indicator dictionary data structure for the patient; responsive to the given span of text matching at least one regular expression and at least one corresponding context, annotate the span of text with a content indicator corresponding to a content category associated with the user specific content indicator dictionary data structure; and suppress or promote content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition.",
    "6. The computer program product of claim 5, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents prescribed targets.",
    "7. The computer program product of claim 6, wherein the metacharacter syntax comprises wildcards.",
    "8. The computer program product of claim 5, wherein the annotator comprises a regular expression processor that translates a regular expression in an implemented syntax into an internal representation that can be executed and matched against a string representing the span of text being searched.",
    "9. An apparatus comprising: at least one processor; and at least one memory coupled to the at least one processor, wherein the at least one memory comprises instructions which, when executed by the at least one processor, cause the at least one processor to implement an annotator for annotating content using context-based surface forms, wherein the instructions cause the processor to: generate a plurality of predefined content indicator data structures, wherein each predefined content indicator data structure comprises a dictionary of context-based surface forms of content indicators corresponding to a corresponding medical condition in a plurality of different medical conditions, wherein each of the context-based surface forms comprise one or more regular expressions and corresponding context that specifies a criterion which, when satisfied, causes an annotator to annotate a corresponding portion of text in which the one or more regular expressions are present; generate a user specific content indicator dictionary (USCID) data structure for a patient at least by selecting a subset of the predefined content indicator data structures that correspond to one or more medical conditions, in the plurality of different medical conditions, that are associated with the patient; receive input content; compare a given span of text in the input content to each regular expression in each context-based surface form of the subset of predefined content indicator data structures in the user specific content indicator dictionary data structure for the patient; responsive to the given span of text matching at least one regular expression and at least one corresponding context, annotate the span of text with a content indicator corresponding to a content category associated with the user specific content indicator dictionary data structure; and suppress or promote content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition.",
    "10. The apparatus of claim 9, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents prescribed targets.",
    "11. The apparatus of claim 9, wherein the annotator comprises a regular expression processor that translates a regular expression in an implemented syntax into an internal representation that can be executed and matched against a string representing the span of text being searched.",
    "12. The method of claim 1, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents context-based surface forms.",
    "13. The method of claim 1, wherein generating the USCID data structure comprises: receiving, by a surface form editor, user input specifying one or more surface forms; generating, by the surface form editor, one or more regular expressions for the one or more surface forms, wherein each regular expression is associated with an annotation type; and storing, by the surface form editor, the one or more regular expressions in a file that is specified during runtime so the surface forms can be processed.",
    "14. The computer program product of claim 5, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents context-based surface forms.",
    "15. The computer program product of claim 5, wherein generating the USCID data structure comprises: receiving, by a surface form editor, user input specifying one or more surface forms; generating, by the surface form editor, one or more regular expressions for the one or more surface forms, wherein each regular expression is associated with an annotation type; and storing, by the surface form editor, the one or more regular expressions in a file that is specified during runtime so the surface forms can be processed.",
    "16. The apparatus of claim 9, wherein the at least one regular expression, to which the given span of text matches, comprises a metacharacter syntax that represents context-based surface forms.",
    "17. The apparatus of claim 9, wherein generating the USCID data structure comprises: receiving, by a surface form editor, user input specifying one or more surface forms; generating, by the surface form editor, one or more regular expressions for the one or more surface forms, wherein each regular expression is associated with an annotation type; and storing, by the surface form editor, the one or more regular expressions in a file that is specified during runtime so the surface forms can be processed.",
    "18. The method of claim 1, wherein suppressing or promoting content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition comprises: in response to the annotated content corresponding to a negative content indicator data structure of the USCID data structure, replacing the annotated content with content corresponding to a positive content indicator data structure of the USCID data structure.",
    "19. The computer program product of claim 5, wherein suppressing or promoting content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition comprises: in response to the annotated content corresponding to a negative content indicator data structure of the USCID data structure, replacing the annotated content with content corresponding to a positive content indicator data structure of the USCID data structure.",
    "20. The apparatus of claim 9, wherein suppressing or promoting content within the annotated content to form modified content that is tailored to the patient based on the at least one medical condition comprises: in response to the annotated content corresponding to a negative content indicator data structure of the USCID data structure, replacing the annotated content with content corresponding to a positive content indicator data structure of the USCID data structure."
  ],
  "description_excerpt": "The present application relates generally to an improved data processing apparatus and method and more specifically to mechanisms for context based surface form generation for cognitive system dictionaries.\n\nIn modern computing environments such as the Internet, computing tools are present to specifically track and analyze a user's interaction with content so that advertisers can target those individuals with advertisements designed to entice the individual to purchase a product or service, view specific content, or the like. For example, these tools may analyze various factors such as search histories, click-throughs, content viewing, electronic product purchases, electronic shopping cart contents, social networking posts, etc. to determine what types of items/services and/or content a user may be interested in and then present to them corresponding advertisements designed to entice them into performing an action to reward the advertiser, e.g., make a purchase, view content, etc.\n\nIn response, other computing tools have been developed to avoid such content, such as pop-up advertisement blockers, adult content blockers, and the like. These computing tools are keyed to broad categories of content, e.g., adult content, or to particular mechanisms for presenting the content, e.g., pop-ups, banner ads, etc.",
  "cpc": [
    "G06F 16/3329",
    "G06F 16/328",
    "G06F 16/3344",
    "G06F 16/335",
    "G06F 40/169",
    "G06F 40/216",
    "G06F 40/237",
    "G06F 40/284",
    "G06F 40/295",
    "G06F 40/30",
    "G16H 15/00",
    "G16H 20/70",
    "G16H 50/20",
    "G16H 70/60"
  ],
  "ipc": [
    "G06F 16/31",
    "G06F 16/33",
    "G06F 16/332",
    "G06F 16/335",
    "G06F 40/169",
    "G16H 15/00"
  ],
  "assignees": [
    "Merative US LP"
  ],
  "inventors": [
    "Robert C. Sizemore",
    "Jennifer L. La Rocca",
    "Sterling R. Smith",
    "Mario J. Lorenzo",
    "Kristin E. McNeil",
    "David B. Werts"
  ],
  "filing_date": "2019-12-02",
  "publication_date": "2023-04-11",
  "grant_date": "2023-04-11",
  "priority_date": "2019-12-02",
  "application_number": "US-201916700218-A",
  "family_id": "76091510",
  "cited_by_count": 1,
  "citations": [
    "US5559693A",
    "US6310629B1",
    "US6236959B1",
    "US7076438B1",
    "US20070050150A1",
    "US9495957B2",
    "US8024173B1",
    "US20140279746A1",
    "US20100131507A1",
    "US20110125734A1",
    "US8515193B1",
    "US20130226843A1",
    "US20140058738A1",
    "US20190342602A1",
    "US10121076B2",
    "US20150100308A1",
    "US20190141411A1",
    "US10121256B2",
    "US20160085742A1",
    "US9754021B2",
    "US20160203267A1",
    "US20160371587A1",
    "US9740956B2",
    "US9616568B1",
    "US20170091164A1",
    "US20170124037A1",
    "WO2017201676A1",
    "US20170358256A1",
    "US20180039699A1",
    "US20180121618A1",
    "US20180143970A1",
    "US20180260680A1",
    "US10268688B2",
    "US20180336318A1",
    "US20190022863A1",
    "US20190080055A1",
    "US20190108912A1",
    "US20190156953A1",
    "US20190156821A1",
    "US20190250895A1",
    "CN108647591A",
    "US10769503B1",
    "CN108765450A",
    "CN109063723A",
    "US11288319B1"
  ]
}

Record 798 of 8,000 in Patents full text (MLC-0201). Request the full dataset.