MLchartDataset catalogue

Patent · US9563920B2 · B2 · US

Method, system and program product for matching of transaction records

(11) Publication number
US9563920B2
(21) Application number
14/175,806
(22) Filing date
2014-02-07
(30) Priority date
2013-03-14
(43) Publication date
2017-02-07
(45) Date of grant
2017-02-07
(51) IPC
G06Q 40/00; G07B 17/00; G07F 19/00
(52) CPC
  • G06Q Information and communication technology [ICT] specially adapted for administrative, commercial, financial, managerial or supervisory purposes; systems or methods specially adapted for administrative, commercial, financial, managerial or supervisory purposes, not otherwise provided for: 40/12
  • G06F Electric digital data processing: 16/24564, 16/285
(73) Assignee
OPERARTIS LLC
(72) Inventors
Tracey Deborah Lall
(54) Title
Method, system and program product for matching of transaction records
(57) Abstract

A method, system and program product comprise accessing a transaction records database. Unmatched records are collected into a first set. The first set at least comprises events and transactions. Probabilities of event matches of transactions originating from an event are calculated. The calculating uses at least defined features and stored probability distributions. A quality value for each of the event matches is calculated. The quality value is at least in part being determined by the probability of the event match. A second set of optimized event matches is determined using at least the quality values. Each of the optimized event matches at least comprises transactions deemed to have been generated by the event.

Full text
View on Google Patents

Claims (32)

  1. A method comprising: accessing, by one or more computer processing units, a database stored in memory of the one or more computer processing units, each entry of the database corresponding to one other entry of the database; collecting unmatched entries of the database into a first set, by the one or more computer processing units, the first set at least comprising database entries for which the corresponding one other entry is unidentified; calculating probabilities of event matches of unmatched entries originating from a single event, by the one or more computer processing units, said calculating comprising, for each entry identified in the first set of collected unmatched entries accessed by the one or more computing processing units from the database: calculating, by the one or more computer processing units, a probability of event matching for each other entry identified in the first set of collected unmatched entries, the probability of event matching identifying a likelihood that said entry and said other entry originated from the same event; for each calculated probability of event matching between pairs of entries identified in the first set of collected unmatched records, calculating a quality value for each of the event matches of the unmatched entries based on the calculated probability of matching for said pair of entries, by the one or more computer processing units; determining a second set of optimized event matches of the unmatched entries using at least the quality values, by the one or more computer processing units, each of the optimized event matches comprising entries deemed to have originated from the same event; modifying, by the one or more computer processing units, each entry in the database corresponding to the entries deemed to have originated from the event to include an identifier for a second entry deemed to have originated from the event; and storing, by the one or more computer processing units, each modified entry to the memory of the one or more computer processing units.
  2. The method as recited in claim 1, in which said step of determining a second set maximizes an overall value of the quality values for the second set and every entry has only one event match.
  3. The method as recited in claim 2, in which the quality value is further at least in part determined by a user defined quality function characterizing an operational cost associated with an event match being incorrect.
  4. The method as recited in claim 1, in which said step of calculating probabilities uses the defined features and probability distributions to determine field range values of an entry where a defined percentage of possible candidate tuples are identified.
  5. The method as recited in claim 2, in which a maximum match optimization for a bipartite graph is used.
  6. The method as recited in claim 1, in which said step of accessing further comprises accessing additional entries which have not yet been processed.
  7. The method as recited in claim 1, in which said step of calculating probabilities of event matches uses a supervised machine learning classifier which has been trained on historical matches which have been validated.
  8. The method as recited in claim 7, in which event matches corresponding to historical matches which have been validated are determined as existing event matches following completion of a period of time for operational review.
  9. The method as recited in claim 7, in which the supervised machine learning classifier comprises a non-parametric bayes classifier.
  10. The method as recited in claim 8, in which the defined features and stored probability distributions are stored as histograms.
  11. The method as recited in claim 1, in which a causal pair probability is calculated from a marginal probability of each causal pair over an entire match set joint probability.
  12. The method as recited in claim 1, in which the defined features and stored probability distributions comprise histograms indexed by key causal characteristics.
  13. The method as recited in claim 5, in which the bipartite graph comprises a set of separate bipartite graphs, each of the separate bipartite graphs comprising a closed set of entries and events which are potentially causally related.
  14. The method as recited in claim 12, in which the defined features and stored probability distributions further comprise historical weighting factors.
  15. The method as recited in claim 1, in which optimized event matches having a quality value below a defined threshold are marked for review.
  16. The method as recited in claim 1, in which the calculated probability for each event match is displayed in a user interface.
  17. A system comprising: a computing device comprising one or more processing units and a memory storage device storing a database comprising a plurality of unmatched entries, each comprising data generated responsive to an event and corresponding to one other unmatched entry, the one or more processing units configured to: retrieve, from the database, stored in the memory storage device of the computing device, the plurality of unmatched entries; calculate, for each entry of the plurality of retrieved unmatched entries, a probability of event matching for each other entry of the plurality of retrieved unmatched entries, the probability of event matching identifying a likelihood that said entry and said other entry originated from the same event store in a data structure in the memory storage device, for each entry of the plurality of retrieved unmatched entries, the calculated probabilities of event matching for each other entry of the plurality of retrieved unmatched entries; calculate, for each calculated probability of event matching stored in the data structure, a quality value of the probability of event matching, each quality value based on the calculated probability of matching for said pair of entries; store each calculated quality value in the memory storage device in association with the corresponding probability of event matching and corresponding pair of transactions entries of the plurality of retrieved unmatched entries; generate a set of optimized matches of entries generated responsive to the same event, based on the calculated quality values for each pair of entries of the plurality of retrieved unmatched entries; modify each entry in the database corresponding to the entries deemed to have been generated by the event to include an identifier for a second entry deemed to have been generated by the event; and store each modified entry to the memory storage device of the computer processing device.
  18. A non-transitory computer-readable storage medium with an executable program stored thereon, wherein the program instructs one or more processors to perform the following steps: accessing a database, each entry of the database corresponding to one other record of the database; collecting unmatched entries into a first set, the first set at least comprising entries for which the corresponding one other entry is unidentified; calculating probabilities of event matches of entries originating from an event, said calculating comprising, for each entry identified in the first set of collected unmatched entries: calculating a probability of event matching for each other entry identified in the first set of collected unmatched entries, the probability of event matching identifying a likelihood that said entry and said other entry originated from the same event; calculating a quality value for each of the event matches of the unmatched entries based on the calculated probability of matching for said pair of entries; determining a second set of optimized event matches using at least the quality values, each of the optimized event matches at least comprising entries deemed to have been generated by the event; modifying each entry in the database corresponding to the entries deemed to have been generated by the event to include an identifier for a second entry deemed to have been generated by the event; and storing each modified record to a memory storage device comprising the database.
  19. The program instructing the processor as recited in claim 18, in which said step of determining a second set maximizes an overall value of the quality values for the second set and every entry is contained in only one event match.
  20. The program instructing the processor as recited in claim 19, in which the quality value is further at least in part determined by a user defined quality function characterizing an operational cost associated with an event match being incorrect.
  21. The program instructing the processor as recited in claim 18, in which said step of calculating probabilities uses the defined features and probability distributions to determine field range values of an entry where a defined percentage of possible candidate tuples are identified.
  22. The program instructing the processor as recited in claim 19, in which a maximum match optimization for a bipartite graph is used.
  23. The program instructing the processor as recited in claim 18, in which said step of accessing further comprises new entries which have not yet been processed.
  24. The program instructing the processor as recited in claim 18, in which said step of calculating probabilities of event matches uses a supervised machine learning classifier which has been trained on historical matches which have been validated.
  25. The program instructing the processor as recited in claim 24, in which event matches corresponding to historical matches which have been validated are determined as existing event matches following completion of a period of time for operational review.
  26. The program instructing the processor as recited in claim 24, in which the supervised machine learning classifier comprises a non-parametric bayes classifier.
  27. The program instructing the processor as recited in claim 25, in which the defined features and stored probability distributions are stored as histograms.
  28. The program instructing the processor as recited in claim 18, in which a causal pair probability is calculated from a marginal probability of each causal pair over an entire match set joint probability.
  29. The program instructing the processor as recited in claim 18, in which the defined features and stored probability distributions comprise histograms indexed by key causal characteristics.
  30. The program instructing the processor as recited in claim 22, in which the bipartite graph comprises a set of separate bipartite graphs, each of the separate bipartite graphs comprising a closed set of entries and events which are potentially causally related.
  31. The method of claim 1, wherein calculating the quality value for each of the event matches further comprises: for each calculated probability of event matching between pairs of entries identified in the first set of collected unmatched records, calculating the quality value of the calculated probability based on (i) the calculated probabilities of matching for said pair of entries, (ii) the calculated probabilities of matching for a first entry of said pair of entries with each other entry identified in the first set, and (iii) the calculated probabilities of matching for a second entry of said pair of entries with each other entry identified in the first set.
  32. The method of claim 1, wherein each entry may be correctly paired with only one other entry, and wherein the calculated quality value represents the probability of matching of said pair of entries being correct in the context of conflicting other possible event matches for a first entry of said pair and other conflicting other possible event matches for a second entry of said pair.

Description

Not applicable.

Not applicable.

Not applicable.

A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or patent disclosure as it appears in the Patent and Trademark Office, patent file or records, but otherwise reserves all copyright rights whatsoever.

One or more embodiments of the invention generally relate to the automated matching of operational transaction records. More particularly, the invention generally relates to a method, apparatus and program for using information derived from validated historical transaction matches to enable the matching of new transactions such that the operational risk associated with any potential mismatches is minimized.

The following is an example of a specific aspect in the prior art that, while expected to be helpful to further educate the reader as to additional aspects of the prior art, is not to be construed as limiting the present invention, or any embodiments thereof, to anything stated or implied therein or inferred thereupon.

By way of educational background, an aspect of the prior art generally useful to be aware of is that in many typical business and financial operational environments matches must be found between business transaction records drawn from two or more different data sources and which have originated from the same business event in order to reconcile said business event with said subsequently created business transaction records.

Citations (8)

  • US5960430A
  • US20070271160A1
  • US7542973B2
  • US20090006380A1
  • US20110004626A1
  • US20120221485A1
  • US20110137928A1
  • US20120197827A1
Record as JSON
{
  "publication_number": "US9563920B2",
  "country": "US",
  "kind": "B2",
  "title": "Method, system and program product for matching of transaction records",
  "abstract": "A method, system and program product comprise accessing a transaction records database. Unmatched records are collected into a first set. The first set at least comprises events and transactions. Probabilities of event matches of transactions originating from an event are calculated. The calculating uses at least defined features and stored probability distributions. A quality value for each of the event matches is calculated. The quality value is at least in part being determined by the probability of the event match. A second set of optimized event matches is determined using at least the quality values. Each of the optimized event matches at least comprises transactions deemed to have been generated by the event.",
  "claims": [
    "1. A method comprising: accessing, by one or more computer processing units, a database stored in memory of the one or more computer processing units, each entry of the database corresponding to one other entry of the database; collecting unmatched entries of the database into a first set, by the one or more computer processing units, the first set at least comprising database entries for which the corresponding one other entry is unidentified; calculating probabilities of event matches of unmatched entries originating from a single event, by the one or more computer processing units, said calculating comprising, for each entry identified in the first set of collected unmatched entries accessed by the one or more computing processing units from the database: calculating, by the one or more computer processing units, a probability of event matching for each other entry identified in the first set of collected unmatched entries, the probability of event matching identifying a likelihood that said entry and said other entry originated from the same event; for each calculated probability of event matching between pairs of entries identified in the first set of collected unmatched records, calculating a quality value for each of the event matches of the unmatched entries based on the calculated probability of matching for said pair of entries, by the one or more computer processing units; determining a second set of optimized event matches of the unmatched entries using at least the quality values, by the one or more computer processing units, each of the optimized event matches comprising entries deemed to have originated from the same event; modifying, by the one or more computer processing units, each entry in the database corresponding to the entries deemed to have originated from the event to include an identifier for a second entry deemed to have originated from the event; and storing, by the one or more computer processing units, each modified entry to the memory of the one or more computer processing units.",
    "2. The method as recited in claim 1, in which said step of determining a second set maximizes an overall value of the quality values for the second set and every entry has only one event match.",
    "3. The method as recited in claim 2, in which the quality value is further at least in part determined by a user defined quality function characterizing an operational cost associated with an event match being incorrect.",
    "4. The method as recited in claim 1, in which said step of calculating probabilities uses the defined features and probability distributions to determine field range values of an entry where a defined percentage of possible candidate tuples are identified.",
    "5. The method as recited in claim 2, in which a maximum match optimization for a bipartite graph is used.",
    "6. The method as recited in claim 1, in which said step of accessing further comprises accessing additional entries which have not yet been processed.",
    "7. The method as recited in claim 1, in which said step of calculating probabilities of event matches uses a supervised machine learning classifier which has been trained on historical matches which have been validated.",
    "8. The method as recited in claim 7, in which event matches corresponding to historical matches which have been validated are determined as existing event matches following completion of a period of time for operational review.",
    "9. The method as recited in claim 7, in which the supervised machine learning classifier comprises a non-parametric bayes classifier.",
    "10. The method as recited in claim 8, in which the defined features and stored probability distributions are stored as histograms.",
    "11. The method as recited in claim 1, in which a causal pair probability is calculated from a marginal probability of each causal pair over an entire match set joint probability.",
    "12. The method as recited in claim 1, in which the defined features and stored probability distributions comprise histograms indexed by key causal characteristics.",
    "13. The method as recited in claim 5, in which the bipartite graph comprises a set of separate bipartite graphs, each of the separate bipartite graphs comprising a closed set of entries and events which are potentially causally related.",
    "14. The method as recited in claim 12, in which the defined features and stored probability distributions further comprise historical weighting factors.",
    "15. The method as recited in claim 1, in which optimized event matches having a quality value below a defined threshold are marked for review.",
    "16. The method as recited in claim 1, in which the calculated probability for each event match is displayed in a user interface.",
    "17. A system comprising: a computing device comprising one or more processing units and a memory storage device storing a database comprising a plurality of unmatched entries, each comprising data generated responsive to an event and corresponding to one other unmatched entry, the one or more processing units configured to: retrieve, from the database, stored in the memory storage device of the computing device, the plurality of unmatched entries; calculate, for each entry of the plurality of retrieved unmatched entries, a probability of event matching for each other entry of the plurality of retrieved unmatched entries, the probability of event matching identifying a likelihood that said entry and said other entry originated from the same event store in a data structure in the memory storage device, for each entry of the plurality of retrieved unmatched entries, the calculated probabilities of event matching for each other entry of the plurality of retrieved unmatched entries; calculate, for each calculated probability of event matching stored in the data structure, a quality value of the probability of event matching, each quality value based on the calculated probability of matching for said pair of entries; store each calculated quality value in the memory storage device in association with the corresponding probability of event matching and corresponding pair of transactions entries of the plurality of retrieved unmatched entries; generate a set of optimized matches of entries generated responsive to the same event, based on the calculated quality values for each pair of entries of the plurality of retrieved unmatched entries; modify each entry in the database corresponding to the entries deemed to have been generated by the event to include an identifier for a second entry deemed to have been generated by the event; and store each modified entry to the memory storage device of the computer processing device.",
    "18. A non-transitory computer-readable storage medium with an executable program stored thereon, wherein the program instructs one or more processors to perform the following steps: accessing a database, each entry of the database corresponding to one other record of the database; collecting unmatched entries into a first set, the first set at least comprising entries for which the corresponding one other entry is unidentified; calculating probabilities of event matches of entries originating from an event, said calculating comprising, for each entry identified in the first set of collected unmatched entries: calculating a probability of event matching for each other entry identified in the first set of collected unmatched entries, the probability of event matching identifying a likelihood that said entry and said other entry originated from the same event; calculating a quality value for each of the event matches of the unmatched entries based on the calculated probability of matching for said pair of entries; determining a second set of optimized event matches using at least the quality values, each of the optimized event matches at least comprising entries deemed to have been generated by the event; modifying each entry in the database corresponding to the entries deemed to have been generated by the event to include an identifier for a second entry deemed to have been generated by the event; and storing each modified record to a memory storage device comprising the database.",
    "19. The program instructing the processor as recited in claim 18, in which said step of determining a second set maximizes an overall value of the quality values for the second set and every entry is contained in only one event match.",
    "20. The program instructing the processor as recited in claim 19, in which the quality value is further at least in part determined by a user defined quality function characterizing an operational cost associated with an event match being incorrect.",
    "21. The program instructing the processor as recited in claim 18, in which said step of calculating probabilities uses the defined features and probability distributions to determine field range values of an entry where a defined percentage of possible candidate tuples are identified.",
    "22. The program instructing the processor as recited in claim 19, in which a maximum match optimization for a bipartite graph is used.",
    "23. The program instructing the processor as recited in claim 18, in which said step of accessing further comprises new entries which have not yet been processed.",
    "24. The program instructing the processor as recited in claim 18, in which said step of calculating probabilities of event matches uses a supervised machine learning classifier which has been trained on historical matches which have been validated.",
    "25. The program instructing the processor as recited in claim 24, in which event matches corresponding to historical matches which have been validated are determined as existing event matches following completion of a period of time for operational review.",
    "26. The program instructing the processor as recited in claim 24, in which the supervised machine learning classifier comprises a non-parametric bayes classifier.",
    "27. The program instructing the processor as recited in claim 25, in which the defined features and stored probability distributions are stored as histograms.",
    "28. The program instructing the processor as recited in claim 18, in which a causal pair probability is calculated from a marginal probability of each causal pair over an entire match set joint probability.",
    "29. The program instructing the processor as recited in claim 18, in which the defined features and stored probability distributions comprise histograms indexed by key causal characteristics.",
    "30. The program instructing the processor as recited in claim 22, in which the bipartite graph comprises a set of separate bipartite graphs, each of the separate bipartite graphs comprising a closed set of entries and events which are potentially causally related.",
    "31. The method of claim 1, wherein calculating the quality value for each of the event matches further comprises: for each calculated probability of event matching between pairs of entries identified in the first set of collected unmatched records, calculating the quality value of the calculated probability based on (i) the calculated probabilities of matching for said pair of entries, (ii) the calculated probabilities of matching for a first entry of said pair of entries with each other entry identified in the first set, and (iii) the calculated probabilities of matching for a second entry of said pair of entries with each other entry identified in the first set.",
    "32. The method of claim 1, wherein each entry may be correctly paired with only one other entry, and wherein the calculated quality value represents the probability of matching of said pair of entries being correct in the context of conflicting other possible event matches for a first entry of said pair and other conflicting other possible event matches for a second entry of said pair."
  ],
  "description_excerpt": "Not applicable.\n\nNot applicable.\n\nNot applicable.\n\nA portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or patent disclosure as it appears in the Patent and Trademark Office, patent file or records, but otherwise reserves all copyright rights whatsoever.\n\nOne or more embodiments of the invention generally relate to the automated matching of operational transaction records. More particularly, the invention generally relates to a method, apparatus and program for using information derived from validated historical transaction matches to enable the matching of new transactions such that the operational risk associated with any potential mismatches is minimized.\n\nThe following is an example of a specific aspect in the prior art that, while expected to be helpful to further educate the reader as to additional aspects of the prior art, is not to be construed as limiting the present invention, or any embodiments thereof, to anything stated or implied therein or inferred thereupon.\n\nBy way of educational background, an aspect of the prior art generally useful to be aware of is that in many typical business and financial operational environments matches must be found between business transaction records drawn from two or more different data sources and which have originated from the same business event in order to reconcile said business event with said subsequently created business transaction records.",
  "cpc": [
    "G06Q 40/12",
    "G06F 16/24564",
    "G06F 16/285"
  ],
  "ipc": [
    "G06Q 40/00",
    "G07B 17/00",
    "G07F 19/00"
  ],
  "assignees": [
    "OPERARTIS LLC"
  ],
  "inventors": [
    "Tracey Deborah Lall"
  ],
  "filing_date": "2014-02-07",
  "publication_date": "2017-02-07",
  "grant_date": "2017-02-07",
  "priority_date": "2013-03-14",
  "application_number": "US-201414175806-A",
  "family_id": "51532512",
  "cited_by_count": 6,
  "citations": [
    "US5960430A",
    "US20070271160A1",
    "US7542973B2",
    "US20090006380A1",
    "US20110004626A1",
    "US20120221485A1",
    "US20110137928A1",
    "US20120197827A1"
  ]
}

Record 4,312 of 8,000 in Patents full text (MLC-0201). Request the full dataset.