Patent · US11487973B2 · B2 · US
Retraining a computer vision model for robotic process automation
- (11) Publication number
- US11487973B2
- (21) Application number
- 16/517,225
- (22) Filing date
- 2019-07-19
- (30) Priority date
- 2019-07-19
- (43) Publication date
- 2022-11-01
- (45) Date of grant
- 2022-11-01
- (51) IPC
- B25J 9/16; G06N 20/00; G06V 30/10; G06V 30/40
- (52) CPC
- G06V Image or video recognition or understanding: 30/40, 10/22, 30/10
- B25J Manipulators; chambers provided with manipulation devices: 9/163, 9/1697
- G06F Electric digital data processing: 18/214, 18/2178, 18/40, 3/14, 30/20, 30/27, 9/452
- G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/6253, 9/6263
- G06N Computing arrangements based on specific computational models: 20/00, 3/08
- G06T Image data processing or generation, in general: 1/60, 2207/20081, 7/11
- (73) Assignee
- UiPath Inc
- (72) Inventors
- Cosmin Voicu
- (54) Title
- Retraining a computer vision model for robotic process automation
- (57) Abstract
A Computer Vision (CV) model generated by a Machine Learning (ML) system may be retrained for more accurate computer image analysis in Robotic Process Automation (RPA). A designer application may receive a selection of a misidentified or non-identified graphical component in an image form a user, determine representative data of an area of the image that includes the selection, and transmit the representative data and the image to an image database. A reviewer may execute the CV model, or cause the CV model to be executed, to confirm that the error exists, and if so, send the image and a correct label to an ML system for retraining. While the CV model is being retrained, an alternative image recognition model may be used to identify the misidentified or non-identified graphical component.
- Full text
- View on Google Patents
Claims (19)
- A non-transitory computer-readable medium storing a computer program, the computer program configured to cause at least one processor to: receive identifications of graphical components within an image from execution of a Computer Vision (CV) model; display the image with the identified graphical components that were identified by the CV model on a visual display; receive a selection of a misidentified or non-identified graphical component in the image; determine representative data of an area of the image that includes the selection; transmit the representative data and the image to an image database; and embed the image and alternative image processing logic in a workflow to identify the misidentified or non-identified graphical component while a retrained CV model is being produced.
- The non-transitory computer-readable medium of claim 1, wherein the computer program is further configured to cause the at least one processor to: receive text information from the image provided by an optical character recognition (OCR) application.
- The non-transitory computer-readable medium of claim 1, wherein the alternative image processing logic comprises an image matching algorithm.
- The non-transitory computer-readable medium of claim 1, wherein the computer program is further configured to cause the at least one processor to: determine the representative data of the area of the image that includes the selection and transmit the image and the selection without providing an indication to a user.
- The non-transitory computer-readable medium of claim 1, wherein the image database stores screenshots as design time images, reported issues, and image matching area selections.
- The non-transitory computer-readable medium of claim 1, wherein the image is from a virtual machine (VM).
- The non-transitory computer-readable medium of claim 1, wherein the representative data comprises coordinates, line segments, or both, that define a shape having an area.
- A computing system, comprising: memory storing machine-readable computer program instructions; and at least one processor configured to execute the machine-readable computer program instructions, the machine-readable computer program instructions configured to cause the at least one processor to: receive a selection of a misidentified or non-identified graphical component in an image, determine representative data of an area of the image that includes the selection, transmit the representative data and the image to an image database for retraining of a Computer Vision (CV) model, receive identifications of graphical components within the image from execution of a retrained CV model, and display the image with the identified graphical components that were identified by the retrained CV model on a visual display.
- The computing system of claim 8, wherein the machine-readable computer program instructions are further configured to cause the at least one processor to: embed the image and alternative image processing logic in a workflow to identify the misidentified or non-identified graphical component while the CV model is being retrained.
- The computing system of claim 9, wherein the alternative image processing logic comprises an image matching algorithm.
- The computing system of claim 8, wherein the machine-readable computer program instructions are configured to cause the at least one processor to: determine the representative data of the area of the image that includes the selection and transmit the image and the selection without providing an indication to a user.
- The computing system of claim 8, wherein the image database stores screenshots as design time images, reported issues, and image matching area selections.
- The computing system of claim 8, wherein the representative data comprises coordinates, line segments, or both, that define a shape having an area.
- A computer-implemented method, comprising: receiving a selection, by a computing system, of a misidentified or non-identified graphical component in an image; determining, by the computing system, representative data of an area of the image that includes the selection; transmitting, by the computing system, the representative data and the image to an image database; and embedding the image and alternative image processing logic in a workflow, by the computing system, to identify the misidentified or non-identified graphical component while a retrained CV model is being produced.
- The computer-implemented method of claim 14, further comprising: receiving, by the computing system, identifications of graphical components within the image from execution of a retrained CV model; and displaying the image, by the computing system, with the identified graphical components that were identified by the retrained CV model on a visual display.
- The computer-implemented method of claim 14, wherein the alternative image processing logic comprises an image matching algorithm.
- The computer-implemented method of claim 14, wherein the computing system is configured to determine the representative data of the area of the image that includes the selection and transmit the image and the selection without providing an indication to a user.
- The computer-implemented method of claim 14, wherein the image database stores screenshots as design time images, reported issues, and image matching area selections.
- The computer-implemented method of claim 14, wherein the representative data comprises coordinates, line segments, or both, that define a shape having an area.
Description
The present invention generally relates to Robotic Process Automation (RPA), and more specifically, to identifying misidentified or non-identified graphical components and retraining a Computer Vision (CV) model for RPA generated by a Machine Learning (ML) system for more accurate computer image analysis.
Currently, training data to automate ML-generated CV model algorithms for recognizing image features for RPA are obtained by generating synthetic data and collecting screenshots (i.e., digital images) of actual user interfaces of various software applications, whether from live applications or the Internet. Synthetic data is data that is produced with the specific purpose of training ML models. This differs from “real” or “organic” data, which is data that already exists and just needs to be collected and labeled. In this case, organic data includes screenshots that are collected through various mechanisms and labeled.
Another source of training data is the screenshots of the application that the user wants to automate. In this approach, if a graphical element of the interface (e.g., a checkbox, a radio button, a text box, etc.) is not being detected by the CV model, the user (e.g., a customer) may select the element that was not identified, create screenshots of the selection, and send the images with the coordinates of the selection to the service provider. However, this approach requires the user to expend the effort to send the images as feedback and report the error. In practice, most users do not do this.
Citations (20)
- JPH0561845U
- JPH11252450A
- US20090063946A1
- US20100138775A1
- US20110047488A1
- US20140068553A1
- EP3161733A1
- AU2014277851A1
- US20170228119A1
- EP3206170A1
- WO2017156628A1
- JP2019512827A
- US20180157386A1
- US20180370029A1
- JP2019008796A
- US20190126463A1
- US20190163499A1
- US20190205363A1
- US20200364485A1
- US11227176B2
Record as JSON
{
"publication_number": "US11487973B2",
"country": "US",
"kind": "B2",
"title": "Retraining a computer vision model for robotic process automation",
"abstract": "A Computer Vision (CV) model generated by a Machine Learning (ML) system may be retrained for more accurate computer image analysis in Robotic Process Automation (RPA). A designer application may receive a selection of a misidentified or non-identified graphical component in an image form a user, determine representative data of an area of the image that includes the selection, and transmit the representative data and the image to an image database. A reviewer may execute the CV model, or cause the CV model to be executed, to confirm that the error exists, and if so, send the image and a correct label to an ML system for retraining. While the CV model is being retrained, an alternative image recognition model may be used to identify the misidentified or non-identified graphical component.",
"claims": [
"1. A non-transitory computer-readable medium storing a computer program, the computer program configured to cause at least one processor to: receive identifications of graphical components within an image from execution of a Computer Vision (CV) model; display the image with the identified graphical components that were identified by the CV model on a visual display; receive a selection of a misidentified or non-identified graphical component in the image; determine representative data of an area of the image that includes the selection; transmit the representative data and the image to an image database; and embed the image and alternative image processing logic in a workflow to identify the misidentified or non-identified graphical component while a retrained CV model is being produced.",
"2. The non-transitory computer-readable medium of claim 1, wherein the computer program is further configured to cause the at least one processor to: receive text information from the image provided by an optical character recognition (OCR) application.",
"3. The non-transitory computer-readable medium of claim 1, wherein the alternative image processing logic comprises an image matching algorithm.",
"4. The non-transitory computer-readable medium of claim 1, wherein the computer program is further configured to cause the at least one processor to: determine the representative data of the area of the image that includes the selection and transmit the image and the selection without providing an indication to a user.",
"5. The non-transitory computer-readable medium of claim 1, wherein the image database stores screenshots as design time images, reported issues, and image matching area selections.",
"6. The non-transitory computer-readable medium of claim 1, wherein the image is from a virtual machine (VM).",
"7. The non-transitory computer-readable medium of claim 1, wherein the representative data comprises coordinates, line segments, or both, that define a shape having an area.",
"8. A computing system, comprising: memory storing machine-readable computer program instructions; and at least one processor configured to execute the machine-readable computer program instructions, the machine-readable computer program instructions configured to cause the at least one processor to: receive a selection of a misidentified or non-identified graphical component in an image, determine representative data of an area of the image that includes the selection, transmit the representative data and the image to an image database for retraining of a Computer Vision (CV) model, receive identifications of graphical components within the image from execution of a retrained CV model, and display the image with the identified graphical components that were identified by the retrained CV model on a visual display.",
"9. The computing system of claim 8, wherein the machine-readable computer program instructions are further configured to cause the at least one processor to: embed the image and alternative image processing logic in a workflow to identify the misidentified or non-identified graphical component while the CV model is being retrained.",
"10. The computing system of claim 9, wherein the alternative image processing logic comprises an image matching algorithm.",
"11. The computing system of claim 8, wherein the machine-readable computer program instructions are configured to cause the at least one processor to: determine the representative data of the area of the image that includes the selection and transmit the image and the selection without providing an indication to a user.",
"12. The computing system of claim 8, wherein the image database stores screenshots as design time images, reported issues, and image matching area selections.",
"13. The computing system of claim 8, wherein the representative data comprises coordinates, line segments, or both, that define a shape having an area.",
"14. A computer-implemented method, comprising: receiving a selection, by a computing system, of a misidentified or non-identified graphical component in an image; determining, by the computing system, representative data of an area of the image that includes the selection; transmitting, by the computing system, the representative data and the image to an image database; and embedding the image and alternative image processing logic in a workflow, by the computing system, to identify the misidentified or non-identified graphical component while a retrained CV model is being produced.",
"15. The computer-implemented method of claim 14, further comprising: receiving, by the computing system, identifications of graphical components within the image from execution of a retrained CV model; and displaying the image, by the computing system, with the identified graphical components that were identified by the retrained CV model on a visual display.",
"16. The computer-implemented method of claim 14, wherein the alternative image processing logic comprises an image matching algorithm.",
"17. The computer-implemented method of claim 14, wherein the computing system is configured to determine the representative data of the area of the image that includes the selection and transmit the image and the selection without providing an indication to a user.",
"18. The computer-implemented method of claim 14, wherein the image database stores screenshots as design time images, reported issues, and image matching area selections.",
"19. The computer-implemented method of claim 14, wherein the representative data comprises coordinates, line segments, or both, that define a shape having an area."
],
"description_excerpt": "The present invention generally relates to Robotic Process Automation (RPA), and more specifically, to identifying misidentified or non-identified graphical components and retraining a Computer Vision (CV) model for RPA generated by a Machine Learning (ML) system for more accurate computer image analysis.\n\nCurrently, training data to automate ML-generated CV model algorithms for recognizing image features for RPA are obtained by generating synthetic data and collecting screenshots (i.e., digital images) of actual user interfaces of various software applications, whether from live applications or the Internet. Synthetic data is data that is produced with the specific purpose of training ML models. This differs from “real” or “organic” data, which is data that already exists and just needs to be collected and labeled. In this case, organic data includes screenshots that are collected through various mechanisms and labeled.\n\nAnother source of training data is the screenshots of the application that the user wants to automate. In this approach, if a graphical element of the interface (e.g., a checkbox, a radio button, a text box, etc.) is not being detected by the CV model, the user (e.g., a customer) may select the element that was not identified, create screenshots of the selection, and send the images with the coordinates of the selection to the service provider. However, this approach requires the user to expend the effort to send the images as feedback and report the error. In practice, most users do not do this.",
"cpc": [
"G06V 30/40",
"B25J 9/163",
"B25J 9/1697",
"G06F 18/214",
"G06F 18/2178",
"G06F 18/40",
"G06F 3/14",
"G06F 30/20",
"G06F 30/27",
"G06F 9/452",
"G06K 9/6253",
"G06K 9/6263",
"G06N 20/00",
"G06N 3/08",
"G06T 1/60",
"G06T 2207/20081",
"G06T 7/11",
"G06V 10/22",
"G06V 30/10"
],
"ipc": [
"B25J 9/16",
"G06N 20/00",
"G06V 30/10",
"G06V 30/40"
],
"assignees": [
"UiPath Inc"
],
"inventors": [
"Cosmin Voicu"
],
"filing_date": "2019-07-19",
"publication_date": "2022-11-01",
"grant_date": "2022-11-01",
"priority_date": "2019-07-19",
"application_number": "US-201916517225-A",
"family_id": "71670084",
"cited_by_count": 2,
"citations": [
"JPH0561845U",
"JPH11252450A",
"US20090063946A1",
"US20100138775A1",
"US20110047488A1",
"US20140068553A1",
"EP3161733A1",
"AU2014277851A1",
"US20170228119A1",
"EP3206170A1",
"WO2017156628A1",
"JP2019512827A",
"US20180157386A1",
"US20180370029A1",
"JP2019008796A",
"US20190126463A1",
"US20190163499A1",
"US20190205363A1",
"US20200364485A1",
"US11227176B2"
]
}
Record 982 of 8,000 in Patents full text (MLC-0201). Request the full dataset.