Patent · US2018157939A1 · A1 · US
System and method for appearance search
- (11) Publication number
- US2018157939A1
- (21) Application number
- 15/832,654
- (22) Filing date
- 2017-12-05
- (30) Priority date
- 2016-12-05
- (43) Publication date
- 2018-06-07
- (51) IPC
- G06N 20/00; G06T 7/70; G06V 10/143; G06V 10/75; G06V 10/764; H04N 21/4223; H04N 21/44; H04N 21/466; G06N 3/04; G06N 3/08
- (52) CPC
- H04N Pictorial communication, e.g. television: 21/44008, 21/4223, 21/44, 21/466, 21/4666
- G06F Electric digital data processing: 18/22, 18/2413
- G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/6215, 9/66
- G06N Computing arrangements based on specific computational models: 20/00, 3/045, 3/0464, 3/08, 3/084, 3/09, 99/005
- G06T Image data processing or generation, in general: 7/70
- G06V Image or video recognition or understanding: 10/143, 10/25, 10/26, 10/454, 10/462, 10/56, 10/75, 10/764, 10/82, 10/955, 20/44, 20/46, 20/52, 30/194
- (73) Assignee
- Avigilon Corp
- (72) Inventors
- Richard Butt; Alexander Chau; Moussa Doumbouya; Levi Glozman; Lu He; Aleksey Lipchin; Shaun P. Marlatt; Sreemanananth Sadanand; Mitul Saha; Mahesh Saptharishi; Yanyan Hu
- (54) Title
- System and method for appearance search
- (57) Abstract
There is provided an appearance search system comprising one or more cameras configured to capture video of a scene, the video having images of objects. The system comprises one or more processors and memory comprising computer program code stored on the memory and configured when executed by the one or more processors to cause the one or more processors to perform a method. The method comprises identifying one or more of the objects within the images of the objects. The method further comprises implementing a learning machine configured to generate signatures of the identified objects and generate a signature of an object of interest. The system further comprises a network configured to send the images of the objects from the camera to the one or more processors. The method further comprises comparing the signatures of the identified objects with the signature of the object of interest to generate similarity scores for the identified objects, and transmitting an instruction for presenting on a display one or more of the images of the objects based on the similarity scores.
- Full text
- View on Google Patents
Claims (1)
- An appearance search system comprising: one or more cameras configured to capture video of a scene, the video having images of objects; one or more processors and memory comprising computer program code stored on the memory and configured when executed by the one or more processors to cause the one or more processors to perform a method comprising: identifying one or more of the objects within the images of the objects; and implementing a learning machine configured to generate signatures of the identified objects and generate a signature of an object of interest; and a network configured to send the images of the objects from the camera to the one or more processors, wherein the method further comprises: comparing the signatures of the identified objects with the signature of the object of interest to generate similarity scores for the identified objects; and transmitting an instruction for presenting on a display one or more of the images of the objects based on the similarity scores. 2. The system of claim 1, further comprising a storage system for storing the generated signatures of the identified objects, and the video. 3. The system of claim 1, wherein the implemented learning machine is a second learning machine, and wherein the identifying is performed by a first learning machine implemented by the one or more processors. 4. The system of claim 3, wherein the first and second learning machines comprise convolutional neural networks. 5. The system of claim 1, wherein the one or more cameras are further configured to filter the images of the objects by classification of the objects, wherein the one or more cameras are further configured to identify one or more of the images comprising human objects, and wherein the network is further configured to send only the identified images to the one or more processors. 6. The system of claim 1, wherein the one or more cameras are further configured to capture the images of the objects using video analytics. 7. The system of claim 1, wherein the images of the objects comprise portions of image frames of the video, and wherein the portions of the image frames comprise first image portions of the image frames, the first image portions including at least the objects. 8. The system of claim 7, wherein the portions of the image frames further comprise second image portions of the image frames, the second image portions being larger than the first image portions. 9. The system of claim 1, wherein the one or more cameras are further configured to generate reference coordinates for allowing extraction from the video of the images of the objects, and wherein the storage system is configured to store the reference coordinates. 10. The system of claim 1, wherein the one or more cameras are further configured to select one or more images from the video captured over a period of time for obtaining one or more of the images of the objects. 11. The system of claim 1, wherein the identifying of the objects comprises outlining the one or more of the objects in the images. 12. The system of claim 1, wherein the identifying comprises: identifying multiple ones of the objects within at least one of the images; and dividing the at least one of images into multiple divided images, each divided image comprising at least a portion of one of the identified objects. 13. The system of claim 12, wherein the method further comprises: for each identified object: determining a confidence level; and if the confidence level does not meet a confidence requirement, then causing the identifying and the dividing to be performed by the first learning machine; or if the confidence level meets the confidence requirement, then causing the identifying and the dividing to be performed by the second learning machine. 14. A computer-readable medium having stored thereon computer program code executable by one or more processors and configured when executed by the one or more processors to cause the one or more processors to perform a method comprising: capturing video of a scene, the video having images of objects; identifying one or more of the objects within the images of the objects; generating, using a learning machine, signatures of the identified objects, and a signature of an object of interest; generating similarity scores for the identified objects by comparing the signatures of the identified objects with the first signature of the object of interest; and presenting on a display one or more of the images of the objects based on the similarity scores. 15. A system comprising: one or more cameras configured to capture video of a scene; and one or more processors and memory comprising computer program code stored on the memory and configured when executed by the one or more processors to cause the one or more processors to perform a method comprising: extracting chips from the video, wherein the chips comprise images of objects; identifying multiple objects within at least one of the chips; and dividing the at least one chip into multiple divided chips, each divided chip comprising at least a portion of one of the identified objects. 16. The system of claim 15, wherein the method further comprises: implementing a learning machine configured to generate signatures of the identified objects and generate a signature of an object of interest. 17. The system of claim 16, wherein the learning machine is a second learning machine, and wherein the identifying and the dividing are performed by a first learning machine implemented by the one or more processors. 18. The system of claim 17, wherein the method further comprises: for each identified object: determining a confidence level; and if the confidence level does not meet a confidence requirement, then causing the identifying and the dividing to be performed by the first learning machine; or if the confidence level meets the confidence requirement, then causing the identifying and the dividing to be performed by the second learning machine. 19. The system of claim 15, wherein the at least one chip comprises at least one padded chip, wherein each padded chip comprises a first image portion of an image frame of the video. 20. The system of claim 19, wherein the at least one chip further comprises at least one non-padded chip, wherein each non-padded chip comprises a second image portion of an image frame of the video, the second image portion being smaller than the first image portion.
Description
The present subject-matter relates to video surveillance, and more particularly to identifying objects of interest in the video of a video surveillance system.
Computer implemented visual object classification, also called object recognition, pertains to the classifying of visual representations of real-life objects found in still images or motion videos captured by a camera. By performing visual object classification, each visual object found in the still images or motion video is classified according to its type (such as, for example, human, vehicle, or animal).
Automated security and surveillance systems typically employ video cameras or other image capturing devices or sensors to collect image data such as video or video footage. In the simplest systems, images represented by the image data are displayed for contemporaneous screening by security personnel and/or recorded for later review after a security breach. In those systems, the task of detecting and classifying visual objects of interest is performed by a human observer. A significant advance occurs when the system itself is able to perform object detection and classification, either partly or completely.
In a typical surveillance system, one may be interested in detecting objects such as humans, vehicles, animals, etc. that move through the environment. However, if for example a child is lost in a large shopping mall, it could be very time consuming for security personnel to manually review video footage for the lost child.
Record as JSON
{
"publication_number": "US2018157939A1",
"country": "US",
"kind": "A1",
"title": "System and method for appearance search",
"abstract": "There is provided an appearance search system comprising one or more cameras configured to capture video of a scene, the video having images of objects. The system comprises one or more processors and memory comprising computer program code stored on the memory and configured when executed by the one or more processors to cause the one or more processors to perform a method. The method comprises identifying one or more of the objects within the images of the objects. The method further comprises implementing a learning machine configured to generate signatures of the identified objects and generate a signature of an object of interest. The system further comprises a network configured to send the images of the objects from the camera to the one or more processors. The method further comprises comparing the signatures of the identified objects with the signature of the object of interest to generate similarity scores for the identified objects, and transmitting an instruction for presenting on a display one or more of the images of the objects based on the similarity scores.",
"claims": [
"1. An appearance search system comprising: one or more cameras configured to capture video of a scene, the video having images of objects; one or more processors and memory comprising computer program code stored on the memory and configured when executed by the one or more processors to cause the one or more processors to perform a method comprising: identifying one or more of the objects within the images of the objects; and implementing a learning machine configured to generate signatures of the identified objects and generate a signature of an object of interest; and a network configured to send the images of the objects from the camera to the one or more processors, wherein the method further comprises: comparing the signatures of the identified objects with the signature of the object of interest to generate similarity scores for the identified objects; and transmitting an instruction for presenting on a display one or more of the images of the objects based on the similarity scores. 2. The system of claim 1, further comprising a storage system for storing the generated signatures of the identified objects, and the video. 3. The system of claim 1, wherein the implemented learning machine is a second learning machine, and wherein the identifying is performed by a first learning machine implemented by the one or more processors. 4. The system of claim 3, wherein the first and second learning machines comprise convolutional neural networks. 5. The system of claim 1, wherein the one or more cameras are further configured to filter the images of the objects by classification of the objects, wherein the one or more cameras are further configured to identify one or more of the images comprising human objects, and wherein the network is further configured to send only the identified images to the one or more processors. 6. The system of claim 1, wherein the one or more cameras are further configured to capture the images of the objects using video analytics. 7. The system of claim 1, wherein the images of the objects comprise portions of image frames of the video, and wherein the portions of the image frames comprise first image portions of the image frames, the first image portions including at least the objects. 8. The system of claim 7, wherein the portions of the image frames further comprise second image portions of the image frames, the second image portions being larger than the first image portions. 9. The system of claim 1, wherein the one or more cameras are further configured to generate reference coordinates for allowing extraction from the video of the images of the objects, and wherein the storage system is configured to store the reference coordinates. 10. The system of claim 1, wherein the one or more cameras are further configured to select one or more images from the video captured over a period of time for obtaining one or more of the images of the objects. 11. The system of claim 1, wherein the identifying of the objects comprises outlining the one or more of the objects in the images. 12. The system of claim 1, wherein the identifying comprises: identifying multiple ones of the objects within at least one of the images; and dividing the at least one of images into multiple divided images, each divided image comprising at least a portion of one of the identified objects. 13. The system of claim 12, wherein the method further comprises: for each identified object: determining a confidence level; and if the confidence level does not meet a confidence requirement, then causing the identifying and the dividing to be performed by the first learning machine; or if the confidence level meets the confidence requirement, then causing the identifying and the dividing to be performed by the second learning machine. 14. A computer-readable medium having stored thereon computer program code executable by one or more processors and configured when executed by the one or more processors to cause the one or more processors to perform a method comprising: capturing video of a scene, the video having images of objects; identifying one or more of the objects within the images of the objects; generating, using a learning machine, signatures of the identified objects, and a signature of an object of interest; generating similarity scores for the identified objects by comparing the signatures of the identified objects with the first signature of the object of interest; and presenting on a display one or more of the images of the objects based on the similarity scores. 15. A system comprising: one or more cameras configured to capture video of a scene; and one or more processors and memory comprising computer program code stored on the memory and configured when executed by the one or more processors to cause the one or more processors to perform a method comprising: extracting chips from the video, wherein the chips comprise images of objects; identifying multiple objects within at least one of the chips; and dividing the at least one chip into multiple divided chips, each divided chip comprising at least a portion of one of the identified objects. 16. The system of claim 15, wherein the method further comprises: implementing a learning machine configured to generate signatures of the identified objects and generate a signature of an object of interest. 17. The system of claim 16, wherein the learning machine is a second learning machine, and wherein the identifying and the dividing are performed by a first learning machine implemented by the one or more processors. 18. The system of claim 17, wherein the method further comprises: for each identified object: determining a confidence level; and if the confidence level does not meet a confidence requirement, then causing the identifying and the dividing to be performed by the first learning machine; or if the confidence level meets the confidence requirement, then causing the identifying and the dividing to be performed by the second learning machine. 19. The system of claim 15, wherein the at least one chip comprises at least one padded chip, wherein each padded chip comprises a first image portion of an image frame of the video. 20. The system of claim 19, wherein the at least one chip further comprises at least one non-padded chip, wherein each non-padded chip comprises a second image portion of an image frame of the video, the second image portion being smaller than the first image portion."
],
"description_excerpt": "The present subject-matter relates to video surveillance, and more particularly to identifying objects of interest in the video of a video surveillance system.\n\nComputer implemented visual object classification, also called object recognition, pertains to the classifying of visual representations of real-life objects found in still images or motion videos captured by a camera. By performing visual object classification, each visual object found in the still images or motion video is classified according to its type (such as, for example, human, vehicle, or animal).\n\nAutomated security and surveillance systems typically employ video cameras or other image capturing devices or sensors to collect image data such as video or video footage. In the simplest systems, images represented by the image data are displayed for contemporaneous screening by security personnel and/or recorded for later review after a security breach. In those systems, the task of detecting and classifying visual objects of interest is performed by a human observer. A significant advance occurs when the system itself is able to perform object detection and classification, either partly or completely.\n\nIn a typical surveillance system, one may be interested in detecting objects such as humans, vehicles, animals, etc. that move through the environment. However, if for example a child is lost in a large shopping mall, it could be very time consuming for security personnel to manually review video footage for the lost child.",
"cpc": [
"H04N 21/44008",
"G06F 18/22",
"G06F 18/2413",
"G06K 9/6215",
"G06K 9/66",
"G06N 20/00",
"G06N 3/045",
"G06N 3/0464",
"G06N 3/08",
"G06N 3/084",
"G06N 3/09",
"G06N 99/005",
"G06T 7/70",
"G06V 10/143",
"G06V 10/25",
"G06V 10/26",
"G06V 10/454",
"G06V 10/462",
"G06V 10/56",
"G06V 10/75",
"G06V 10/764",
"G06V 10/82",
"G06V 10/955",
"G06V 20/44",
"G06V 20/46",
"G06V 20/52",
"G06V 30/194",
"H04N 21/4223",
"H04N 21/44",
"H04N 21/466",
"H04N 21/4666"
],
"ipc": [
"G06N 20/00",
"G06T 7/70",
"G06V 10/143",
"G06V 10/75",
"G06V 10/764",
"H04N 21/4223",
"H04N 21/44",
"H04N 21/466",
"G06N 3/04",
"G06N 3/08"
],
"assignees": [
"Avigilon Corp"
],
"inventors": [
"Richard Butt",
"Alexander Chau",
"Moussa Doumbouya",
"Levi Glozman",
"Lu He",
"Aleksey Lipchin",
"Shaun P. Marlatt",
"Sreemanananth Sadanand",
"Mitul Saha",
"Mahesh Saptharishi",
"Yanyan Hu"
],
"filing_date": "2017-12-05",
"publication_date": "2018-06-07",
"priority_date": "2016-12-05",
"application_number": "US-201715832654-A",
"family_id": "62243913",
"cited_by_count": 121
}
Record 3,446 of 8,000 in Patents full text (MLC-0201). Request the full dataset.