Patent · US10133933B1 · B1 · US
Item put and take detection using image recognition
- (11) Publication number
- US10133933B1
- (21) Application number
- 15/907,112
- (22) Filing date
- 2018-02-27
- (30) Priority date
- 2017-08-07
- (43) Publication date
- 2018-11-20
- (45) Date of grant
- 2018-11-20
- (51) IPC
- G06V 10/764
- (52) CPC
- H04N Pictorial communication, e.g. television: 17/002, 23/60, 23/90
- G06F Electric digital data processing: 18/24143
- G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/00718, 9/00771
- G06T Image data processing or generation, in general: 2207/10016, 2207/20084, 2210/12
- G06V Image or video recognition or understanding: 10/454, 10/764, 10/82, 20/41, 20/52
- (73) Assignee
- Standard Cognition Corp
- (72) Inventors
- Jordan E. Fisher; Daniel L. Fischetti; Brandon L. Ogle; John F. Novak; Kyle E. Dorman; Kenneth S. Kihara; Juan C. Lasheras
- (54) Title
- Item put and take detection using image recognition
- (57) Abstract
Systems and techniques are provided for tracking puts and takes of inventory items by subjects in an area of real space. A plurality of cameras with overlapping fields of view produce respective sequences of images of corresponding fields of view in the real space. A processing system is coupled to the system. In one embodiment, the processing system comprises image recognition engines receiving corresponding sequences of images from the plurality of cameras. The image recognition engines process the images in the corresponding sequences to identify subjects represented in the images and generate classifications of the identified subjects. The system processes the classifications of identified subjects for sets of images in the sequences of images to detect takes and puts of inventory items on shelves by identified subjects.
- Full text
- View on Google Patents
Claims (27)
- A system for tracking puts and takes of inventory items by subjects in an area of real space, comprising: a plurality of cameras, cameras in the plurality of cameras producing respective sequences of images of corresponding fields of view in the real space, the field of view of each camera overlapping with the field of view of at least one other camera in the plurality of cameras; a processing system coupled to the plurality of cameras, the processing system including a plurality of image recognition engines, receiving corresponding sequences of images from the plurality of cameras, image recognition engines in the plurality of image recognition engines processing the images in the corresponding sequences to identify subjects represented in the images; and logic to process sets of images in the sequences of images that include the identified subjects to detect takes of inventory items by identified subjects and puts of inventory items on shelves by identified subjects, wherein the logic to process sets of images includes: for identified subjects, logic to process images to generate classifications of the images of the identified subjects, the classifications including whether the identified subject is holding an inventory item, a first nearness classification indicating a location of a hand of the identified subject relative to a shelf, a second nearness classification indicating a location a hand of the identified subject relative to the identified subject.
- The system of claim 1, wherein the second nearness classification indicates a location of a hand of the identified subject relative to a body of the identified subject, and the generated classifications include a third nearness classification indicating a location of a hand of the identified subject relative to a basket associated with an identified subject.
- The system of claim 1, including logic to perform time sequence analysis over the classifications of images to detect said takes and said puts by the identified subjects.
- The system of claim 1, wherein the logic to process sets of images includes: for identified subjects, logic to identify bounding boxes of data representing hands in images in the sets of images of the identified subjects, and to process data in the bounding boxes to generate classifications of data within the bounding boxes for the identified subjects.
- The system of claim 3, wherein the classifications include whether the identified subject is holding an inventory item, a first nearness classification indicating a location of a hand of the identified subject relative to a shelf, a second nearness classification indicating a location a hand of the identified subject relative to a body of the identified subject, a third nearness classification indicating a location a hand of the identified subject relative to a basket associated with an identified subject, and an identifier of a likely inventory item.
- The system of claim 3, including logic to perform time sequence analysis over the classifications of data within the bounding boxes in the sets of images to detect said takes and said puts by the identified subjects.
- The system of claim 1, wherein the logic to process sets of images comprises convolutional neural networks.
- The system of claim 1, wherein cameras in the plurality of cameras are configured to generate synchronized sequences of images.
- The system of claim 1, wherein the plurality of cameras comprise cameras disposed over and having fields of view encompassing respective parts of the area in real space.
- The system of claim 1, including logic responsive to the detected takes and puts, to generate log data structures including a list of inventory items for identified subjects.
- A method for tracking puts and takes of inventory items by subjects in an area of real space, the method including: using a plurality of cameras to produce respective sequences of images of corresponding fields of view in the real space, the field of view of each camera overlapping with the field of view of at least one other camera in the plurality of cameras; receiving corresponding sequences of images from the plurality of cameras, processing the images in the corresponding sequences using image recognition engines in a plurality of image recognition engines and identifying subjects represented in the images wherein the plurality of image recognition engines are part of a processing system coupled to the plurality of cameras; and processing sets of images in the sequences of images that include the identified subjects to detect takes of inventory items by identified subjects and puts of inventory items on shelves by identified subjects, wherein the processing sets of images includes: for identified subjects, generating classifications of the images of the identified subjects, the classifications including whether the identified subject is holding an inventory item, a first nearness classification indicating a location of a hand of the identified subject relative to a shelf, a second nearness classification indicating a location of a hand of the identified subject relative to a body of the identified subject, a third nearness classification indicating a location a hand of the identified subject relative to a basket associated with an identified subject, and an identifier of a likely inventory item.
- The method of claim 11, including performing time sequence analysis over the classifications of images to detect said takes and said puts by the identified subjects.
- The method of claim 11, wherein the processing sets of images includes: for identified subjects, identifying bounding boxes of data representing hands in images in the sets of images of the identified subjects, and processing data in the bounding boxes to generate classifications of data within the bounding boxes for the identified subjects.
- The method of claim 13, including performing time sequence analysis over the classifications of data within the bounding boxes in the sets of images to detect said takes and said puts by the identified subjects.
- The method of claim 11, including circular buffers coupled to cameras in the plurality of cameras to store sets of images in the sequences of images from the plurality of cameras.
- The method of claim 11, including processing sets of images using convolutional neural networks.
- The method of claim 11, wherein cameras in the plurality of cameras are configured to generate synchronized sequences of images.
- The method of claim 11, wherein the plurality of cameras comprise cameras disposed over and having fields of view encompassing respective parts of the area in real space.
- The method of claim 11, including responsive to the detected takes and puts, generating a log data structure including a list of inventory items for each identified subject.
- A system for tracking puts and takes of inventory items by subjects in an area of real space, comprising: a plurality of cameras, cameras in the plurality of cameras producing respective sequences of images of corresponding fields of view in the real space, the field of view of each camera overlapping with the field of view of at least one other camera in the plurality of cameras; a processing system coupled to the plurality of cameras, the processing system including: first image recognition engines, receiving the sequences of images from the plurality of cameras, which process images to generate first data sets that identify subjects and locations of the identified subjects in the real space; logic to process the first data sets to specify bounding boxes which include images of hands of identified subjects in images in the sequences of images; second image recognition engines, receiving the sequences of images from the plurality of cameras, which process the specified bounding boxes in the images to generate a classification of hands of the identified subjects, the classification including whether the identified subject is holding an inventory item, a first nearness classification indicating a location of a hand of the identified subject relative to a shelf, a second nearness classification indicating a location of a hand of the identified subject relative to the identified subject, and an identifier of a likely inventory item; and logic to process the classifications of hands for sets of images in the sequences of images of identified subjects to detect takes of inventory items by identified subjects and puts of inventory items on shelves by identified subjects; and logic to generate a log data structure including a list of inventory items for each identified subject.
- The system of claim 20, including circular buffers coupled to cameras in the plurality of cameras to store sets of images in the sequences of images from the plurality of cameras for processing by corresponding ones of the first and second image recognition engines.
- The system of claim 20, wherein the first data sets comprise for each identified subject sets of candidate joints having coordinates in real space.
- The system of claim 20, wherein the logic to process the first data sets to specify bounding boxes specifies bounding boxes based on locations of joints in the sets of candidate joints for each subject.
- The system of claim 20, wherein the second image recognition engines comprise convolutional neural networks.
- The system of claim 20, wherein the logic to process the classifications of bounding boxes comprise convolutional neural networks.
- The system of claim 20, wherein cameras in the plurality of cameras are configured to generate synchronized sequences of images.
- The system of claim 20, wherein the plurality of cameras comprise cameras disposed over and having fields of view encompassing respective parts of the area in real space.
Description
This application is a continuation-in-part of U.S. patent application Ser. No. 15/847,796 filed 19 Dec. 2017 (now U.S. Pat. No. 10,055,853); and benefit is claimed of U.S. Provisional Patent Application No. 62/542,077 filed 7 Aug. 2017. Both applications are incorporated herein by reference.
A portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure as it appears in the Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever.
A computer program listing appendix (Copyright, Standard Cognition, Inc.) submitted electronically via the EFS-Web in ASCII text accompanies this application and is incorporated by reference. The name of the ASCII text file is “STCG_Computer_Program_Appx” created on 10 Jan. 2018 and is 21,742 bytes.
The present invention relates to systems that identify and track puts and takes of items by subjects in real space.
A difficult problem in image processing arises when images from multiple cameras disposed over large spaces are used to identify and track actions of subjects.
Tracking actions of subjects within an area of real space, such as a shopping store, present many technical challenges. For example, consider such an image processing system deployed in a shopping store with multiple customers moving in aisles between the shelves and open spaces within the shopping store. Customers take items from shelves and put those in their respective shopping carts or baskets.
Citations (35)
- US6154559A
- WO2000021021A1
- US7050624B2
- US9582891B2
- WO2002059836A2
- WO2002059836A3
- US20030107649A1
- EP1574986B1
- US20180070056A1
- US20080159634A1
- US20090083815A1
- US8577705B1
- US9449233B2
- US20120159290A1
- US20130156260A1
- JP2013196199A
- JP2014089626A
- US20140282162A1
- US20150019391A1
- US20150039458A1
- US9881221B2
- US9536177B2
- US20150206188A1
- US20170116473A1
- US20150262116A1
- US20160125245A1
- US20180025175A1
- CN104778690B
- US9911290B1
- WO2017151241A2
- US20170278255A1
- US20170309136A1
- US20170323376A1
- US20180165728A1
- US10055853B1
Record as JSON
{
"publication_number": "US10133933B1",
"country": "US",
"kind": "B1",
"title": "Item put and take detection using image recognition",
"abstract": "Systems and techniques are provided for tracking puts and takes of inventory items by subjects in an area of real space. A plurality of cameras with overlapping fields of view produce respective sequences of images of corresponding fields of view in the real space. A processing system is coupled to the system. In one embodiment, the processing system comprises image recognition engines receiving corresponding sequences of images from the plurality of cameras. The image recognition engines process the images in the corresponding sequences to identify subjects represented in the images and generate classifications of the identified subjects. The system processes the classifications of identified subjects for sets of images in the sequences of images to detect takes and puts of inventory items on shelves by identified subjects.",
"claims": [
"1. A system for tracking puts and takes of inventory items by subjects in an area of real space, comprising: a plurality of cameras, cameras in the plurality of cameras producing respective sequences of images of corresponding fields of view in the real space, the field of view of each camera overlapping with the field of view of at least one other camera in the plurality of cameras; a processing system coupled to the plurality of cameras, the processing system including a plurality of image recognition engines, receiving corresponding sequences of images from the plurality of cameras, image recognition engines in the plurality of image recognition engines processing the images in the corresponding sequences to identify subjects represented in the images; and logic to process sets of images in the sequences of images that include the identified subjects to detect takes of inventory items by identified subjects and puts of inventory items on shelves by identified subjects, wherein the logic to process sets of images includes: for identified subjects, logic to process images to generate classifications of the images of the identified subjects, the classifications including whether the identified subject is holding an inventory item, a first nearness classification indicating a location of a hand of the identified subject relative to a shelf, a second nearness classification indicating a location a hand of the identified subject relative to the identified subject.",
"2. The system of claim 1, wherein the second nearness classification indicates a location of a hand of the identified subject relative to a body of the identified subject, and the generated classifications include a third nearness classification indicating a location of a hand of the identified subject relative to a basket associated with an identified subject.",
"3. The system of claim 1, including logic to perform time sequence analysis over the classifications of images to detect said takes and said puts by the identified subjects.",
"4. The system of claim 1, wherein the logic to process sets of images includes: for identified subjects, logic to identify bounding boxes of data representing hands in images in the sets of images of the identified subjects, and to process data in the bounding boxes to generate classifications of data within the bounding boxes for the identified subjects.",
"5. The system of claim 3, wherein the classifications include whether the identified subject is holding an inventory item, a first nearness classification indicating a location of a hand of the identified subject relative to a shelf, a second nearness classification indicating a location a hand of the identified subject relative to a body of the identified subject, a third nearness classification indicating a location a hand of the identified subject relative to a basket associated with an identified subject, and an identifier of a likely inventory item.",
"6. The system of claim 3, including logic to perform time sequence analysis over the classifications of data within the bounding boxes in the sets of images to detect said takes and said puts by the identified subjects.",
"7. The system of claim 1, wherein the logic to process sets of images comprises convolutional neural networks.",
"8. The system of claim 1, wherein cameras in the plurality of cameras are configured to generate synchronized sequences of images.",
"9. The system of claim 1, wherein the plurality of cameras comprise cameras disposed over and having fields of view encompassing respective parts of the area in real space.",
"10. The system of claim 1, including logic responsive to the detected takes and puts, to generate log data structures including a list of inventory items for identified subjects.",
"11. A method for tracking puts and takes of inventory items by subjects in an area of real space, the method including: using a plurality of cameras to produce respective sequences of images of corresponding fields of view in the real space, the field of view of each camera overlapping with the field of view of at least one other camera in the plurality of cameras; receiving corresponding sequences of images from the plurality of cameras, processing the images in the corresponding sequences using image recognition engines in a plurality of image recognition engines and identifying subjects represented in the images wherein the plurality of image recognition engines are part of a processing system coupled to the plurality of cameras; and processing sets of images in the sequences of images that include the identified subjects to detect takes of inventory items by identified subjects and puts of inventory items on shelves by identified subjects, wherein the processing sets of images includes: for identified subjects, generating classifications of the images of the identified subjects, the classifications including whether the identified subject is holding an inventory item, a first nearness classification indicating a location of a hand of the identified subject relative to a shelf, a second nearness classification indicating a location of a hand of the identified subject relative to a body of the identified subject, a third nearness classification indicating a location a hand of the identified subject relative to a basket associated with an identified subject, and an identifier of a likely inventory item.",
"12. The method of claim 11, including performing time sequence analysis over the classifications of images to detect said takes and said puts by the identified subjects.",
"13. The method of claim 11, wherein the processing sets of images includes: for identified subjects, identifying bounding boxes of data representing hands in images in the sets of images of the identified subjects, and processing data in the bounding boxes to generate classifications of data within the bounding boxes for the identified subjects.",
"14. The method of claim 13, including performing time sequence analysis over the classifications of data within the bounding boxes in the sets of images to detect said takes and said puts by the identified subjects.",
"15. The method of claim 11, including circular buffers coupled to cameras in the plurality of cameras to store sets of images in the sequences of images from the plurality of cameras.",
"16. The method of claim 11, including processing sets of images using convolutional neural networks.",
"17. The method of claim 11, wherein cameras in the plurality of cameras are configured to generate synchronized sequences of images.",
"18. The method of claim 11, wherein the plurality of cameras comprise cameras disposed over and having fields of view encompassing respective parts of the area in real space.",
"19. The method of claim 11, including responsive to the detected takes and puts, generating a log data structure including a list of inventory items for each identified subject.",
"20. A system for tracking puts and takes of inventory items by subjects in an area of real space, comprising: a plurality of cameras, cameras in the plurality of cameras producing respective sequences of images of corresponding fields of view in the real space, the field of view of each camera overlapping with the field of view of at least one other camera in the plurality of cameras; a processing system coupled to the plurality of cameras, the processing system including: first image recognition engines, receiving the sequences of images from the plurality of cameras, which process images to generate first data sets that identify subjects and locations of the identified subjects in the real space; logic to process the first data sets to specify bounding boxes which include images of hands of identified subjects in images in the sequences of images; second image recognition engines, receiving the sequences of images from the plurality of cameras, which process the specified bounding boxes in the images to generate a classification of hands of the identified subjects, the classification including whether the identified subject is holding an inventory item, a first nearness classification indicating a location of a hand of the identified subject relative to a shelf, a second nearness classification indicating a location of a hand of the identified subject relative to the identified subject, and an identifier of a likely inventory item; and logic to process the classifications of hands for sets of images in the sequences of images of identified subjects to detect takes of inventory items by identified subjects and puts of inventory items on shelves by identified subjects; and logic to generate a log data structure including a list of inventory items for each identified subject.",
"21. The system of claim 20, including circular buffers coupled to cameras in the plurality of cameras to store sets of images in the sequences of images from the plurality of cameras for processing by corresponding ones of the first and second image recognition engines.",
"22. The system of claim 20, wherein the first data sets comprise for each identified subject sets of candidate joints having coordinates in real space.",
"23. The system of claim 20, wherein the logic to process the first data sets to specify bounding boxes specifies bounding boxes based on locations of joints in the sets of candidate joints for each subject.",
"24. The system of claim 20, wherein the second image recognition engines comprise convolutional neural networks.",
"25. The system of claim 20, wherein the logic to process the classifications of bounding boxes comprise convolutional neural networks.",
"26. The system of claim 20, wherein cameras in the plurality of cameras are configured to generate synchronized sequences of images.",
"27. The system of claim 20, wherein the plurality of cameras comprise cameras disposed over and having fields of view encompassing respective parts of the area in real space."
],
"description_excerpt": "This application is a continuation-in-part of U.S. patent application Ser. No. 15/847,796 filed 19 Dec. 2017 (now U.S. Pat. No. 10,055,853); and benefit is claimed of U.S. Provisional Patent Application No. 62/542,077 filed 7 Aug. 2017. Both applications are incorporated herein by reference.\n\nA portion of the disclosure of this patent document contains material that is subject to copyright protection. The copyright owner has no objection to the facsimile reproduction by anyone of the patent document or the patent disclosure as it appears in the Patent and Trademark Office patent file or records, but otherwise reserves all copyright rights whatsoever.\n\nA computer program listing appendix (Copyright, Standard Cognition, Inc.) submitted electronically via the EFS-Web in ASCII text accompanies this application and is incorporated by reference. The name of the ASCII text file is “STCG_Computer_Program_Appx” created on 10 Jan. 2018 and is 21,742 bytes.\n\nThe present invention relates to systems that identify and track puts and takes of items by subjects in real space.\n\nA difficult problem in image processing arises when images from multiple cameras disposed over large spaces are used to identify and track actions of subjects.\n\nTracking actions of subjects within an area of real space, such as a shopping store, present many technical challenges. For example, consider such an image processing system deployed in a shopping store with multiple customers moving in aisles between the shelves and open spaces within the shopping store. Customers take items from shelves and put those in their respective shopping carts or baskets.",
"cpc": [
"H04N 17/002",
"G06F 18/24143",
"G06K 9/00718",
"G06K 9/00771",
"G06T 2207/10016",
"G06T 2207/20084",
"G06T 2210/12",
"G06V 10/454",
"G06V 10/764",
"G06V 10/82",
"G06V 20/41",
"G06V 20/52",
"H04N 23/60",
"H04N 23/90"
],
"ipc": [
"G06V 10/764"
],
"assignees": [
"Standard Cognition Corp"
],
"inventors": [
"Jordan E. Fisher",
"Daniel L. Fischetti",
"Brandon L. Ogle",
"John F. Novak",
"Kyle E. Dorman",
"Kenneth S. Kihara",
"Juan C. Lasheras"
],
"filing_date": "2018-02-27",
"publication_date": "2018-11-20",
"grant_date": "2018-11-20",
"priority_date": "2017-08-07",
"application_number": "US-201815907112-A",
"family_id": "64176574",
"cited_by_count": 305,
"citations": [
"US6154559A",
"WO2000021021A1",
"US7050624B2",
"US9582891B2",
"WO2002059836A2",
"WO2002059836A3",
"US20030107649A1",
"EP1574986B1",
"US20180070056A1",
"US20080159634A1",
"US20090083815A1",
"US8577705B1",
"US9449233B2",
"US20120159290A1",
"US20130156260A1",
"JP2013196199A",
"JP2014089626A",
"US20140282162A1",
"US20150019391A1",
"US20150039458A1",
"US9881221B2",
"US9536177B2",
"US20150206188A1",
"US20170116473A1",
"US20150262116A1",
"US20160125245A1",
"US20180025175A1",
"CN104778690B",
"US9911290B1",
"WO2017151241A2",
"US20170278255A1",
"US20170309136A1",
"US20170323376A1",
"US20180165728A1",
"US10055853B1"
]
}
Record 3,168 of 8,000 in Patents full text (MLC-0201). Request the full dataset.