MLchartDataset catalogue

Patent · US10628708B2 · B2 · US

Utilizing a deep neural network-based model to identify visually similar digital images based on user-selected visual attributes

(11) Publication number
US10628708B2
(21) Application number
15/983,949
(22) Filing date
2018-05-18
(30) Priority date
2018-05-18
(43) Publication date
2020-04-21
(45) Date of grant
2020-04-21
(51) IPC
G06T 7/73; G06V 10/764; G06V 10/771; G06V 20/00
(52) CPC
  • G06V Image or video recognition or understanding: 20/00, 10/454, 10/761, 10/764, 10/771, 10/82
  • G06F Electric digital data processing: 16/532, 16/5854, 18/211, 18/214, 18/2148, 18/22, 18/2413, 18/41
  • G06K Graphical data reading; presentation of data; record carriers; handling record carriers: 9/52, 9/6254, 9/6257
  • G06N Computing arrangements based on specific computational models: 3/045
  • G06T Image data processing or generation, in general: 2207/20081, 2207/20084, 7/75
(73) Assignee
Adobe Inc
(72) Inventors
Zhe Lin; Xiaohui SHEN; Mingyang Ling; Jianming Zhang; Jason Kuen; Brett Butterfield
(54) Title
Utilizing a deep neural network-based model to identify visually similar digital images based on user-selected visual attributes
(57) Abstract

The present disclosure relates to systems, methods, and non-transitory computer readable media for utilizing a deep neural network-based model to identify similar digital images for query digital images. For example, the disclosed systems utilize a deep neural network-based model to analyze query digital images to generate deep neural network-based representations of the query digital images. In addition, the disclosed systems can generate results of visually-similar digital images for the query digital images based on comparing the deep neural network-based representations with representations of candidate digital images. Furthermore, the disclosed systems can identify visually similar digital images based on user-defined attributes and image masks to emphasize specific attributes or portions of query digital images.

Full text
View on Google Patents

Claims (20)

  1. In a digital medium environment for searching digital images, a non-transitory computer readable medium for matching digital images based on visual similarity comprising instructions that, when executed by a processor, cause a computer device to: receive, from a user client device, a user selection of a query digital image and at least one of spatial selectivity, image composition, or object count to use to identify similar digital images; receive, from the user client device, an image mask that indicates a portion of the query digital image to emphasize in identifying similar digital images and an image mask weight corresponding to the image mask; utilize a deep neural network-based model to generate a set of deep features for the query digital image; generate, based on weighting the set of deep features of the query digital image utilizing the image mask and the image mask weight and in accordance with the user selection, a deep neural network-based representation of the query digital image by utilizing one or more of a spatial selectivity algorithm, an image composition algorithm, or an object count algorithm; and based on the deep neural network-based representation of the query digital image, identify, from a digital image database, a similar digital image for the query digital image.
  2. The non-transitory computer readable medium of claim 1, further comprising instructions that, when executed by the processor, cause the computer device to train the deep neural network-based model to generate deep features for digital images.
  3. The non-transitory computer readable medium of claim 1, wherein the instructions, when executed by the processor, cause the computer device to generate the deep neural network-based representation by modifying the set of deep features by utilizing a deep neural network-based representation comprising at least one of the spatial selectivity algorithm, the image composition algorithm, or the object count algorithm.
  4. The non-transitory computer readable medium of claim 3, further comprising instructions that, when executed by the processor, cause the computer device to receive, from the user client device: an additional query digital image and an additional image mask that indicates a portion of the additional query digital image; and an indication to invert the additional image mask to emphasize portions of the additional query digital image that are outside the additional image mask in identifying similar digital images.
  5. The non-transitory computer readable medium of claim 4, further comprising instructions that, when executed by the processor, cause the computer device to generate, based on receiving the additional query digital image, the additional image mask, and the indication to invert the additional image mask, an additional deep neural network-based representation for the additional query digital image based on weighting the portion of the additional query digital image outside of the additional image mask.
  6. The non-transitory computer readable medium of claim 5, wherein weighting the portion of the additional query digital image outside of the additional image mask comprises applying weights to features of a set of features that correspond to the portion of the additional query digital image outside of the additional image mask.
  7. The non-transitory computer readable medium of claim 1, further comprising instructions that, when executed by the processor, cause the computer device to: determine similarity scores for a plurality of digital images from the digital image database; and rank the plurality of digital images based on the determined similarity scores.
  8. The non-transitory computer readable medium of claim 1, further comprising instructions that, when executed by the processor, cause the computer device to: receive, from the user client device, a selection of a second query digital image; utilize the deep neural network-based model to generate a second set of deep features for the second query digital image; generate, based on the second set of deep features of the second query digital image, a multi-query vector representation that represents a composite of the query digital image and the second query digital image; and identify a similar digital image for the multi-query vector representation from within the digital image database.
  9. The non-transitory computer readable medium of claim 8, further comprising instructions that, when executed by the processor, cause the computer device to: receive, in relation to the second query digital image, a second user selection of at least one of spatial selection, image composition, or object count; and wherein the instructions cause the computer device to generate, based on the second user selection, the multi-query vector representation by utilizing at least one of the spatial selectivity algorithm, the image composition algorithm, or the object count algorithm.
  10. The non-transitory computer readable medium of claim 8, further comprising instructions that, when executed by the processor, cause the computer device to receive, from the user client device: a second image mask that indicates a portion of the second query digital image to emphasize in identifying similar digital images.
  11. The non-transitory computer readable medium of claim 10, further comprising instructions that, when executed by the processor, cause the computer device to modify, in response to receiving the second image mask, the multi-query vector representation to emphasize the portion of the second query digital image indicated by the second image mask.
  12. In a digital medium environment for searching for digital images, a system for matching digital images based on visual similarity comprising: a processor; and a non-transitory computer readable medium comprising instructions that, when executed by the processor, cause the system to: receive, from a user client device, a user selection of a query digital image and at least one of spatial selectivity, image composition, or object count to use to identify similar digital images; receive, from the user client device, an image mask that indicates a portion of the query digital image to emphasize in identifying similar digital images and an image mask weight corresponding to the image mask; utilize a deep neural network-based model to generate a set of deep features for the query digital image; generate, based on weighting the set of deep features of the query digital image utilizing the image mask and the image mask weight and in accordance with the user selection, a deep neural network-based representation of the query digital image by utilizing one or more of a spatial selectivity algorithm, an image composition algorithm, or an object count algorithm; determine, based on the deep neural network-based representation of the query digital image, similarity scores for a plurality of digital images within a digital image database; and identify, based on the deep neural network-based representation of the query digital image and further based on the similarity scores of the plurality of digital images, a similar digital image for the query digital image from the digital image database.
  13. The system of claim 12, wherein the instructions, when executed by the processor, cause the system to determine the similarity scores by comparing deep features of each of the plurality of digital images with deep features of the query digital image.
  14. The system of claim 13, further comprising instructions that, when executed by the processor, cause the system to: receive, from the user client device: an additional query digital image and an additional image mask that indicates a portion of the additional query digital image; and an indication to invert the additional image mask to emphasize portions of the additional query digital image that are outside the additional image mask in identifying similar digital images; and generate, in response to receiving the indication to invert the additional image mask, an additional deep neural network-based representation for the additional query digital image based on weighting the portion of the additional query digital image outside of the additional image mask.
  15. The system of claim 12, further comprising instructions that, when executed by the processor, cause the system to: receive, from the user client device, a selection of a second query digital image; utilize the deep neural network-based model to generate a second set of deep features for the second query digital image; generate, based on the second set of deep features of the second query digital image, a multi-query vector representation that represents a composite of the query digital image and the second query digital image; determine, based on the deep neural network-based representation of the second query digital image, similarity scores for plurality of digital images within a multi-query digital image database; and identify a similar digital image for the multi-query vector representation from within the digital image database.
  16. The system of claim 15, wherein the instructions cause the system to determine the similarity scores by comparing deep features of each of the plurality of digital images with deep features of the multi-query vector representation.
  17. The system of claim 16, further comprising instructions that, when executed by the processor, cause the system to: receive, from the user client device: a second image mask that indicates a portion of the second query digital image to further emphasize in identifying similar digital images; and a second image mask weight to apply to the second image mask of the second query digital image; and modify the multi-query vector representation to apply the second image mask weight to the portion of the second query digital image indicated by the second image mask.
  18. In a digital medium environment for searching digital images, a computer-implemented method for matching digital images based on visual similarity comprising: receiving, from a user client device, a user selection of a query digital image and at least one of spatial selectivity, image composition, or object count; and a step for identifying a similar digital image for the query digital image based on the user selection.
  19. The computer-implemented method of claim 18, further comprising receiving, from the user client device, a second query digital image.
  20. The computer-implemented method of claim 19, further comprising a step for identifying a similar digital image for a compound feature vector representing visual attributes for the query digital image and visual attributes for the second query digital image.

Description

Advancements in computing devices and image analysis techniques have led to a variety of innovations in identifying digital images that are visually similar. For example, image analysis systems are now able to analyze high-resolution digital images to identify objects within the images and search through terabytes of information stored in digital image databases to identify other digital images that depict the same or similar objects.

Despite these advances however, conventional image analysis systems continue to suffer from a number of disadvantages, particularly in the accuracy and flexibility of identifying similar digital images. For instance, while conventional image analysis systems can identify the same objects in two different digital images, these systems often disregard other aspects of the images (e.g., backgrounds, spatial arrangement of objects, and other visual attributes of the images). Indeed, because conventional image analysis systems often rely solely on semantic content to classify images based on various image tags, these systems are too object-focused in their analysis. As a result, conventional image analysis systems often produce inaccurate results when determining the visual similarity of two images. This is a particularly significant problem because, due to this inaccuracy, users of conventional image analysis systems are often required to spend an inordinate amount of time performing excessive user actions searching through match results before locating desirable image matches.

Citations (5)

  • US8165407B1
  • US20100111396A1
  • US20140108016A1
  • US20170097948A1
  • US20170262479A1
Record as JSON
{
  "publication_number": "US10628708B2",
  "country": "US",
  "kind": "B2",
  "title": "Utilizing a deep neural network-based model to identify visually similar digital images based on user-selected visual attributes",
  "abstract": "The present disclosure relates to systems, methods, and non-transitory computer readable media for utilizing a deep neural network-based model to identify similar digital images for query digital images. For example, the disclosed systems utilize a deep neural network-based model to analyze query digital images to generate deep neural network-based representations of the query digital images. In addition, the disclosed systems can generate results of visually-similar digital images for the query digital images based on comparing the deep neural network-based representations with representations of candidate digital images. Furthermore, the disclosed systems can identify visually similar digital images based on user-defined attributes and image masks to emphasize specific attributes or portions of query digital images.",
  "claims": [
    "1. In a digital medium environment for searching digital images, a non-transitory computer readable medium for matching digital images based on visual similarity comprising instructions that, when executed by a processor, cause a computer device to: receive, from a user client device, a user selection of a query digital image and at least one of spatial selectivity, image composition, or object count to use to identify similar digital images; receive, from the user client device, an image mask that indicates a portion of the query digital image to emphasize in identifying similar digital images and an image mask weight corresponding to the image mask; utilize a deep neural network-based model to generate a set of deep features for the query digital image; generate, based on weighting the set of deep features of the query digital image utilizing the image mask and the image mask weight and in accordance with the user selection, a deep neural network-based representation of the query digital image by utilizing one or more of a spatial selectivity algorithm, an image composition algorithm, or an object count algorithm; and based on the deep neural network-based representation of the query digital image, identify, from a digital image database, a similar digital image for the query digital image.",
    "2. The non-transitory computer readable medium of claim 1, further comprising instructions that, when executed by the processor, cause the computer device to train the deep neural network-based model to generate deep features for digital images.",
    "3. The non-transitory computer readable medium of claim 1, wherein the instructions, when executed by the processor, cause the computer device to generate the deep neural network-based representation by modifying the set of deep features by utilizing a deep neural network-based representation comprising at least one of the spatial selectivity algorithm, the image composition algorithm, or the object count algorithm.",
    "4. The non-transitory computer readable medium of claim 3, further comprising instructions that, when executed by the processor, cause the computer device to receive, from the user client device: an additional query digital image and an additional image mask that indicates a portion of the additional query digital image; and an indication to invert the additional image mask to emphasize portions of the additional query digital image that are outside the additional image mask in identifying similar digital images.",
    "5. The non-transitory computer readable medium of claim 4, further comprising instructions that, when executed by the processor, cause the computer device to generate, based on receiving the additional query digital image, the additional image mask, and the indication to invert the additional image mask, an additional deep neural network-based representation for the additional query digital image based on weighting the portion of the additional query digital image outside of the additional image mask.",
    "6. The non-transitory computer readable medium of claim 5, wherein weighting the portion of the additional query digital image outside of the additional image mask comprises applying weights to features of a set of features that correspond to the portion of the additional query digital image outside of the additional image mask.",
    "7. The non-transitory computer readable medium of claim 1, further comprising instructions that, when executed by the processor, cause the computer device to: determine similarity scores for a plurality of digital images from the digital image database; and rank the plurality of digital images based on the determined similarity scores.",
    "8. The non-transitory computer readable medium of claim 1, further comprising instructions that, when executed by the processor, cause the computer device to: receive, from the user client device, a selection of a second query digital image; utilize the deep neural network-based model to generate a second set of deep features for the second query digital image; generate, based on the second set of deep features of the second query digital image, a multi-query vector representation that represents a composite of the query digital image and the second query digital image; and identify a similar digital image for the multi-query vector representation from within the digital image database.",
    "9. The non-transitory computer readable medium of claim 8, further comprising instructions that, when executed by the processor, cause the computer device to: receive, in relation to the second query digital image, a second user selection of at least one of spatial selection, image composition, or object count; and wherein the instructions cause the computer device to generate, based on the second user selection, the multi-query vector representation by utilizing at least one of the spatial selectivity algorithm, the image composition algorithm, or the object count algorithm.",
    "10. The non-transitory computer readable medium of claim 8, further comprising instructions that, when executed by the processor, cause the computer device to receive, from the user client device: a second image mask that indicates a portion of the second query digital image to emphasize in identifying similar digital images.",
    "11. The non-transitory computer readable medium of claim 10, further comprising instructions that, when executed by the processor, cause the computer device to modify, in response to receiving the second image mask, the multi-query vector representation to emphasize the portion of the second query digital image indicated by the second image mask.",
    "12. In a digital medium environment for searching for digital images, a system for matching digital images based on visual similarity comprising: a processor; and a non-transitory computer readable medium comprising instructions that, when executed by the processor, cause the system to: receive, from a user client device, a user selection of a query digital image and at least one of spatial selectivity, image composition, or object count to use to identify similar digital images; receive, from the user client device, an image mask that indicates a portion of the query digital image to emphasize in identifying similar digital images and an image mask weight corresponding to the image mask; utilize a deep neural network-based model to generate a set of deep features for the query digital image; generate, based on weighting the set of deep features of the query digital image utilizing the image mask and the image mask weight and in accordance with the user selection, a deep neural network-based representation of the query digital image by utilizing one or more of a spatial selectivity algorithm, an image composition algorithm, or an object count algorithm; determine, based on the deep neural network-based representation of the query digital image, similarity scores for a plurality of digital images within a digital image database; and identify, based on the deep neural network-based representation of the query digital image and further based on the similarity scores of the plurality of digital images, a similar digital image for the query digital image from the digital image database.",
    "13. The system of claim 12, wherein the instructions, when executed by the processor, cause the system to determine the similarity scores by comparing deep features of each of the plurality of digital images with deep features of the query digital image.",
    "14. The system of claim 13, further comprising instructions that, when executed by the processor, cause the system to: receive, from the user client device: an additional query digital image and an additional image mask that indicates a portion of the additional query digital image; and an indication to invert the additional image mask to emphasize portions of the additional query digital image that are outside the additional image mask in identifying similar digital images; and generate, in response to receiving the indication to invert the additional image mask, an additional deep neural network-based representation for the additional query digital image based on weighting the portion of the additional query digital image outside of the additional image mask.",
    "15. The system of claim 12, further comprising instructions that, when executed by the processor, cause the system to: receive, from the user client device, a selection of a second query digital image; utilize the deep neural network-based model to generate a second set of deep features for the second query digital image; generate, based on the second set of deep features of the second query digital image, a multi-query vector representation that represents a composite of the query digital image and the second query digital image; determine, based on the deep neural network-based representation of the second query digital image, similarity scores for plurality of digital images within a multi-query digital image database; and identify a similar digital image for the multi-query vector representation from within the digital image database.",
    "16. The system of claim 15, wherein the instructions cause the system to determine the similarity scores by comparing deep features of each of the plurality of digital images with deep features of the multi-query vector representation.",
    "17. The system of claim 16, further comprising instructions that, when executed by the processor, cause the system to: receive, from the user client device: a second image mask that indicates a portion of the second query digital image to further emphasize in identifying similar digital images; and a second image mask weight to apply to the second image mask of the second query digital image; and modify the multi-query vector representation to apply the second image mask weight to the portion of the second query digital image indicated by the second image mask.",
    "18. In a digital medium environment for searching digital images, a computer-implemented method for matching digital images based on visual similarity comprising: receiving, from a user client device, a user selection of a query digital image and at least one of spatial selectivity, image composition, or object count; and a step for identifying a similar digital image for the query digital image based on the user selection.",
    "19. The computer-implemented method of claim 18, further comprising receiving, from the user client device, a second query digital image.",
    "20. The computer-implemented method of claim 19, further comprising a step for identifying a similar digital image for a compound feature vector representing visual attributes for the query digital image and visual attributes for the second query digital image."
  ],
  "description_excerpt": "Advancements in computing devices and image analysis techniques have led to a variety of innovations in identifying digital images that are visually similar. For example, image analysis systems are now able to analyze high-resolution digital images to identify objects within the images and search through terabytes of information stored in digital image databases to identify other digital images that depict the same or similar objects.\n\nDespite these advances however, conventional image analysis systems continue to suffer from a number of disadvantages, particularly in the accuracy and flexibility of identifying similar digital images. For instance, while conventional image analysis systems can identify the same objects in two different digital images, these systems often disregard other aspects of the images (e.g., backgrounds, spatial arrangement of objects, and other visual attributes of the images). Indeed, because conventional image analysis systems often rely solely on semantic content to classify images based on various image tags, these systems are too object-focused in their analysis. As a result, conventional image analysis systems often produce inaccurate results when determining the visual similarity of two images. This is a particularly significant problem because, due to this inaccuracy, users of conventional image analysis systems are often required to spend an inordinate amount of time performing excessive user actions searching through match results before locating desirable image matches.",
  "cpc": [
    "G06V 20/00",
    "G06F 16/532",
    "G06F 16/5854",
    "G06F 18/211",
    "G06F 18/214",
    "G06F 18/2148",
    "G06F 18/22",
    "G06F 18/2413",
    "G06F 18/41",
    "G06K 9/52",
    "G06K 9/6254",
    "G06K 9/6257",
    "G06N 3/045",
    "G06T 2207/20081",
    "G06T 2207/20084",
    "G06T 7/75",
    "G06V 10/454",
    "G06V 10/761",
    "G06V 10/764",
    "G06V 10/771",
    "G06V 10/82"
  ],
  "ipc": [
    "G06T 7/73",
    "G06V 10/764",
    "G06V 10/771",
    "G06V 20/00"
  ],
  "assignees": [
    "Adobe Inc"
  ],
  "inventors": [
    "Zhe Lin",
    "Xiaohui SHEN",
    "Mingyang Ling",
    "Jianming Zhang",
    "Jason Kuen",
    "Brett Butterfield"
  ],
  "filing_date": "2018-05-18",
  "publication_date": "2020-04-21",
  "grant_date": "2020-04-21",
  "priority_date": "2018-05-18",
  "application_number": "US-201815983949-A",
  "family_id": "65996870",
  "cited_by_count": 7,
  "citations": [
    "US8165407B1",
    "US20100111396A1",
    "US20140108016A1",
    "US20170097948A1",
    "US20170262479A1"
  ]
}

Record 2,205 of 8,000 in Patents full text (MLC-0201). Request the full dataset.