DETAILED ACTION
This action is in response to amendments filed 05/01/2026. Claims 1, 3-14, and 17-23 are pending and have been examined.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of the first paragraph of 35 U.S.C. 112(a):
(a) IN GENERAL.—The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor or joint inventor of carrying out the invention.
The following is a quotation of the first paragraph of pre-AIA 35 U.S.C. 112:
The specification shall contain a written description of the invention, and of the manner and process of making and using it, in such full, clear, concise, and exact terms as to enable any person skilled in the art to which it pertains, or with which it is most nearly connected, to make and use the same, and shall set forth the best mode contemplated by the inventor of carrying out his invention.
Claims 1, 3-14, and 17-23 are rejected under 35 U.S.C. 112(a) or 35 U.S.C. 112 (pre-AIA ), first paragraph, as failing to comply with the written description requirement. The claim(s) contains subject matter which was not described in the specification in such a way as to reasonably convey to one skilled in the relevant art that the inventor or a joint inventor, or for applications subject to pre-AIA 35 U.S.C. 112, the inventor(s), at the time the application was filed, had possession of the claimed invention.
Claim 1 recites “where each of the support images have at least the first label or the second label in common with the query image by randomly selecting each of the plurality of support images from an image group stored in an image database” in its sixth limitation. While the instant specification discloses a set of support images having at least one label in common with the query image (paragraph [0096] of the published application), it fails to disclose a scenario where every single support image has at least one label in common with the query image. Thus, claim 1 contains new matter not described in the instant specification, and fails to comply with the written description requirement. This deficiency is present in substantially similar independent claims 12 and 13, and is inherited by all dependent claims.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claims 1, 4, 6, 8, 11-14, 17-19, and 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Karlinsky (RepMet: Representative-based metric learning for classification and few-shot object detection, 2019, CVPR 2019 pp. 5192-5201) in view of Hu (COMPUTER VISION IMAGE FEATURE IDENTIFICATION VIA MULTI-LABEL FEW-SHOT MODEL, filed 11/27/2019, US 2021/0142054 A1), and further in view of Bayar (A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer, published 2016, IH&MMSec '16: Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security pp. 5-10), and Hwang (ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION, published 12/19/2019, ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION.
Regarding claim 1, Karlinsky discloses [a] learning system for assigning multiple labels to a query image, comprising at least one processor configured to:
receive the query image: “Each episode, for the case of the n-shot, m-way few-shot detection task, contains random n training examples for each of the m randomly chosen classes, and 10 ·m random query images containing one or more instances belonging to these classes” (Karlinsky, page 5202, left column, paragraph 3)
determine a first label for the query image; determine a second label for the query image:
“For few-shot object detection, we build upon modern approaches … to generate regions of interest, and a classifier ‘head’ that classifies these ROIs into one of the object categories or a background region. The input to this subnet are the feature vectors pooled from the ROIs, and the class posteriors (class label[s]) for a given ROI are computed by comparing its embedding vector to the set of representatives for each category.” (Karlinsky, page 5198, left column, paragraph 1)
“we will refer to the input of the subnet as a single (pooled) feature vector X ∈ Rf computed by the backbone for the given image (query image) (or ROI).” (Karlinsky, page 5199, right column, paragraph 4)
PNG
media_image1.png
415
426
media_image1.png
Greyscale
”Figure 2. Overview of our approach. (a) Train time: backbone, embedding space, and mixture models for the classes are learned jointly” (Karlinsky, page 5198, Figure 2). An instance where the query image is mapped to a first label (dog) and a second label (a bicycle).
input the query image into a learning model and obtain an output of the learning model: “We propose a subnet architecture and corresponding losses that allow us to train a DML embedding jointly with the multi-modal mixture distribution used for computing the class posterior in the resulting embedding space. This subnet then becomes a DML-based classifier head, which can be attached on top of a classification or a detection backbone. It is important to note that our DML-subnet (learning model) is trained jointly with the feature producing backbone. The architecture of the proposed subnet is depicted in Figure 3” (Karlinsky, page 5199, right column, paragraph 3)“we will refer to the input of the subnet (learning model) as a single (pooled) feature vector X ∈
R
f
computed by the backbone for the given image (query image) (or ROI)” (Karlinsky, page 5199, right column, paragraph 4); “For a given image (or detector ROI) and its corresponding embedding vector E, our network (learning model) computes a matrix of N × K distances
d
i
j
E
=
d
(
E
,
R
i
j
)
between E and the representatives
R
i
j
. These distances are used to compute the probability of the given image (or ROI) in each mode j of each class i:
PNG
media_image2.png
71
258
media_image2.png
Greyscale
” (Karlinsky, page 5200, left column, paragraph 2); “we do not learn the mixing coefficients and set the discriminative class posterior (output) to be:
PNG
media_image3.png
57
563
media_image3.png
Greyscale
” (Karlinsky, page 5200, left column, paragraph 3)
calculate a first loss by calculating a difference between the output of the learning model and a target output: “Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image representation) computed by the backbone for the given image (query image) (or ROI)” (Karlinsky, page 5199, right column, paragraph 4); “Having P(C = i|X) and P(B|X) (output of the learning model) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone). The first loss is the regular cross-entropy (CE) with the ground truth labels (target output) given for the image (or ROI) corresponding to X” (Karlinsky, page 5200, right column, paragraph 3). As one of ordinary skill in the art would know, cross-entropy determines a difference between an observed output and a desired target output. Note for further mappings below, “query image” and values derived from the query image, such as regions of interest and pooled feature vectors computed from the image, are used somewhat interchangeably.
acquire support data comprising a plurality of support images:
PNG
media_image1.png
415
426
media_image1.png
Greyscale
”Train time: backbone, embedding space and mixture models for the classes are learned jointly, class representatives (support images) are mixture mode centers in the embedding space” (Karlinsky, page 5198, Figure 2). Representatives, each a mode of a distribution representing an image class, can be represented as images, as seen in Figure 2.
“We represent each class by a mixture model with multiple modes, and consider the centers of these modes as the representative vectors (support images) for the class. Unlike previous methods, we simultaneously learn (acquire) the embedding space, backbone network parameters, and the representative vectors of the training categories, in a single end-to-end training process.” (Karlinsky, page 5197, right column, paragraph 2).
… where each of the support images have at least the first label or the second label in common with the query image:
PNG
media_image1.png
415
426
media_image1.png
Greyscale
(Karlinsky, page 5198, Figure 2). In this instance, each representative (support image) represents one label shared with the query image (which has three labels: dog, bike, and truck).
acquire a feature amount of the query image
PNG
media_image4.png
315
524
media_image4.png
Greyscale
(Karlinsky, page 5200, Figure 3). The DML embedding module creates an embedding vector (feature amount) of the feature vector input (query image).
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module, which consists of a few fully connected (FC) layers with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
(feature amount)” (Karlinsky, page 5199, right column, paragraph 4). As indicated by paragraph [0087] of the published application, a feature amount can be an embedding vector.
… and a feature amount of each of the plurality of support images corresponding to the query image: “As an additional set of trained parameters, we hold a set of ‘representatives’
R
i
j
∈
R
e
(feature amount of each of the plurality of support images). Each vector
R
i
j
represents the center of the j-th mode of the learned discriminative mixture distribution in the embedding space, for the i-th class out of the total of N classes. We assume a fixed number of K modes (peaks) in the distribution of each class, so 1 ≤ j ≤ K” (Karlinsky, page 5200, left column, paragraph 1). Each support image has a label corresponding to the query image.
…which are calculated based on a parameter of the learning model:
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module, which consists of a few fully connected (FC) layers (parameter[s]) with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
(feature amount of the query image)” (Karlinsky, page 5199, right column, paragraph 4).
“In our implementation, the representatives (feature amount of each of plurality of support images) are realized as weights (parameter[s]) of an FC layer of size N ·K · e receiving a fixed scalar input 1. The output of this layer is reshaped to an N× K×e tensor. During training, this simple construction flows the gradients to the weights of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2);
calculate a second loss by calculating a contrastive loss between the feature amount of the query image and the feature amount of each of the plurality of support images such that as the contrastive loss becomes larger, the second loss becomes larger: “For a given image (or detector ROI) and its corresponding embedding vector E, our network computes a matrix of N × K distances
d
i
j
(E) = d(E,
R
i
j
) between E (feature amount of the query image) and the representatives
R
i
j
(feature amount of the plurality of support images)” (Karlinsky, page 5200, left column, paragraph 2); “we use a sum of two losses to train our model ... The other (second loss / contrastive loss) is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
where i∗ is the correct class index for the current example” (Karlinsky, page 5200, right column, paragraph 3). d, a measure of distance between the feature amounts of the query data and feature amounts of the support data, is used to calculate this second loss. Since the distance function is minimized over all representatives of all classes, distance between the query embedding and each of the plurality of support images must be calculated.
adjust the parameter based on the first loss and the second loss: “During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
Karlinsky relates to few-shot learning for multi-label image processing systems and is analogous to the claimed invention.
While Karlinsky fails to disclose the further limitations of the claim, Hu teaches [a] learning system for assigning multiple labels to a query image, comprising at least one processor:
“[0012] Embodiments disclosed herein include a computer vision model that identifies a combination of graphic elements present in a query image based on a support set of images that include other various combinations of the graphic features.” (Hu, [0012])
“Software or firmware to implement the techniques introduced here may be stored on a machine-readable storage medium and may be executed by one or more general-purpose or special-purpose programmable microprocessors. A "machine-readable medium", as the term is used herein, includes any mechanism that can store information in a form accessible by a machine (a machine may be, for example, a computer, network device, cellular phone, personal digital assistant (PDA), manufacturing tool, any device with one or more processors, etc.)” (Hu, [0058]).
Hu relates to few-shot learning for image processing systems and is analogous to the claimed invention. Karlinsky teaches a method of training an image classifier with few-shot learning. The claimed invention improves upon this method by executing it with computer processors. Hu teaches a method of training an image classifier with few-shot learning which can be executed on computer processors, applicable to Karlinsky. A person of ordinary skill in the art would have recognized that running Karlinsky’s method on Hu’s processor hardware would lead to the predictable result of the method being performed by a computer as-described, and would improve the known device by allowing it to be used to process real data on a computer (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
While Karlinsky and Hu fail to disclose the further limitations of the claim, Bayar discloses a system, wherein each of the first label and the second label of the query image are related to an edited feature of the query image: “In this paper, we proposed a novel CNN-based universal forgery detection technique that can automatically learn how to detect different image manipulations (edit[s]). To prevent the CNN from learning features that represent an image's content, we proposed a new form of convolutional specifically designed to suppress an image's content and learn manipulation detection features. We accomplished this by specifically constraining this new convolutional layer to learn prediction error filters. Through a series of experiments, we demonstrated that our CNN-based universal forensic approach can automatically learn how to detect multiple image manipulations (multiple label[s]) without relying on pre-selected features or any preprocessing.” (Bayar, page 7, left column, paragraph 4).
Bayar relates to machine learning for image analysis and is analogous to the claimed invention. The existing combination teaches a system that generates losses based on similarities of query images and support images, and based on model output vs. intended output. Bayar teaches a method of detecting edits in images. It would have been obvious to one of ordinary skill in the art to combine the existing combination with Bayar by detecting edits in the query images with Bayar’s method. This would achieve the predictable result of identifying whether and how a query image has been manipulated, with the system of the existing combination method performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
While Bayar fails to disclose the further limitations of the claim, Hwang discloses a method of randomly selecting each of the plurality of support images from an image group stored in an image database:
“As shown, the query image and the support image may be images of the database” (Hwang, [0253])
“A predetermined number of the query images and the support images are randomly classified in the database DB image including the previously acquired images and the currently acquired images” (Hwang, [0254])
“Hereinafter, the learning process using the object image will be described in detail using FIGS. 9 to 12. FIG. 9 shows one example of images in database DB. FIG. 10 shows a query image and FIG. 11 shows one example of a support image” (Hwang, [0251])
PNG
media_image6.png
226
625
media_image6.png
Greyscale
(Hwang, Figure 9)
PNG
media_image7.png
362
345
media_image7.png
Greyscale
(Hwang, Figure 10)
PNG
media_image8.png
362
349
media_image8.png
Greyscale
(Hwang, Figure 11)
Hwang relates to few-shot learning with support data randomly sampled from a database and is analogous to the claimed invention. The existing combination teaches a system that generates a set of support images. The claimed invention improves upon this method by randomly sampling support images from a database. Hwang teaches a method of randomly sampling support images from a database, applicable to the existing combination. A person of ordinary skill in the art would have recognized that storing support image data on a database and retrieving it with random sampling would lead to the predictable result of increasing accessibility of the support dataset and create variability through randomness during training, and would improve the known device by reducing overfit during training and enabling collating of support data from multiple sources (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
Regarding claim 4, the rejection of claim 1 is incorporated. Karlinsky, in combination with Hu, further teaches a system, wherein the at least one processor is configured to
calculate a total loss based on the first loss and the second loss: (Karlinsky) “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
adjust the parameter based on the total loss: (Karlinsky) “During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2).
Regarding claim 6, the rejection of claim 1 is incorporated. Karlinsky, in combination with Hu further teaches a method, wherein
the learning model is configured to recognize three or more labels, which can be arranged in unique combinations: (Karlinsky) “we trained our DML classifier on the ImageNet Attributes dataset defined in [25], which contains 116236 images from 90 classes (labels)” (Karlinsky, page 5201, right column, paragraph 4). In this instance, a combination of 90 classes is being used for training.
for each unique combination of the three or more labels, a data set which includes the query image and the support data exists: (Karlinsky)
PNG
media_image9.png
563
957
media_image9.png
Greyscale
(Karlinsky, page 5200, Figure 3). This is part of a diagram showing Karlinsky’s training architecture. As shown here, an input (query image) and a set of representatives (support data) exist in the training scheme, supporting a variable number of labels (N classes).
PNG
media_image10.png
724
782
media_image10.png
Greyscale
“Overview of our approach. (a) Train time: backbone, embedding space and mixture models for the classes are learned jointly, class representatives are mixture mode centers in the embedding space” (Karlinsky, page 5198, Figure 2). Figure 2(a) shows the process of mapping classes to an embedding space during training. This training methodology supports at least one unique combination of three labels (Bike class, dog class, and truck class).
the at least one processor is configured to
calculate, for each unique combination of the three or more labels, the first loss based on the query image corresponding to the combination: (Karlinsky) “Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI)” (Karlinsky, page 5199, right column, paragraph 4); “The first loss is the regular cross-entropy (CE) with the ground truth labels given for the image (or ROI) corresponding to X” (Karlinsky, page 5200, right column, paragraph 3).
acquire, for each unique combination of the three or more labels, the feature amount of the query image corresponding to the combination: (Karlinsky)
PNG
media_image9.png
563
957
media_image9.png
Greyscale
(Karlinsky, page 5200, Figure 3). The DML embedding module creates an embedding vector (feature amount) of the feature vector input (query image). This corresponds to the N classes of the representative set, against which it will be compared.
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query data) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module, which consists of a few fully connected (FC) layers with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
} (feature amount), where commonly embedding size e ≪ f” (Karlinsky, page 5199, right column, paragraph 4). As indicated by page 25 line 23 to page 26 line 3 of the instant Specification, a feature amount can be an embedding vector.
acquire, for each combination of the three or more labels, ... the feature amount of the support images corresponding to the combination: (Karlinsky) “As an additional set of trained parameters, we hold a set of ‘representatives’
R
i
j
∈
R
e
(feature amount). Each vector
R
i
j
represents the center of the j-th mode of the learned discriminative mixture distribution (support image) in the embedding space, for the i-th class out of the total of N classes. We assume a fixed number of K modes (peaks) in the distribution of each class, so 1 ≤ j ≤ K” (Karlinsky, page 5200, left column, paragraph 1).
calculate, for each unique combination of the three or more labels, the second loss based on the feature amount of the query image corresponding to the combination and the feature amount of the support images corresponding to the combination: (Karlinsky) “For a given image (or detector ROI) and its corresponding embedding vector E, our network computes a matrix of N × K distances
d
i
j
(E) = d(E,
R
i
j
) between E (feature amount of the query data) and the representatives
R
i
j
(feature amount of the support images)” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone) ... The other (second loss) is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
” (Karlinsky, page 5200, right column, paragraph 3). d, a measure of distance between the feature amounts of the query data and feature amounts of the support data, is used to calculate this second loss.
adjust the parameter based on the first loss and the second loss which are calculated for each unique combination of the three or more labels: (Karlinsky) “During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
Regarding claim 8, the rejection of claim 1 is incorporated. Karlinsky, in combination with Hu, further teaches a system, wherein the at least one processor is configured to acquire the second loss based on the feature amount of the query image, the feature amount of the support data, and a coefficient corresponding to a label similarity between the query image and the support images: (Karlinsky) “For a given image (or detector ROI) and its corresponding embedding vector E, our network computes a matrix of N × K distances
d
i
j
(E) = d(E,
R
i
j
) (label similarity) between E (feature amount of the query image) and the representatives
R
i
j
(feature amount of the support images)” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone) ... The other (second loss) is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
where
i
*
is the correct class index for the current example and | · |+ is the ReLU function.” (Karlinsky, page 5200, right column, paragraph 3). d measures the distance between the query embedding and a support embedding. It’s minimized for classes that include the query and maximized for classes that don’t, thus it can be considered a measure of label similarity.
Regarding claim 11, the rejection of claim 1 is incorporated. Karlinsky further teaches a method, wherein
the learning model is a model for recognizing an object included in an image: “we demonstrate the effectiveness of our approach on the problem of few-shot object detection, by incorporating the proposed DML architecture as a classification head into a standard object detection model. We achieve the best results on the ImageNet-LOC dataset” (Karlinsky, page 5197, Abstract).
the query image is a multi-label query image:
PNG
media_image1.png
415
426
media_image1.png
Greyscale
”Figure 2. Overview of our approach. (a) Train time: backbone, embedding space, and mixture models for the classes are learned jointly” (Karlinsky, page 5198, Figure 2). An instance where the query image is mapped to a first label (dog) and a second label (a bicycle).
the support data is a support image corresponding to the multi-label query image: ”Train time: backbone, embedding space and mixture models for the classes are learned jointly, class representatives (support images) are mixture mode centers in the embedding space” (Karlinsky, page 5198, Figure 2); “We represent each class by a mixture model with multiple modes, and consider the centers of these modes as the representative vectors (support images) for the class. Unlike previous methods, we simultaneously learn (acquire) the embedding space, backbone network parameters, and the representative vectors of the training categories, in a single end-to-end training process.” (Karlinsky, page 5197, right column, paragraph 2).
Regarding claim 12, Karlinsky teaches a learning method, comprising:
receiving the query image: “Each episode, for the case of the n-shot, m-way few-shot detection task, contains random n training examples for each of the m randomly chosen classes, and 10 ·m random query images containing one or more instances belonging to these classes” (Karlinsky, page 5202, left column, paragraph 3)
determining a first label for the query image; determining a second label for the query image:
“For few-shot object detection, we build upon modern approaches … to generate regions of interest, and a classifier ‘head’ that classifies these ROIs into one of the object categories or a background region. The input to this subnet are the feature vectors pooled from the ROIs, and the class posteriors (class label[s]) for a given ROI are computed by comparing its embedding vector to the set of representatives for each category.” (Karlinsky, page 5198, left column, paragraph 1)
“we will refer to the input of the subnet as a single (pooled) feature vector X ∈ Rf computed by the backbone for the given image (query image) (or ROI).” (Karlinsky, page 5199, right column, paragraph 4)
PNG
media_image1.png
415
426
media_image1.png
Greyscale
”Figure 2. Overview of our approach. (a) Train time: backbone, embedding space, and mixture models for the classes are learned jointly” (Karlinsky, page 5198, Figure 2). An instance where the query image is mapped to a first label (dog) and a second label (a bicycle).
inputting the query image into a learning model and obtain an output of the learning model: “We propose a subnet architecture and corresponding losses that allow us to train a DML embedding jointly with the multi-modal mixture distribution used for computing the class posterior in the resulting embedding space. This subnet then becomes a DML-based classifier head, which can be attached on top of a classification or a detection backbone. It is important to note that our DML-subnet (learning model) is trained jointly with the feature producing backbone. The architecture of the proposed subnet is depicted in Figure 3” (Karlinsky, page 5199, right column, paragraph 3)“we will refer to the input of the subnet (learning model) as a single (pooled) feature vector X ∈
R
f
computed by the backbone for the given image (query image) (or ROI)” (Karlinsky, page 5199, right column, paragraph 4); “For a given image (or detector ROI) and its corresponding embedding vector E, our network (learning model) computes a matrix of N × K distances
d
i
j
E
=
d
(
E
,
R
i
j
)
between E and the representatives
R
i
j
. These distances are used to compute the probability of the given image (or ROI) in each mode j of each class i:
PNG
media_image2.png
71
258
media_image2.png
Greyscale
” (Karlinsky, page 5200, left column, paragraph 2); “we do not learn the mixing coefficients and set the discriminative class posterior (output) to be:
PNG
media_image3.png
57
563
media_image3.png
Greyscale
” (Karlinsky, page 5200, left column, paragraph 3)
calculate a first loss by calculating a difference between the output of the learning model and a target output: “Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image representation) computed by the backbone for the given image (query image) (or ROI)” (Karlinsky, page 5199, right column, paragraph 4); “Having P(C = i|X) and P(B|X) (output of the learning model) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone). The first loss is the regular cross-entropy (CE) with the ground truth labels (target output) given for the image (or ROI) corresponding to X” (Karlinsky, page 5200, right column, paragraph 3). As one of ordinary skill in the art would know, cross-entropy determines a difference between an observed output and a desired target output. Note for further mappings below, “query image” and values derived from the query image, such as regions of interest and pooled feature vectors computed from the image, are used somewhat interchangeably.
acquire support data comprising a plurality of support images:
PNG
media_image1.png
415
426
media_image1.png
Greyscale
”Train time: backbone, embedding space and mixture models for the classes are learned jointly, class representatives (support images) are mixture mode centers in the embedding space” (Karlinsky, page 5198, Figure 2). Representatives, each a mode of a distribution representing an image class, can be represented as images, as seen in Figure 2.
“We represent each class by a mixture model with multiple modes, and consider the centers of these modes as the representative vectors (support images) for the class. Unlike previous methods, we simultaneously learn (acquire) the embedding space, backbone network parameters, and the representative vectors of the training categories, in a single end-to-end training process.” (Karlinsky, page 5197, right column, paragraph 2).
… where each of the support images have at least the first label or the second label in common with the query image:
PNG
media_image1.png
415
426
media_image1.png
Greyscale
(Karlinsky, page 5198, Figure 2). In this instance, each representative (support image) represents one label shared with the query image (which has three labels: dog, bike, and truck).
acquiring a feature amount of the query image
PNG
media_image4.png
315
524
media_image4.png
Greyscale
(Karlinsky, page 5200, Figure 3). The DML embedding module creates an embedding vector (feature amount) of the feature vector input (query image).
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module, which consists of a few fully connected (FC) layers with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
(feature amount)” (Karlinsky, page 5199, right column, paragraph 4). As indicated by paragraph [0087] of the published application, a feature amount can be an embedding vector.
… and a feature amount of each of the plurality of support images corresponding to the query image: “As an additional set of trained parameters, we hold a set of ‘representatives’
R
i
j
∈
R
e
(feature amount of each of the plurality of support images). Each vector
R
i
j
represents the center of the j-th mode of the learned discriminative mixture distribution in the embedding space, for the i-th class out of the total of N classes. We assume a fixed number of K modes (peaks) in the distribution of each class, so 1 ≤ j ≤ K” (Karlinsky, page 5200, left column, paragraph 1). Each support image has a label corresponding to the query image.
…which are calculated based on a parameter of the learning model:
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module, which consists of a few fully connected (FC) layers (parameter[s]) with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
(feature amount of the query image)” (Karlinsky, page 5199, right column, paragraph 4).
“In our implementation, the representatives (feature amount of each of plurality of support images) are realized as weights (parameter[s]) of an FC layer of size N ·K · e receiving a fixed scalar input 1. The output of this layer is reshaped to an N× K×e tensor. During training, this simple construction flows the gradients to the weights of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2);
calculating a second loss by calculating a contrastive loss between the feature amount of the query image and the feature amount of each of the plurality of support images such that as the contrastive loss becomes larger, the second loss becomes larger: “For a given image (or detector ROI) and its corresponding embedding vector E, our network computes a matrix of N × K distances
d
i
j
(E) = d(E,
R
i
j
) between E (feature amount of the query image) and the representatives
R
i
j
(feature amount of the plurality of support images)” (Karlinsky, page 5200, left column, paragraph 2); “we use a sum of two losses to train our model ... The other (second loss / contrastive loss) is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
where i∗ is the correct class index for the current example” (Karlinsky, page 5200, right column, paragraph 3). d, a measure of distance between the feature amounts of the query data and feature amounts of the support data, is used to calculate this second loss. Since the distance function is minimized over all representatives of all classes, distance between the query embedding and each of the plurality of support images must be calculated.
adjusting the parameter based on the first loss and the second loss: “During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
Karlinsky relates to few-shot learning for multi-label image processing systems and is analogous to the claimed invention.
While Karlinsky fails to disclose the further limitations of the claim, Bayar discloses a system, wherein each of the first label and the second label of the query image are related to an edited feature of the query image: “In this paper, we proposed a novel CNN-based universal forgery detection technique that can automatically learn how to detect different image manipulations (edit[s]). To prevent the CNN from learning features that represent an image's content, we proposed a new form of convolutional specifically designed to suppress an image's content and learn manipulation detection features. We accomplished this by specifically constraining this new convolutional layer to learn prediction error filters. Through a series of experiments, we demonstrated that our CNN-based universal forensic approach can automatically learn how to detect multiple image manipulations (multiple label[s]) without relying on pre-selected features or any preprocessing.” (Bayar, page 7, left column, paragraph 4).
Bayar relates to machine learning for image analysis and is analogous to the claimed invention. The existing combination teaches a system that generates losses based on similarities of query images and support images, and based on model output vs. intended output. Bayar teaches a method of detecting edits in images. It would have been obvious to one of ordinary skill in the art to combine the existing combination with Bayar by detecting edits in the query images with Bayar’s method. This would achieve the predictable result of identifying whether and how a query image has been manipulated, with the system of the existing combination method performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
While Bayar fails to disclose the further limitations of the claim, Hwang discloses a method of randomly selecting each of the plurality of support images from an image group stored in an image database:
“As shown, the query image and the support image may be images of the database” (Hwang, [0253])
“A predetermined number of the query images and the support images are randomly classified in the database DB image including the previously acquired images and the currently acquired images” (Hwang, [0254])
“Hereinafter, the learning process using the object image will be described in detail using FIGS. 9 to 12. FIG. 9 shows one example of images in database DB. FIG. 10 shows a query image and FIG. 11 shows one example of a support image” (Hwang, [0251])
PNG
media_image6.png
226
625
media_image6.png
Greyscale
(Hwang, Figure 9)
PNG
media_image7.png
362
345
media_image7.png
Greyscale
(Hwang, Figure 10)
PNG
media_image8.png
362
349
media_image8.png
Greyscale
(Hwang, Figure 11)
Hwang relates to few-shot learning with support data randomly sampled from a database and is analogous to the claimed invention. The existing combination teaches a system that generates a set of support images. The claimed invention improves upon this method by randomly sampling support images from a database. Hwang teaches a method of randomly sampling support images from a database, applicable to the existing combination. A person of ordinary skill in the art would have recognized that storing support image data on a database and retrieving it with random sampling would lead to the predictable result of increasing accessibility of the support dataset and create variability through randomness during training, and would improve the known device by reducing overfit during training and enabling collating of support data from multiple sources (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
Regarding claim 13, Karlinsky teaches program instructions to:
receive the query image: “Each episode, for the case of the n-shot, m-way few-shot detection task, contains random n training examples for each of the m randomly chosen classes, and 10 ·m random query images containing one or more instances belonging to these classes” (Karlinsky, page 5202, left column, paragraph 3)
determine a first label for the query image; determine a second label for the query image:
“For few-shot object detection, we build upon modern approaches … to generate regions of interest, and a classifier ‘head’ that classifies these ROIs into one of the object categories or a background region. The input to this subnet are the feature vectors pooled from the ROIs, and the class posteriors (class label[s]) for a given ROI are computed by comparing its embedding vector to the set of representatives for each category.” (Karlinsky, page 5198, left column, paragraph 1)
“we will refer to the input of the subnet as a single (pooled) feature vector X ∈ Rf computed by the backbone for the given image (query image) (or ROI).” (Karlinsky, page 5199, right column, paragraph 4)
PNG
media_image1.png
415
426
media_image1.png
Greyscale
”Figure 2. Overview of our approach. (a) Train time: backbone, embedding space, and mixture models for the classes are learned jointly” (Karlinsky, page 5198, Figure 2). An instance where the query image is mapped to a first label (dog) and a second label (a bicycle).
input the query image into a learning model and obtain an output of the learning model: “We propose a subnet architecture and corresponding losses that allow us to train a DML embedding jointly with the multi-modal mixture distribution used for computing the class posterior in the resulting embedding space. This subnet then becomes a DML-based classifier head, which can be attached on top of a classification or a detection backbone. It is important to note that our DML-subnet (learning model) is trained jointly with the feature producing backbone. The architecture of the proposed subnet is depicted in Figure 3” (Karlinsky, page 5199, right column, paragraph 3)“we will refer to the input of the subnet (learning model) as a single (pooled) feature vector X ∈
R
f
computed by the backbone for the given image (query image) (or ROI)” (Karlinsky, page 5199, right column, paragraph 4); “For a given image (or detector ROI) and its corresponding embedding vector E, our network (learning model) computes a matrix of N × K distances
d
i
j
E
=
d
(
E
,
R
i
j
)
between E and the representatives
R
i
j
. These distances are used to compute the probability of the given image (or ROI) in each mode j of each class i:
PNG
media_image2.png
71
258
media_image2.png
Greyscale
” (Karlinsky, page 5200, left column, paragraph 2); “we do not learn the mixing coefficients and set the discriminative class posterior (output) to be:
PNG
media_image3.png
57
563
media_image3.png
Greyscale
” (Karlinsky, page 5200, left column, paragraph 3)
calculate a first loss by calculating a difference between the output of the learning model and a target output: “Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image representation) computed by the backbone for the given image (query image) (or ROI)” (Karlinsky, page 5199, right column, paragraph 4); “Having P(C = i|X) and P(B|X) (output of the learning model) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone). The first loss is the regular cross-entropy (CE) with the ground truth labels (target output) given for the image (or ROI) corresponding to X” (Karlinsky, page 5200, right column, paragraph 3). As one of ordinary skill in the art would know, cross-entropy determines a difference between an observed output and a desired target output. Note for further mappings below, “query image” and values derived from the query image, such as regions of interest and pooled feature vectors computed from the image, are used somewhat interchangeably.
acquire support data comprising a plurality of support images:
PNG
media_image1.png
415
426
media_image1.png
Greyscale
”Train time: backbone, embedding space and mixture models for the classes are learned jointly, class representatives (support images) are mixture mode centers in the embedding space” (Karlinsky, page 5198, Figure 2). Representatives, each a mode of a distribution representing an image class, can be represented as images, as seen in Figure 2.
“We represent each class by a mixture model with multiple modes, and consider the centers of these modes as the representative vectors (support images) for the class. Unlike previous methods, we simultaneously learn (acquire) the embedding space, backbone network parameters, and the representative vectors of the training categories, in a single end-to-end training process.” (Karlinsky, page 5197, right column, paragraph 2).
… where each of the support images have at least the first label or the second label in common with the query image:
PNG
media_image1.png
415
426
media_image1.png
Greyscale
(Karlinsky, page 5198, Figure 2). In this instance, each representative (support image) represents one label shared with the query image (which has three labels: dog, bike, and truck).
acquire a feature amount of the query image
PNG
media_image4.png
315
524
media_image4.png
Greyscale
(Karlinsky, page 5200, Figure 3). The DML embedding module creates an embedding vector (feature amount) of the feature vector input (query image).
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module, which consists of a few fully connected (FC) layers with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
(feature amount)” (Karlinsky, page 5199, right column, paragraph 4). As indicated by paragraph [0087] of the published application, a feature amount can be an embedding vector.
… and a feature amount of each of the plurality of support images corresponding to the query image: “As an additional set of trained parameters, we hold a set of ‘representatives’
R
i
j
∈
R
e
(feature amount of each of the plurality of support images). Each vector
R
i
j
represents the center of the j-th mode of the learned discriminative mixture distribution in the embedding space, for the i-th class out of the total of N classes. We assume a fixed number of K modes (peaks) in the distribution of each class, so 1 ≤ j ≤ K” (Karlinsky, page 5200, left column, paragraph 1). Each support image has a label corresponding to the query image.
…which are calculated based on a parameter of the learning model:
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module, which consists of a few fully connected (FC) layers (parameter[s]) with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
(feature amount of the query image)” (Karlinsky, page 5199, right column, paragraph 4).
“In our implementation, the representatives (feature amount of each of plurality of support images) are realized as weights (parameter[s]) of an FC layer of size N ·K · e receiving a fixed scalar input 1. The output of this layer is reshaped to an N× K×e tensor. During training, this simple construction flows the gradients to the weights of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2);
calculate a second loss by calculating a contrastive loss between the feature amount of the query image and the feature amount of each of the plurality of support images such that as the contrastive loss becomes larger, the second loss becomes larger: “For a given image (or detector ROI) and its corresponding embedding vector E, our network computes a matrix of N × K distances
d
i
j
(E) = d(E,
R
i
j
) between E (feature amount of the query image) and the representatives
R
i
j
(feature amount of the plurality of support images)” (Karlinsky, page 5200, left column, paragraph 2); “we use a sum of two losses to train our model ... The other (second loss / contrastive loss) is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
where i∗ is the correct class index for the current example” (Karlinsky, page 5200, right column, paragraph 3). d, a measure of distance between the feature amounts of the query data and feature amounts of the support data, is used to calculate this second loss. Since the distance function is minimized over all representatives of all classes, distance between the query embedding and each of the plurality of support images must be calculated.
adjust the parameter based on the first loss and the second loss: “During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
Karlinsky relates to few-shot learning for multi-label image processing systems and is analogous to the claimed invention.
While Karlinsky fails to disclose the further limitations of the claim, Hu teaches [a] non-transitory computer-readable information storage medium for storing a program for causing a computer to: “FIG. 8 is a high-level block diagram showing an example of a processing device 800 that can represent a system to run any of the methods/algorithms described above” (Hu, [0054]); “Physical and functional components ( e.g., devices, engines, modules, and data repositories, etc.) associated with processing device 800 can be implemented as circuitry, firmware, software, other executable instructions, or any combination thereof ... the functional components described can be implemented as instructions on a tangible storage memory capable of being executed by a processor or other integrated circuit chip ( e.g., software, software libraries, application program interfaces, etc.). The tangible storage memory can be computer readable data storage. The tangible storage memory may be volatile or non-volatile memory. In some embodiments, the volatile memory may be considered "non-transitory" in the sense that it is not a transitory signal. Memory space and storages described in the figures can be implemented with the tangible storage memory as well, including volatile or nonvolatile memory” (Hu, [0059]).
Hu relates to few-shot learning for image processing systems and is analogous to the claimed invention. Karlinsky teaches a method of training an image classifier with few-shot learning. The claimed invention improves upon this method by storing it as instructions on non-transitory computer media. Hu teaches a method of training an image classifier with few-shot learning which can be encoded in instructions and stored in non-transitory computer media, applicable to Karlinsky. A person of ordinary skill in the art would have recognized that storing Karlinsky’s method on Hu’s computer media would lead to the predictable result of the method being executed in volatile memory or stored to be executed at a later time in non-volatile memory, and would improve the known device by allowing it to be used to process real data on a computer (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
While Karlinsky and Hu fail to disclose the further limitations of the claim, Bayar discloses a system, wherein each of the first label and the second label of the query image are related to an edited feature of the query image: “In this paper, we proposed a novel CNN-based universal forgery detection technique that can automatically learn how to detect different image manipulations (edit[s]). To prevent the CNN from learning features that represent an image's content, we proposed a new form of convolutional specifically designed to suppress an image's content and learn manipulation detection features. We accomplished this by specifically constraining this new convolutional layer to learn prediction error filters. Through a series of experiments, we demonstrated that our CNN-based universal forensic approach can automatically learn how to detect multiple image manipulations (multiple label[s]) without relying on pre-selected features or any preprocessing.” (Bayar, page 7, left column, paragraph 4).
Bayar relates to machine learning for image analysis and is analogous to the claimed invention. The existing combination teaches a system that generates losses based on similarities of query images and support images, and based on model output vs. intended output. Bayar teaches a method of detecting edits in images. It would have been obvious to one of ordinary skill in the art to combine the existing combination with Bayar by detecting edits in the query images with Bayar’s method. This would achieve the predictable result of identifying whether and how a query image has been manipulated, with the system of the existing combination method performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
While Bayar fails to disclose the further limitations of the claim, Hwang discloses a method of randomly selecting each of the plurality of support images from an image group stored in an image database:
“As shown, the query image and the support image may be images of the database” (Hwang, [0253])
“A predetermined number of the query images and the support images are randomly classified in the database DB image including the previously acquired images and the currently acquired images” (Hwang, [0254])
“Hereinafter, the learning process using the object image will be described in detail using FIGS. 9 to 12. FIG. 9 shows one example of images in database DB. FIG. 10 shows a query image and FIG. 11 shows one example of a support image” (Hwang, [0251])
PNG
media_image6.png
226
625
media_image6.png
Greyscale
(Hwang, Figure 9)
PNG
media_image7.png
362
345
media_image7.png
Greyscale
(Hwang, Figure 10)
PNG
media_image8.png
362
349
media_image8.png
Greyscale
(Hwang, Figure 11)
Hwang relates to few-shot learning with support data randomly sampled from a database and is analogous to the claimed invention. The existing combination teaches a system that generates a set of support images. The claimed invention improves upon this method by randomly sampling support images from a database. Hwang teaches a method of randomly sampling support images from a database, applicable to the existing combination. A person of ordinary skill in the art would have recognized that storing support image data on a database and retrieving it with random sampling would lead to the predictable result of increasing accessibility of the support dataset and create variability through randomness during training, and would improve the known device by reducing overfit during training and enabling collating of support data from multiple sources (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
Regarding claim 14, the rejection of claim 1 is incorporated. Hu further discloses a system, wherein the learning model recognizes multi-label data of the query image:
PNG
media_image1.png
415
426
media_image1.png
Greyscale
”Figure 2. Overview of our approach. (a) Train time: backbone, embedding space, and mixture models for the classes are learned jointly” (Karlinsky, page 5198, Figure 2). An instance where the query image is mapped to a first label (dog) and a second label (a bicycle).
Regarding claim 17, the rejection of claim 6 is incorporated. Karlinsky, in combination with Hu, further teaches a method, wherein:
wherein the learning model is configured to recognize three or more labels, which can be arranged in unique combinations: (Karlinsky) “we trained our DML classifier on the ImageNet Attributes dataset defined in [25], which contains 116236 images from 90 classes (labels)” (Karlinsky, page 5201, right column, paragraph 4). In this instance, a combination of 90 classes is being used for training.
wherein, for every unique combination of the three or more labels, the data set which includes the query image and the support data exists: (Karlinsky)
PNG
media_image9.png
563
957
media_image9.png
Greyscale
(Karlinsky, page 5200, Figure 3). This is part of a diagram showing Karlinsky’s training architecture. As shown here, an input (query image) and a set of representatives (support data) exist in the training scheme, supporting one unique combination of N labels.
PNG
media_image10.png
724
782
media_image10.png
Greyscale
“Overview of our approach. (a) Train time: backbone, embedding space and mixture models for the classes are learned jointly, class representatives are mixture mode centers in the embedding space” (Karlinsky, page 5198, Figure 2). Figure 2(a) shows the process of mapping classes to an embedding space during training. This training methodology supports at least one unique combination of three labels (Bike class, dog class, and truck class).
wherein the at least one processor is configured to
calculate, for every unique combination of the three or more labels, the first loss based on the query image corresponding to the combination: (Karlinsky) “Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query data representation) computed by the backbone for the given image (or ROI)” (Karlinsky, page 5199, right column, paragraph 4); “The first loss is the regular cross-entropy (CE) with the ground truth labels given for the image (or ROI) corresponding to X” (Karlinsky, page 5200, right column, paragraph 3).
acquire, for every unique combination of the three or more labels, the feature amount of the query image corresponding to the combination: (Karlinsky)
PNG
media_image9.png
563
957
media_image9.png
Greyscale
(Karlinsky, page 5200, Figure 3). The DML embedding module creates an embedding vector (feature amount) of the feature vector input (query image).
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module, which consists of a few fully connected (FC) layers with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
(feature amount)” (Karlinsky, page 5199, right column, paragraph 4). As indicated by paragraph [0087] of the published application, a feature amount can be an embedding vector.
acquire, for every combination of the three or more labels, … the feature amount of the support images corresponding to the combination: (Karlinsky) “As an additional set of trained parameters, we hold a set of ‘representatives’
R
i
j
∈
R
e
(feature amount of each of the plurality of images of the support images). Each vector
R
i
j
represents the center of the j-th mode of the learned discriminative mixture distribution in the embedding space, for the i-th class out of the total of N classes. We assume a fixed number of K modes (peaks) in the distribution of each class, so 1 ≤ j ≤ K” (Karlinsky, page 5200, left column, paragraph 1). Each support image has a label corresponding to the multi-label query image.
calculate, for every unique combination of the three or more labels, the second loss based on the feature amount of the query image corresponding to the combination and the feature amount of the support images corresponding to the combination: (Karlinsky) “For a given image (or detector ROI) and its corresponding embedding vector E, our network computes a matrix of N × K distances
d
i
j
(E) = d(E,
R
i
j
) between E (feature amount of the multi-label query image) and the representatives
R
i
j
(feature amount of the plurality of images of the support images)” (Karlinsky, page 5200, left column, paragraph 2); “we use a sum of two losses to train our model ... The other (second loss) is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
where i∗ is the correct class index for the current example” (Karlinsky, page 5200, right column, paragraph 3).
adjust[ing] the parameter based on the first loss and the second loss which are calculated for every unique combination of the three or more labels: (Karlinsky) “During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
Regarding claim 18, the rejection of claim 1 is incorporated. Bayar further discloses a system, wherein the learning system makes a determination as to if the query image was edited: “In this paper, we proposed a novel CNN-based universal forgery detection technique that can automatically learn how to detect different image manipulations (edit[s]). To prevent the CNN from learning features that represent an image's content, we proposed a new form of convolutional specifically designed to suppress an image's content and learn manipulation detection features. We accomplished this by specifically constraining this new convolutional layer to learn prediction error filters. Through a series of experiments, we demonstrated that our CNN-based universal forensic approach can automatically learn how to detect multiple image manipulations (multiple label[s]) without relying on pre-selected features or any preprocessing.” (Bayar, page 7, left column, paragraph 4).
Bayar relates to machine learning for image analysis and is analogous to the claimed invention. The existing combination teaches a system that generates losses based on similarities of query images and support images, and based on model output vs. intended output. Bayar teaches a method of detecting edits in images. It would have been obvious to one of ordinary skill in the art to combine The existing combination by detecting edits in the query images with Bayar’s method. This would achieve the predictable result of identifying whether and how a query image has been manipulated, with the system of the existing combination method performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
Regarding claim 19, the rejection of claim 1 is incorporated. Bayar further discloses a system, wherein the learning system makes a determination as to how the query image was edited:
“In this paper, we proposed a novel CNN-based universal forgery detection technique that can automatically learn how to detect different image manipulations (edit[s]). To prevent the CNN from learning features that represent an image's content, we proposed a new form of convolutional specifically designed to suppress an image's content and learn manipulation detection features. We accomplished this by specifically constraining this new convolutional layer to learn prediction error filters. Through a series of experiments, we demonstrated that our CNN-based universal forensic approach can automatically learn how to detect multiple image manipulations (multiple label[s]) without relying on pre-selected features or any preprocessing.” (Bayar, page 7, left column, paragraph 4).
“In our first set of experiments, we trained different CNNs to detect each of the four manipulations discussed in Section 5.1. Each CNN corresponds to a binary classifier that detects one type of possible image operation with the same architecture outlined in Section 4.” (Bayar, page 6, left column, paragraph 5)
Bayar relates to machine learning for image analysis and is analogous to the claimed invention. The existing combination teaches a system that generates losses based on similarities of query images and support images, and based on model output vs. intended output. Bayar teaches a method of detecting edits in images. It would have been obvious to one of ordinary skill in the art to combine the existing combination and Hu by detecting edits in the query images with Bayar’s method. This would achieve the predictable result of identifying whether and how a query image has been manipulated, with the system of the existing combination method performing the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
Regarding claim 21, the rejection of claim 1 is incorporated. Karlinsky further discloses a learning system, wherein the at least one processor is configured to adjust the parameter based on the first loss and the second loss using an inverse propagation method or a gradient descent method such that the first loss and the second loss become smaller:
“During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
“we use a sum of two losses to train our model (DML subnet + backbone). The first loss is the regular cross-entropy (CE) with the ground truth labels given for the image (or ROI) corresponding to X. The other is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
where i∗ is the correct class index for the current example” (Karlinsky, page 5200, right column, paragraph 3).
Regarding claim 22, the rejection of claim 1 is incorporated. Karlinsky further discloses a system, wherein the at least one processor is configured to:
acquire the feature amount of the query image as an embedded vector by inputting the query image into the learning model:
PNG
media_image4.png
315
524
media_image4.png
Greyscale
(Karlinsky, page 5200, Figure 3). The DML embedding module (learning model component) creates an embedding vector (feature amount) of the feature vector input (query image).
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module (learning model component), which consists of a few fully connected (FC) layers with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
(embedded vector / feature amount)” (Karlinsky, page 5199, right column, paragraph 4). As indicated by paragraph [0087] of the published application, a feature amount can be an embedding vector.
acquire the feature amounts of each of the plurality of support images as a plurality of embedded vectors by inputting each of the plurality of support images into the learning model:
PNG
media_image11.png
551
659
media_image11.png
Greyscale
(Karlinsky, Figure 3)
“As an additional set of trained parameters, we hold a set of ‘representatives’
R
i
j
∈
R
e
(feature amount of each of the plurality of support images). Each vector
R
i
j
represents the center of the j-th mode of the learned discriminative mixture distribution in the embedding space, for the i-th class out of the total of N classes. We assume a fixed number of K modes (peaks) in the distribution of each class, so 1 ≤ j ≤ K” (Karlinsky, page 5200, left column, paragraph 1). Representatives output from the FC layer are embedding vectors in the same space as the embedded query images.
“For a given image (or detector ROI) and its corresponding embedding vector E, our network computes a matrix of N × K distances dij(E) = d(E,Rij) between E and the representatives Rij . These distances are used to compute the probability of the given image (or ROI) in each mode j of each class i” (Karlinsky, page 5200, left column, paragraph 2). Embedded representatives are input to the distance computational module of the model to be compared with embedded query images.
calculate the contrastive loss based on a Euclidean distance between the feature amount of the query image and the feature amount of each of the plurality of support images:
“In DML, the metric being learned is usually implemented as an L2 distance (Euclidean distance) between the samples in an embedding space generated by a neural network” (Karlinsky, page 5199, left column, paragraph 2)
“For a given image (or detector ROI) and its corresponding embedding vector E, our network computes a matrix of N × K distances
d
i
j
(E) = d(E,
R
i
j
) between E (feature amount of the query image) and the representatives
R
i
j
(feature amount of the plurality of support images)” (Karlinsky, page 5200, left column, paragraph 2)
“we use a sum of two losses to train our model ... The other (contrastive loss) is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
where i∗ is the correct class index for the current example” (Karlinsky, page 5200, right column, paragraph 3)
Claim 3 is rejected under 35 U.S.C. 103 as being unpatentable over Karlinsky (RepMet: Representative-based metric learning for classification and few-shot object detection, 2019, CVPR 2019 pp. 5192-5201) in view of Hu (COMPUTER VISION IMAGE FEATURE IDENTIFICATION VIA MULTI-LABEL FEW-SHOT MODEL, filed 11/27/2019, US 2021/0142054 A1), and further in view of Bayar (A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer, published 2016, IH&MMSec '16: Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security pp. 5-10), Hwang (ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION, published 12/19/2019, US 2019/0382941 A1), and Liu (Transductive Prototypical Network For Few-Shot Classification, published 9/30/2020, 2020 IEEE International Conference on Image Processing (ICIP), pp. 1671-1675).
Regarding claim 3, the rejection of claim 1 is incorporated. While Karlinsky fails to disclose the further limitations of the claim, Liu, in combination with Hu, discloses a system, wherein the at least one processor is configured to:
calculate an average feature amount based on the feature amount of each of the plurality of support images:
(Liu) “Specifically, we construct each episode by randomly sampling N classes from Ctrain and K labeled samples per class as the support set S” (Liu, page 2, left column, paragraph 2)
(Liu) “it computes the mean vector (average feature amount) of embedded features (feature amount) per class from the support set (support data) as the corresponding prototype representation So, the prototype of the class j can be expressed using the following formulation:
PNG
media_image12.png
91
331
media_image12.png
Greyscale
” (Liu, page 2, right column, paragraph 1)
“We evaluate our Td-PN approach compared with the recent state-of-the-art methods on two datasets, miniImageNet [3] and tieredImageNet [13]” (Liu, page 3, right column, paragraph 1); “The miniImageNet dataset is a subset of ILSVRC-12 [1], which is the most popular few-shot learning benchmark. It contains 100 classes randomly selected from ILSVRC-12 with 600 images per class” (Liu, page 3, right column, paragraph 2). Liu’s system is compatible with image classes.
acquire the second loss based on the feature amount of the query image and the average feature amount:
“Specifically, we construct each episode by randomly sampling N classes from Ctrain and K labeled samples per class as the support set S … while a portion of the rest samples from the same N classes as the query set, denoted as Q” (Liu, page 2, left column, paragraph 2).
“In order to find out the nearest class prototype for each query sample, Prototypical Network [6] calculates a soft assignment of embedded features in the query set to each class prototype (average feature amount) as follows
PNG
media_image13.png
80
419
media_image13.png
Greyscale
Where A ∈
R
T
×
N
is an assignment matrix, and
d
(
f
θ
x
~
i
,
p
j
)
denotes the euclidean distance between the embedded feature of the query sample i (feature amount of the query image) and the prototype of the class j (average feature amount)” (Liu, page 3, left column, paragraph 1).
“we design a weighted contrastive loss (second loss) to obtain a classifying-friendly (discriminative) feature embedding space. The designed loss conduces to select top k confident embedded features from the query set for each class in the next subsection. Specifically, if the query sample i belongs to the class j, we expect the value of the soft assignment Aij in the equation (2) to be close to 1; otherwise, we expect the value of the soft assignment Aij to be close to 0. Therefore, the contrastive loss function can be computed as:
PNG
media_image14.png
148
630
media_image14.png
Greyscale
” (Liu, page 3, left column, paragraph 1).
Examiner’s note: While Liu does not disclose the usage of multi-label data or a processor, these attributes are already present in The existing combination (see claim 1), and would simply necessitate using multi-label query data for Liu’s loss, and running Liu’s method on a generic computer processor.
Liu relates to few-shot learning for image classification and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified The existing combination to use distances to average support data instead of distances to direct support data points in loss, as disclosed by Liu. Prototype distances are a simple and efficient implementation of few-shot learning. In particular, Liu’s method avoids overfitting common on prototypical network settings, and performs better than similar state-of-the-art approaches. See Liu, page 1, right column, paragraphs 2-3 and page 4, right column, paragraph 1.
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Karlinsky (RepMet: Representative-based metric learning for classification and few-shot object detection, 2019, CVPR 2019 pp. 5192-5201) in view of Hu (COMPUTER VISION IMAGE FEATURE IDENTIFICATION VIA MULTI-LABEL FEW-SHOT MODEL, filed 11/27/2019, US 2021/0142054 A1), and further in view of Bayar et al. (A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer, published 2016, IH&MMSec '16: Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security pp. 5-10), Hwang (ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION, published 12/19/2019, US 2019/0382941 A1), and Gunel (SUPERVISED CONTRASTIVE LEARNING FOR PRE-TRAINED LANGUAGE MODEL FINE-TUNING, November 2020, arXiv:2011.01403v2).
Regarding claim 5, the rejection of claim 4 is incorporated. Gunel, in combination with Hu, discloses a system, wherein the at least one processor is configured to
calculate the total loss based on the first loss, the second loss, and a weighting coefficient specified by a creator: (Gunel) “
λ
is a scalar weighting hyperparameter (weighting coefficient) that we tune for each downstream task. The loss is given by the following formulas:
PNG
media_image15.png
194
742
media_image15.png
Greyscale
The overall loss is a weighted average of CE (first loss) and the SCL loss (second loss), as given in equation (1)” (Gunel, page 2, paragraph 5).
Gunel relates to few-shot learning for classification and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified The existing combination to use a combined loss function, as disclosed by Gunel. Adding a contrastive loss term to a cross-entropy loss can improve performance in few-shot learning models, and makes the model more robust to noise in the training data as well as imbuing it with more generalization ability to related tasks. See (Gunel, page 2, paragraph 2) and (Gunel, page 8, paragraph 2).
Claim 7 is rejected under 35 U.S.C. 103 as being unpatentable over Karlinsky (RepMet: Representative-based metric learning for classification and few-shot object detection, 2019, CVPR 2019 pp. 5192-5201) in view of Hu (COMPUTER VISION IMAGE FEATURE IDENTIFICATION VIA MULTI-LABEL FEW-SHOT MODEL, filed 11/27/2019, US 2021/0142054 A1), and further in view of Bayar (A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer, published 2016, IH&MMSec '16: Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security pp. 5-10), Hwang (ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION, published 12/19/2019, US 2019/0382941 A1), and Chen (Exploring Simple Siamese Representation Learning, published November 2020, arXiv:2011.10566v1).
Regarding claim 7, the rejection of claim 1 is incorporated. Karlinsky, in combination with Hu, further teaches a method, wherein:
the query image is input to a first learning model: (Karlinsky)
PNG
media_image9.png
563
957
media_image9.png
Greyscale
(Karlinsky, page 5200, Figure 3). The pooled feature vector (query image) is input into the DML embedding module of the model.
“Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module (first learning model), which consists of a few fully connected (FC) layers with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
}, where commonly embedding size e ≪ f” (Karlinsky, page 5199, right column, paragraph 4).
the support data is input to a second learning model: (Karlinsky) “As an additional set of trained parameters, we hold a set of ‘representatives’
R
i
j
∈
R
e
. Each vector
R
i
j
represents the center of the j-th mode of the learned discriminative mixture distribution (support data) in the embedding space, for the i-th class out of the total of N classes” (Karlinsky, page 5200, left column, paragraph 1); “In our implementation, the representatives are realized as weights of an FC layer (second learning model) of size N ·K · e receiving a fixed scalar input 1. The output of this layer is reshaped to an N× K×e tensor” (Karlinsky, page 5200, left column, paragraph 2).
the at least one processor is configured to
calculate the first loss based on the parameter of the first learning model: “Having P(C = i|X) and P(B|X) (output of the first learning model) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone). The first loss is the regular cross-entropy (CE) with the ground truth labels given for the image (or ROI) corresponding to X” (Karlinsky, page 5200, right column, paragraph 3). The first loss is based on the output of the first learning model, and is necessarily based on its parameters.
acquire the feature amount of the query image calculated based on the parameter of the first learning model: “Batch-training is used, but for simplicity we will refer to the input of the subnet as a single (pooled) feature vector X ∈
R
f
(query image) computed by the backbone for the given image (or ROI) ... We first employ a DML embedding module (first learning model), which consists of a few fully connected (FC) layers with batch normalization (BN) and ReLU nonlinearity (2-3 such layers in our experiments). The output of the embedding module is a vector E = E(X) ∈
R
e
} (feature amount), where commonly embedding size e ≪ f” (Karlinsky, page 5199, right column, paragraph 4). Outputs of a learning model are based on parameters of that model.
acquire the feature amount of the ... support data calculated based on the parameter of the second learning model: “As an additional set of trained parameters, we hold a set of ‘representatives’
R
i
j
∈
R
e
(feature amount). Each vector
R
i
j
represents the center of the j-th mode of the learned discriminative mixture distribution (support data) in the embedding space, for the i-th class out of the total of N classes. We assume a fixed number of K modes (peaks) in the distribution of each class, so 1 ≤ j ≤ K” (Karlinsky, page 5200, left column, paragraph 1); “In our implementation, the representatives are realized as weights (parameter[s]) of an FC layer (second learning model) of size N ·K · e receiving a fixed scalar input 1. The output of this layer is reshaped to an N× K×e tensor. During training, this simple construction flows the gradients to the weights of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2). Outputs of a learning model are based on parameters of that model.
adjust each of the parameter of the first learning model and the parameter of the second learning model: “In our implementation, the representatives are realized as weights of an FC layer (second learning model) of size N ·K · e receiving a fixed scalar input 1. The output of this layer is reshaped to an N× K×e tensor. During training, this simple construction flows the gradients to the weights (parameter of the second learning model) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2).
While Karlinsky, Hu, and Bayar fail to disclose the further limitations of the claim, Chen teaches a method, wherein a parameter of the first learning model and a parameter of the second learning model are shared: “Siamese networks are weight-sharing neural networks applied on two or more inputs. They are natural tools for comparing (including but not limited to “contrasting”) entities” (X Chen, page 1, left column, paragraph 1).
Chen relates to n-shot learning for image classification and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified The existing combination to share parameters between two embedding modules, as disclosed by Chen. Siamese networks such as these can learn meaningful representations with minimal data, not requiring negative sample pairs, large batches, or momentum episodes. They can also model invariance with respect to complicated transformations / augmentations. See (Chen, page 1, Abstract) and (Chen, page 2, left column, paragraph 1).
Claim 9 is rejected under 35 U.S.C. 103 as being unpatentable over Karlinsky (RepMet: Representative-based metric learning for classification and few-shot object detection, 2019, CVPR 2019 pp. 5192-5201) in view of Hu (COMPUTER VISION IMAGE FEATURE IDENTIFICATION VIA MULTI-LABEL FEW-SHOT MODEL, filed 11/27/2019, US 2021/0142054 A1), and further in view of Bayar (A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer, published 2016, IH&MMSec '16: Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security pp. 5-10), Hwang (ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION, published 12/19/2019, US 2019/0382941 A1), and, and Rios (Few-Shot and Zero-Shot Multi-Label Learning for Structured Label Spaces, 2018, Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 3132–3142).
Regarding claim 9, the rejection of claim 1 is incorporated. Hu discloses a system with at least one processor configured: “Software or firmware to implement the techniques introduced here may be stored on a machine-readable storage medium and may be executed by one or more general-purpose or special-purpose programmable microprocessors. A "machine-readable medium", as the term is used herein, includes any mechanism that can store information in a form accessible by a machine (a machine may be, for example, a computer, network device, cellular phone, personal digital assistant (PDA), manufacturing tool, any device with one or more processors, etc.)” (Hu, [0058]).
Hu relates to few-shot learning for image processing systems and is analogous to the claimed invention. The existing combination teaches a method of training an image classifier with few-shot learning. The claimed invention improves upon this method by executing it with computer processors. Hu teaches a method of training an image classifier with few-shot learning which can be executed on computer processors, applicable to The existing combination. A person of ordinary skill in the art would have recognized that running The existing combination’s method on Hu’s processor hardware would lead to the predictable result of the method being performed by a computer as-described, and would improve the known device by allowing it to be used to process real data on a computer (MPEP 2143 I. (D) Applying a known technique to a known device (method, or product) ready for improvement to yield predictable results).
While Karlinsky, Hu, and Bayar fail to disclose the further limitations of the claim, Rios teaches a system, able to acquire the query image and the support images from a data group having a long-tail distribution for multi-labels: “Rubin et al. (2012) refer to datasets that have long-tail frequency distributions as “power-law datasets”. Methods that predict in- frequent labels fall under the paradigm of few-shot classification which refers to supervised methods in which only a few examples, typically between 1 and 5, are available in the training dataset for each label ... time. In this paper, we explore both of these issues, long documents and power-law datasets, with an emphasis on analyzing the few- and zero-shot aspects of large-scale multi-label problems.” (Rios, page 1, right column, paragraph 1).
Rios relates to few-shot learning for classification and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified The existing combination to acquire data from long-tail distributions, as disclosed by Rios. Doing so would build a model capable of resisting data sparsity, particularly for datasets with a large number of labels. See (Rios, page 1, right column, paragraph 1).
Claim 10 is rejected under 35 U.S.C. 103 as being unpatentable over Karlinsky (RepMet: Representative-based metric learning for classification and few-shot object detection, 2019, CVPR 2019 pp. 5192-5201) in view of Hu (COMPUTER VISION IMAGE FEATURE IDENTIFICATION VIA MULTI-LABEL FEW-SHOT MODEL, filed 11/27/2019, US 2021/0142054 A1), and further in view of Bayar (A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer, published 2016, IH&MMSec '16: Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security pp. 5-10), Hwang (ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION, published 12/19/2019, US 2019/0382941 A1), and, and Szegedy (Going deeper with convolutions, published 2014, arXiv:1409.4842v1).
Regarding claim 10, the rejection of claim 1 is incorporated. Karlinsky, in combination with Hu, further teaches a method, wherein
the learning model is configured such that a last layer of a model which has learned another label other than a plurality of labels to be recognized is replaced with a layer corresponding to the plurality of labels: (Karlinsky)
“We propose a subnet architecture and corresponding losses that allow us to train a DML embedding jointly with the multi-modal mixture distribution used for computing the class posterior in the resulting embedding space. This subnet then becomes a DML-based classifier head, which can be attached on top of a classification or a detection backbone. It is important to note that our DML-subnet is trained jointly with the feature producing backbone.” (Karlinsky, page 5199, right column, paragraph 3). The few-shot classifier taught by Karlinsky is attached to the head of an image classifier backbone for initial feature processing.
PNG
media_image16.png
171
1058
media_image16.png
Greyscale
(Karlinsky, page 5201, figure 4). InceptionV3 can perform initial feature extraction of images, acting as a backbone of Karlinsky’s classifier.
“For the DML-based classification experiments, we used the InceptionV3 backbone (model), attaching the proposed DML subnet to the layer before its last FC layer (last layer)” (Karlinsky, page 5201, left column, paragraph 1). In Karlinsky’s experiments, the backbone is an Inception network.
“We tested our approach on a set of fine-grained classification datasets, widely adopted in the state-of-the-art DML classification works: Stanford Dogs, Oxford-IIIT Pet, Oxford 102 Flowers, and ImageNet Attributes” (Karlinsky, page 5201, right column, paragraph 3). Karlinsky trained and tested his classifier using these four image datasets.
the at least one processor is configured to calculate the first loss based on the output of the learning model in which the last layer is replaced with the layer corresponding to the plurality of labels, and based on the target output: (Karlinsky) “Having P(C = i|X) and P(B|X) (output of the learning model) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone). The first loss is the regular cross-entropy (CE) with the ground truth labels given for the image (or ROI) corresponding to X” (Karlinsky, page 5200, right column, paragraph 3).
While Karlinsky, Hu, and Bayar fail to disclose the further limitations of the claim, Szegedy teaches a method of constructing a model which has learned another label other than a plurality of labels to be recognized:
“We propose a deep convolutional neural network architecture codenamed Inception (model) ... One particular incarnation used in our submission for ILSVRC14 is called GoogLeNet, a 22 layers deep network, the quality of which is assessed in the context of classification and detection” (Szegedy, page 1, Abstract).
“The ILSVRC 2014 classification challenge involves the task of classifying the image into one of 1000 leaf-node categories in the Imagenet (another label [set]) hierarchy. There are about 1.2 million images for training, 50,000 for validation and 100,000 images for testing. Each image is associated with one ground truth category, and performance is measured based on the highest scoring classifier predictions” (Szegedy, page 8, paragraph 4). The Examiner notes Karlinsky discloses that their model is an Inception model (Karlinsky, page 5201, left column, paragraph 1) Szegedy discloses training an Inception model using ImageNet image set (Szegedy, page 8, paragraph 4) Karlinsky further discloses further training the Inception model using Stanford Dogs, Oxford-IIIT Pet and Oxford Flowers Image sets (Karlinsky, page 5201, right column, paragraph 3). Accordingly, the combination of Karlinsky and Szegedy discloses a model that has learned another label other than the initial labels.
“We independently trained 7 versions of the same GoogLeNet model (including one wider version), and performed ensemble prediction with them” (Szegedy, page 8, paragraph 5).
Szegedy relates to image classification with neural networks and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified The existing combination to use Inception as the backbone network, as disclosed by Szegedy. The Inception architecture is particularly useful as the base model of an object detection network, and is easily adaptable or fine-tuned to different label sets. See (Szegedy, page 4, paragraph 2) and (Szegedy, page 6, paragraph 3).
Claim 20 is rejected under 35 U.S.C. 103 as being unpatentable over Karlinsky (RepMet: Representative-based metric learning for classification and few-shot object detection, 2019, CVPR 2019 pp. 5192-5201) in view of Hu (COMPUTER VISION IMAGE FEATURE IDENTIFICATION VIA MULTI-LABEL FEW-SHOT MODEL, filed 11/27/2019, US 2021/0142054 A1), and further in view of Bayar (A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer, published 2016, IH&MMSec '16: Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security pp. 5-10), Hwang (ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION, published 12/19/2019, US 2019/0382941 A1), and, and Huang et al. (SYSTEMS AND METHODS FOR AUTOMATED VIDEO CLASSIFICATION, filed 12/27/2018, US 10,922,548 B1), hereafter referred to as Huang.
Regarding claim 20, the rejection of claim 19 is incorporated. While the currently cited prior art fails to disclose the further limitations of the claim, Huang discloses a system, wherein in determining how the query image was edited, the learning system determines if digital text was added to the query image:
“In general, a meme video may include text overlaid on image or video content. The text may be added after the image and/or video content has been captured, such that the text does not belong to the originally captured image or video content, but is added after-the-fact. In various embodiments, the meme video detector module 602 can identify meme videos based on identification of such synthetically added text. As shown in the example of FIG. 6A, the meme video detector module 602 can include a dynamic region detection module 604 and a synthetic text detection module 606.” (Huang, column 22, paragraph 2).
PNG
media_image17.png
498
744
media_image17.png
Greyscale
(Huang, Figure 6B)
“FIGS. 6B and 6C illustrate example scenarios 620, 640 that illustrate functionality of the meme video detector module 602, according to an embodiment of the present technology. In the example scenario 620 depicted in FIG. 6B, a video frame 622 from a meme video is depicted. The 35 meme video includes a video of a couple dancing (arrow 626) overlaid with static, synthetic text 624 which reads "When you thought it was Thursday and realize it's actually Friday." (Huang, column 23, paragraph 4)
Huang relates to machine learning for image analysis and is analogous to the claimed invention. The existing combination teaches a system that determines image manipulations in images. Huang teaches a system to detect the presence of synthetic text in images. It would have been obvious to one of ordinary skill in the art to combine the existing combination with Huang by detecting synthetic text in the query images with Huang’s system. This would achieve the predictable result of identifying whether a query image has had text added after being taken, with the system of the existing combination and the system of Huang working the same together as they did separately. (MPEP 2143 I. (A) Combining prior art elements according to known methods to yield predictable results).
Claim 23 is rejected under 35 U.S.C. 103 as being unpatentable over Karlinsky (RepMet: Representative-based metric learning for classification and few-shot object detection, 2019, CVPR 2019 pp. 5192-5201) in view of Hu (COMPUTER VISION IMAGE FEATURE IDENTIFICATION VIA MULTI-LABEL FEW-SHOT MODEL, filed 11/27/2019, US 2021/0142054 A1), and further in view of Bayar (A Deep Learning Approach to Universal Image Manipulation Detection Using a New Convolutional Layer, published 2016, IH&MMSec '16: Proceedings of the 4th ACM Workshop on Information Hiding and Multimedia Security pp. 5-10), Hwang (ARTIFICIAL INTELLIGENCE WASHING MACHINE PROVIDING AUTOMATIC WASHING COURSE BASED ON VISION, published 12/19/2019, US 2019/0382941 A1), and Min (NETWORK REPARAMETERIZATION FOR NEW CLASS CATEGORIZATION, published 3/26/2020, US 2020/0097757 A1).
Regarding claim 23, the rejection of claim 22 is incorporated. Karlinsky further discloses a system, wherein … the at least one processor is configured to:
calculate the first loss, calculate the second loss, and adjust the parameter for each episode of the series:
“During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
“we use a sum of two losses to train our model (DML subnet + backbone). The first loss is the regular cross-entropy (CE) with the ground truth labels given for the image (or ROI) corresponding to X. The other (second loss) is intended to ensure there is at least
α
margin between the distance of E to the closest representative of the correct class, and the distance of E to the closest representative of a wrong class:
PNG
media_image5.png
65
447
media_image5.png
Greyscale
where i∗ is the correct class index for the current example” (Karlinsky, page 5200, right column, paragraph 3).
calculate the first loss as a multi-label cross-entropy loss between the output of the learning model and the target output such that a large first loss indicates a low accuracy of the learning model: “Having P(C = i|X) and P(B|X) (output of the learning model) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone). The first loss is the regular cross-entropy (CE) with the ground truth labels (target output) given for the image (or ROI) corresponding to X” (Karlinsky, page 5200, right column, paragraph 3). Since this is being calculated on multi-label data, it’s a multi-label cross entropy loss.
adjust the parameter based on the first loss and the second loss by executing a learning of the learning model: “During training, this simple construction flows the gradients to the weights (parameter[s]) of the FC layer and learns the representatives” (Karlinsky, page 5200, left column, paragraph 2); “Having P(C = i|X) and P(B|X) computed in the network, we use a sum of two losses to train our model (DML subnet + backbone)” (Karlinsky, page 5200, right column, paragraph 3).
While Karlinsky fails to disclose the further limitations of the claim, Min discloses a system, wherein
the learning system uses few-shot learning to train the learning model using a series of episodes:
“Embodiments of the present invention can be flexibly applied to both ZSL and FSL (few-shot learning), where the exemplar information about unseen are provided in the form of the semantic attributes or one/a few labeled samples, respectively” (Min, [0026])
“Since fθ (learning model component) is trained as a standard multiclass classification task to distinguish all classes within the training set, the resultant feature extractor is supposed to be able to generate more discriminative feature representation for images of new classes than that generated by a model trained in episode-based fashion where the model is trained to distinguish several classes within mini-batches. Meanwhile, gϕ (learning model component) is trained in episode-based fashion by constantly sampling new classes and minimizing the classification loss using the weights generated by gϕ. After training, whenever some new classes come (e.g., in a query Q), along with supporting information in the form of either attribute vectors (ZLS) or few-labeled samples (FSL) (samples for few-shot learning)” (Min, [0036])
each episode of the series corresponds to a different combination of labels: “ In each episode, some randomly sampled classes (labels) are selected and serve as a NCC task for the model” (Min, [0021])
each episode of the series comprises at least one episode query image and at least one episode support image: “weight generation network gϕ is trained in the episode-based manner to enable gϕ to grasp enough knowledge of classifying new classes based on one/a few labeled samples. In details, during the training, we keep randomly sampling from Dt={Xt,Yt} FSL tasks, each of which includes a support set and a query image set.” (Min, [0052])
wherein the at least one processor is configured to: calculate the first loss, calculate the second loss, and adjust the parameter for each episode of the series: “gϕ is trained (adjustment of the parameter[s]) in episode-based fashion by constantly sampling new classes and minimizing the classification loss (first loss) using the weights generated by gϕ” (Min, [0036]); “Through minimizing the least square embedding loss (second loss), the visual-attribute relationship can be established” (Min, [0048])
Min relates to episodic few-shot learning for image classification and is analogous to the claimed invention. It would have been obvious to one of ordinary skill in the art before the effective filing date of the claimed invention to have modified the existing combination to train the few-shot classification system episodically, as disclosed by Min. Episodic training in this context results in a model that’s adaptive to new classes, generalizing well to new classes during testing. Additionally, Min’s implementation of this type of training avoids limited discriminative ability due to neglect of global information, overcoming a common error with this type of methodology. See Min, [0021], [0026].
Response to Arguments
The following responses address arguments and remarks made in the instant remarks dated 05/01/2026.
Claim Interpretation
In light of the instant amendments, limitations are no longer found to be contingent.
Objections
Previous objections to the claims are withdrawn in light of the instant amendments.
112(a) Rejections
New rejections under 35 U.S.C. 112(a) have been made in light of the amendments.
112(b) Rejections
Previous rejections under 35 U.S.C. 112(b) have been withdrawn in light of the instant amendments.
101 Rejections
On pages 12-13 of the instant remarks, the Applicant argues that the claimed invention improves on an existing technological process:
“Applicant respectfully submits that claim 1 as amended resolves a technical problem and
improves an existing technological process. These improvements are discussed, for example, in
paragraphs [0041], [0063] and [0092] of the published application which provide a detailed
discussion of how the technical problem is overcome including a specific application of few-shot
learning based on contrastive learning for handling machine learning based multi-labeling systems
As discussed during the interview, applicant amended claim 1 to clarify that the
improvements are recited in the claim. Specifically, as suggested by the examiner, applicant
amended claim 1 to clarify the calculation of the first loss and the calculation of the second loss”
The Applicant’s arguments above have been fully considered and are persuasive. Rejections under 35 U.S.C. 101 have been withdrawn.
103 Rejections
On pages 15-16 of the instant remarks, the Applicant argues that Karlinsky’s representatives are not equivalent to the claimed support images:
“First, applicant submits that the cited references do not disclose ""acquire support data
comprising a plurality of support images .... "
In the rejection, the examiner cites the representative of Karlinsky as the support data and
indicates that the representatives "can be represented as images" on page 36 of the rejection. This
is inconsistent with the disclosure of Karlinsky. That is, there is no disclosure in Karlinsky that
representatives are images, nor is there any inherent or necessary implication that the
representative vectors correspond to images. Therefore, interpreting the representatives as
"support images" is not supported by the reference.
For example, as explained by Karlinsky on 5198 and quoted by the examiner, "class
representatives are mixture mode centers in the embedding space." (Emphasis added). A mixture
mode center is a mean vector or specific component within mixture model
Similarly, as explained by Karlinsky on 5197 and also quoted by the examiner, "[w]e
represent each class by a mixture model with multiple modes, and consider the centers of these
modes as the representative vectors for the class." (Emphasis added).
Nowhere in Karlinsky is there any disclosure that the representatives are images.”
Regarding the Applicant’s arguments above, the Examiner respectfully disagrees. As one of ordinary skill in the art would know, a mode of continuous data from a multimodal distribution is typically defined as a range of data values having a local maximum frequency within a frequency distribution or local maximum probability within a probability distribution. With that in mind, the “center of a mode”, equivalent to “representative vectors” disclosed by Karlinsky (Karlinsky, page 5197, right column, paragraph 2), could be reasonably interpreted as either the median or mean of the mode’s range of data values.
The probability distributions calculated by Karlinsky are made over a dataset of images, each individual datapoint being an image. Thus, the median of the mode would be equivalent to the original image corresponding to the middle of that data range. Undeniably, an original image from the dataset constitutes an image. The mean of the mode would be equivalent to an average image of the interval, created by summing all the images within it together and normalizing the sum by the count of images. The Examiner maintains that an average of images can itself be considered an image, even if it’s not represented in the original image dataset.
Karlinsky makes it clear the original images from the dataset and representative vectors are somewhat interchangeable in their interpretation and use. While the representatives are created by fully-connected layers of a neural network (FC layer with weights trained to produce useful representatives, see Karlinsky, page 5200, left column, paragraph 2), they each have the same dimensionality as the embedded original images (Karlinsky, page 5199, right column, paragraph 4 to page 5200, left column, paragraph 2), they seem to be visualizable as graphic images (Karlinsky, Figure 2), and representatives that are compared to query images to determine classes during training time are replaced with embedded actual images for the same purpose during test time (Karlinsky, 5200, left column, paragraph 2).
Thus, Karlinsky discloses “acquire support data comprising a plurality of support images” in the form of image class representatives (Karlinsky, page 5197, right column, paragraph 2 & page 5198, Figure 2). See the 103 rejections section for more detail. No rejections are withdrawn on these grounds.
On pages 17-18 of the instant remarks, the Applicant argues that the cited references fail to disclose support images with first or second labels in common with the query image:
“Second, applicant submits that the cited references do not disclose "acquire support data
comprising ... where each of the support images have at least the first label or the second label in
common with the query image."
…
Accordingly, the representatives cannot be the recited support data because not only do the
representatives not comprise a plurality of images as discussed above, but there is no common
label for images with the acquired representatives as each representative applies to a single distinct
"label."
Moreover, the M training examples are not the cited representatives and are not selected
from these representatives. Instead, the M training examples are training examples for each of the
categories used for training the machine learning model. As explained, for example, at 5202-5203,
the training examples are obtained from ImageNet-LOC datasets and are previously unseen by the
learning model.”
Regarding the Applicant’s assertion that the relied upon prior art fails to disclose, as amended, labels being shared between every support image and the query image, the Examiner respectfully disagrees. While the N-way M-shot episodes of Karlinsky previously mapped to an older version of the claim may no longer be applicable to the amended limitation, Karlinsky still discloses acquiring support data, “where each of the support images have at least the first label or the second label in common with the query image”, as claimed in the instant amendments.
As established in arguments above, Karlinsky’s representatives are sufficient to disclose “support images”. Karlinsky, in Figure 2, demonstrates an example where each class representative has one label of a multi-label query image, thus disclosing the amended claim limitation. No rejections are withdrawn on these grounds. See the 103 rejections section for more information.
On pages 19-20 of the instant remarks, the Applicant argues that the cited references fail to disclose acquiring support data with random selection:
“Third, applicant submits that the cited references do not disclose "acquire support data ...
by randomly selecting each of the plurality of support images from an image group stored in an
image database."
…
For the features of now-canceled claim 16 which recited a similar random selection feature,
the examiner cited Karlinsky at 5198, 5200, and 5201 which discuss both the representatives and
the training data
The only portion of the cited sections that discuss random selection, however, are the
training batches. Nowhere does it indicate that the representatives are randomly selected. As
noted above, the training batches cannot be the recited plurality of support images because they
are not used to calculate the second loss as recited in the claim.
Moreover, applicant submits that the cited references do not the disclose selection from an
image group as recited and discussed in paragraph [0060] of the published application.
…
Accordingly, as discussed during the interview, Karlinsky only discloses randomly
selecting the training batch which included a representative for each randomly selected class. Thus,
the system randomly selects classes (and therefore one representative for each class) and does not
disclose randomly selecting a plurality of images for each class, let alone selecting them from a
group of images (because the representatives are not images).”
The Applicant’s arguments above, with respect to the rejection(s) of the amended independent claims, have been fully considered and are persuasive. Therefore, the rejection has been withdrawn. However, upon further consideration, a new ground(s) of rejection is made in view of Hwang. See the 103 rejections section for more detail.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure:
van den Berg (Mode (Statistics) - Simple Tutorial, published June 2020, retrieved from https://www.spss-tutorials.com/mode-in-statistics/ or https://web.archive.org/web/20201202222050/https://www.spss-tutorials.com/mode-in-statistics/) summarizes common definitions of modes in data
Fu (Transductive Multi-label Zero-shot Learning, 2015, arXiv:1503.07790v1) teaches a method of synthesizing a multi-label dataset containing all possible combinations of labels, applicable to zero-shot learning
Saliou (US 20190073565 A1) teaches a method of using CNNs with shared parameters to map multiple inputs to the same embedding space
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Aaron P Gormley whose telephone number is (571)272-1372. The examiner can normally be reached Monday - Friday 12:00 PM - 8:00 PM EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle T Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AG/Examiner, Art Unit 2148
/MICHELLE T BECHTOLD/Supervisory Patent Examiner, Art Unit 2148