DETAILED ACTION
[1] Remarks
I. The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
II. Claims 1-20 are pending and have been examined, where claims 1-20 is/are rejected. Explanations will be provided below.
III. Inventor and/or assignee search were performed and determined no double patenting rejection(s) is/are necessary.
IV. Patent eligibility (updated in 2019) shown by the following: Claims 1-20 pass patent eligibility test because there is/are no limitation or a combination of limitations amounting to an abstract idea. Also, the following limitation or the combinations of the limitations: “the set of query embeddings, using a classification subnetwork of the object detection neural network to generate, for each object embedding, a respective classification score distribution over the set of query embeddings, wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood that the region of the input image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding” effects a transformation or a reduction of a particular article to a different state or thing / adds a specific limitation(s) other than what is well-understood, routine and conventional in the field, or adding unconventional steps that confine the claim to a particular useful application and providing improvements to the technical field of
Deep learning, which recite additional elements that integrate the judicial exception into a practical application and amounting significant more.
[2] Claim Interpretation
The following is a quotation of 35 U.S.C. 112(f):
(f) Element in Claim for a Combination. – An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
The following is a quotation of pre-AIA 35 U.S.C. 112, sixth paragraph:
An element in a claim for a combination may be expressed as a means or step for performing a specified function without the recital of structure, material, or acts in support thereof, and such claim shall be construed to cover the corresponding structure, material, or acts described in the specification and equivalents thereof.
Use of the word “means” (or “step for”) in a claim with functional language creates a rebuttable presumption that the claim element is to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is invoked is rebutted when the function is recited with sufficient structure, material, or acts within the claim itself to entirely perform the recited function. Absence of the word “means” (or “step for”) in a claim creates a rebuttable presumption that the claim element is not to be treated in accordance with 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph). The presumption that 35 U.S.C. 112(f) (pre-AIA 35 U.S.C. 112, sixth paragraph) is not invoked is rebutted when the claim element recites function but fails to recite sufficiently definite structure, material or acts to perform that function.
Claim elements in this application that use the word “means” (or “step for”) are presumed to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action. Similarly, claim elements that do not use the word “means” (or “step for”) are presumed not to invoke 35 U.S.C. 112(f) except as otherwise indicated in an Office action.
Claim(s) 12-16 are not interpreted under 35 U.S.C. 112(f) or pre-AIA U.S.C. 112 6th paragraph because of the following reason(s): limitations are modified by sufficient structure or material for performing the claimed function.
Claim(s) 2-11 and 17-21 do not require 35 U.S.C. 112(f) or pre-AIA U.S.C. 112 6th paragraph interpretation because they are method claims and / or they are CRM claims.
Upon examination of the specification and claims, the examiner has determined, under the best understanding of the scope of the claim(s), rejection(s) under 35 U.S.C. 112(a)/(b) is not necessitated because of the following reasons: sufficient support are provided in the written description / drawings of the invention.
[3] Grounds of Rejection
Double Patenting
The nonstatutory double patenting rejection is based on a judicially created doctrine grounded in public policy (a policy reflected in the statute) so as to prevent the unjustified or improper timewise extension of the "right to exclude" granted by a patent and to prevent possible harassment by multiple assignees. See In re Goodman, 11 F.3d 1046, 29 USPQ2d 2010 (Fed. Cir. 1993); In re Longi, 759 F.2d 887, 225 USPQ 645 (Fed. Cir. 1985); In re Van Ornum, 686 F.2d 937, 214 USPQ 761 (CCPA 1982); In re Vogel, 422 F.2d 438, 164 USPQ 619 (CCPA 1970); and In re Thorington, 418 F.2d 528, 163 USPQ 644 (CCPA 1969).
A timely filed terminal disclaimer in compliance with 37 CFR 1.321(c) may be used to overcome an actual or provisional rejection based on a nonstatutory double patenting ground provided the conflicting application or patent is shown to be commonly owned with this application. See 37 CFR 1.130(b).
Effective January 1, 1994, a registered attorney or agent of record may sign a terminal disclaimer. A terminal disclaimer signed by the assignee must fully comply with 37 CFR 3.73(b).
Claim 2 is rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over claim 1 of U.S. Patent No. 11,928,854. The conflicting claims are not identical because patent claim 1 requires the additional step of “processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings, wherein the image encoding subnetwork comprises one or more self-attention neural network layers (underlined portion)”, not required by claim 2. However, the conflicting claims are not patentably distinct from each other because: Claim 2 and patent claim 1 recite common subject matter; whereby claim 2, which recites the open-ended transitional phrase “comprising”, does not preclude the additional elements recited by patent claim 1, and whereby the elements of claim 2 are fully anticipated by patent claim 1.
Claim 12 is rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over claim 14 of U.S. Patent No. 11,928,854. The conflicting claims are not identical because patent claim 14 requires the additional step of “processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings, wherein the image encoding subnetwork comprises one or more self-attention neural network layers (underlined portion)”, not required by claim 12. However, the conflicting claims are not patentably distinct from each other because:
Claim 12 and patent claim 14 recite common subject matter;
whereby claim 12, which recites the open-ended transitional phrase “comprising”, does not preclude the additional elements recited by patent claim 14, and whereby the elements of claim 12 are fully anticipated by patent claim 14.
Claim 17 is rejected under the judicially created doctrine of obviousness-type double patenting as being unpatentable over claim 15 of U.S. Patent No. 11,928,854. The conflicting claims are not identical because patent claim 15 requires the additional step of “processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings, wherein the image encoding subnetwork comprises one or more self-attention neural network layers (underlined portion)”, not required by claim 17. However, the conflicting claims are not patentably distinct from each other because:
Claim 17 and patent claim 15 recite common subject matter;
whereby claim 17, which recites the open-ended transitional phrase “comprising”, does not preclude the additional elements recited by patent claim 15, and whereby the elements of claim 17 are fully anticipated by patent claim 15.
Note: All claims will be indicated rejected because all independent claims are rejected under Double Patenting.
[4] Allowable Subject Matter
Claims 2-21 are allowable / patentable if applicant(s) overcome double patenting rejections. The following is an examiner’s statement of reasons for allowance by comparing claims to closest references. The references are divided into primary and secondary, where primary would have been utilized in a USC 102 or main USC 103 reference and secondary would have been utilized a secondary USC 103 reference, but these references do not cover enough of the claim’s scope to warrant a rejection.
Primary reference, Minderer “Simple Open-Vocabulary Object Detection with Vision Transformers” (publication date: 05/06/2022) discloses a method performed by one or more computers, the method comprising:
obtaining: (i) an image, and (ii) a set of one or more query embeddings, wherein each query embedding represents a respective category of object (see figure 1, illustration below, image of giraffes and its text are input to the transformer);
processing the image and the set of query embeddings using an object detection neural network to generate object detection data for the image (see figure 1, MLP is read as the neural network), comprising:
processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings, wherein the image encoding subnetwork comprises one or more self-attention neural network layers (see figure 1, the image is divided into four equal regions, each sub regions is read as its own self attention, see “Vision Transformer Encoder” and “Text Transformer Encoder”):
PNG
media_image1.png
467
1472
media_image1.png
Greyscale
processing each object embedding using a localization subnetwork of the object detection neural network to generate localization data defining a corresponding region of the image (see figure 1, output of MLP is read as localization data); and
processing: (i) the set of object embeddings, and (ii) the set of query embeddings, using a classification subnetwork of the object detection neural network to generate, for each object embedding, a respective classification score distribution over the set of query embeddings (see figure 1, illustration below the output of Linear Projections is read as classification scores),
wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood that the text corresponding to the category represented by the query embedding (see figure 1 illustration below, the embedding corresponds to the likelihood of text input):
PNG
media_image2.png
623
1519
media_image2.png
Greyscale
.
Minderer is silent in disclosing but not a likelihood that the region of the image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding. Also, Minderer does not qualify as a prior art because its publication date (May 12, 2022) is after the earliest priority of the current application (05/06/2022).
Primary reference, Wang (US 20230019211) discloses a method performed by one or more computers, the method comprising: obtaining: (i) an image (see figure 2, 120), and (ii) a set of one or more query embeddings, wherein each query embedding represents a respective category of object (see figure 2, 140);
processing the image and the set of query embeddings using an object detection neural network to generate object detection data for the image (see figure 2, 104 receives text embedding, also see illustrations below, see paragraph 63, trains one or more neural networks from data comprising a mix of paired and unpaired data), comprising:
processing the image using an image encoding subnetwork of the object detection neural network to generate a set of object embeddings (see paragraph 70, a pre-training framework obtains or otherwise receives as input an image 120, calculates an image embedding 122, which is input to a multi-scale image encoder 124 to calculate an image embedding 126, see paragraph 215, a CNN include a region-based or regional convolutional neural networks, RCNNs, and Fast RCNNs, e.g., as used for object detection or other type of CNN, 124 is read as the subnetwork, 122 is read as the image embedding);
PNG
media_image3.png
395
897
media_image3.png
Greyscale
processing each object embedding using a localization subnetwork of the object detection neural network to generate localization data defining a corresponding region of the image (see paragraph 245, a CNN for emergency vehicle detection and identification may use data from microphones 1496 to detect and identify emergency vehicle sirens); and
processing: (i) the set of object embeddings (see figure 1, 108), and (ii) the set of query embeddings, using a classification subnetwork of the object detection neural network to generate, for each object embedding, a respective classification score distribution over the set of query embeddings (see paragraph 97, a softmax function is a generalization of a logistic function to multiple dimensions, a softmax function normalizes inputs into a probability distribution comprising probabilities), but is silent in disclosing wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood that the region of the image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding.
Secondary reference, Lund (US 20200372610) discloses using a classification subnetwork of the object detection neural network to generate a respective classification score distribution over the set of query embeddings (see figure 3 and figure 5B), but not for each object embedding:
PNG
media_image4.png
417
1465
media_image4.png
Greyscale
;
wherein the respective classification score distribution for each of the
PNG
media_image5.png
264
844
media_image5.png
Greyscale
.
Lund is silent in disclosing wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood that the region of the image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding.
Secondary reference, Ermolov “Hyperbolic Vision Transformers: Combining Improvements in Metric Learning” discloses a method performed by one or more computers, the method comprising: obtaining (1) an image (see figure 1 illustrations below), and (ii) a set of one or more query embeddings, wherein each query embedding represents a respective category of object (see figure 1 illustrations below). Ermolov is silent in disclosing processing the image and the set of query embeddings using an object detection neural network to generate object detection data for the image.
Secondary reference, AGARWAL (US 20190080225) discloses wherein the respective classification score distribution for each of the object embeddings defines, for each query embedding, a likelihood paragraph 36, a view to force the network to learn better separation of the embeddings, query embeddings, the above loss may be increased slightly for all predictions, irrespective of whether the prediction is right or wrong, a square-root of all the probabilities in the prediction distribution Pi and then re-normalize to obtain the new probability distribution Qi, FIG. 4 is a graphical representation illustrating a predicted Probability Distribution (P), new probability distribution obtained after square-root and normalization of P, and T is the target distribution):
PNG
media_image6.png
308
712
media_image6.png
Greyscale
.
AGARWAL is silent in disclosing a likelihood that the region of the image corresponding to the object embedding depicts an object that is included in the category represented by the query embedding. AGARWAL’s network is not a CNN used for object detection.
Wang, Lund, Ermolov and AGARWAL, taken alone or in combination with each other, are silent in disclosing all the limitations of claims.
CONTACT INFORMATION
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEX LIEW (duty station is located in New York City) whose telephone number is (571)272-8623 (FAX 571-273-8623), cell (917)763-1192 or email alexa.liew@uspto.gov. Please note the examiner cannot reply through email unless an internet communication authorization is provided by the applicant. The examiner can be reached anytime.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, MISTRY ONEAL R, can be reached on (313)446-4912. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative or access to the automated information system, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/ALEX KOK S LIEW/Primary Examiner, Art Unit 2674 Telephone: 571-272-8623
Date: 9/3/26