DETAILED ACTION
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Response to Amendment
In response to applicant’s amendment received on 6/9/26, all requested changes to the claims have been entered. Claims 1-20 were previously pending. Claims 21 and 22 have been added. Claims 8 and 9 have been cancelled. Claims 1-7 and 10-22 are currently pending. The amendments have resolved the pending 112(b) rejections which are herein withdrawn.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1, 7 and 15 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument. The new grounds of rejection being necessitated by the amendment. Additionally, the arguments made with regards to the motivation to combine Lu and Lee are also moot in view of the new grounds of rejection presented below.
Claim Rejections - 35 USC § 102
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
Claims 1, 7 and 15 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by US 2017/0206465 to Jin et al. (“Jin”).
Regarding claim 1, Jin discloses a computer-implemented method comprising:
partitioning, using a trained image segmentation model, an input image into a plurality of patches (Fig. 6, element 606; paragraphs 49, 52, 67, 92, wherein the region proposal technique such as a CNN (i.e. “trained image segmentation model”) partitions an image into a plurality of regions/patches);
generating, using a vision transformer model, a plurality of patch embeddings, each patch embedding comprising a multidimensional numerical representation of a patch in the plurality of patches (paragraphs 49, 52 and 67, wherein the region proposal technique, which corresponds to the broadest reasonable interpretation of a “vision transformer model”, which the Examiner is not interpreting to be a true vision transformer (ViT) as is known in the art but rather any type of broader “model” (imitation/emulation) thereof, which generates d-dimensional feature vectors (i.e. “multidimensional numerical representation”) for each region/patch);
generating, using a trained patch-label similarity model, a plurality of word embeddings corresponding to the plurality of patch embeddings, each word embedding comprising a word or phrase label for the corresponding patch (Fig. 3; Fig. 6, element 608; paragraphs 57, 58, 67 and 93, wherein a trained embedding function (i.e. patch-label similarity model) is used to generate a word/text embedding corresponding to each region/patch); and
generating, using a trained label prediction model and the plurality of word embeddings, a text label corresponding to the input image (Fig. 2; Fig. 6, element 610; paragraphs 62-65, 69, 94 and 95, wherein the multi-instance embedding module (MIE) (i.e. “label prediction model”) is trained to generate, based on the word/text embedding associated with each region/patch, a plurality of text labels corresponding to the overall image).
Regarding claim 7, please refer to the rejection of claim 1 above. Jin further discloses a computer program product comprising one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by a processor to cause the processor to perform operations disclosed in claim 1 (Fig. 1; paragraphs 23).
Regarding claim 15, please refer to the rejection of claim 1 above. Jin further discloses a computer system comprising a processor and one or more computer readable storage media, and program instructions collectively stored on the one or more computer readable storage media, the program instructions executable by the processor to cause the processor to perform operations the operations of claim 1 (Fig. 1; paragraphs 23).
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 2-4, 10-12 and 16-18 are rejected under 35 U.S.C. 103 as being unpatentable over US 2017/0206465 to Jin et al. (“Jin”) in view of US 2020/0097604 to Lee et al. (“Lee”).
Regarding claim 2, Jin discloses the computer-implemented method of claim 1, wherein the embedding function of the MIE (i.e. patch-label similarity model) is trained and stores a pair-wise similarity score between patch and word embeddings (paragraph 55).
However, Jin does not disclose expressly wherein the trained patch-label similarity model comprises a similarity matrix, a cell of the similarity matrix storing a pair-wise similarity score between a patch embedding and a word embedding.
Lee discloses a process for image captioning comprising partitioning an image into patches/regions, generating region/patch vectors and comparing the region/patch vectors to word vectors in a first stage attention mechanism that include determining a pair-wise similarity score between the vectors and storing each score in a matrix (Fig. 2, element 214; paragraphs 53-56).
Jin & Lee are combinable because they are from the same art of image captioning.
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention to incorporate the technique of using an attention mechanism to generate a similarity matrix where each cell stores a pair-wise similarity score between a patch vector/embedding and a word vector/embedding, as taught by Lee, into the image caption process of Jin.
The suggestion/motivation for doing so would have been to provide improved techniques in determining which regions correspond to which words by determining the biggest region-word pair response for the matching process (Lee, paragraphs 01 and 58).
Therefore, it would have been obvious to combine Lee with Jin to obtain the invention as specified in claim 2.
Regarding claim 3, the combination of Jin and Lee discloses the computer-implemented method of claim 2, wherein the pair-wise similarity score between a patch embedding and a word embedding is computed by analyzing a plurality of training images and corresponding training image captions (Jin, paragraphs 53-61; Lee, paragraphs 18-19).
Regarding claim 4, the combination of Lu and Lee discloses the computer-implemented method of claim 3, further comprising:
partitioning a training image in the plurality of training images into a plurality of training patches; and generating, using the vision transformer model, a plurality of training patch embeddings, each training patch embedding comprising a multidimensional numerical representation of a training patch in the plurality of training patches (Lee, paragraphs 38 and 39. Jin, paragraphs 49-52, wherein a region proposal technique is used to partition and generate the region/patch vectors/embeddings for training images).
Regarding claims 10-12, please refer to the rejections of claims 2-4, respectively, above.
Regarding claims 16-18, please refer to the rejections of claims 2-4, respectively, above.
Claims 21 and 22 are rejected under 35 U.S.C. 103 as being unpatentable over US 2017/0206465 to Jin et al. (“Jin”) in view of US 2023/0103305 to Xu et al. (“Xu”).
Regarding claim 21, Jin discloses the computer-implemented method of claim 1.
Jin does not disclose expressly generating, from a training image caption corresponding to a training image, a caption graph, each node in the caption graph representing an object described in the training image caption and each edge in the caption graph representing a relationship between two objects; generating a plurality of node embeddings corresponding to nodes in the caption graph; and computing a pair-wise similarity score between a patch embedding and a word embedding derived from the node embeddings.
Xu discloses a process for training a scene graph generation neural network for use in image captioning (paragraph 01) that comprises generating, from a training image caption corresponding to a training image, a caption graph, each node in the caption graph representing an object described in the training image caption and each edge in the caption graph representing a relationship between two objects (Figs. 3 and 4; paragraphs 46, 51, 53, 56 and 57, wherein a label graph (406) (i.e. caption graph) is generated from an image description (302, 402) (i.e. training image caption) associated with a training image (300, 400), each node of the label/caption graph representing an object and each edge a relationship between them);
generating a plurality of node embeddings corresponding to nodes in the caption graph (Figs 3 and 4; paragraphs 51-52 and 58-62, wherein the label embedding model generates label graph embedding (i.e. node embedding) from the label/caption graph); and
computing a pair-wise similarity score between a patch embedding and a word embedding derived from the node embeddings (Figs. 3 and 4; paragraphs 52-54, 62 and 63, wherein graph matching computes a one-to-one similarity mapping between visual embeddings (i.e. patch embedding) and the label graph embeddings (i.e. word embeddings) derived from the node embeddings).
Jin & Xu are combinable because they are from the same art of image processing, specifically related to image captioning.
It would have been obvious to a person of ordinary skill in the art, before the effective filing date of the claimed invention to incorporate the technique of generating, from a training image caption corresponding to a training image, a caption graph, generating a plurality of node embeddings corresponding to nodes in the caption graph, and computing a pair-wise similarity score between a patch embedding and a word embedding derived from the node embeddings, as taught by Xu, into the image caption process disclosed by Jin.
The suggestion/motivation for doing so would have been provide efficient, flexible and accurate semantic scene graph generation for use in image captioning (Xu, paragraph 01-02).
Therefore, it would have been obvious to combine Xu with Jin to obtain the invention as specified in claim 21.
Claim 22 recites substantially similar limitations to those of claim 21 and is therefore rejected for the same reasoning indicated above with regards to claim 21.
Allowable Subject Matter
Claims 5, 6, 13, 14, 19 and 20 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. See attached PTO-892.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AARON W CARTER whose telephone number is (571)272-7445. The examiner can normally be reached 8am - 5pm (Mon - Fri).
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AARON W CARTER/Primary Examiner, Art Unit 2661