Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
DETAILED ACTION
RESPONSE TO ARGUMENTS
Claim Objections
In view of applicant amendment and claim construction guidance, claims 1 & 11-12 objection is withdrawn.
Allowable Subject Matter
Claims 6 & 8-10 remain objected to allowable subject matter.
Prior Art Rejection
After carefully reviewing applicant amendments, prior art guidance and claim limitations, examiner respectfully disagrees.
Argument 1
Applicant submits Yu describes “a quantization system that maps image tokens to entries in a single, global quantization codebook,” and does not disclose a codebook organized on a per object basis, separate groups of embeddings for respective objects, or an architecture in which a model selects an embedding from a group associated with the particular object depicted in the image.
In response to applicant's argument that the references fail to show certain features of applicant’s invention, it is noted that the features upon which applicant relies (i.e., [ embeddings be stored in separate object specific codebooks, that a separate codebook be maintained for each respective object or that the codebook itself be organized according to the object identity.]) are not recited in the rejected claim(s). Although the claims are interpreted in light of the specification, limitations from the specification are not read into the claims. See In re Van Geuns, 988 F.2d 1181, 26 USPQ2d 1057 (Fed. Cir. 1993).
In view of above arguments, examiner submits rejection is sufficient and respectfully maintained.
Argument 2
Applicant submits Yu’s [0019-0020] disclosing a class label or natural language conditioning input to a code prediction model is not itself an “object associated embedding” and that Yu lacks a data structure organized according to object identity.
In response, examiner submits the rejection does not require that the auxiliary conditioning data of Yu [0019-0020] itself constitute the entirety of the claimed embedding. Claim 1 does not require that the embedding itself contain an express object identifier, claim 1 also does not require a separate object specific codebook.
In view of above arguments, examiner submits rejection is sufficient and respectfully maintained.
Argument 3
Applicants arguments filed on (06/22/2026) have been fully considered but are deemed moot in view of new grounds of rejection. Due to the variation in claim scope via amendments a new ground of rejection is proper.
ALLOWABLE SUBJECT MATTER
Claims 2-3, 6 & 8-10 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
CLAIM REJECTIONS - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claims 1, 4-5, 7 & 11-14 are rejected under 35 U.S.C. 103 as being unpatentable over YU (WO 2023/059699) in view of He et al. (U.S. Publication 2021/0334475)
As to claims 1 & 11-12, YU discloses providing embeddings that are associated with objects ([0019-0020] discloses an auxiliary conditioning data that comprises a class label descriptive of a desired class of the synthesized image, that auxiliary conditioning data can also be natural language text token. [0084] discloses the auxiliary conditioning data comprises a class label descriptive of a desired class of the synthesized image. [0108] discloses separate embedding layers can be learned from scratch for class id token and image tokens); providing a token that represents a part of a digital image that depicts at least a part of an object ([0070] discloses the computing system can process the plurality of input image patches 12 with a machine learned image encoder 16 to generate a plurality of image tokens 18 in a latent space. The plurality of images token 18 can correspond to the plurality of input image patches 12.), wherein the token represents a part with a lower resolution than a resolution of the pixel of the part ([0093] discloses an example encoder of ViT-VQGAN first maps 8x8 non-overlapping image patches into image tokens. Encoding a 256x256 resolution image into a 32x32 = 1024 token sequence); selecting with a model an embedding of the embeddings to represent the token in a representation of the digital image ([0113] discloses a visual codebook is learned that snaps a patch embedding to its nearest codebook entry which is a learned and index-able location in the overall latent space. These entries can be thought of as visual word types, and the appearance of any of these words in a patch in a given image is thus an image token. [0097] discloses the quantized codebook index is determined by looking up the codebook vector closest to the input features ze(x) in the terms of the Euclidean distance.);
YU is silent to determining with the model a reconstruction of the token that represents the object depending on the representation of the digital image; and determining, depending on a difference between the token and the reconstruction of the token, a parameter that defines at least one of the embeddings and/or a parameter that defines the model.
Yu compares predicted codes against quantized codes and modifies model parameters based upon the resulting loss, Yu does not disclose the predicted code as a reconstruction of the corresponding original token and the comparison as the claimed difference between that token and its reconstruction.
However, He discloses determining with the model a reconstruction of the token ([0081] discloses a standard BERT pre-training consists of applying final hidden vectors from the final encoding layer 915, the hidden vectors corresponding to the masked tokens, to an output softmax 917 over vocabulary to reconstruct the masked tokens. [0083] discloses the objective of the computer system is to reconstruct corrupted tokens. [0084] discloses an enhanced mask decoder (EMD) which is a task specific decoder designed to reconstruct masked token of an MLM.) that represents the object depending on the representation of the digital image; and determining, depending on a difference between the token and the reconstruction of the token ([0081] discloses after reconstructing the masked tokens the computing system then trains and updates the model based on the accuracy of the predictions. [0083] discloses reconstruction based on a formula according to whether the reconstructed token equals the original token. Accordingly, He discloses determining correspondence/difference between an original token and the reconstruction of that token an training/updating the model depending on that result.), a parameter that defines at least one of the embeddings and/or a parameter that defines the model.
It would have been obvious to one of ordinary skill in the art at the time of effective filing to modify YU’s disclosure to include the above limitations in order to better reconstruct tokens, improve pre-training convergence, and improve self-training of the model.
As to claim 4, YU in view of He discloses everything as disclosed in claim 1. In addition, YU discloses determining the token depending on the pixel of the part of the digital image using an encoder. ([0008] discloses processing by the computing system, the plurality of input image patches with a machine learned image encoder to generate a plurality of image tokens in a latent space.)([0070, 0093])
As to claim 5, YU in view of He discloses everything as disclosed in claim 1. In addition, YU discloses determining, using a decoder: (i) a reconstruction of the digital image that depicts a reconstruction of the object, or (ii) a digital image that depicts a reconstruction of the object depending on the reconstruction of the token that represents the object. ([0016] discloses auto aggressively predicting, by the computing system using a machine learned code prediction model, a plurality of predicted codes (i.e. reconstructions of the token) from the quantizqation codebook based at least in part on one or more of the plurality of quantized codes; and processing, by the computing system, the plurality of predicted codes with a machine learned image decoder to generate a plurality of synthesized image patches that form a synthesized image (e.g. reconstruction of the digital image))([0072, 0093, 0119])
As to claim 7, YU in view of He discloses everything as disclosed in claim 5. In addition, YU discloses (i) providing an embedding for the object for the reconstruction of the part of the digital image, determining a token that represents the object for the reconstruction of the digital image depending on the embedding, and determining the reconstruction of the digital image depending on the token that represents the object for the reconstruction of the digital image; or (ii) providing an embedding for an object for the digital image, determining a token that represents the object for the digital image depending on the embedding, and determining the digital image depending on the token that represents the object for the digital image. ([0016] discloses auto aggressively predicting, by the computing system using a machine learned code prediction model, a plurality of predicted codes (i.e. reconstructions of the token) from the quantization codebook based at least in part on one or more of the plurality of quantized codes; and processing, by the computing system, the plurality of predicted codes with a machine learned image decoder to generate a plurality of synthesized image patches that form a synthesized image (e.g. reconstruction of the digital image))([0093, 0113, 0119])
As to claim 13, YU in view of He discloses everything as disclosed in claim 1. In addition, YU discloses the parameter that defines at least one of the embeddings. ([0074, 0097-0098])
YU in view of He is silent to determining, depending on the difference between the token and the reconstruction of the token.
However, He discloses determining, depending on the difference between the token and the reconstruction of the token ([0081, 0083])
It would have been obvious to one of ordinary skill in the art at the time of effective filing to modify YU in view of He’s disclosure to include the above limitations in order to better reconstruct the tokens, improve pre-training convergence, and improve the learned token representations.
As to claim 14, YU in view of He discloses everything as disclosed in claim 1. In addition, YU discloses determining, depending on the difference between the token and the reconstruction of the token, the parameter that defines the model. (Yu’s [0033] & He [0081, 0083])
CONCLUSION
No prior art has been found for claims 2-3, 6 & 8-10 in their current form.
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to Stephen P Coleman whose telephone number is (571)270-5931. The examiner can normally be reached Monday-Thursday 8AM-5PM.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Andrew Moyer can be reached at (571) 272-9523. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
Stephen P. Coleman
Primary Examiner
Art Unit 2675
/STEPHEN P COLEMAN/Primary Examiner, Art Unit 2675