Prosecution Insights
Last updated: October 02, 2026
Application No. 18/879,621

MODEL TRAINING METHOD, IMAGE PROCESSING METHOD, ELECTRONIC DEVICE AND STORAGE MEDIUM

Non-Final OA §103
Filed
Dec 27, 2024
Priority
Jun 28, 2022 — CN 202210754292.2 +1 more
Examiner
SOFRONIOU, MICHAEL MARIO
Art Unit
Tech Center
Assignee
Lemon Inc.
OA Round
1 (Non-Final)
100%
Grant Probability
Favorable
1-2
OA Rounds
6m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 100% — above average
100%
Career Allowance Rate
3 granted / 3 resolved
+40.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
2y 3m
Avg Prosecution
20 currently pending
Career history
22
Total Applications
across all art units

Statute-Specific Performance

§101
5.6%
-34.4% vs TC avg
§103
42.3%
+2.3% vs TC avg
§102
12.7%
-27.3% vs TC avg
§112
33.8%
-6.2% vs TC avg
Black line = Tech Center average estimate • Based on career data from 3 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Preliminary Amendment In response to applicant’s preliminary amendment received on 12/27/2024, all requested changes to the claims, specification & abstract have been entered. Claim(s) 1-14 were previously pending. Claim(s) 15-24 have been added. Claim(s) 9-10 & 13-14 have been cancelled. Claim(s) 1-8,11-12 and 15-24 are currently pending. Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The information disclosure statement(s) (IDS) submitted on 02/03/2025 & 06/05/2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Specification The title of the invention is not descriptive. A new title is required that is clearly indicative of the invention to which the claims are directed. While this application is directed to training a model, a title that reflects the field of endeavor for the type of training to be conducted would better embody the claimed invention as a whole. The invention is directed towards improvements in masked autoencoders via incorporation of a fusion feature prediction for applications in computer vision for downstream classification tasks. A title that reflects these features or other features that highlight the inventive concept would be appropriate. The disclosure is objected to because of the following informalities: [¶0004] recites “graphics relationship between a current input image and of other image by comparing…”. The examiner believes this was intended to recite “between a current input image and another image” [¶0047] recites “However, there is also an association relationship between regions inside the image, buy the model obtained…”. The examiner believes this was intended to recite “…, but the model obtained …” Appropriate correction is required. Claim Objections Claims 7, 23 & 24 are objected to because of the following informalities: Claim 7 recites “wherein the acquiring a target image feature in a target image, comprises…”. The limitation of the “target image” was previously introduced in claim 1, which claim 7 depends on. The examiner believes this claim should instead recite “wherein acquiring the target image feature in the target image, comprises…” Claims 23 & 24 both recite “the non-transient computer readable storage medium…”. Claim 23 (and claim 24, by extension of its dependence on claim 23) depend on claim 12, which first introduces this element as “A non-transitory computer readable storage medium”. The examiner understands these to be substantially identical elements, but advises applicant to use consistent terminology among claim terms in the same chain of claim dependency. Appropriate correction is required. Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claim(s) 1-3, 5-7, 11-12, 17-18, 20-24 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al ("Masked Autoencoders are Scalable Vision Learners", Dec 19, 2021, arXiv), hereinafter referred to as “He”, in view of Chen et al ("Context Autoencoder for Self-Supervised Representation Learning", Feb 07 2022 (v1), arXiv), hereinafter referred to as “Chen”, further in view of Zhou et al (US 2023/0376729 A1), hereinafter referred to as “Zhou”. Regarding claim 1, He teach A model training method (He: the masked autoencoder (MAE) training method outlined in Figure 1 [Sec 1. Introduction - ¶09-10 (pg. 2)]), comprising: acquiring a first sample image, wherein the first sample image comprises a first image block which is uncovered and a second image block, which is covered (He: an input image (i.e., a first sample image) is divided into a subset of non-overlapping patches and a subset of those patches are masked (i.e., covered), resulting in both visible, unmasked patches and invisible, masked patches [Sec 3. Approach - Subsec: Masking - ¶01-02; Figs. 1 & 2]); processing the first sample image through a first model, to obtain a first image feature corresponding to the first image block (He: a ViT is applied to visible, unmasked patches resulting in embedded patches (i.e., a first image feature) with additional positional embeddings [Sec 3. Approach - Subsec: MAE encoder - ¶01; Fig. 1]); reconstructing the second image block according to the first image feature, to obtain a first image (He: a MAE decoder utilizes the embedded patches (again, the first image feature) and mask tokens embedding information of masked regions to then yield a reconstructed input image (i.e., a first image) [Sec 3. Approach - Subsec: MAE decoder - ¶01-02 & Subsec: Reconstruction target - ¶01-02; Fig. 1]), and updating a model parameter of the first model, according to the first image, the second image block, (He: a loss function computes the mean square error (MSE) between the reconstructed image (i.e., first image) and the original input image (i.e., a first sample image) on only the masked patches (i.e., the second image block) to update the decoder layers [Sec 3. Approach - Subsec: Reconstruction target - ¶01-02; Fig. 1]). He fails to recite determining a fusion prediction feature of the first and second image block, acquiring a target image after preprocessing the first sample image and using these to update the model parameter. Chen, however, is analogous art pertinent to the field of endeavor of the present application and describes a context autoencoder (CAE) that utilizes a context regressor to predict masked patch representations from visible patches and masked queries. More specifically, Chen teaches and determining a fusion prediction feature of the first image block and the second image block, according to the first image feature; (Chen: the latent contextual regressor H predicts latent representations of masked patches Zm (i.e., a fusion prediction feature) from the latent representations of visible patches Zv (i.e., a first image feature) [Sec 2.1. Architecture - Subsec: Latent contextual regressor - ¶01-02; Figs. 1 & 2]) and updating a model parameter of the first model, according to the first image, the second image block, the fusion prediction feature (Chen: the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - 01-04; eq. 1; Fig. 2]). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the latent contextual regressor utilized by Chen with the base masked autoencoder framework of He provides context to inform what masked regions would be based on their visible counterparts, leading to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. Neither He nor Chen explain acquiring a target image feature after preprocessing the first sample image and utilizing it to update model parameters. Zhou, on the other hand, is analogous art pertinent to the field of endeavor and teach a method for training neural network utilizing a transformed autoencoder (TAE) that utilizes transformed views of an input image. Zhou teach acquiring a target image feature in a target image, wherein the target image is an image after preprocessing of the first sample image (Zhou: in the TAE approach a an input image 203 which may be a crop of a larger image (i.e., a target image after preprocessing of the first sample image – the examiner notes here that, that cropping is described of preprocessing of the first sample image in accordance with [¶0073] of the specification of the instant application) is encoded into a set of latent representations 304 (i.e., a target image feature) [¶0045-49; Figs. 3 & 5]); and updating a model parameter of the first model, according to (Zhou: a loss between the reconstructed target y derived from features of the transformed input image x (i.e., a target image feature) is calculated to measure the discrepancy between the reconstruction and ground truth y' to update the decoder [¶0051-58; eqs. 2-3]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the transform view training framework of Zhou with the combined MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. Considering claim 2, He in view of Chen, further in view of Zhou teach The method according to claim 1 (outlined above), wherein, the updating a model parameter of the first model, according to the first image, the second image block, the fusion prediction feature and the target image feature, comprises: determining a first loss function of the first model, according to the first image and the second image block; updating the model parameter of the first model, according to the first loss function (He: a loss function computes the mean square error (MSE) between the reconstructed image (i.e., first image) and the original input image (i.e., a first sample image) on only the masked patches (i.e., the second image block) to update the decoder layers [Sec 3. Approach - Subsec: Reconstruction target - ¶01-02; Fig. 1]). He, however, fails to teach determining a second loss function according to a fusion prediction feature and target image feature. Chen, per contra, teach determining a second loss function of the first model, according to the fusion prediction feature updating the model parameter of the first model, according to (Chen: the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - ¶01-04; eq. 1; Fig. 2]). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the loss driven by the predicted latent representations of masked patches from Chen in tandem with the loss taught by He optimizes masked region reconstruction, leading to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. He in view of Chen, however, fail to teach determining the second loss function according to the target image feature. Zhou, on the other hand, teach determining a second loss function of the first model, according to updating the model parameter of the first model, according to (Zhou: a loss between the reconstructed target y derived from features of the transformed input image x (i.e., a target image feature) is calculated to measure the discrepancy between the reconstruction and ground truth y' to update the decoder [¶0051-58; eqs. 2-3]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [¶0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the loss from transformed views of Zhou with the loss functions taught by the MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. With respect to claim 3, He in view of Chen, further in view of Zhou teach The method according to claim 2 (detailed above), wherein, the determining a second loss function of the first model, according to the fusion prediction feature and the target image feature, comprises: acquiring a second sample image, and determining a second image feature of the second sample image, wherein the second sample image is an image other than the first sample image (Chen: the CAE architecture is trained using the imageNet-1K dataset comprising a plurality of images for model training as illustrated in Fig. 6 [Sec 4.1 Implementation - 01-02], visible features Zv (second image features) are determined from visible patches Xv via the encoder F [Sec 2.1 - Architecture - ¶01-02; Fig. 2]); and determining the second loss function according to the fusion prediction feature, the target image feature, (Chen: the latent contextual regressor H predicts latent representations of masked patches Zm (i.e., a fusion prediction feature) from the latent representations of visible patches Zv (i.e., a second image feature) [Sec 2.1. Architecture - Subsec: Latent contextual regressor - 01-02; Figs 1 & 2], wherein the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - ¶01-04; eq. 1; Fig. 2] – the examiner notes that here, given that the latent representations of masked patches (the fusion prediction feature) are derived from the latent representations of visual features (a second image feature), the loss calculated using the latent representations of masked patches is inherently determined using the second image features). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the loss driven by the predicted latent representations of masked patches from Chen in tandem with the loss taught by He optimizes masked region reconstruction, leading to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. Chen again fails to disclose using a target image to determine the second loss. Zhou, however, teach determining the second loss function according to (Zhou: a loss between the reconstructed target y derived from features of the transformed input image x (i.e., a target image feature) is calculated to measure the discrepancy between the reconstruction and ground truth y' to update the decoder [¶0051-58; eqs. 2-3]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [¶0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the loss from transformed views of Zhou with the loss functions taught by the MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. Turning to claim 5, He in view of Chen, further in view of Zhou The method according to claim 2 (described previously), wherein, the updating the model parameter of the first model, according to the first loss function and the second loss function, comprises: acquiring a first weight corresponding to the first loss function and a second weight corresponding to the second loss function; and updating the model parameter of the first model according to the first loss function, the first weight, the second loss function and the second weight (Chen: the whole loss is calculated as a weighted sum of a decoding, cross-entropy loss ly and alignment loss lz in the form of mean-square error (MSE) according to a weighting parameter λ [Sec 2.2 Objective Function - ¶01-04] - the examiner notes that while only one weighting factor, λ , is recited by Chen, it would be obvious to one of ordinary skill in the art to implement a similar weighting factor for the decoding loss ly via routine optimization and experimentation (see MPEP § 2144.05(II))). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation -¶01-03]. One of ordinary skill in the art would recognize the advantage of utilizing Chen's predicted latent representations of masked patches Zm with the base masked autoencoder framework of He provides context to inform what masked regions would be based on their visible counterparts. This leads to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. As for claim 6, He in view of Chen, further in view of Zhou The method according to claim 1 (described previously), wherein, the determining a fusion prediction feature of the first image block and the second image block, according to the first image feature, comprises: acquiring a first vector (Chen: generating mask queries Qm [Sec 2.1 Architecture - Subsec: Latent contextual regressor - ¶01- 02; Figs. 1-2]); fusing the first vector and the first image feature, to obtain a fusion vector; and obtaining the fusion prediction feature according to the fusion vector (Chen: the latent contextual regressor H predicts (i.e., fuses) latent representations of masked patches Zm (i.e., a fusion prediction feature) from the latent representations of visible patches Zv (i.e., a first image feature) and mask queries Qm (i.e., a first vector) [Sec 2.1. Architecture - Subsec: Latent contextual regressor - ¶01-02; Figs 1 & 2] – the examiner notes that the latent representations of masked patches Zm, are already in vector form, and thus also serve as the fusion vector(s)). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize the advantage of utilizing Chen's latent contextual regressor to generate latent representations of masked patches Zm from mask queries Qm and latent representations of visible patches Zv to inform overall image reconstruction during model training. This leads to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. Concerning claim 7, He in view of Chen, further in view of Zhou The method according to claim 1 (described previously), wherein, the acquiring a target image feature in a target image, comprises: determining a first region in the target image, according to a position of the second image block in the first sample image (Zhou: the positions of masked regions (i.e., a first region) are encoded into mask tokens for input image 203 (i.e., the target image) [¶0046; Fig. 3]); performing offset processing on the first region, to obtain a second region (Zhou: training images may be a crop (i.e., a type of offset processing) of a larger original image, to form the transformed crop (i.e., a second region) [¶0075]); and determining an image feature corresponding to an image block within the second region in the target image as the target image feature (Zhou: transformed crops of the original images can then be resized back to the training size and run through TAE as the training image to encode latent representations 304 of the transformed crop [¶0045-49; Figs. 3 & 5]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [¶0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the transform view training framework of Zhou with the combined MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. Regarding claim 11, He teach An electronic device, comprising: a processor and a memory (He: MAE training was conducted in 128 TPU-v3 cores [Sec 4.2 Comparisons with Previous Results - ¶01-04; Table 2] – the examiner notes that each TPU (tensor processing unit) acts as electronic device, containing an associated processor and memory); wherein, the memory stores computer execution instructions; and the processor executes the computer execution instructions stored in the memory, so that the processor executes a model training method (He: the masked autoencoder (MAE) training method outlined in Figure 1 [Sec 1. Introduction - ¶09-10 (pg. 2)]), which comprises: acquiring a first sample image, wherein the first sample image comprises a first image block which is uncovered and a second image block, which is covered (He: an input image (i.e., a first sample image) is divided into a subset of non-overlapping patches and a subset of those patches are masked (i.e., covered), resulting in both visible, unmasked patches and invisible, masked patches [Sec 3. Approach - Subsec: Masking - ¶01-02; Figs. 1 & 2]); processing the first sample image through a first model, to obtain a first image feature corresponding to the first image block (He: a ViT is applied to visible, unmasked patches resulting in embedded patches (i.e., a first image feature) with additional positional embeddings [Sec 3. Approach - Subsec: MAE encoder - ¶01; Fig. 1]); reconstructing the second image block according to the first image feature, to obtain a first image (He: a MAE decoder utilizes the embedded patches (again, the first image feature) and mask tokens embedding information of masked regions to then yield a reconstructed input image (i.e., a first image) [Sec 3. Approach - Subsec: MAE decoder - ¶01-02 & Subsec: Reconstruction target - ¶01-02; Fig. 1]), and updating a model parameter of the first model, according to the first image, the second image block, (He: a loss function computes the mean square error (MSE) between the reconstructed image (i.e., first image) and the original input image (i.e., a first sample image) on only the masked patches (i.e., the second image block) to update the decoder layers [Sec 3. Approach - Subsec: Reconstruction target - ¶01-02; Fig. 1]). He fails to recite determining a fusion prediction feature of the first and second image block, acquiring a target image after preprocessing the first sample image and using these to update the model parameter. Chen, however, teach and determining a fusion prediction feature of the first image block and the second image block, according to the first image feature; (Chen: the latent contextual regressor H predicts latent representations of masked patches Zm (i.e., a fusion prediction feature) from the latent representations of visible patches Zv (i.e., a first image feature) [Sec 2.1. Architecture - Subsec: Latent contextual regressor - ¶01-02; Figs. 1 & 2]) and updating a model parameter of the first model, according to the first image, the second image block, the fusion prediction feature (Chen: the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - 01-04; eq. 1; Fig. 2]). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the latent contextual regressor utilized by Chen with the base masked autoencoder framework of He provides context to inform what masked regions would be based on their visible counterparts, leading to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. Neither He nor Chen explain acquiring a target image feature after preprocessing the first sample image and utilizing it to update model parameters. Zhou, on the other hand, teach acquiring a target image feature in a target image, wherein the target image is an image after preprocessing of the first sample image (Zhou: in the TAE approach a an input image 203 which may be a crop of a larger image (i.e., a target image after preprocessing of the first sample image – the examiner notes here that, that cropping is described of preprocessing of the first sample image in accordance with [¶0073] of the specification of the instant application) is encoded into a set of latent representations 304 (i.e., a target image feature) [¶0045-49; Figs. 3 & 5]); and updating a model parameter of the first model, according to (Zhou: a loss between the reconstructed target y derived from features of the transformed input image x (i.e., a target image feature) is calculated to measure the discrepancy between the reconstruction and ground truth y' to update the decoder [¶0051-58; eqs. 2-3]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the transform view training framework of Zhou with the combined MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. With respect to claim 12, He in view of Chen teach A non-transitory computer readable storage medium, having computer execution instructions stored therein, wherein, the processor, upon executing the computer execution instructions (He: MAE training was conducted in 128 TPU-v3 cores [Sec 4.2 Comparisons with Previous Results - ¶01-04; Table 2] – the examiner notes that each TPU (tensor processing unit) contains an associated processor and memory for storing and executing computer instructions), implements the model training method according to claim 1 (as described previously). Turning to claim 17, He in view of Chen, further in view of Zhou The electronic device according to claim 11 (described previously), wherein, the updating a model parameter of the first model, according to the first image, the second image block, the fusion prediction feature and the target image feature, comprises: determining a first loss function of the first model, according to the first image and the second image block; updating the model parameter of the first model, according to the first loss function (He: a loss function computes the mean square error (MSE) between the reconstructed image (i.e., first image) and the original input image (i.e., a first sample image) on only the masked patches (i.e., the second image block) to update the decoder layers [Sec 3. Approach - Subsec: Reconstruction target - ¶01-02; Fig. 1]). He, however, fails to teach determining a second loss function according to a fusion prediction feature and target image feature. Chen, per contra, teach determining a second loss function of the first model, according to the fusion prediction feature updating the model parameter of the first model, according to (Chen: the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - ¶01-04; eq. 1; Fig. 2]). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the loss driven by the predicted latent representations of masked patches from Chen in tandem with the loss taught by He optimizes masked region reconstruction, leading to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. He in view of Chen, however, fail to teach determining the second loss function according to the target image feature. Zhou, on the other hand, teach determining a second loss function of the first model, according to updating the model parameter of the first model, according to (Zhou: a loss between the reconstructed target y derived from features of the transformed input image x (i.e., a target image feature) is calculated to measure the discrepancy between the reconstruction and ground truth y' to update the decoder [¶0051-58; eqs. 2-3]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [¶0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the loss from transformed views of Zhou with the loss functions taught by the MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. Concerning claim 18, He in view of Chen, further in view of Zhou teach The electronic device according to claim 17 (outlined above), wherein, the determining a second loss function of the first model, according to the fusion prediction feature and the target image feature, comprises: acquiring a second sample image, and determining a second image feature of the second sample image, wherein the second sample image is an image other than the first sample image (Chen: the CAE architecture is trained using the imageNet-1K dataset comprising a plurality of images for model training as illustrated in Fig. 6 [Sec 4.1 Implementation - 01-02] visible features Zv (second image features) are determined from visible patches Xv via the encoder F [Sec 2.1 - Architecture - ¶01-02; Fig. 2]); and determining the second loss function according to the fusion prediction feature, the target image feature, (Chen: the latent contextual regressor H predicts latent representations of masked patches Zm (i.e., a fusion prediction feature) from the latent representations of visible patches Zv (i.e., a second image feature) [Sec 2.1. Architecture - Subsec: Latent contextual regressor - 01-02; Figs 1 & 2], wherein the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - ¶01-04; eq. 1; Fig. 2] – the examiner notes that here, given that the latent representations of masked patches (the fusion prediction feature) are derived from the latent representations of visual features (a second image feature), the loss calculated using the latent representations of masked patches is inherently determined using the second image features). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the loss driven by the predicted latent representations of masked patches from Chen in tandem with the loss taught by He optimizes masked region reconstruction, leading to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. Chen again fails to disclose using a target image to determine the second loss. Zhou, however, teach determining the second loss function according to (Zhou: a loss between the reconstructed target y derived from features of the transformed input image x (i.e., a target image feature) is calculated to measure the discrepancy between the reconstruction and ground truth y' to update the decoder [¶0051-58; eqs. 2-3]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [¶0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the loss from transformed views of Zhou with the loss functions taught by the MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. As for claim 20, He in view of Chen, further in view of Zhou teach The electronic device according to claim 17 (outlined previously), wherein, the updating the model parameter of the first model, according to the first loss function and the second loss function, comprises: acquiring a first weight corresponding to the first loss function and a second weight corresponding to the second loss function; and updating the model parameter of the first model according to the first loss function, the first weight, the second loss function and the second weight (Chen: the whole loss is calculated as a weighted sum of a decoding, cross-entropy loss ly and alignment loss lz in the form of mean-square error (MSE) according to a weighting parameter λ [Sec 2.2 Objective Function - ¶01-04] - the examiner notes that while only one weighting factor, λ , is recited by Chen, it would be obvious to one of ordinary skill in the art to implement a similar weighting factor for the decoding loss ly via routine optimization and experimentation (see MPEP § 2144.05(II))). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation -¶01-03]. One of ordinary skill in the art would recognize the advantage of utilizing Chen's predicted latent representations of masked patches Zm with the base masked autoencoder framework of He provides context to inform what masked regions would be based on their visible counterparts. This leads to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. With regard to claim 21, He in view of Chen, further in view of Zhou teach The electronic device according to claim 11 (described previously), wherein, the determining a fusion prediction feature of the first image block and the second image block, according to the first image feature, comprises: acquiring a first vector (Chen: generating mask queries Qm [Sec 2.1 Architecture - Subsec: Latent contextual regressor - ¶01- 02; Figs. 1-2]); fusing the first vector and the first image feature, to obtain a fusion vector; and obtaining the fusion prediction feature according to the fusion vector (Chen: the latent contextual regressor H predicts (i.e., fuses) latent representations of masked patches Zm (i.e., a fusion prediction feature) from the latent representations of visible patches Zv (i.e., a first image feature) and mask queries Qm (i.e., a first vector) [Sec 2.1. Architecture - Subsec: Latent contextual regressor - ¶01-02; Figs 1 & 2] – the examiner notes that the latent representations of masked patches Zm, are already in vector form, and thus also serve as the fusion vector(s)). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize the advantage of utilizing Chen's latent contextual regressor to generate latent representations of masked patches Zm from mask queries Qm and latent representations of visible patches Zv to inform overall image reconstruction during model training. This leads to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. Considering claim 22, He in view of Chen, further in view of Zhou teach The electronic device according to claim 11, wherein, the acquiring a target image feature in a target image, comprises: determining a first region in the target image, according to a position of the second image block in the first sample image (Zhou: the positions of masked regions (i.e., a first region) are encoded into mask tokens for input image 203 (i.e., the target image) [¶0046; Fig. 3]); performing offset processing on the first region, to obtain a second region (Zhou: training images may be a crop (i.e., a type of offset processing) of a larger original image, to form the transformed crop (i.e., a second region) [¶0075]); and determining an image feature corresponding to an image block within the second region in the target image as the target image feature (Zhou: transformed crops of the original images can then be resized back to the training size and run through TAE as the training image to encode latent representations 304 of the transformed crop [¶0045-49; Figs. 3 & 5]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [¶0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the transform view training framework of Zhou with the combined MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. Turning to claim 23, He in view of Chen teach The non-transient computer readable storage medium according to claim 12 (described previously), wherein, the updating a model parameter of the first model, according to the first image, the second image block, the fusion prediction feature and the target image feature, comprises: updating the model parameter of the first model, according to the first loss function (He: a loss function computes the mean square error (MSE) between the reconstructed image (i.e., first image) and the original input image (i.e., a first sample image) on only the masked patches (i.e., the second image block) to update the decoder layers [Sec 3. Approach - Subsec: Reconstruction target - ¶01-02; Fig. 1]). He, however, fails to teach determining a second loss function according to a fusion prediction feature and target image feature. Chen, per contra, teach determining a second loss function of the first model, according to the fusion prediction feature updating the model parameter of the first model, according to (Chen: the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - ¶01-04; eq. 1; Fig. 2]). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the loss driven by the predicted latent representations of masked patches from Chen in tandem with the loss taught by He optimizes masked region reconstruction, leading to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. He in view of Chen, however, fail to teach determining the second loss function according to the target image feature. Zhou, on the other hand, teach determining a second loss function of the first model, according to updating the model parameter of the first model, according to (Zhou: a loss between the reconstructed target y derived from features of the transformed input image x (i.e., a target image feature) is calculated to measure the discrepancy between the reconstruction and ground truth y' to update the decoder [¶0051-58; eqs. 2-3]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [¶0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the loss from transformed views of Zhou with the loss functions taught by the MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. With respect to claim 24, He in view of Chen teach The non-transient computer readable storage medium according to claim 23 (described above), wherein, the determining a second loss function of the first model, according to the fusion prediction feature and the target image feature, comprises: acquiring a second sample image, and determining a second image feature of the second sample image, wherein the second sample image is an image other than the first sample image (Chen: the CAE architecture is trained using the imageNet-1K dataset comprising a plurality of images for model training as illustrated in Fig. 6 [Sec 4.1 Implementation - 01-02] visible features Zv (second image features) are determined from visible patches Xv via the encoder F [Sec 2.1 - Architecture - ¶01-02; Fig. 2]); and determining the second loss function according to the fusion prediction feature, the target image feature, (Chen: the latent contextual regressor H predicts latent representations of masked patches Zm (i.e., a fusion prediction feature) from the latent representations of visible patches Zv (i.e., a second image feature) [Sec 2.1. Architecture - Subsec: Latent contextual regressor - 01-02; Figs 1 & 2], wherein the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - ¶01-04; eq. 1; Fig. 2] – the examiner notes that here, given that the latent representations of masked patches (the fusion prediction feature) are derived from the latent representations of visual features (a second image feature), the loss calculated using the latent representations of masked patches is inherently determined using the second image features). Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the loss driven by the predicted latent representations of masked patches from Chen in tandem with the loss taught by He optimizes masked region reconstruction, leading to overall superior downstream performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. Chen again fails to disclose using a target image to determine the second loss. Zhou, however, teach determining the second loss function according to (Zhou: a loss between the reconstructed target y derived from features of the transformed input image x (i.e., a target image feature) is calculated to measure the discrepancy between the reconstruction and ground truth y' to update the decoder [¶0051-58; eqs. 2-3]). Zhou further states that their transformed autoencoder (TAE) is compatible with other self-supervised learning methods such as MAEs and that integrating the transformed image reconstruction task in TAE with an MAE-like framework improves overall performance [¶0080]. It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to utilize the loss from transformed views of Zhou with the loss functions taught by the MAE-CAE architecture taught by He in view of Chen to improve overall image reconstruction during training. Claim(s) 8, 15 & 16 is/are rejected under 35 U.S.C. 103 as being unpatentable over He et al ("Masked Autoencoders are Scalable Vision Learners", Dec 19, 2021, arXiv), hereinafter referred to as “He”, in view of Chen et al ("Context Autoencoder for Self-Supervised Representation Learning", Feb 07 2022 (v1), arXiv), hereinafter referred to as “Chen”. Regarding claim 8, He teach An image processing method (He: image processing methods using the pre-trained model are detailed in Sec 5. Transfer Learning Experiments [with particular attention to the classification tasks [¶06; table 6]), comprising: processing a plurality of images through a first model, to obtain a plurality of image features corresponding to the plurality of images (He: a ViT is applied to visible, unmasked patches resulting in embedded patches (i.e., an image feature) with additional positional embeddings [Sec 3. Approach - Subsec: MAE encoder - ¶01; Fig. 1], which is trained across a plurality of images in the IN1K dataset [Sec 5. Transfer Learning Experiments - ¶01-07; table 6]), wherein the first model is a model obtained through training of reconstructed image comparative learning (He: a loss function computes the mean square error (MSE) between the reconstructed image (i.e., first image) and the original input image (i.e., a first sample image) on only the masked patches (i.e., the second image block) to update the decoder layers [Sec 3. Approach - Subsec: Reconstruction target - ¶01-02; Fig. 1] – the examiner notes that here, the loss between the reconstructed and original input image informs the comparative learning used during model training); and classifying the plurality of images, according to the plurality of image features (He: the pre-trained model is utilized for downstream classification tasks using the embedded patches and compared to other computer vision architectures [Sec 5. Transfer Learning - ¶06; table 6]). He, however, fails to disclose utilizing predictive feature comparative learning for model training. Chen, per contra, is analogous art pertinent to the field of endeavor of the present application and disclose a context autoencoder (CAE) that utilizes a context regressor to predict masked patch representations from visible patches and masked queries. Chen teach wherein the first model is a model obtained through training of reconstructed image comparative learning combined with predicted feature comparative learning (Chen: the latent contextual regressor H predicts latent representations of masked patches Zm (i.e., a fusion prediction feature) from the latent representations of visible patches Zv (i.e., a first image feature) [Sec 2.1. Architecture - Subsec: Latent contextual regressor - ¶01-02; Figs. 1 & 2], the latent representations of masked patches Zm (i.e., a fusion prediction feature) inform an alignment loss l2 of the decoder [Sec 2.2 Objective Function - ¶01-04; eq. 1; Fig. 2]) Chen compares their context autoencoder to human recognition, wherein the latent contextual regressor H forms a plausible hypothesis for what the masked (i.e., covered) patches of the image are based on the positional relationship and classification of visible (i.e., visible) portions [Sec 3.1 Analysis - Subsec: Intuitive Interpretation - ¶01-03]. One of ordinary skill in the art would recognize that implementing the latent contextual regressor utilized by Chen with the base masked autoencoder framework of He provides context to inform what masked regions would be based on their visible counterparts during model training. This leads to overall superior downstream classification performance after training compared to the base MAE framework alone [Sec 4.4 Downstream Tasks - ¶01-02; Table 3]. With respect to claim 15, He in view of Chen teach An electronic device, comprising: a processor and a memory; wherein, the memory stores computer execution instructions; and the processor executes the computer execution instructions stored in the memory (He: MAE training was conducted in 128 TPU-v3 cores [Sec 4.2 Comparisons with Previous Results - ¶01-04; Table 2] – the examiner notes that each TPU (tensor processing unit) acts as electronic device, containing an associated processor and memory), so that the processor executes the image processing method according to claim 8 (as described above). 16. A non-transitory computer readable storage medium, having computer execution instructions stored therein, wherein, the processor, upon executing the computer execution instructions (He: MAE training was conducted in 128 TPU-v3 cores [Sec 4.2 Comparisons with Previous Results - ¶01-04; Table 2] – the examiner notes that each TPU (tensor processing unit) acts as electronic device, containing an associated processor and memory for storing and executing computer instructions), implements the image processing method according to claim 8 (described previously). Allowable Subject Matter Claims 4 & 19 are objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims. Regarding claim 4, the primary reason for the indication of allowable subject matter is that the prior art fail to teach or reasonably suggest: “acquiring a first similarity between the fusion prediction feature and the target image feature;” The closest cited prior art, Chen, disclose generating a latent representation of masked patches Zm (i.e., a fusion prediction feature) but the alignment loss is dictated by the similarity between the latent representations of masked patches Zm and representations of masked patches Z - m. The next closes prior art, Zhou, teach obtaining features transformed image (i.e., a target image feature), wherein the transformed image is generated via a transformation of the original training image (in the form of a cro), and determines a loss based on the transformed image features and the reconstructed image features. Neither of these references describe a similarity between the fusion prediction feature and target image, nor do they provide a reasonable motivation for combining the individual elements they teach to determine a first similarity between those two features. Claim 19, which recites substantially identical subject matter, is objected to for the same analysis applied to claim 4. Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure: Chen et al (US 2024/0221128 A1) describe training a masked autoencoder network utilizing a feature prediction output for image inpainting. Caron et al (“Emerging Properties in Self-Supervised Vision Transformers”, 2021, arXiv) detail a self-supervised training method, deemed “DINO” for adaptive image classification. Zhou et al (“iBOT: Image BERT Pre-Training with Online Tokenizer”, 2022, arXiv), detail a self-supervised framework for performing masked predictions using an online tokenizer for semantic segmentation. Wu et al (“Object-wise Masked Autoencoders for Fast Pre-training, 2022. arXiv) describe an augmentation to the base MAE framework that encodes embeddings for objects in images. Any inquiry concerning this communication or earlier communications from the examiner should be directed to Michael M. Sofroniou whose telephone number is (571)272-0287. The examiner can normally be reached M-F: 8:30 AM - 5:00 PM. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, John M. Villecco can be reached at (571) 272-7319. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /MICHAEL M SOFRONIOU/Examiner, Art Unit 2661 /JOHN VILLECCO/Supervisory Patent Examiner, Art Unit 2661
Read full office action

Prosecution Timeline

Dec 27, 2024
Application Filed
Aug 25, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12711652
IMAGE PROCESSING APPARATUS, IMAGE PICKUP APPARATUS, CONTROL METHOD FOR IMAGE PROCESSING APPARATUS, AND STORAGE MEDIUM CAPABLE OF NOTIFYING USER OF BLUR INFORMATION, OR OF DISPLAYING INDICATOR
2y 8m to grant Granted Aug 18, 2026
Study what changed to get past this examiner. Based on 1 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
100%
Grant Probability
99%
With Interview (+0.0%)
2y 3m (~6m remaining)
Median Time to Grant
Low
PTA Risk
Based on 3 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month