Prosecution Insights
Last updated: October 02, 2026
Application No. 19/102,347

METHOD, APPARATUS, DEVICE AND MEDIUM FOR OBJECT PROCESSING BASED ON PRE-TRAINING AND TWO-PHASE DEPLOYMENT

Non-Final OA §102§103§112
Filed
Feb 07, 2025
Priority
Apr 06, 2023 — CN 202310363149.5 +1 more
Examiner
ELLIOTT, JORDAN MCKENZIE
Art Unit
Tech Center
Assignee
Lemon Inc.
OA Round
1 (Non-Final)
47%
Grant Probability
Moderate
1-2
OA Rounds
1y 3m
Est. Remaining
51%
With Interview

Examiner Intelligence

Grants 47% of resolved cases
47%
Career Allowance Rate
15 granted / 32 resolved
-13.1% vs TC avg
Minimal +4% lift
Without
With
+4.2%
Interview Lift
resolved cases with interview
Typical timeline
2y 11m
Avg Prosecution
25 currently pending
Career history
69
Total Applications
across all art units

Statute-Specific Performance

§101
7.9%
-32.1% vs TC avg
§103
55.8%
+15.8% vs TC avg
§102
25.2%
-14.8% vs TC avg
§112
11.2%
-28.8% vs TC avg
Black line = Tech Center average estimate • Based on career data from 32 resolved cases

Office Action

§102 §103 §112
DETAILED ACTION Claims 1-6 and 13-24 are pending in this application. In view of the preliminary amendment filed on 03/06/2026, claims 7-12 are cancelled, claims 13 and 14 are amended and claims 15-24 are newly added. Claims 1-16 and 13-24 have been examined under the priority date of 04/06/2023 in accordance with applicant’s claim for foreign priority. Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Priority Receipt is acknowledged of certified copies of papers required by 37 CFR 1.55. Information Disclosure Statement The information disclosure statement (IDS) submitted on 02/26/2025 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(d): (d) REFERENCE IN DEPENDENT FORMS.—Subject to subsection (e), a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. The following is a quotation of pre-AIA 35 U.S.C. 112, fourth paragraph: Subject to the following paragraph [i.e., the fifth paragraph of pre-AIA 35 U.S.C. 112], a claim in dependent form shall contain a reference to a claim previously set forth and then specify a further limitation of the subject matter claimed. A claim in dependent form shall be construed to incorporate by reference all the limitations of the claim to which it refers. Claim 16 is rejected under 35 U.S.C. 112(d) or pre-AIA 35 U.S.C. 112, 4th paragraph, as being of improper dependent form for failing to further limit the subject matter of the claim upon which it depends, or for failing to include all the limitations of the claim upon which it depends. Claim 16 recites the limitation of an “electronic device according to claim 14”, however claim 14 does not include an electronic device. Therefore, the claim references a limitation which has not been previously introduced in the claims from which it depends. Applicant may cancel the claim(s), amend the claim(s) to place the claim(s) in proper dependent form, rewrite the claim(s) in independent form, or present a sufficient showing that the dependent claim(s) complies with the statutory requirements. Claim Rejections - 35 USC § 102 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention. Claims 1-2, 5-6, 13-15, 18-20 and 23-24 are rejected under 35 U.S.C. 102(a)(1) as being anticipated by Pyzow (US 20230065870 A1, filed 08-30-2022). Regarding claim 1 Pyzow discloses; A method for object processing, comprising: in response to receiving a predetermined operation by a user on a first selection control for pre-trained at least one generic model presented in a user interface, selecting the at least one generic model (Pyzow, [0042] the system has a user interface, [0044] where the user interface presents the user with clustering parameters to be inputted to a machine learning model, where the user is also presented with “mode” selection options which are used to generate a machine learning model according to the parameters (generic model)); PNG media_image1.png 74 346 media_image1.png Greyscale (Pyzow, [0042]) PNG media_image2.png 158 342 media_image2.png Greyscale PNG media_image3.png 236 348 media_image3.png Greyscale (Pyzow, [0044]) at least acquiring at least one generic feature that is generated by the at least one generic model and associated with a sample of an object of a target category among a plurality of categories (Pyzow, [0323] the first trained model (generic model) may parse input data objects (first object sample) and generate an output, where [0324] the output is a second set of features generated from the first features/data sets, where the features are descriptive fields which indicate object types from multiple types (categories), where [0051] the generated second features may correspond to an object where [0060] the feature is displayed to an interface as belonging to a cluster/category after the model outputs this); PNG media_image4.png 38 342 media_image4.png Greyscale PNG media_image5.png 278 342 media_image5.png Greyscale (Pyzow, [0324]) PNG media_image6.png 162 342 media_image6.png Greyscale (Pyzow, [0051]) and training an individual model for processing the object of the target category at least based at least on the at least one generic feature and annotation information of the sample of the object of the target category (Pyzow, [0325] a second machine learning model (individual model) may be trained based on the second features/feature output from the first model (generic model) and output a ranking or modification of the input object (sample of the object) as shown in figure 16, where [0072] the clustering and image feature classification can be updated by the user via the user interface on the images using annotations as shown in figure 8B, 834B, “these annotations” are used in training the machine learning models, where according to paragraph [0322] the methods shown in figure 8A-8C are used in training and feature generation and modification for training the first and second machine learning models). PNG media_image7.png 224 344 media_image7.png Greyscale (Pyzow, [0325]) PNG media_image8.png 310 344 media_image8.png Greyscale PNG media_image9.png 202 344 media_image9.png Greyscale (Pyzow, [0072]) PNG media_image10.png 408 784 media_image10.png Greyscale (Pyzow, figure 8B) PNG media_image11.png 830 690 media_image11.png Greyscale (Pyzow, figure 16) Regarding claim 2 Pyzow discloses; The method according to claim 1, wherein at least acquiring the at least one generic feature comprises: in response to receiving a predetermined operation by the user on a first selection control corresponding to a target generic model of the at least one generic model (Pyzow, [0056] the user is prompted to select at least one clustering model to add the pipeline, further the user may as select at least one characteristic of a model as present to select or indicate the one or more machine learning model structures (selecting a generic target model)), presenting, in the user interface, a second selection control for at least one intermediate feature of the target generic model (Pyzow, [0057] the user interface has at least one feature impact model generated which presents at least one feature generated by the model as well as the clusters generated by the first model/target generic model, where [0065] the user interface can prompt the user to modify at least one parameter of the clusters for the intermediate feature, then a modified feature is output from the model (target generic model), [0121]-[0122] the system generates intermediate inputs from a parent task to be input into subsequent models for subsequent modeling tasks indicating that the modified or selected features and inputs are used as intermediate inputs as shown in figure 11); PNG media_image12.png 830 554 media_image12.png Greyscale (Pyzow, Figure 11 emphasis added) and in response to receiving a predetermined operation by the user on the second selection control, acquiring the generic feature and at least one intermediate feature that are generated by the target generic model and that are associated with the sample of the object of the target category (Pyzow, [0057] the user interface has at least one feature impact model generated which presents at least one feature generated by the model as well as the clusters generated by the first model/target generic model, where [0065] the user interface can prompt the user to modify at least one parameter of the clusters for the intermediate feature, then a modified feature is output from the target generic model, [0121]-[0122] the system generates intermediate inputs from a parent task to be input into subsequent models for subsequent modeling tasks indicating that the modified or selected features and inputs are used as intermediate inputs as shown in figure 11, [0325] a second machine learning model (individual model) may be trained based on the second features/feature output from the first model (generic feature) and output a ranking or modification of the input object (sample of the object) as shown in figure 16, where [0072] the clustering and image feature classification can be updated by the user via the user interface on the images using annotations as shown in figure 8B, 834B, “these annotations” are used in training the machine learning models, where according to paragraph [0322] the methods shown in figure 8A-8C are used in training and feature generation and modification for training the first and second machine learning models, therefore both a generic feature and an intermediate feature are generated). Regarding claim 5 Pyzow discloses; The method according to claim 1, wherein the object comprises an image (Pyzow, [0050] the model is trained using data structures which comprise images and text associated with the images), and the method further comprises: acquiring image information and text information of a plurality of training image samples (Pyzow, [0050] the first data structure acquired corresponds to text, and the second corresponds to an image); PNG media_image13.png 98 304 media_image13.png Greyscale PNG media_image14.png 30 304 media_image14.png Greyscale (Pyzow, [0050]) and pre-training a target generic model of the at least one generic model based on a matching degree between the image information and the text information (Pyzow, [0069] the image and text data may be clustered by the cluster insights model (target generic model) to generate text classification showing how text inputs correspond to objects appearing in the images (matching degree), where the model has been trained with these inputs, where according to [0066] the image classification presentation portion is an output of the cluster insights model (target generic model)). PNG media_image15.png 252 304 media_image15.png Greyscale PNG media_image16.png 264 314 media_image16.png Greyscale (Pyzow, [0069] emphasis added) Regarding claim 6 the combination of Pyzow and Yu teaches; The method according to claim 5, wherein the target generic model comprises an image encoder and a text encoder (Pyzow, [0046] the model structure includes both a text and an image processor (text and image encoders) as shown in figure 2, where the image processor (240 in fig. 2) can be a pre-trained squeeze net, which can be used as an image encoder and the text processor (242 in fig. 2) can be a Truncated SVD, which also acts as an encoder), PNG media_image17.png 248 304 media_image17.png Greyscale (Pyzow, [0046]) PNG media_image18.png 310 714 media_image18.png Greyscale (Pyzow, figure 2, emphasis added) and pre-training the target generic model comprises: generating, by the image encoder, a plurality of image features based on the image information of the plurality of training image samples (Pyzow, [0050] the data processing system generates first and second features using the feature generator (includes the image encoder), in which the second features correspond to images, where according to paragraph [0054] the data processing system uses the structure as specified in figure 2 (includes the image and text encoders respectively) to generate the features, [0055] where the features are extracted from the text and image data structures respectively indicating the corresponding encoders are used for each); generating, by the text encoder, a plurality of text features based on the text information of the plurality of training image samples (Pyzow, [0050] the data processing system generates first and second features using the feature generator (includes the image encoder), in which the first features correspond to text features/information, where according to paragraph [0054] the data processing system uses the structure as specified in figure 2 (includes the image and text encoders respectively) to generate the features, [0055] where the features are extracted from the text and image data structures respectively indicating the corresponding encoders are used for each); determining a matching degree between a respective image feature of the plurality of image features and a respective text feature of the plurality of text features (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked); and training the image encoder and the text encoder to increase a matching degree between an image feature and a corresponding text feature of a target training image sample among the plurality of training image samples (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked and then the model (including both the text and vision encoders in the structure) may be retrained to increase correlation as described in [0072]), PNG media_image19.png 278 304 media_image19.png Greyscale PNG media_image20.png 180 308 media_image20.png Greyscale (Pyzow, [0070]) and to reduce a matching degree between the image feature of the target image sample and a text feature of another training image sample among the plurality of training image samples (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked and then the model (including both the text and vision encoders in the structure) may be retrained to increase correlation as described in [0072]), given that the ranking and thresholding is updated when the model is retrained to improve accurate correlation, this would also reduce the correlation of the text features which do not correspond to the image). Regarding claim 13 Pyzow discloses; An electronic device, comprising: at least one processor (Pyzow, [0076] the system has at least one processor included in a computer); and at least one memory, wherein the at least one memory is coupled to the at least one processor and stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the electronic device to perform acts comprising (Pyzow, [0078] the system has one or more processors couples to a memory for executing computer program instructions to carry out the method described): in response to receiving a predetermined operation by a user on a first selection control for pre- trained at least one generic model presented in a user interface, selecting the at least one generic model (Pyzow, [0042] the system has a user interface, [0044] where the user interface presents the user with clustering parameters to be inputted to a machine learning model, where the user is also presented with “mode” selection options which are used to generate a machine learning model according to the parameters (generic model)); at least acquiring at least one generic feature that is generated by the at least one generic model and associated with a sample of an object of a target category among a plurality of categories (Pyzow, [0323] the first trained model (generic model) may parse input data objects (first object sample) and generate an output, where [0324] the output is a second set of features generated from the first features/data sets, where the features are descriptive fields which indicate object types from multiple types (categories), where [0051] the generated second features may correspond to an object where [0060] the feature is displayed to an interface as belonging to a cluster/category after the model outputs this); and training an individual model for processing the object of the target category at least based at least on the at least one generic feature and annotation information of the sample of the object of the target category (Pyzow, [0325] a second machine learning model (individual model) may be trained based on the second features/feature output from the first model (generic feature) and output a ranking or modification of the input object (sample of the object) as shown in figure 16, where [0072] the clustering and image feature classification can be updated by the user via the user interface on the images using annotations as shown in figure 8B, 834B, “these annotations” are used in training the machine learning models, where according to paragraph [0322] the methods shown in figure 8A-8C are used in training and feature generation and modification for training the first and second machine learning models). Regarding claim 14 Pyzow discloses; A non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, acts comprising (Pyzow, [0078] the system has one or more processors couples to a memory for executing computer program instructions to carry out the method described): in response to receiving a predetermined operation by a user on a first selection control for pre- trained at least one generic model presented in a user interface, selecting the at least one generic model (Pyzow, [0042] the system has a user interface, [0044] where the user interface presents the user with clustering parameters to be inputted to a machine learning model, where the user is also presented with “mode” selection options which are used to generate a machine learning model according to the parameters (generic model)); at least acquiring at least one generic feature that is generated by the at least one generic model and associated with a sample of an object of a target category among a plurality of categories (Pyzow, [0323] the first trained model (generic model) may parse input data objects (first object sample) and generate an output, where [0324] the output is a second set of features generated from the first features/data sets, where the features are descriptive fields which indicate object types from multiple types (categories), where [0051] the generated second features may correspond to an object where [0060] the feature is displayed to an interface as belonging to a cluster/category after the model outputs this); and training an individual model for processing the object of the target category at least based at least on the at least one generic feature and annotation information of the sample of the object of the target category (Pyzow, [0325] a second machine learning model (individual model) may be trained based on the second features/feature output from the first model (generic feature) and output a ranking or modification of the input object (sample of the object) as shown in figure 16, where [0072] the clustering and image feature classification can be updated by the user via the user interface on the images using annotations as shown in figure 8B, 834B, “these annotations” are used in training the machine learning models, where according to paragraph [0322] the methods shown in figure 8A-8C are used in training and feature generation and modification for training the first and second machine learning models). Regarding claim 15 Pyzow discloses; The electronic device according to claim 13, wherein at least acquiring the at least one generic feature comprises: in response to receiving a predetermined operation by the user on a first selection control corresponding to a target generic model of the at least one generic model (Pyzow, [0056] the user is prompted to select at least one clustering model to add the pipeline, further the user may as select at least one characteristic of a model as present to select or indicate the one or more machine learning model structures (selecting a generic target model)), presenting, in the user interface, a second selection control for at least one intermediate feature of the target generic model (Pyzow, [0057] the user interface has at least one feature impact model generated which presents at least one feature generated by the model as well as the clusters generated by the first model/target generic model, where [0065] the user interface can prompt the user to modify at least one parameter of the clusters for the intermediate feature, then a modified feature is output from the target generic model, [0121]-[0122] the system generates intermediate inputs from a parent task to be input into subsequent models for subsequent modeling tasks indicating that the modified or selected features and inputs are used as intermediate inputs as shown in figure 11); and in response to receiving a predetermined operation by the user on the second selection control, acquiring the generic feature and at least one intermediate feature that are generated by the target generic model and that are associated with the sample of the object of the target category (Pyzow, [0057] the user interface has at least one feature impact model generated which presents at least one feature generated by the model as well as the clusters generated by the first model/target generic model, where [0065] the user interface can prompt the user to modify at least one parameter of the clusters for the intermediate feature, then a modified feature is output from the target generic model, [0121]-[0122] the system generates intermediate inputs from a parent task to be input into subsequent models for subsequent modeling tasks indicating that the modified or selected features and inputs are used as intermediate inputs as shown in figure 11, [0325] a second machine learning model (individual model) may be trained based on the second features/feature output from the first model (generic feature) and output a ranking or modification of the input object (sample of the object) as shown in figure 16, where [0072] the clustering and image feature classification can be updated by the user via the user interface on the images using annotations as shown in figure 8B, 834B, “these annotations” are used in training the machine learning models, where according to paragraph [0322] the methods shown in figure 8A-8C are used in training and feature generation and modification for training the first and second machine learning models, therefore both a generic feature and an intermediate feature are generated). Regarding claim 18 Pyzow discloses; The electronic device according to claim 13, wherein the object comprises an image (Pyzow, [0050] the model is trained using data structures which comprise images and text associated with the images), and the acts further comprise: acquiring image information and text information of a plurality of training image samples (Pyzow, [0050] the first data structure acquired corresponds to text, and the second corresponds to an image); and pre-training a target generic model of the at least one generic model based on a matching degree between the image information and the text information (Pyzow, [0069] the image and text data may be clustered by the cluster insights model (target generic model) to generate text classification showing how text inputs correspond to objects appearing in the images (matching degree), where the model has been trained with these inputs, where according to [0066] the image classification presentation portion is an output of the cluster insights model (target generic model)). Regarding claim 19 Pyzow discloses; The electronic device according to claim 18, wherein the target generic model comprises an image encoder and a text encoder (Pyzow, [0046] the model structure includes both a text and an image processor (text and image encoders) as shown in figure 2, where the image processor (240 in fig. 2) can be a pre-trained squeeze net, which can be used as an image encoder and the text processor (242 in fig. 2) can be a Truncated SVD, which also acts as an encoder) and pre-training the target generic model comprises: generating, by the image encoder, a plurality of image features based on the image information of the plurality of training image samples (Pyzow, [0050] the data processing system generates first and second features using the feature generator (includes the image encoder), in which the second features correspond to images, where according to paragraph [0054] the data processing system uses the structure as specified in figure 2 (includes the image and text encoders respectively) to generate the features, [0055] where the features are extracted from the text and image data structures respectively indicating the corresponding encoders are used for each); generating, by the text encoder, a plurality of text features based on the text information of the plurality of training image samples (Pyzow, [0050] the data processing system generates first and second features using the feature generator (includes the image encoder), in which the first features correspond to text features/information, where according to paragraph [0054] the data processing system uses the structure as specified in figure 2 (includes the image and text encoders respectively) to generate the features, [0055] where the features are extracted from the text and image data structures respectively indicating the corresponding encoders are used for each); determining a matching degree between a respective image feature of the plurality of image features and a respective text feature of the plurality of text features (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked); and training the image encoder and the text encoder to increase a matching degree between an image feature and a corresponding text feature of a target training image sample among the plurality of training image samples (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked and then the model (including both the text and vision encoders in the structure) may be retrained to increase correlation as described in [0072]), and to reduce a matching degree between the image feature of the target image sample and a text feature of another training image sample among the plurality of training image samples (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked and then the model (including both the text and vision encoders in the structure) may be retrained to increase correlation as described in [0072]), given that the ranking and thresholding is updated when the model is retrained to improve accurate correlation, this would also reduce the correlation of the text features which do not correspond to the image). Regarding claim 20, Pyzow discloses; The non-transitory computer-readable storage medium according to claim 14, wherein at least acquiring the at least one generic feature comprises: in response to receiving a predetermined operation by the user on a first selection control corresponding to a target generic model of the at least one generic model (Pyzow, [0056] the user is prompted to select at least one clustering model to add the pipeline, further the user may as select at least one characteristic of a model as present to select or indicate the one or more machine learning model structures (selecting a generic target model)), presenting, in the user interface, a second selection control for at least one intermediate feature of the target generic model (Pyzow, [0057] the user interface has at least one feature impact model generated which presents at least one feature generated by the model as well as the clusters generated by the first model/target generic model, where [0065] the user interface can prompt the user to modify at least one parameter of the clusters for the intermediate feature, then a modified feature is output from the target generic model, [0121]-[0122] the system generates intermediate inputs from a parent task to be input into subsequent models for subsequent modeling tasks indicating that the modified or selected features and inputs are used as intermediate inputs as shown in figure 11); and in response to receiving a predetermined operation by the user on the second selection control, acquiring the generic feature and at least one intermediate feature that are generated by the target generic model and that are associated with the sample of the object of the target category (Pyzow, [0057] the user interface has at least one feature impact model generated which presents at least one feature generated by the model as well as the clusters generated by the first model/target generic model, where [0065] the user interface can prompt the user to modify at least one parameter of the clusters for the intermediate feature, then a modified feature is output from the target generic model, [0121]-[0122] the system generates intermediate inputs from a parent task to be input into subsequent models for subsequent modeling tasks indicating that the modified or selected features and inputs are used as intermediate inputs as shown in figure 11, [0325] a second machine learning model (individual model) may be trained based on the second features/feature output from the first model (generic feature) and output a ranking or modification of the input object (sample of the object) as shown in figure 16, where [0072] the clustering and image feature classification can be updated by the user via the user interface on the images using annotations as shown in figure 8B, 834B, “these annotations” are used in training the machine learning models, where according to paragraph [0322] the methods shown in figure 8A-8C are used in training and feature generation and modification for training the first and second machine learning models, therefore both a generic feature and an intermediate feature are generated). Regarding claim 23 Pyzow discloses; The non-transitory computer-readable storage medium according to claim 14, wherein the object comprises an image (Pyzow, [0050] the model is trained using data structures which comprise images and text associated with the images), and the acts further comprise: acquiring image information and text information of a plurality of training image samples (Pyzow, [0050] the first data structure acquired corresponds to text, and the second corresponds to an image); and pre-training a target generic model of the at least one generic model based on a matching degree between the image information and the text information (Pyzow, [0069] the image and text data may be clustered by the cluster insights model (target generic model) to generate text classification showing how text inputs correspond to objects appearing in the images (matching degree), where the model has been trained with these inputs, where according to [0066] the image classification presentation portion is an output of the cluster insights model (target generic model)). Regarding claim 24 Pyzow discloses; The non-transitory computer-readable storage medium according to claim 23, wherein the target generic model comprises an image encoder and a text encoder (Pyzow, [0046] the model structure includes both a text and an image processor (text and image encoders) as shown in figure 2, where the image processor (240 in fig. 2) can be a pre-trained squeeze net, which can be used as an image encoder and the text processor (242 in fig. 2) can be a Truncated SVD, which also acts as an encoder), and pre-training the target generic model comprises: generating, by the image encoder, a plurality of image features based on the image information of the plurality of training image samples (Pyzow, [0050] the data processing system generates first and second features using the feature generator (includes the image encoder), in which the second features correspond to images, where according to paragraph [0054] the data processing system uses the structure as specified in figure 2 (includes the image and text encoders respectively) to generate the features, [0055] where the features are extracted from the text and image data structures respectively indicating the corresponding encoders are used for each); generating, by the text encoder, a plurality of text features based on the text information of the plurality of training image samples (Pyzow, [0050] the data processing system generates first and second features using the feature generator (includes the image encoder), in which the first features correspond to text features/information, where according to paragraph [0054] the data processing system uses the structure as specified in figure 2 (includes the image and text encoders respectively) to generate the features, [0055] where the features are extracted from the text and image data structures respectively indicating the corresponding encoders are used for each); determining a matching degree between a respective image feature of the plurality of image features and a respective text feature of the plurality of text features (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked); and training the image encoder and the text encoder to increase a matching degree between an image feature and a corresponding text feature of a target training image sample among the plurality of training image samples (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked and then the model (including both the text and vision encoders in the structure) may be retrained to increase correlation as described in [0072]), and to reduce a matching degree between the image feature of the target image sample and a text feature of another training image sample among the plurality of training image samples (Pyzow, [0072] the correlation of text features with image features can be ranked using scalar values and thresholds can be used to modify the clusters and correlations using retraining or training of the model such that the thresholds are satisfied, [0069] correlation thresholds between the text and the images must be met, these are ranked, therefore the correlation may be ranked and then the model (including both the text and vision encoders in the structure) may be retrained to increase correlation as described in [0072]), given that the ranking and thresholding is updated when the model is retrained to improve accurate correlation, this would also reduce the correlation of the text features which do not correspond to the image). Claim Rejections - 35 USC § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. Claims 3-4, 16-17 and 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Pyzow (US 20230065870 A1, filed 08-30-2022) in view of Yu (CN 115131635 A, published 09-30-2022). Regarding claim 3 Pyzow fails to disclose; The method according to claim 2, wherein training the individual model comprises: generating a fusion feature based on the at least one intermediate feature and the generic feature; and training the individual model based on the fusion feature and the annotation information. However, in the same field of endeavor, Yu teaches; wherein training the individual model comprises: generating a fusion feature based on the at least one intermediate feature and the generic feature (Yu, [0008] a neural network model (generic model) is obtained, and a first sample feature is extracted from the sample image, where this is related to the category of the image (generic feature), [0009] the neural network is trained (target generic model), and a first memory feature is generated (intermediate feature), these two features, the first sample feature (generic feature) and the first memory feature (intermediate feature) are fused to form an output prediction image (fusion feature)); PNG media_image21.png 288 592 media_image21.png Greyscale (Yu, [0007]- [0009]) and training the individual model based on the fusion feature and the annotation information (Yu, [0010] the neural network model is trained based on the output predicted image/images (fusion feature) to obtain a trained image generation model (individual model), [0025] annotation results are acquired for each prediction image, therefore they would be included with the prediction image (fusion result) for training the image generation model (individual model)). PNG media_image22.png 96 580 media_image22.png Greyscale (Yu, [0025]) The combination of Pyzow and Yu would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the addition of the fusion feature results improves the trained models because the generation of the fused feature prediction image to train or update the models would allow the model to continuously learn new features and improve the content of the generated images or results. (Yu, [0043]) Regarding claim 4 The combination of Pyzow and Yu teaches; The method according to claim 1, wherein training the individual model comprises: generating, by the individual model, a processing result for the sample of the object of the target category based on the at least one generic feature (Yu, [0012] a reference image (sample) and the image generation model (individual model) are obtained, [0013] the image generation model (individual model) extracts target features from the reference image based on the category of the image (sample and target category information), [0014] the image generate model then selects target memory features which are features previously learned (generic feature) and [0015] the target memory feature (generic feature) and the target features (sample and target category information) are fused to obtained a fusion result (processing result)); PNG media_image23.png 310 588 media_image23.png Greyscale PNG media_image24.png 58 582 media_image24.png Greyscale (Yu, [0012]- [0015]) and training the individual model based on the processing result and the annotation information (Yu, [0182] the adjustment module is used to adjust the model to obtain or update the image generation model by using prediction images (obtained in the same way as the target images, analogous to the processing result), [0184] annotation results are acquired for each image [00184] the annotation results and the predicted images/processing results are used to adjust the model to obtain an adjusted generation model). PNG media_image25.png 326 588 media_image25.png Greyscale (Yu, [0182]- [0184]) The combination of Pyzow and Yu would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the use of the fused processing result, continuous adjusting and retraining allows for diversity in the content generated by the model (Yu, [0155]) Regarding claim 16 the combination of Pyzow and Yu teaches; The electronic device according to claim 14, wherein training the individual model comprises: generating a fusion feature based on the at least one intermediate feature and the generic feature (Yu, [0008] a neural network model (generic model) is obtained, and a first sample feature is extracted from the sample image, where this is related to the category of the image (generic feature), [0009] the neural network is trained (target generic model), and a first memory feature is generated (intermediate feature), these two features, the first sample feature (generic feature) and the first memory feature (intermediate feature) are fused to form an output prediction image (fusion feature); and training the individual model based on the fusion feature and the annotation information (Yu, [0010] the neural network model is trained based on the output predicted image/images (fusion feature) to obtain a trained image generation model (individual model), [0025] annotation results are acquired for each prediction image, therefore they would be included with the prediction image (fusion result) for training the image generation model (individual model)). The combination of Pyzow and Yu would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the addition of the fusion feature results improves the trained models because the generation of the fused feature prediction image to train or update the models would allow the model to continuously learn new features and improve the content of the generated images or results. (Yu, [0043]) Regarding claim 17 the combination of Pyzow and Yu teaches; The electronic device according to claim 13, wherein training the individual model comprises: generating, by the individual model, a processing result for the sample of the object of the target category based on the at least one generic feature (Yu, [0012] a reference image (sample) and the image generation model (individual model) are obtained, [0013] the image generation model (individual model) extracts target features from the reference image based on the category of the image (sample and target category information), [0014] the image generate model then selects target memory features which are features previously learned (generic feature) and [0015] the target memory feature (generic feature) and the target features (sample and target category information) are fused to obtained a fusion result (processing result)); and training the individual model based on the processing result and the annotation information (Yu, [0182] the adjustment module is used to adjust the model to obtain or update the image generation model by using prediction images (obtained in the same way as the target images, analogous to the processing result), [0184] annotation results are acquired for each image [00184] the annotation results and the predicted images/processing results are used to adjust the model to obtain an adjusted generation model). The combination of Pyzow and Yu would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the use of the fused processing result, continuous adjusting and retraining allows for diversity in the content generated by the model (Yu, [0155]) Regarding claim 21 the combination of Pyzow and Yu teaches; The non-transitory computer-readable storage medium according to claim 14, wherein training the individual model comprises: generating a fusion feature based on the at least one intermediate feature and the generic feature(Yu, [0008] a neural network model (generic model) is obtained, and a first sample feature is extracted from the sample image, where this is related to the category of the image (generic feature), [0009] the neural network is trained (target generic model), and a first memory feature is generated (intermediate feature), these two features, the first sample feature (generic feature) and the first memory feature (intermediate feature) are fused to form an output prediction image (fusion feature); and training the individual model based on the fusion feature and the annotation information (Yu, [0010] the neural network model is trained based on the output predicted image/images (fusion feature) to obtain a trained image generation model (individual model), [0025] annotation results are acquired for each prediction image, therefore they would be included with the prediction image (fusion result) for training the image generation model (individual model)). The combination of Pyzow and Yu would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the addition of the fusion feature results improves the trained models because the generation of the fused feature prediction image to train or update the models would allow the model to continuously learn new features and improve the content of the generated images or results. (Yu, [0043]) Regarding claim 22 the combination of Pyzow and Yu teaches; The non-transitory computer-readable storage medium according to claim 14, wherein training the individual model comprises: generating, by the individual model, a processing result for the sample of the object of the target category based on the at least one generic feature (Yu, [0012] a reference image (sample) and the image generation model (individual model) are obtained, [0013] the image generation model (individual model) extracts target features from the reference image based on the category of the image (sample and target category information), [0014] the image generate model then selects target memory features which are features previously learned (generic feature) and [0015] the target memory feature (generic feature) and the target features (sample and target category information) are fused to obtained a fusion result (processing result)); and training the individual model based on the processing result and the annotation information (Yu, [0182] the adjustment module is used to adjust the model to obtain or update the image generation model by using prediction images (obtained in the same way as the target images, analogous to the processing result), [0184] annotation results are acquired for each image [00184] the annotation results and the predicted images/processing results are used to adjust the model to obtain an adjusted generation model). The combination of Pyzow and Yu would have been obvious to one of ordinary skill in the art prior to the effective filing date of the presently claimed invention. The motivation for the combination lies in that the use of the fused processing result, continuous adjusting and retraining allows for diversity in the content generated by the model (Yu, [0155]) Conclusion The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. For a listing of analogous prior art as determined by the examiner, please see the attached PTOL-892 Notice of References Cited form. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JORDAN M ELLIOTT whose telephone number is (703)756-5463. The examiner can normally be reached M-F 8AM-5PM ET. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Emily Terrell can be reached at (571) 270-3717. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /J.M.E./Examiner, Art Unit 2666 /EMILY C TERRELL/Supervisory Patent Examiner, Art Unit 2666
Read full office action

Prosecution Timeline

Feb 07, 2025
Application Filed
Sep 22, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705715
Tone Mapping for Preserving Contrast of Fine Features in an Image
3y 6m to grant Granted Aug 11, 2026
Patent 12682611
IMAGE ACQUISITION MODEL TRAINING METHOD AND APPARATUS, IMAGE DETECTION METHOD AND APPARATUS, AND DEVICE
3y 1m to grant Granted Jul 14, 2026
Patent 12573117
METHOD AND DEVICE FOR DEEP LEARNING-BASED PATCHWISE RECONSTRUCTION FROM CLINICAL CT SCAN DATA
3y 1m to grant Granted Mar 10, 2026
Patent 12475998
SYSTEMS AND METHODS OF ADAPTIVELY GENERATING FACIAL DEVICE SELECTIONS BASED ON VISUALLY DETERMINED ANATOMICAL DIMENSION DATA
2y 2m to grant Granted Nov 18, 2025
Patent 12450918
AUTOMATIC LANE MARKING EXTRACTION AND CLASSIFICATION FROM LIDAR SCANS
3y 1m to grant Granted Oct 21, 2025
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
47%
Grant Probability
51%
With Interview (+4.2%)
2y 11m (~1y 3m remaining)
Median Time to Grant
Low
PTA Risk
Based on 32 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month