Prosecution Insights
Last updated: September 01, 2026
Application No. 18/467,096

APPARATUS, METHOD, AND COMPUTER PROGRAM FOR TRANSFER LEARNING

Non-Final OA §102§103§112
Filed
Sep 14, 2023
Priority
Oct 06, 2022 — EU 22200122.4
Examiner
GONZALES, VINCENT
Art Unit
2124
Tech Center
2100 — Computer Architecture & Software
Assignee
Nokia Corporation
OA Round
1 (Non-Final)
78%
Grant Probability
Favorable
1-2
OA Rounds
5m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 78% — above average
78%
Career Allowance Rate
419 granted / 534 resolved
+23.5% vs TC avg
Moderate +11% lift
Without
With
+11.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 5m
Avg Prosecution
16 currently pending
Career history
556
Total Applications
across all art units

Statute-Specific Performance

§101
21.1%
-18.9% vs TC avg
§103
41.6%
+1.6% vs TC avg
§102
13.8%
-26.2% vs TC avg
§112
14.3%
-25.7% vs TC avg
Black line = Tech Center average estimate • Based on career data from 534 resolved cases

Office Action

§102 §103 §112
CTNF 18/467,096 CTNF 86904 DETAILED ACTION This action is written in response to the application filed 9/14/23. Subject Matter Eligibility In determining whether the claims are subject matter eligible, the examiner has considered and applied guidance from MPEP § 2106. The examiner finds that the independent claims are directed to the practical application of providing for transfer learning for neural network models. Furthermore, the combination of steps performed in the recited method cannot be practically performed as a mental process. Claim Rejections - 35 USC § 112 07-30-02 The following is a quotation of the second paragraph of 35 U.S.C. 112: (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. Claims 2 and 13-14 are rejected under 35 U.S.C. 112(b), as being indefinite for failing to particularly point out and distinctly claim the subject matter which applicant regards as the invention. Claim 2 recites “An apparatus as claimed in claim 1, wherein the intermediate layer is a layer of the pre-trained neural network node model that is performed prior to an aggregation of the time series data over a time dimension.” A layer of a neural network is an abstract data structure, and it is not clear how the verb ‘performed’ applies here. Accordingly, the claim is indefinite. Claim 13 recites an apparatus [comprising] “receive, from an apparatus, a request for a first plurality of embeddings” and “signal said first plurality of embeddings to the apparatus.” The language of the claim as a whole is unclear and ambiguous. For example: The apparatus is described in terms of its functionality, but without any explicit means for language. The claimed apparatus seems to receive information (a request) from itself . Likewise, the claimed apparatus seems to signal information (the embeddings) to itself . Because it is not clear which of the above interpretations is applicable, the term is ambiguous, and consequently a person of ordinary skill would not be able to understand the scope of the claim with reasonable certainty. Therefore the claim is indefinite. Claim 14 inherits these deficiencies from claim 13. Claim Rejections - 35 USC § 102 07-07-aia AIA 07-07 The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action: A person shall be entitled to a patent unless – 07-08-aia AIA (a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale or otherwise available to the public before the effective filing date of the claimed invention. 07-09-fti 07-09 (b) the invention was patented or described in a printed publication in this or a foreign country or in public use or on sale in this country, more than one year prior to the date of application for patent in the United States. 07-15 AIA Claim s 13-14 are rejected under 35 U.S.C. 102( a)(1 ) as being anticipated by Kiros (US 2019/0258713 A1) . Regarding claim 13, Kiros discloses an apparatus for a network node comprising access to a pre-trained neural network node model, the apparatus further comprises: receive, from an apparatus, a request for a first plurality of embeddings associated with an intermediate layer of the neural network node model; and [0029] “As another example, the system can combine the embeddings in the data set 110 with embeddings from a different data set (the alternative data set 118) that has been generated using a different technique in order to provide task-specific embeddings in response to received requests.” See also [0003] describing NN architecture: “Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.” signal said first plurality of embeddings to the apparatus. Id. Regarding claim 14, Kiros discloses the further limitation wherein the request further comprises unlabelled input data. [0003] “Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.” Claim Rejections - 35 USC § 103 07-20 AIA The following is a quotation of 35 U.S.C. 103(a) which forms the basis for all obviousness rejections set forth in this Office action: (a) A patent may not be obtained though the invention is not identically disclosed or described as set forth in section 102 of this title, if the differences between the subject matter sought to be patented and the prior art are such that the subject matter as a whole would have been obvious at the time the invention was made to a person having ordinary skill in the art to which said subject matter pertains. Patentability shall not be negatived by the manner in which the invention was made. The following are the references relied upon in the rejections below: Kiros (US 2019/0258713 A1) Lundgaard (US 2021/0141995 A1) O’Shea (US 2018/0367192 A1) Profentzas (Profentzas, Christos, Magnus Almgren, and Olaf Landsiedel. "MicroTL: Transfer learning on low-power IoT devices." 2022 IEEE 47th Conference on Local Computer Networks (LCN). IEEE. Published online 26 August 2022.) 07-21-aia AIA Claim s 1-2, 7-12, 15-16 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Lundgaard and Profentzas . Regarding claims 1, 15 and 20, Lundgaard discloses an apparatus, comprising: at least one processor; and [0025] ‘processors’. at least one memory storing instructions that, when executed by the at least one processor, cause the apparatus at least to: [0084] “volatile and non-volatile physical storage mediums”, ‘memory’. obtain, from a pre-trained neural network node model, a first plurality of embeddings associated with an intermediate layer of the neural network node model; … [0028] “For both textual and image data, pre-trained deep learning models may be used to produce embedding representations, which quantitatively describe data in the activation space of the pre-trained network. … Such embeddings may be created by passing the image date and/or textual data as input to a pre-trained model and using hidden layer activations of the network as embeddings.” use the value of the first number of resources to determine a number of averaging functions to be performed, by the device, over a time dimension for each channel of the pre-trained neural network node model; [0027] “Typically, when pre-trained models are not performing well, fine-tuning the pre-trained model is performed to improve performance.” [0032] “By using the activation maps, the server may generate a resulting image embedding by performing a global average pooling on a convolutional layer of the CNN to determine a value for every channel. For example, convolutional layer may be a final convolutional layer, and/or one or more layers that precede the final convolutional layer.” transform the first plurality of embeddings into a second plurality of embeddings by performing said number of averaging functions for the each channel; and Id. cause the device to train a device specific neural network model using the second plurality of embeddings. [0053] “In this example, an embedding vector is generated with 1792 elements to represent each image. In cases where both image and textual data are available, image and textual embeddings may be created separately and concatenated together before being passed as input to the downstream model.” [0054] “The downstream classification model may be a deep, fully-connected network, which may accept a fixed size input and may output a probability distribution over possible classes. … The downstream model may be smaller than most pre-trained models used for transfer learning, and it may be retrained quickly and at a low computational cost.” Profentzas discloses the following further limitation which Lundgaard does not disclose: obtain a value of a first number of resources available on a device for fine-tuning and/or retraining at least part of the pre-trained neural network node model; P. 4, second col., “We implement MicroTL in C, using the Arm-gcc 9.2.1. We evaluate MicroTL on nRF-52840-DK board featuring: a 32-bit ARM Cortex-M4 with an FPU at 64 MHz, a DSP co-processor, 256 KB of RAM, and 1MB KB of flash. The board has wireless communication capability with Bluetooth Low Energy (BLE), Thread, and Zigbee. The chip of this board is widely used on low-power IoT applications like wearable and smart-watches [20].” P. 4, table 1 (reproduced below). PNG media_image1.png 342 780 media_image1.png Greyscale P. 7, second col., “MicroTL tailors transfer learning on low-power IoT devices running on 32-64 MHz and KBs of RAM by addressing their recourse constraints, without needing extra parameters for training and by dynamically utilizing personalized data.” At the time of filing, it would have been obvious to a person of ordinary skill to apply the technique disclosed by Profentzas for tailoring a transfer learning architecture for a low-powered edge device to the Lundgaard system because this would provide for effective learning at the edge node without the need to transfer training data, thereby providing privacy. Regarding claims 2 and 16, Profentzas discloses the further limitation wherein the intermediate layer is a layer of the pre-trained neural network node model that is performed prior to an aggregation of the time series data over a time dimension. P. 4, table 1 (reproduced supra ), listing layers which are “Pre-trained … [on the] HAR-UCI-3a” dataset. P. 5, first col., “HAR-UCI is based on time-series of sensors (accelerometer and gyroscope) captured by smartphone”. Regarding claim 7, Profentzas discloses the further limitation wherein the obtaining of the first plurality of embeddings further comprises: signal, to a network node, a request for said first plurality of embeddings, wherein said request comprises unlabelled input data; and P. 5, table II, illustrating test set results. receive said first plurality of embeddings from the network node. Id. The classification results are determined by the edge device networks after additional local training. Regarding claim 8, Profentzas discloses the further limitation further comprising: cause the device to run the trained device-specific neural network model using a second plurality of input data in order to output at least one inference; and P. 5, first col., describing a classification task using HAR-UCI data. use said inference to identify at least a type of data. Id. Regarding claim 9, Profentzas discloses the further limitation wherein the device-specific neural network model relates to recognizing audio data, the second plurality of input data comprises an audio sample, and wherein the identifying at least one type of data comprises identifying of different types of audio signals within the audio sample. P. 1, abstract, “Deep Neural Networks (DNNs) on IoT devices are becoming readily available for classification tasks using sensor data like images and audio .” Regarding claim 10, Profentzas discloses the further limitation wherein the device-specific neural network model relates to recognizing activity data, the second plurality of input data comprises activity data produced when a user performs at least one type of activity, and wherein the identifying at least one type of data comprises identifying at least one activity from said activity data. P. 5, first col., describing a classification task using HAR-UCI data. Regarding claim 11, Profentzas discloses the further limitation wherein the intermediate layer further comprises a last high dimensional layer prior to a penultimate layer of the neural network node model. P. 4, table 1 (reproduced supra): both HAR-UCI-3a and CIFAR-7 neural network models have the claimed architecture feature. Regarding claim 12, Profentzas discloses the further limitation wherein the pre-trained neural network node model is pre-trained on time-series data, the time series data comprising a first plurality of input data, each of said first plurality of input data comprising a tensor having associated sets of sample values, timestep values, and channel values. P. 5, first col., “HAR-UCI. The data-set is a human activity recognition data set. HAR-UCI is based on time-series of sensors (accelerometer and gyroscope) captured by smartphone (Samsung Galaxy S II), and it has six classes (activities): 1) walking, 2) walking upstairs, 3) walking downstairs, 4) sitting, 5) standing, 6) laying. The data type is a time series of signals needing preprocessing before use. Our experiments parse the data per 128 time-step and apply min-max normalization over the nine-axis (final dimension 128x9), leading to 7,300 training samples and 3,000 testing samples. We use the following three classes as the source domain (HAR-UCI-3a): walking upstairs, walking downstairs, and standing. HAR-UCI-3a has 3,650 samples for training and 1,500 for testing. We use the other three classes as the target domain: walking, sitting, and lying. We named the data set as HAR-UCI-3b, and it has 3,650 samples for training and 1,500 for testing.” Regarding claim 19, Profentzas discloses the further limitation comprising: cause the device to run the trained device-specific neural network model using a second plurality of input data in order to output at least one inference; and P. 5, table II, illustrating test results for two classification tasks. use said inference to identify at least a type of data. Id. , illustrating results for two classification tasks . 07-21-aia AIA Claim s 3-4 and 17 are rejected under 35 U.S.C. 103 as being unpatentable over Lundgaard, Profentzas and O’Shea . Regarding claim 3, O’Shea discloses the following further limitation which neither Lundgaard/Profentzas discloses wherein said averaging functions comprise one or more of power means functions. [0082] “power means amplitude”. At the time of filing, it would have been obvious to a person of ordinary skill to apply a power means function (as taught by O’Shea) to the Lundgaard/Profentzas system because this would provide a meaningful summary statistic for the underlying data, providing for more efficient transmission of data. Regarding claim 4, O’Shea discloses the further limitation wherein the one or more power means functions are a fractional subset of a plurality of power means functions, each of the plurality of power means functions associated with respective priority for selection, and wherein the transforming further comprises: select said one or more of power means functions from the plurality of power means functions in order descending from highest priority to lowest priority; and [0026] “Further, each of the K different power mean values is one selected from the group”. generate, for each of said one or more of power means functions and each of said first plurality of embeddings, said second plurality of embeddings. Id. See also [0006] and [0098] describing embeddings. Regarding claim 17, O’Shea discloses the following further limitation which neither Lundgaard/Profentzas discloses wherein said averaging functions comprise one or more of power means functions. [0082] “power means amplitude”. The obviousness analysis of claim 3 applies equally here . 07-21-aia AIA Claim 18 is rejected under 35 U.S.C. 103 as being unpatentable over Lundgaard, Profentzas and Kiros . Regarding claim 18, Kiros discloses the further limitation which neither Lundgaard/Profentzas discloses wherein the obtaining of the first plurality of embeddings further comprises: signal, to a network node, a request for said first plurality of embeddings, wherein said request comprises unlabelled input data; and [0029] “As another example, the system can combine the embeddings in the data set 110 with embeddings from a different data set (the alternative data set 118) that has been generated using a different technique in order to provide task-specific embeddings in response to received requests.” See also [0003] describing NN architecture: “Neural networks are machine learning models that employ one or more layers of nonlinear units to predict an output for a received input. Some neural networks include one or more hidden layers in addition to an output layer. The output of each hidden layer is used as input to the next layer in the network, i.e., the next hidden layer or the output layer. Each layer of the network generates an output from a received input in accordance with current values of a respective set of parameters.” receive said first plurality of embeddings from the network node. Id. Claim Objections and Allowable Subject Matter Claims 5-6 are allowable over the prior art, but are objected to as depending from a rejected parent claim. Additional Relevant Prior Art The following references were identified by the Examiner as being relevant to the disclosed invention, but are not relied upon in any particular prior art rejection: Choo discloses a transfer learning system comprising a technique for measuring the suitability of each of a plurality of source models for the task at hand. (US 2020/0134469 A1) Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Vincent Gonzales whose telephone number is (571) 270-3837. The examiner can normally be reached on Monday-Friday 7 a.m. to 4 p.m. MT. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Miranda Huang, can be reached at (571) 270-7092. Information regarding the status of an application may be obtained from the USPTO Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. /Vincent Gonzales/Primary Examiner, Art Unit 2124 Application/Control Number: 18/467,096 Page 2 Art Unit: 2124 Application/Control Number: 18/467,096 Page 3 Art Unit: 2124 Application/Control Number: 18/467,096 Page 4 Art Unit: 2124 Application/Control Number: 18/467,096 Page 5 Art Unit: 2124 Application/Control Number: 18/467,096 Page 6 Art Unit: 2124 Application/Control Number: 18/467,096 Page 7 Art Unit: 2124 Application/Control Number: 18/467,096 Page 8 Art Unit: 2124 Application/Control Number: 18/467,096 Page 9 Art Unit: 2124 Application/Control Number: 18/467,096 Page 10 Art Unit: 2124 Application/Control Number: 18/467,096 Page 11 Art Unit: 2124 Application/Control Number: 18/467,096 Page 12 Art Unit: 2124
Read full office action

Prosecution Timeline

Sep 14, 2023
Application Filed
Jun 01, 2026
Non-Final Rejection mailed — §102, §103, §112 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12705472
Prefetching Weights For Use In A Neural Network Processor
2y 9m to grant Granted Aug 11, 2026
Patent 12675989
FUSION MODEL TRAINING USING DISTANCE METRICS
2y 3m to grant Granted Jul 07, 2026
Patent 12651182
IDENTIFYING TRAITS OF PARTITIONED GROUP FROM IMBALANCED DATASET
4y 11m to grant Granted Jun 09, 2026
Patent 12639623
FAIR SELECTIVE CLASSIFICATION VIA A VARIATIONAL MUTUAL INFORMATION UPPER BOUND FOR IMPOSING SUFFICIENCY
4y 4m to grant Granted May 26, 2026
Patent 12619681
SYSTEMS AND METHODS FOR DOMAIN-AWARE CLASSIFICATION OF UNLABELED DATA
4y 7m to grant Granted May 05, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
78%
Grant Probability
90%
With Interview (+11.2%)
3y 5m (~5m remaining)
Median Time to Grant
Low
PTA Risk
Based on 534 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month