Prosecution Insights
Last updated: October 02, 2026
Application No. 18/274,531

DISTRIBUTED MACHINE LEARNING WITH NEW LABELS USING HETEROGENEOUS LABEL DISTRIBUTION

Non-Final OA §103
Filed
Jul 27, 2023
Priority
Jan 29, 2021 — nonprovisional of PCTIN2021050097
Examiner
BARRETT, RYAN S
Art Unit
2148
Tech Center
2100 — Computer Architecture & Software
Assignee
Telefonaktiebolaget LM Ericsson
OA Round
2 (Non-Final)
66%
Grant Probability
Favorable
2-3
OA Rounds
1m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 66% — above average
66%
Career Allowance Rate
281 granted / 429 resolved
+10.5% vs TC avg
Strong +41% interview lift
Without
With
+41.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
16 currently pending
Career history
444
Total Applications
across all art units

Statute-Specific Performance

§101
10.4%
-29.6% vs TC avg
§103
38.2%
-1.8% vs TC avg
§102
11.4%
-28.6% vs TC avg
§112
8.7%
-31.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 429 resolved cases

Office Action

§103
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the Amendment filed on 6/24/2026. Claims 1-3, 6, 9-17, 25, and 31 are pending in the case. Claims 4-5, 7-8, and 18 have been cancelled. Claims 1, 9, 15, and 25 are independent claims. Response to Arguments Applicant’s amendments regarding the objections are persuasive. These objections are respectfully withdrawn. Applicant’s amendments regarding 35 U.S.C. § 112 rejections are persuasive. These rejections are respectfully withdrawn. Applicant’s arguments regarding 35 U.S.C. § 101 rejections are persuasive. These rejections are respectfully withdrawn. Applicant’s arguments regarding 35 U.S.C. § 103 rejections are persuasive. New rejections based on a different version of the same reference appear below. Claim Interpretation Claims 2 and 16 recite “for training local ML models” which appears to be an intended use rather than a positive claim limitation. Claim 11 recites “artificial neural network (ANN)” which is an umbrella term encompassing the other recited alternatives. Claim Rejections - 35 U.S.C. § 103 In the event the determination of the status of the application as subject to AIA 35 U.S.C. §§ 102 and 103 (or as subject to pre-AIA 35 U.S.C. §§ 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. The following is a quotation of 35 U.S.C. § 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102 of this title, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 C.F.R. § 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. § 102(b)(2)(C) for any potential 35 U.S.C. § 102(a)(2) prior art against the later invention. Claims 1-3, 9-12, 15-17, 25, and 31 are rejected under 35 U.S.C. § 103 as being unpatentable over Kairouz et al. (“Advances and Open Problems in Federated Learning,” 10 December 2019, https://arxiv.org/abs/1912.04977v1, hereinafter Kairouz) in view of Yang et al. (“Federated Machine Learning: Concept and Applications,” 13 February 2019, https://arxiv.org/abs/1902.04885, hereinafter Yang) and Nayak et al. (“Data Impressions: Mining Deep Models to Extract Samples for Data-free Applications,” 15 January 2021, https://arxiv.org/abs/2101.06069v1, hereinafter Nayak). As to independent claim 1, Kairouz teaches a method for distributed machine learning (ML) at a central computing device, the method comprising: providing a first dataset (“public datasets,” page 28 section “3.3.1 Personalization via Featurization” line 10) including a first set of labels (“additional data or metadata might need to be maintained, e.g. user interaction data to provide labels for a supervised learning task,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 4-5) to a plurality of local computing devices including a first local computing device and a second local computing device (“the clients (e.g. an app running on mobile phones) are instrumented to store locally (with limits on time and quantity) the necessary training data,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 1-2); receiving, from the first local computing device, a first set of ML model probabilities values from training a first local ML model using the first set of labels (“The server collects an aggregate of the device updates,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “4. Aggregation” line 1); receiving, from the second local computing device, a second set of ML model probabilities values from training a second local ML model using the first set of labels (“The server collects an aggregate of the device updates,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “4. Aggregation” line 1); generating a third set of ML model probabilities values [] (“The server locally updates the shared model based on the aggregated update computed from the clients that participated in the current round,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “5. Model update” lines 1-2); training a global ML model using the generated second set of data impressions (“A server (service provider) orchestrates the training process, by repeating the following steps until training is stopped,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph 2 lines 1-2). Kairouz does not appear to expressly teach a method comprising one or more labels different from any label in the first set of labels. Yang teaches a method comprising one or more labels different from any label in the first set of labels (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space. Consider two institutions, one is a bank located in China, and the other is an e-commerce company located in the United States. Due to geographical restrictions, the user groups of the two institutions have a small intersection. On the other hand, due to the different businesses, only a small portion of the feature space from both parties overlaps. In this case, transfer learning [50] techniques can be applied to provide solutions for the entire sample and feature space under a federation (Figure2c). Specially, a common representation between the two feature space [sic] is learned using the limited common sample sets and later applied to obtain predictions for samples with only one-side features,” page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-9). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the machine learning of Kairouz to comprise the different labels of Yang. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely making use of different data sets (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space,” Yang page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-1). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). Kairouz/Yang does not appear to expressly teach a method comprising: generating a weights matrix using the received first set of ML model probabilities values and the received second set of ML model probabilities values; generating a first set of data impressions using the generated third set of ML model probabilities values, wherein the first set of data impressions includes data impressions for each of the one or more labels []; generating a second set of data impressions by clustering using the generated first set of data impressions for each of the one or more labels []; and [generating a third set of ML model probabilities values] by sampling using the generated weights matrix. Nayak teaches a method comprising: generating a weights matrix using the received first set of ML model probabilities values and the received second set of ML model probabilities values (“We compute a normalized class similarity matrix (C) using the weights W connecting the final (softmax) and the pre-final layers. The element C(i; j) of this matrix denotes the visual similarity between the categories i and i in [0, 1]. Thus, a row ck of the class similarity matrix (C) gives the similarity of class k with each of the K categories (including itself),” page 3 section “3.1 Modelling the Data in Softmax Space” paragraph 5 lines 2-8); generating a first set of data impressions using the generated third set of ML model probabilities values, wherein the first set of data impressions includes data impressions for each of the one or more labels [] (“Corresponding to each sampled softmax vector y i k , we can craft a Data Impression x - i k , for which the Trained network predicts a similar softmax output,” page 4 section “3.2 Crafting Data Impressions via Dirichlet Sampling” lines 9-11); generating a second set of data impressions by clustering using the generated first set of data impressions for each of the one or more labels [] (“They optimize a random noise in the input space till it results in a one-hot vector (softmax) output. This means, their optimization to craft the representative samples would expect a one-hot vector in the output space. Hence, they call the reconstructions Class Impressions. Our reconstruction (eq. (2)) is inspired from this, though we model the output space utilizing the class similarities perceived by the Teacher model. Because of this, we argue that our modelling is closer to the original distribution and results in better patterns in the reconstructions, calling them Data Impressions of the Teacher model,” page 7 section “4.1.4 Class Versus Data Impressions” paragraph 2 lines 2-12); and generating a third set of ML model probabilities values by sampling using the generated weights matrix (“actual sampling of the probability vectors happen from p(s) = Dir(K; β × α),” page 4 column right lines 12-13; “we generate a diverse set of pseudo training examples that can provide with enough information to train the Student model via Dirichlet sampling,” page 5 section “4.1.1 Zero-Shot Knowledge Distillation” lines 10-12). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the machine learning of Kairouz/Yang to comprise the data impressions of Nayak. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely “to provide a competitive performance” (Nayak page 7 line 2). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). As to dependent claim 2, the rejection of claim 1 is incorporated. Kairouz/Yang/Nayak further teaches a method comprising: generating a fourth set (“repeating the following steps until training is stopped,” page 8 section “1.1.2 A Typical Federated Training Process” Kairouz paragraph 2 lines 1-2) of ML model probabilities values by averaging using the first set of data impressions and the second set of data impressions (“β intuitively models the spread of the Dirichlet distribution and acts as a scaling parameter atop α to yield the final concentration parameter (prior),” Nayak page 4 column right lines 13-15) for each label of the first set of labels and the one or more labels different from any label in the first set of labels (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space. Consider two institutions, one is a bank located in China, and the other is an e-commerce company located in the United States. Due to geographical restrictions, the user groups of the two institutions have a small intersection. On the other hand, due to the different businesses, only a small portion of the feature space from both parties overlaps. In this case, transfer learning [50] techniques can be applied to provide solutions for the entire sample and feature space under a federation (Figure2c). Specially, a common representation between the two feature space [sic] is learned using the limited common sample sets and later applied to obtain predictions for samples with only one-side features,” Yang page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-9); and providing the generated fourth set (“repeating the following steps until training is stopped,” page 8 section “1.1.2 A Typical Federated Training Process” Kairouz paragraph 2 lines 1-2) of ML model probabilities values to the plurality of local computing devices, including the first local computing device and the second local computing device, for training local ML models (“additional data or metadata might need to be maintained, e.g. user interaction data to provide labels for a supervised learning task,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 4-5). As to dependent claim 3, the rejection of claim 1 is incorporated. Kairouz/Yang/Nayak further teaches a method wherein the received first set of ML model probabilities values and the received second set of ML model probabilities values are one of: Softmax values (“We compute a normalized class similarity matrix (C) using the weights W connecting the final (softmax) and the pre-final layers. The element C(i; j) of this matrix denotes the visual similarity between the categories i and i in [0, 1]. Thus, a row ck of the class similarity matrix (C) gives the similarity of class k with each of the K categories (including itself),” page 3 section “3.1 Modelling the Data in Softmax Space” Nayak paragraph 5 lines 2-8), sigmoid values, and Dirichlet values (“actual sampling of the probability vectors happen from p(s) = Dir(K; β × α),” Nayak page 4 column right lines 12-13). As to independent claim 9, Kairouz teaches a method for distributed machine learning (ML) learning at a local computing device, the method comprising: receiving a first dataset (“public datasets,” page 28 section “3.3.1 Personalization via Featurization” line 10) including a first set of labels (“additional data or metadata might need to be maintained, e.g. user interaction data to provide labels for a supervised learning task,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 4-5); training a local ML model [] (“Each selected device locally computes an update to the model,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “3. Client computation” line 1); and providing the generated set of ML model probabilities values to a central computing device (“The server collects an aggregate of the device updates,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “4. Aggregation” line 1). Kairouz does not appear to expressly teach a method comprising generating a second dataset including the first set of labels from the received first dataset and one or more labels different from any label in the first set of labels. Yang teaches a method comprising generating a second dataset including the first set of labels from the received first dataset and one or more labels different from any label in the first set of labels (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space. Consider two institutions, one is a bank located in China, and the other is an e-commerce company located in the United States. Due to geographical restrictions, the user groups of the two institutions have a small intersection. On the other hand, due to the different businesses, only a small portion of the feature space from both parties overlaps. In this case, transfer learning [50] techniques can be applied to provide solutions for the entire sample and feature space under a federation (Figure2c). Specially, a common representation between the two feature space [sic] is learned using the limited common sample sets and later applied to obtain predictions for samples with only one-side features,” page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-9). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the machine learning of Kairouz to comprise the different labels of Yang. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely making use of different data sets (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space,” Yang page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-1). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). Kairouz/Yang does not appear to expressly teach a method comprising: generating a weights matrix using the one or more labels []; and generating a set of ML model probabilities values by using the generated weights matrix and trained local ML model. Nayak teaches a method comprising: generating a weights matrix using the one or more labels [] (“We compute a normalized class similarity matrix (C) using the weights W connecting the final (softmax) and the pre-final layers. The element C(i; j) of this matrix denotes the visual similarity between the categories i and i in [0, 1]. Thus, a row ck of the class similarity matrix (C) gives the similarity of class k with each of the K categories (including itself),” page 3 section “3.1 Modelling the Data in Softmax Space” paragraph 5 lines 2-8); and generating a set of ML model probabilities values by using the generated weights matrix and trained local ML model (“actual sampling of the probability vectors happen from p(s) = Dir(K; β × α),” page 4 column right lines 12-13; “we generate a diverse set of pseudo training examples that can provide with enough information to train the Student model via Dirichlet sampling,” page 5 section “4.1.1 Zero-Shot Knowledge Distillation” lines 10-12). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the machine learning of Kairouz/Yang to comprise the data impressions of Nayak. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely “to provide a competitive performance” (Nayak page 7 line 2). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). As to dependent claim 10, the rejection of claim 9 is incorporated. Kairouz/Yang/Nayak further teaches a method wherein the received first data set is a public dataset (“public datasets,” Kairouz page 28 section “3.3.1 Personalization via Featurization” line 10) and the generated second dataset is a private dataset (“keeping data private at each site,” Kairouz page 15 paragraph 7 lines 3-4). As to dependent claim 11, the rejection of claim 9 is incorporated. Kairouz/Yang/Nayak further teaches a method wherein the local ML model is one of: a convolutional neural network (CNN), a artificial neural network (ANN), and a recurrent neural network (RNN) (“neural networks,” Kairouz page 15 line 4). As to dependent claim 12, the rejection of claim 9 is incorporated. Kairouz/Yang/Nayak further teaches a method comprising: receiving a set of ML model probabilities values from the central computing device representing (“additional data or metadata might need to be maintained, e.g. user interaction data to provide labels for a supervised learning task,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 4-5) an averaging using a first set of data impressions and a second set of data impressions (“β intuitively models the spread of the Dirichlet distribution and acts as a scaling parameter atop α to yield the final concentration parameter (prior),” Nayak page 4 column right lines 13-15) for each label of the first set of labels and the one or more labels different from any label in the first set of labels (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space. Consider two institutions, one is a bank located in China, and the other is an e-commerce company located in the United States. Due to geographical restrictions, the user groups of the two institutions have a small intersection. On the other hand, due to the different businesses, only a small portion of the feature space from both parties overlaps. In this case, transfer learning [50] techniques can be applied to provide solutions for the entire sample and feature space under a federation (Figure2c). Specially, a common representation between the two feature space [sic] is learned using the limited common sample sets and later applied to obtain predictions for samples with only one-side features,” Yang page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-9); and training the local ML model using the received set of ML model probabilities values (“Each selected device locally computes an update to the model,” Kairouz page 8 section “1.1.2 A Typical Federated Training Process” paragraph “3. Client computation” line 1). As to independent claim 15, Kairouz teaches a central computing device comprising: a memory (“central server,” page 1 section “Abstract” line 2); and a processor (“central server,” page 1 section “Abstract” line 2) coupled to the memory, wherein the processor is configured to: provide a first dataset (“public datasets,” page 28 section “3.3.1 Personalization via Featurization” line 10) including a first set of labels (“additional data or metadata might need to be maintained, e.g. user interaction data to provide labels for a supervised learning task,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 4-5) to plurality of local computing devices including a first local computing device and a second local computing device (“the clients (e.g. an app running on mobile phones) are instrumented to store locally (with limits on time and quantity) the necessary training data,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 1-2); receive, from the first local computing device, a first set of ML model probabilities values from training a first local ML model using the first set of labels (“The server collects an aggregate of the device updates,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “4. Aggregation” line 1); receive, from the second local computing device, a second set of ML model probabilities values from training a second local ML model using the first set of labels (“The server collects an aggregate of the device updates,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “4. Aggregation” line 1); generate a third set of ML model probabilities values [] (“The server locally updates the shared model based on the aggregated update computed from the clients that participated in the current round,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “5. Model update” lines 1-2); train a global ML model using the generated second set of data impressions (“A server (service provider) orchestrates the training process, by repeating the following steps until training is stopped,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph 2 lines 1-2). Kairouz does not appear to expressly teach a device comprising one or more labels different from any label in the first set of labels. Yang teaches a device comprising one or more labels different from any label in the first set of labels (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space. Consider two institutions, one is a bank located in China, and the other is an e-commerce company located in the United States. Due to geographical restrictions, the user groups of the two institutions have a small intersection. On the other hand, due to the different businesses, only a small portion of the feature space from both parties overlaps. In this case, transfer learning [50] techniques can be applied to provide solutions for the entire sample and feature space under a federation (Figure2c). Specially, a common representation between the two feature space [sic] is learned using the limited common sample sets and later applied to obtain predictions for samples with only one-side features,” page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-9). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the machine learning of Kairouz to comprise the different labels of Yang. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely making use of different data sets (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space,” Yang page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-1). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). Kairouz/Yang does not appear to expressly teach a device configured to: generate a weights matrix using the received first set of ML model probabilities values and the received second set of ML model probabilities values; generate a first set of data impressions using the generated third set of ML model probabilities values, wherein the first set of data impressions includes data impressions for each of the one or more labels []; generate a second set of data impressions by clustering using the generated first set of data impressions for each of the one or more labels []; and [generate a third set of ML model probabilities values] by sampling using the generated weights matrix. Nayak teaches a device configured to: generate a weights matrix using the received first set of ML model probabilities values and the received second set of ML model probabilities values (“We compute a normalized class similarity matrix (C) using the weights W connecting the final (softmax) and the pre-final layers. The element C(i; j) of this matrix denotes the visual similarity between the categories i and i in [0, 1]. Thus, a row ck of the class similarity matrix (C) gives the similarity of class k with each of the K categories (including itself),” page 3 section “3.1 Modelling the Data in Softmax Space” paragraph 5 lines 2-8); generate a first set of data impressions using the generated third set of ML model probabilities values, wherein the first set of data impressions includes data impressions for each of the one or more labels [] (“Corresponding to each sampled softmax vector y i k , we can craft a Data Impression x - i k , for which the Trained network predicts a similar softmax output,” page 4 section “3.2 Crafting Data Impressions via Dirichlet Sampling” lines 9-11); generate a second set of data impressions by clustering using the generated first set of data impressions for each of the one or more labels [] (“They optimize a random noise in the input space till it results in a one-hot vector (softmax) output. This means, their optimization to craft the representative samples would expect a one-hot vector in the output space. Hence, they call the reconstructions Class Impressions. Our reconstruction (eq. (2)) is inspired from this, though we model the output space utilizing the class similarities perceived by the Teacher model. Because of this, we argue that our modelling is closer to the original distribution and results in better patterns in the reconstructions, calling them Data Impressions of the Teacher model,” page 7 section “4.1.4 Class Versus Data Impressions” paragraph 2 lines 2-12); and generate a third set of ML model probabilities values by sampling using the generated weights matrix (“actual sampling of the probability vectors happen from p(s) = Dir(K; β × α),” page 4 column right lines 12-13; “we generate a diverse set of pseudo training examples that can provide with enough information to train the Student model via Dirichlet sampling,” page 5 section “4.1.1 Zero-Shot Knowledge Distillation” lines 10-12). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the machine learning of Kairouz/Yang to comprise the data impressions of Nayak. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely “to provide a competitive performance” (Nayak page 7 line 2). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). As to dependent claim 16, the rejection of claim 15 is incorporated. Kairouz/Yang/Nayak further teaches a device wherein the processor is further configured to: generate a fourth set (“repeating the following steps until training is stopped,” page 8 section “1.1.2 A Typical Federated Training Process” Kairouz paragraph 2 lines 1-2) of ML model probabilities values by averaging using the generated first set of data impressions and the generated second set of data impressions (“β intuitively models the spread of the Dirichlet distribution and acts as a scaling parameter atop α to yield the final concentration parameter (prior),” Nayak page 4 column right lines 13-15) for each label of the first set of labels and the one or more labels different from any label in the first set of labels (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space. Consider two institutions, one is a bank located in China, and the other is an e-commerce company located in the United States. Due to geographical restrictions, the user groups of the two institutions have a small intersection. On the other hand, due to the different businesses, only a small portion of the feature space from both parties overlaps. In this case, transfer learning [50] techniques can be applied to provide solutions for the entire sample and feature space under a federation (Figure2c). Specially, a common representation between the two feature space [sic] is learned using the limited common sample sets and later applied to obtain predictions for samples with only one-side features,” Yang page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-9); and provide the generated fourth set (“repeating the following steps until training is stopped,” page 8 section “1.1.2 A Typical Federated Training Process” Kairouz paragraph 2 lines 1-2) of ML model probabilities values to a plurality of local computing devices, including the first local computing device and the second local computing device, for training local ML models (“additional data or metadata might need to be maintained, e.g. user interaction data to provide labels for a supervised learning task,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 4-5). As to dependent claim 17, the rejection of claim 15 is incorporated. Kairouz/Yang/Nayak further teaches a device wherein the received first set of ML model probabilities values and the received second set of ML model probabilities values are one of: Softmax values (“We compute a normalized class similarity matrix (C) using the weights W connecting the final (softmax) and the pre-final layers. The element C(i; j) of this matrix denotes the visual similarity between the categories i and i in [0, 1]. Thus, a row ck of the class similarity matrix (C) gives the similarity of class k with each of the K categories (including itself),” page 3 section “3.1 Modelling the Data in Softmax Space” Nayak paragraph 5 lines 2-8), sigmoid values, and Dirichlet values (“actual sampling of the probability vectors happen from p(s) = Dir(K; β × α),” Nayak page 4 column right lines 12-13). As to independent claim 25, Kairouz teaches a local computing device comprising: a memory (“mobile phones,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” line 1); and a processor (“mobile phones,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” line 1) coupled to the memory, wherein the processor is configured to: receive a first dataset (“public datasets,” page 28 section “3.3.1 Personalization via Featurization” line 10) including a first set of labels (“additional data or metadata might need to be maintained, e.g. user interaction data to provide labels for a supervised learning task,” page 7 section “1.1.1 The Lifecycle of a Model in Federated Learning” paragraph “2. Client instrumentation” lines 4-5); train a local ML model [] (“Each selected device locally computes an update to the model,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “3. Client computation” line 1); and provide the generated set of model probabilities values to a central computing device (“The server collects an aggregate of the device updates,” page 8 section “1.1.2 A Typical Federated Training Process” paragraph “4. Aggregation” line 1). Kairouz does not appear to expressly teach a device configured to generate a second dataset including the first set of labels from the received first dataset and one or more labels different from any label in the first set of labels. Yang teaches a device configured to generate a second dataset including the first set of labels from the received first dataset and one or more labels different from any label in the first set of labels (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space. Consider two institutions, one is a bank located in China, and the other is an e-commerce company located in the United States. Due to geographical restrictions, the user groups of the two institutions have a small intersection. On the other hand, due to the different businesses, only a small portion of the feature space from both parties overlaps. In this case, transfer learning [50] techniques can be applied to provide solutions for the entire sample and feature space under a federation (Figure2c). Specially, a common representation between the two feature space [sic] is learned using the limited common sample sets and later applied to obtain predictions for samples with only one-side features,” page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-9). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the machine learning of Kairouz to comprise the different labels of Yang. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely making use of different data sets (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space,” Yang page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-1). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). Kairouz/Yang does not appear to expressly teach a device configured to: generate a weights matrix using the one or more labels []; and generate a set of model probabilities values by using the generated weights matrix and trained local ML model. Nayak teaches a device configured to: generate a weights matrix using the one or more labels [] (“We compute a normalized class similarity matrix (C) using the weights W connecting the final (softmax) and the pre-final layers. The element C(i; j) of this matrix denotes the visual similarity between the categories i and i in [0, 1]. Thus, a row ck of the class similarity matrix (C) gives the similarity of class k with each of the K categories (including itself),” page 3 section “3.1 Modelling the Data in Softmax Space” paragraph 5 lines 2-8); and generate a set of model probabilities values by using the generated weights matrix and trained local ML model (“actual sampling of the probability vectors happen from p(s) = Dir(K; β × α),” page 4 column right lines 12-13; “we generate a diverse set of pseudo training examples that can provide with enough information to train the Student model via Dirichlet sampling,” page 5 section “4.1.1 Zero-Shot Knowledge Distillation” lines 10-12). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the machine learning of Kairouz/Yang to comprise the data impressions of Nayak. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely “to provide a competitive performance” (Nayak page 7 line 2). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). As to dependent claim 31, the rejection of claim 1 is incorporated. Kairouz/Yang/Nayak further teaches a computer program product comprising a non-transitory computer-readable medium storing a computer program comprising instructions which when executed by processing circuitry causes the processing circuitry to perform the method of claim 1 (“central server,” Kairouz page 1 section “Abstract” line 2). Claim 6 is rejected under 35 U.S.C. § 103 as being unpatentable over Kairouz in view of Yang, Nayak, and Muraoka et al. (US 2019/0095525 A1, hereinafter Muraoka). As to dependent claim 6, the rejection of claim 1 is incorporated. Kairouz/Yang/Nayak further teaches a method comprising the generated first set of data impressions (“Corresponding to each sampled softmax vector y i k , we can craft a Data Impression x - i k , for which the Trained network predicts a similar softmax output,” Nayak page 4 section “3.2 Crafting Data Impressions via Dirichlet Sampling” lines 9-11) for each of the one or more labels different from any label in the first set of labels (“Federated Transfer Learning applies to the scenarios that the two data sets differ not only in samples but also in feature space. Consider two institutions, one is a bank located in China, and the other is an e-commerce company located in the United States. Due to geographical restrictions, the user groups of the two institutions have a small intersection. On the other hand, due to the different businesses, only a small portion of the feature space from both parties overlaps. In this case, transfer learning [50] techniques can be applied to provide solutions for the entire sample and feature space under a federation (Figure2c). Specially, a common representation between the two feature space [sic] is learned using the limited common sample sets and later applied to obtain predictions for samples with only one-side features,” Yang page 12:7 section “2.3.3 Federated Transfer Learning (FTL)” lines 1-9). Kairouz/Yang/Nayak does not appear to expressly teach a method wherein clustering [] is according to a k-medoids clustering algorithm and uses the elbow method to determine the number of clusters k. Muraoka teaches a method wherein clustering [] is according to a k-medoids clustering algorithm and uses the elbow method to determine the number of clusters k (“Any known clustering algorithms such as aggregative hierarchical clustering (including group average method) and non-hierarchical clustering (such as k-means, k-medoids, x-means, etc.) can be applied to feature vectors of images. When an algorithm such as k-means, which has a fixed number of clusters as a parameter, is used, the appropriate number of the cluster can be determined by using any known criteria used in elbow method, silhouette method, etc.,” paragraph 0031 lines 1-9). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the clustering of Kairouz/Yang/Nayak to comprise the k-medoids of Muraoka. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely clustering via k-medoids (“Any known clustering algorithms such as aggregative hierarchical clustering (including group average method) and non-hierarchical clustering (such as k-means, k-medoids, x-means, etc.) can be applied to feature vectors of images,” Muraoka paragraph 0031 lines 1-5). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). Claims 13-14 are rejected under 35 U.S.C. § 103 as being unpatentable over Kairouz in view of Yang, Nayak, and Moreno et al. (US 2018/0114137 A1, hereinafter Moreno). As to dependent claim 13, the rejection of claim 1 is incorporated. Kairouz/Yang/Nayak does not appear to expressly teach a method wherein the plurality of local computing devices, including the first local computing device and the second local computing device, comprises a plurality of radio network nodes which are configured to classify an alarm type using the trained local ML models. Moreno teaches a method wherein the plurality of local computing devices, including the first local computing device and the second local computing device, comprises a plurality of radio network nodes (“the node 130 includes (1) a power solution 128 (See FIG. 10) that can be a battery, a generator or a composition of both; (2) a low-power wireless network module 126 responsible for receiving and transmitting data,” paragraph 0046 lines 4-7) which are configured to classify an alarm type (“input signals (in this case, coming from sensors) will be transformed into output signals (ex: fire alarm, crop time, irrigate now, etc.),” paragraph 0051 lines 11-13) using the trained local ML models (“the node monitors data from each valid sensor (604). Then, a valid neuromorphic program (NP) processes the input data 606. Thereafter, the node checks whether there is data in the ‘alert’ class 608. If yes, then the system triggers an alert process 616,” paragraph 0056 lines 2-6). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the local computing devices of Kairouz/Yang/Nayak to comprise the radio network nodes of Moreno. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely classifying an alarm type using the trained local ML models (“the node monitors data from each valid sensor (604). Then, a valid neuromorphic program (NP) processes the input data 606. Thereafter, the node checks whether there is data in the ‘alert’ class 608. If yes, then the system triggers an alert process 616,” Moreno paragraph 0056 lines 2-6). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). As to dependent claim 14, the rejection of claim 1 is incorporated. Kairouz/Yang/Nayak does not appear to expressly teach a method wherein the plurality of local computing devices, including the first local computing device and the second local computing device, comprises a plurality of wireless sensor devices which are configured to classify an alarm type using the trained local ML models. Moreno teaches a method wherein the plurality of local computing devices, including the first local computing device and the second local computing device, comprises a plurality of wireless sensor devices (“In addition to the components from the first level (sensors 122 and actuators 124) 120, the node 130 includes (1) a power solution 128 (See FIG. 10) that can be a battery, a generator or a composition of both; (2) a low-power wireless network module 126 responsible for receiving and transmitting data,” paragraph 0046 lines 2-7) which are configured to classify an alarm type (“input signals (in this case, coming from sensors) will be transformed into output signals (ex: fire alarm, crop time, irrigate now, etc.),” paragraph 0051 lines 11-13) using the trained local ML models (“the node monitors data from each valid sensor (604). Then, a valid neuromorphic program (NP) processes the input data 606. Thereafter, the node checks whether there is data in the ‘alert’ class 608. If yes, then the system triggers an alert process 616,” paragraph 0056 lines 2-6). Accordingly, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify the local computing devices of Kairouz/Yang/Nayak to comprise the wireless sensor devices of Moreno. (1) The Examiner finds that the prior art included each claim element listed above, although not necessarily in a single prior art reference, with the only difference between the claimed invention and the prior art being the lack of actual combination of the elements in a single prior art reference. (2) The Examiner finds that one of ordinary skill in the art could have combined the elements as claimed by known software development methods, and that in combination, each element merely performs the same function as it does separately. (3) The Examiner finds that one of ordinary skill in the art would have recognized that the results of the combination were predictable, namely classifying an alarm type using the trained local ML models (“the node monitors data from each valid sensor (604). Then, a valid neuromorphic program (NP) processes the input data 606. Thereafter, the node checks whether there is data in the ‘alert’ class 608. If yes, then the system triggers an alert process 616,” Moreno paragraph 0056 lines 2-6). Therefore, the rationale to support a conclusion that the claim would have been obvious is that the combining prior art elements according to known methods to yield predictable results to one of ordinary skill in the art. See MPEP § 2143(I)(A). Conclusion The prior art made of record and not relied upon is considered pertinent to Applicant’s disclosure: US 2014/0058763 A1 disclosing k-means clustering for machine learning Applicant is required under 37 C.F.R. § 1.111(c) to consider these references fully when responding to this action. It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331, 1332-33, 216 U.S.P.Q. 1038, 1039 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 U.S.P.Q. 275, 277 (C.C.P.A. 1968)). In the interests of compact prosecution, Applicant is invited to contact the examiner via electronic media pursuant to USPTO policy outlined MPEP § 502.03. All electronic communication must be authorized in writing. Applicant may wish to file an Internet Communications Authorization Form PTO/SB/439. Applicant may wish to request an interview using the Interview Practice website: http://www.uspto.gov/patent/laws-and-regulations/interview-practice. Applicant is reminded Internet e-mail may not be used for communication for matters under 35 U.S.C. § 132 or which otherwise require a signature. A reply to an Office action may NOT be communicated by Applicant to the USPTO via Internet e-mail. If such a reply is submitted by Applicant via Internet e-mail, a paper copy will be placed in the appropriate patent application file with an indication that the reply is NOT ENTERED. See MPEP § 502.03(II). Any inquiry concerning this communication or earlier communications from the examiner should be directed to Ryan Barrett whose telephone number is 571 270 3311. The examiner can normally be reached 9:00am to 5:30pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, Applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor Michelle Bechtold can be reached at 571 431 0762. The fax phone number for the organization where this application or proceeding is assigned is 571 273 8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /Ryan Barrett/ Primary Examiner, Art Unit 2148
Read full office action

Prosecution Timeline

Jul 27, 2023
Application Filed
Mar 25, 2026
Non-Final Rejection mailed — §103
Jun 24, 2026
Response Filed
Aug 21, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748970
NEURAL NETWORK MODEL TRAINING METHOD, DATA PROCESSING METHOD, AND APPARATUS
3y 2m to grant Granted Sep 29, 2026
Patent 12749020
REINFORCEMENT LEARNING FOR MACHINE LEARNING MODELS USING DYNAMIC CONFIDENCE THRESHOLDS
2y 10m to grant Granted Sep 29, 2026
Patent 12743660
RANKED PRUNING OF DATA SET TO TRAIN MACHINE LEARNING MODEL MODELS
3y 7m to grant Granted Sep 22, 2026
Patent 12743187
Display Method and Related Apparatus
2y 3m to grant Granted Sep 22, 2026
Patent 12731074
SYSTEM AND METHOD FOR A PERSISTENT AND PERSONALIZED DATASET SOLUTION FOR IMPROVING GUEST INTERACTION WITH AN INTERACTIVE AREA
3y 8m to grant Granted Sep 08, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
66%
Grant Probability
99%
With Interview (+41.2%)
3y 3m (~1m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 429 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month