Prosecution Insights
Last updated: October 01, 2026
Application No. 18/663,656

MODEL TRAINING METHOD AND COMMUNICATION APPARATUS

Non-Final OA §103
Filed
May 14, 2024
Priority
Nov 15, 2021 — CN 202111350005.3 +1 more
Examiner
COLEMAN, PAUL
Art Unit
Tech Center
Assignee
Huawei Technologies Co., Ltd.
OA Round
1 (Non-Final)
65%
Grant Probability
Favorable
1-2
OA Rounds
1y 4m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 65% — above average
65%
Career Allowance Rate
17 granted / 26 resolved
+5.4% vs TC avg
Strong +47% interview lift
Without
With
+47.4%
Interview Lift
resolved cases with interview
Typical timeline
3y 8m
Avg Prosecution
15 currently pending
Career history
39
Total Applications
across all art units

Statute-Specific Performance

§101
31.3%
-8.7% vs TC avg
§103
47.0%
+7.0% vs TC avg
§102
4.2%
-35.8% vs TC avg
§112
16.9%
-23.1% vs TC avg
Black line = Tech Center average estimate • Based on career data from 26 resolved cases

Office Action

§103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . Information Disclosure Statement The information disclosure statements (IDS) submitted on September 16, 2024, and March 12, 2025 are in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statements are being considered by the examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1-2, 7-8, and 14-15 are rejected under 35 U.S.C. 103 as being unpatentable over Dinesh C. Verma et al. (US2020/0219014A1) in view of Tao Lin et al. (Don’t Use Large Mini-Batches, Use Local SGD). Regarding claim 1, Verma in view of Lin, teaches a model training method comprising: “performing, by an i t h device in a k t h group of devices, n * M i t h device completes a model parameter exchange with at least one other device in the k t h group of devices k t h group of devices, M is greater than or equal to 2, ” – Verma teaches this limitation in part. Verma teaches a plurality of system nodes, each having a respective model: “The terms "a plurality" may be understood to include any integer number greater than or equal to two” (Verma, pg. 2, ¶ [0019]) The plurality of participating system nodes corresponds to the claimed K t h group of devices, an individual system node corresponds to the claimed i t h device, and the quantity of system nodes corresponds to claimed M . Verma further teaches local model training by each individual system node. Specifically: “Each agent uses the mini-batches to train its local AI model.” (Verma, pg. 4, ¶ [0045]) Verma also teaches movement of model parameters among the participating nodes: “the neural network model parameters can be moved and the neural networks at each system node can be trained using local data.” (Verma, pg. 3, ¶ [0035]) “The neural network model parameters are moved until each agent has the opportunity to train each neural network model with its own local data.” (Verma, pg. 3, ¶ [0035]) Verma further teaches an ordered transfer between participating system nodes: “a permutation order for moving the neural network models between system nodes for training.” (Verma, pg. 4, ¶ [0042]) After local training, each agent may: “pass their respective model along to the next system node in the permutation order.” (Verma, pg. 4, ¶ [0044]) Verma expressly permits that transfer to occur directly between participating devices: “the agent can transfer the model to the next system node in a peer to peer manner” (Verma, pg. 4, ¶ [0046]) Thus, Verma teaches an i t h device in a group of at least two devices performing model training and completing model/model-parameter exchange with another participating device of the same group. However, Verma does not expressly teach that the i t h device performs specifically: “ n * M times of” model training; that the parameter exchange occurs specifically: “every M times of model training”; or that “n is an integer”. Accordingly, Verma teaches a permutation and complete training rounds, but does not express the number of local model-training operations using the claimed n * M relationship or expressly equate an exchange interval to the number M of participating devices. “sending, by the i t h device, a model M i , n * M to a target device, wherein the model M i , n * M is obtained by the i t h device – Verma teaches this limitation in part. Verma teaches that, after an agent performs local model training, the resulting model is passed to the next participating system node in the permutation order: ”The mini-batches enables each agent to train its local model with same data classes and pass their respective model along to the next system node in the permutation order.” (Verma, pg. 4, ¶ [0044]) Verma further teaches: “The fusion server can then direct the model over to the next system node in the permutation sequence.” (Verma, pg. 4, ¶ [0045]) Verma further expressly permits the trained model to be sent directly from the training device to that next system node: “the agent can transfer the model to the next system node in a peer to peer manner” (Verma, pg. 4, ¶ [0046]) Thus, the next system node in Verma’s permutation corresponds to the claimed target device, and Verma teaches sending, by the i t h device, the resulting trained model to that target device. Verma does not, however, expressly establish that the resulting model is obtained specifically “by completing the n*M times of model training”. Verma does not teach these limitations and/or portions of: “n*M times of” model training; that the parameter exchange occurs specifically: “every M times of model training”; “n is an integer”. “by completing the n*M times of model training.” Lin, however, teaches these remaining limitations and/or portions of: “n*M times of” – Lin teaches that each worker performs H sequential loop SGD updates between synchronization rounds: “In local SGD, each worker k ∈ K evolves a local model by performing sequential SGD updates … before communication (synchronization by averaging) among the workers.” (Lin, pg. 1, §§ Introduction – Local SGD) And further describes the local model as existing after: “ t global synchronization rounds and subsequent h ∈ H local steps” (Lin, pg. 1, §§ Introduction – Local SGD) Lin’s Algorithm 1 likewise defines: “input: number of synchronization steps T ; input: number of local steps H ; input: number of nodes K .” (Lin, pg. 18, Algorithm 1) For each synchronization step t , each node performs the local-update loop: “for h := 1, …, H do” (Lin, pg. 18, Algorithm 1) before the synchronization operation is performed. Thus, after T synchronization intervals, each node has performed T * H local model-training operations. Lin further expressly discloses a post-local SGD configuration: “A5: post-local SGD ( K = 16 , H = 16 )” (Lin, pg. 2, Figure 1) Designating Lin’s number of nodes K as the claimed quantity of devices M , Lin’s disclosed configuration has: H = K = M = 16 . Therefore, after an integer number T of synchronization intervals, each worker performs: T * H = T * M model training operations. Designating Lin’s synchronization-round count T as the claimed n , Lin thereby teaches “n*M times of” model training. “every M times of model training” – Lin expressly teaches that synchronization occurs after every H local SGD updates: “each worker k ∈ K evolves a local model by performing sequential SGD updates … before communication (synchronization by averaging) among the workers.” (Lin, pg. 1, §§ Introduction – Local SGD) Lin’s Algorithm 1 makes the periodic relationship explicit. For each synchronization round, every node performs: “for h := 1, …, H do” (Lin, pg. 18, Algorithm 1) Including computation of the gradient and an update of the local model. Only after the H local updates are completed does Lin perform: “all-reduce aggregation of the gradients” (Lin, pg. 18, Algorithm 1) and obtain a: “new global (synchronized) model … for all K nodes” (Lin, pg. 18, Algorithm 1) Lin expressly discloses the post-local SGD configuration: “A5: post-local SGD ( K = 16 , H = 16 )” (Lin, pg. 2, Figure 1) Thus, in Lin’s disclosed embodiment: K = 16 is the number of participating workers/nodes; H = 16 is the number of model-training updates performed before synchronization. Designating Lin’s K as claimed M , Lin teaches H = M = 16 . Accordingly, each participating worker completes synchronization after every 16 local model-training updates, i.e., “every M times of model training”. “n is an integer” – Lin’s Local SGD algorithm defines the training process using a discrete number of synchronization steps: “input: number of synchronization steps T ;” (Lin, pg. 18, Algorithm 1) and then iterates: “for t := 1, …, T do” (Lin, pg. 18, Algorithm 1) Thus, T is a count of synchronization rounds and is necessarily an integer quantity. As discussed above, each synchronization interval comprises H local training updates. Designating Lin’s synchronization-round count T as the claimed n , Lin therefore teaches “n is an integer”. “by completing the n*M times of model training.” – Lin teaches this limitation. Lin’s Algorithm 1 expressly updates each local model during each one of the H local training steps: “update the local model” (Lin, pg. 18, Algorithm 1) This update is within the h   : =   1 ,   … ,   H loop, after which synchronization occurs. Therefore, after n synchronization intervals, each worker’s trained model has resulted from completion of n * H local model-training operations. In the expressly disclosed K = 16 ,   H = 16 configuration, where H = K = M , that model has been obtained after: n * H   =   n * M local model-training operations. Verma, as discussed above, teaches sending the resulting trained model to the fusion server (see Verma, pg. 4, ¶ [0045]). Accordingly, Verma in view of Lin teaches a model sent by the i t h device to the target device, “wherein the model M i , n * M is obtained by the i t h device by completing the n*M times of model training”. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Verma’s distributed model-training method to use Lin’s post-local SGD synchronization schedule. As established above, Verma performs distributed stochastic-gradient-based model training at multiple system nodes and permits the nodes to perform local model training between synchronization events. Lin teaches a known Local SGD technique in which workers perform multiple local model-training updates before synchronization, thereby reducing the number of communication rounds and improving communication efficiency. One of ordinary skill in the art therefore would have had reason to employ Lin’s post-local SGD synchronization schedule in Verma’s distributed training system to reduce synchronization and communication overhead while allowing the participating devices to continue training locally between synchronization events. Applying Lin’s expressly disclosed K = 16 ,     H = 16 configuration to Verma’s participating training devices would have predictably caused each of 16 participating devices to perform 16 local model-training operations before synchronization. With Lin’s worker count K corresponding to claimed M , the resulting system would therefore synchronize “every M times of model training”. Thus, n successive synchronization intervals, each comprising M local model-training operations, result in n * M model-training operations. Regarding claim 2, Verma in view of Lin, teach the model training method according to claim 1, wherein the target device: “ k t h ” “ k t h ” “or the target device specified by a device other than the devices of the devices k t h group of devices.” – Verma teaches this limitation. Verma teaches that the participating system nodes are separate from a fusion server. In the disclosed architecture, the: “system includes a sequence of interconnected systems nodes and agents that in electrical connection with a fusion server.” (Verma, pg. 3, ¶ [0033]) Verma further states that the fusion server: “determines a random permutation order between the agents” (Verma, pg. 3, ¶ [0033]) and: “mediating a transfer of models based upon the permutation order” (Verma, pg. 3, ¶ [0033]) Verma more specifically teaches that the fusion server generates instructions identifying the order in which the models are to move among the participating system nodes: “the fusion server creates a set of control and coordination instructions for each agent at the system nodes” (Verma, pg. 4, ¶ [0042]) and those instructions: “may include a permutation order for moving the neural network models between system nodes for training.” (Verma, pg. 4, ¶ [0042]) After an agent trains its model, Verma teaches that: “The fusion server can then direct the model over to the next system node in the permutation sequence.” (Verma, pg. 4, ¶ [0045]) Verma further expressly permits the model to be sent directly by the training device to that specified next system node: “the agent can transfer the model to the next system node in a peer to peer manner” (Verma, pg. 4, ¶ [0046]) Thus, in this embodiment, the next system node corresponds to the claimed target device, and the fusion server, which is separate from the participating system nodes of the k t h group, corresponds to “a device other than the devices of the k t h group of devices.” By determining the permutation order and directing the model to the next system node in that permutation sequence, the fusion server specifies the target device. Because claim 2 recites the three target-device conditions in the alternative, Verma’s disclosure of the third alternative is sufficient to meet the additional limitation of claim 2. Accordingly, claim 2 would have been obvious over Verma in view of Lin for the same reasons set forth above with respect to claim 1, with Verma further teaching the additional limitation of claim 2. Regarding claim 7 Claim 7 recites substantially the same functional limitations already analyzed above with respect to corresponding method claim 1, but recast in communication-apparatus form. Accordingly, the teachings of the respective references identified above with respect to claim 1 are equally applicable to the corresponding functional limitations of claim 7. In particular, the processor-executed operations of claim 7 correspond to the respective steps of claim 1, including performing n * M times of model training, completing model-parameter exchange with at least one other device every M times of model training, wherein M is the quality of devices in the k t h group, and sending the model M i , n * M to the target device after completion of the n * M times of model training. In addition, Verma expressly teaches a distributed-learning system comprising a processor communicatively coupled to a memory and configured to perform the disclosed distributed-learning operations. Accordingly, implementing the corresponding method operations of claim 1 by the processor and memory recited in claim 7 is expressly contemplated by Verma. Accordingly, the Verma-Lin combination applied to claim 1 likewise teaches or suggests the corresponding limitations of claim 7. Regarding claim 8 Claim 8 recites substantially the same functional limitations already analyzed above with respect to corresponding method claim 2, but recast in communication-apparatus form. Accordingly, the teachings of the respective references identified above with respect to claim 2 are equally applicable to the corresponding functional limitations of claim 8. In particular, Verma teaches that the target device is specified by a device other than the devices of the k t h group of devices. As discussed above with respect to claim 2, the next system node in Verma’s permutation sequence corresponds to the claimed target device, while the fusion server, which is separate from the participating system nodes of the k t h group, corresponds to the device other than the devices of the k t h group. By determining the permutation order and directing the model to the next system node in that permutation sequence, the fusion server specifies the target device. Accordingly, the Verma-Lin combination applied to claim 2 likewise teaches or suggests the corresponding limitations of claim 8. Claims 3 and 9 are rejected under 35 U.S.C. 103 as being unpatentable over Verma in view of Lin and further in view of Deepak Narayanan et al. (PipeDream: Generalized Pipeline Parallelism for DNN Training). Regarding claim 3, Verma in view of Lin and further in view of Narayanan, teach the model training method according to claim 1, wherein the performing, by the i t h device in the k t h group of devices, n*M times of model training comprises: “for a j t h time of model training in the n*M times of model training, i t h ( i - 1 ) t h ( i - 1 ) t h ” – Verma teaches this limitation in part. Verma teaches repeated local model training and sequential transfer among participating system nodes: ”The mini-batches enables each agent to train its local model with same data classes and pass their respective model along to the next system node in the permutation order.” (Verma, pg. 4, ¶ [0044]) Thus, Verma teaches successive model-training operations involving adjacent devices in an ordered sequence. Verma does not teach these limitations and/or portions of: “receiving, by the i t h device, a result obtained through inference by an ( i - 1 ) t h device from the ( i - 1 ) t h device;” “determining, by the i t h , a first gradient and a second gradient based on the received result, wherein the first gradient is for updating a model M i , j - 1 , the second gradient is for updating a model M i - 1 , j - 1 , the model M i , j - 1 is obtained by the i t h device by completing a ( j - 1 ) t h time of model training, and the model M i - 1 , j - 1 is obtained by the ( i - 1 ) t h device by completing the ( j - 1 ) t h time of model training;” ”training, by the i t h device, the model M i , j - 1 based on the first gradient.” Narayanan, however, teaches these limitations and/or portions of: “receiving, by the i t h device, a result obtained through inference by an ( i - 1 ) t h device from the ( i - 1 ) t h device;” – Narayanan teaches or renders obvious this limitation. Narayanan teaches partitioning a DNN into stages executed on successive workers: “Each stage is mapped to a separate GPU that performs the forward pass (and backward pass) for all layers in that stage.” (Narayanan, pg. 4, § 3) Upon completing its forward computation: “each stage asynchronously sends the output activations to the next stage” (Narayanan, pg. 4-5, § 3) Narayanan expressly teaches that a stage receives intermediate data from the prior stage: “Upon receiving intermediate data from the prior stage” (Narayanan, pg. 8, § 4) The stage receiving intermediate data from the prior stage copies that data to GPU memory for processing. Thus, designating consecutive workers as the ( i - 1 ) t h and i t h devices, the i t h device receives the output activations produced by the ( i - 1 ) t h device’s forward computation. Under the broadest reasonable interpretation, those activations are a result obtained through inference by the proceeding device. “determining, by the i t h , a first gradient and a second gradient based on the received result, wherein the first gradient is for updating a model M i , j - 1 , the second gradient is for updating a model M i - 1 , j - 1 , the model M i , j - 1 is obtained by the i t h device by completing a ( j - 1 ) t h time of model training, and the model M i - 1 , j - 1 is obtained by the ( i - 1 ) t h device by completing the ( j - 1 ) t h time of model training;” – Narayanan teaches that DNN training is bidirectional: “a forward pass through the computation graph is followed by a backward pass that uses state and intermediate data computed during the forward pass.” (Narayanan, pg. 1, § Abstract) Thus, the backward computation performed by the i t h stage uses the forward/intermediate result received from the preceding stage. Narayanan further teaches that (during the corresponding backward pass): “The same weight version is then used to compute the weight update and upstream weight gradient” (Narayanan, pg. 7, § 3.3) Narayanan therefore teaches that the i t h stage’s backward computation produces both (1) information for updating its locally maintained model parameters and (2) an upstream gradient for propagation to the preceding stage. Narayanan expressly teaches: “On completing its backward pass, each stage asynchronously sends the gradient to the previous stage” (Narayanan, pg. 5, § 3) The upstream gradient transmitted to the ( i - 1 ) t h stage is used by that stage to continue back-propagation and determine the local parameter gradient by which the ( i - 1 ) t h stage’s model is updated. Thus, the local parameter gradient at the i t h stage corresponds to the claimed first gradient, and the upstream gradient corresponds to the claimed second gradient. Narayanan further teaches successive local model states: “Each backward pass in a stage results in weight updates; the next forward pass uses the latest version of weights available” (Narayanan, pg. 2, § 1) And, as training proceeds, Narayanan’s PipeDream: “applies updates to the most recent parameter version” (Narayanan, pg. 8, § 4) Accordingly, designating the model state produced by the preceding ( j - 1 ) t h training/update operation at the i t h stage as M i , j - 1 , and the corresponding preceding model state at the ( i - 1 ) t h stage as M i - 1 , j - 1 , Narayanan teaches or renders obvious the recited model-state relationship. ”training, by the i t h device, the model M i , j - 1 based on the first gradient.” – Narayanan teaches that: “Each backward pass in a stage results in weight updates; the next forward pass uses the latest version of weights available” (Narayanan, pg. 2, § 1) As discussed above, the weight update is obtained from the local parameter gradient computed during the i t h worker’s backward pass. Thus, the i t h worker trains/updates its model based on the claimed first gradient. It would have been obvious to one of skill in the art before the effective filing date to employ Narayanan’s pipeline-parallel DNN training technique in Verma’s distributed neural-network training system. As established above, Verma distributes neural-network training among interconnected system nodes. Narayanan teaches partitioning DNN training among successive workers such that forward activations are communicated downstream and backward gradients are communicated upstream. Narayanan identifies reduced inter-worker communication as a technical benefit: each pipeline worker communicates only relevant activations and gradients with a successive worker, resulting in communication reduction exceeding 85% for certain models (see Narayanan, pg. 5, § 3). One of ordinary skill therefore would have had reason to use Narayanan’s pipeline technique in Verma’s distributed training system to reduce inter-device communication and distribute neural-network computations among the participating devices. The modification would predictably cause successive devices to exchange forward activations and backward gradients, with each device computing a local parameter gradient to update its model and propagating an upstream gradient to the preceding device. Regarding claim 9 Claim 9 recites substantially the same functional limitations already analyzed above with respect to corresponding method claim 3, but recast in communication-apparatus form. Accordingly, the teachings of the respective references identified above with respect to claim 3 are equally applicable to the corresponding functional limitations of claim 9. In particular, Narayanan teaches successive workers communicating forward-pass output activations and backward-pass gradients, with each worker determining and applying a local weight update and communicating an upstream gradient to the preceding worker. Thus, the same teachings relied upon for receiving an inference result from the ( i - 1 ) t h device, determining the first and second gradients, associating those gradients with the respective preceding model states, and training the i t h -device model based on the first gradient apply equally to the apparatus of claim 9. Narayanan expressly describes output activations and gradients being communicated between successive workers and local parameter versions being updated during training. Accordingly, the Verma-Lin-Narayanan combination applied to claim 3 likewise teaches or suggests the corresponding limitations of claim 9. Claims 4 and 10 are rejected under 35 U.S.C. 103 as being unpatentable over Verma in view of Lin in further view of Narayanan and further in view of Greg Storm et al. (WO2021/226302A1). Regarding claim 4, Verma in view of Lin in further view of Narayanan and further in view of Storm, teach the model training method according to claim 3, wherein “the first gradient is determined based on the received result 1 s t i = M ;” – Verma does not teach this limitation. Narayanan, however, teaches or renders obvious this limitation in part. Narayanan teaches: “The last stage starts the backward pass on a minibatch immediately after the forward pass completes.” (Narayanan, pg. 5, § 3) Narayanan further explains that DNN training is bidirectional, wherein the backward pass: “uses state and intermediate data computed during the forward pass.” (Narayanan, pg. 1, § Abstract) Thus, designing the last/output stage as the claimed M t h device, Narayanan teaches or renders obvious that, when i = M , the local parameter gradient corresponding to the claimed first gradient is determined based on the received forward/intermediate result. “and the first gradient is determined based on the second gradient transmitted by an ( i + 1 ) t h device in response to determining i ≠ M .” – Verma does not teach this limitation. Narayanan, however, teaches or renders obvious this limitation. Narayanan expressly teaches: “On completing its backward pass, each stage asynchronously sends the gradient to the previous stage” (Narayanan, pg. 5, § 3) Thus, for a non-final stage i , the gradient generated by succeeding stage ( i + 1 ) is transmitted backward to stage i . Narayanan further teaches: “Each backward pass in a stage results in weight updates” (Narayanan, pg. 2, § 1) and that the backward pass computes the gradients used for the corresponding model update. Accordingly, the gradient transmitted from stage ( i + 1 ) is used in stage i ’s backward computation to determine the local parameter gradient by which stage i ’s model is updated. The transmitted upstream gradient corresponds to the claimed second gradient, and the local parameter gradient determined at stage i corresponds to the claimed first gradient. Narayanan does not teach these limitations and/or portions of: “a label received from a 1 s t device” Storm, however, teaches these limitations and/or portions of: “a label received from a 1 s t device” – Storm teaches split neural-network training in which a client-side portion performs an initial forward computation and sends the resulting activations to a server-side portion. Specifically, Storm teaches that the claim: “The clients 206,208, 210 in their respective turn do a forward step on the model A and sends the output of A (i.e., activations at S only or SA1 (206B), SA2 (208B), SAN 210B)) to the server 202 in addition to the required labels.” (Storm, col. 5-6, lines 64-67, 1) Thus, Storm expressly teaches transmitting a trained label from the device performing the input-side forward computation to the device performing the output-side computation, where the received activation and label are used for loss calculation and backpropagation. Applying Storm’s label-transfer teaching to Narayanan’s pipeline, the first/input-stage device corresponds to the claimed 1 s t device, and the last/output-stage device corresponds to the claimed M t h device. The resulting arrangement therefore provides the M t h device with: “a label received from a 1 s t device” for determining the first gradient based on the received result and the corresponding label. It would have been obvious to one of ordinary skill in the art before the effective filing date to provide the label associated with a training sample from Narayanan’s first/input-stage device to its final/output-stage device, as taught by Storm. As established above, Narayanan teaches that the final stage receives the forward result and initiates the backward computation. Storm expressly teaches a known split-learning implementation in which the input-side device sends the forward activation in addition to the required label to the output-side device, which calculates the loss and performs back-propagation. One of ordinary skill therefore would have had reason to make the corresponding training label available to Narayanan’s final/output stage so that the final stage could calculate the supervised loss associated with the received forward result and determine the gradient imitating back-propagation. The modification would predictably preserve Narayanan’s pipeline partitioning and backward-gradient propagation while supplying the final stage with the supervision information associated with the training sample. Regarding claim 10 Claim 10 recites substantially the same functional limitations already analyzed above with respect to corresponding method claim 4, but recast in communication-apparatus form. Accordingly, the teachings of the respective references identified above with respect to claim 4 are equally applicable to the corresponding functional limitations of claim 10. In particular, the previously identified teachings concerning determining the first gradient, when i=M, based on the received inference result and the label associated with the training data, and determining the first gradient, when i ≠ M , based on the second gradient transmitted by the ( i + 1 ) t h device, are equally applicable to the apparatus configured to perform those same operations. Accordingly, the Verma-Lin-Narayanan-Storm combination applied to claim 4 likewise teaches or suggests the corresponding limitations of claim 10. Claims 5 and 11 are rejected under 35 U.S.C. 103 as being unpatentable over Verma in view of Lin and further in view of Abhijit Guha Roy et al. (BrainTorrent: A Peer-to-Peer Environment for Decentralized Federated Learning). Regarding claim 5, Verma in view of Lin and further in view of Roy, teach the model training method according to claim 1, further comprising: “ i t h device completing the model parameter exchange with the at least one other device in the k t h group of devices, i t h k t h ” – Verma teaches this limitation in part. As established with respect to claim 1, Verma teaches model/model-parameter transfer between devices. In particular, Verma teaches: “the agent can transfer the model to the next system node in a peer to peer manner” (Verma, pg. 4, ¶ [0046]) Thus, Verma teaches an i t h device completing a model-parameter exchange with another device. Verma also expressly teaches determining a quantity of samples locally maintained at each system node. Figure 5 provides: “Determine number of data points at each system node” (Verma, FIG. 5, step 501) Consistently, Verma teaches: “Each agent computes statistics of data distribution stored at its respective system node. The agents share these statistics with the fusion server” (Verma, pg. 3, ¶ [0034]) Thus, Verma teaches both peer-to-peer model-parameter transfer and a sample quantity corresponding to data locally stored at each participating device. Verma does not, however, teach that, in connection with completion of the peer model-parameter exchange, the i t h device exchanges its locally stored sample quantity with that other peer device. Verma does not teach these limitations and/or portions of: “in response to … exchanging, by the i t h device, a locally stored sample quantity with the at least one other device in the k t h group of devices.“ Roy, however, teaches these limitations and/or portions of: “in response to … exchanging, by the i t h device, a locally stored sample quantity with the at least one other device in the k t h group of devices.“ – Roy expressly defines a peer’s local training-sample quantity. Roy considers N centers C 1 , … , C N , wherein each C i has: “training data D i = x i , y i , … , ( x a i , y a i with a i labeled samples.” (Roy, pg. 3, § 2) Thus, a i expressly represents the quantity of training samples locally maintained by client C i . Roy then teaches a serverless peer-to-peer environment in which: “all clients C i i = 1 N in this environment are connected directly in a peer-to-peer fashion” (Roy, pg. 4, § 2.2) And, when another client has an updated model: “All clients C j with updates, i.e., v o l d j < v n e w j , send their weights W j and the training sample size a j to C i .” (Roy, pg. 4, § 2.2) Roy’s Algorithm 1 confirms the corresponding receiving operation: “Receive updated W j and a j from C j ;” (Roy, pg. 5, Algorithm 1) Thus, designating C j as the claimed i t h device and C i as the claimed ”at least one other device”, Roy teaches a device exchanging its model weights and its locally stored training-sample quantity a j to another peer device as part of the corresponding peer model-update exchange. Under the broadest reasonable interpretation of the claim, the recited “in response to … completing the model parameter exchange” does not require a separate sample-quantity transmission that begins only after the model-parameter transmission has temporally terminated. Rather, the responsive relationship encompasses exchanging the sample quantity in connection with the corresponding model-parameter exchange. Roy expressly satisfies that relationship by transmitting the updated weights W j and associated training-sample quantity a j together in the same peer update. It would have been obvious to one of ordinary skill in the art before the effective filing date to modify Verma’s peer-to-peer model-parameter exchange to additionally exchange the transmitted device’s locally stored sample quantity, as taught by Roy. Verma already determines data statistics at the respective system nodes, including the “number of data points at each system node”, and uses data-distribution information to coordinate its distributed-training process. Verma also expressly permits direct peer-to-peer transfer of the trained model between system nodes. Roy demonstrates a known peer-to-peer distributed-training implementation in which a device communicates its updated model weights together with the quantity of training samples locally maintained by that device. Roy further associates the sample quantity a j with the corresponding model W j during processing of the received peer update. One of ordinary skill therefore would have had reason to communicate Verma’s already-maintained local sample quantity with the corresponding peer-transferred model so that information concerning the amount of local training data associated with that model is available to the receiving device. The modification would predictably provide model-associated training-data information during Verma’s existing peer model exchange using Roy’s known technique of transmitting the model weights and corresponding local sample quantity together. Regarding claim 11 Claim 11 recites substantially the same functional limitations already analyzed above with respect to corresponding method claim 5, but recast in communication-apparatus form. Accordingly, the teachings of the respective references identified above with respect to claim 5 are equally applicable to the corresponding functional limitations of claim 11. In particular, Roy teaches a peer-to-peer federated-learning arrangement in which a client having an updated model sends both its model weights and its local training sample size a j to another for model aggregation. Thus, the same teaching relied upon for exchanging a locally stored sample quantity when model-parameter exchange occurs applies equally to the apparatus of claim 11. Accordingly, the Verma-Lin-Roy combination applied to claim 5 likewise teaches or suggests the corresponding limitations of claim 11. Claims 6, 12, 14, and 16 are rejected under 35 U.S.C. 103 as being unpatentable over Verma in view of Lin and further in view of Srinivas Sridharan et al. (US20190205745A1). Regarding claim 6, Verma in view of Lin and further in view of Sridharan, teach the model training method according to claim 1, further comprising: “for a next time of training following the n * M time of model training, obtaining, by the i t h device, information about a model M r from the target device, wherein the M r is r t h M i , n * M k , r ∈ 1 , M , the model M i , n * M k is obtained by the i t h device in the k t h group of devices by completing the n * M times of model training, i M k K .” – Verma teaches this limitation in part. Verma teaches that an agent trains its local AI model and may: “send the resulting AI model to the fusion server.” (Verma, pg. 4, ¶ [0045]) Verma further states: “At the end of the permutation order, the trained model is sent to the fusion server.” (Verma, pg. 4, ¶ [0045]) Thus, in this embodiment, the fusion server corresponds to the claimed target device that receives the model resulting from the preceding model-training operations. Verma further teaches: “the fusion server creates a set of control and coordination instructions for each agent at the system nodes” (Verma, pg. 4, ¶ [0042]) And that the fused model may be formed: “by averaging the weights of all of the AI models provided by the system node” (Verma, pg. 4, ¶ [0047]) After which the system nodes are directed: “to further train the fused AI-model using new mini-batches.” (Verma, pg. 4, ¶ [0047]) Figure 4 and ¶ [0048] provide the same iterative operation. Verma teaches that: “The fusion server can collect intelligence models trained on incremental mini-batches from each system node and average their weights to create a fused artificial intelligence model.” (Verma, pg. 6, ¶ [0048]) If additional training remains, the fusion server: “can direct the fused intelligence model weights to each system node to create next mini-batch, update the model weights” (Verma, pg. 5, ¶ [0048]) Thus, Verma teaches that, following the preceding training, the i t h device obtains from the target/fusion server information about a fused model for a next time of training, and that the fused model is obtained by the target device through fusion of trained model parameters. Verma does not, however, teach the claimed indexed inter-group arrangement in which there are M corresponding model positions within each of K groups and a respective r t h fused model is obtained from corresponding trained models across the groups. Verma does not teach these limitations and/or portions of: “an r t h model … inter-group … r ∈ 1 , M … i traverses from 1 to M , and k traverses from 1 to K .” Sridharan, however, teaches these limitations and/or portions of: “an r t h model” – Sridharan renders obvious this portion. Sridharan teaches combined model and data parallelism in which: “each computational node includes multiple GPUs.” (Sridharan, pg. 19, ¶ [0191]) Sridharan further teaches: “Each node can have a complete instance of the model with separate GPUs within each node are used to train different portions of the model.” (Sridharan, pg. 19, ¶ [0191]) Thus, Sridharan teaches multiple replicated complete models, with each complete model divided into a plurality of corresponding model portions trained by respective GPUs. Sridharan further teaches data-parallel parameter averaging in which: “the global parameters (e.g., weights, biases)” (Sridharan, pg. 19, ¶ [0190]) are set to the average of the parameters from the respective nodes, and teaches that parameter averaging uses a central parameter server maintaining the parameter data. (Sridharan, pg. 19, ¶ [0190]). When Sridharan’s disclosed parameter averaging is applied to its disclosed combined model/data-parallel arrangement, corresponding model portions in the replicated complete models are predictably synchronized or averaged with one another. The result is a respective fused/global model portion corresponding to each model-partition position. Accordingly, designating one such resulting corresponding fused model portion as the claimed r t h model renders obvious “an r t h model”. “inter-group” – Sridharan expressly teaches grouping worker nodes for distributed model training. In particular, Sridharan teaches that a communication framework may establish multiple groups of worker nodes for exchange of model-data parameters, including gradient updates: “can include multiple sets of worker nodes” (Sridharan, pg. 23, ¶ [0223]) “the communications framework can establish multiple groups of worker nodes.” (Sridharan, pg. 23, ¶ [0223]) Sridharan specifically teaches a first group 2210 and a second group 2230 and states: “One or more parameter servers 2220 can then be instantiated.” (Sridharan, pg. 23, ¶ [0223]) Sridharan further expressly states that the parameter servers may be: “configured to enable efficient inter-group communication between the various groups 2210, 2230.” (Sridharan, pg. 23, ¶ [0223]) Figure 22 correspondingly depicts first and second worker groups 2210 and 2230 communicating through parameter server 2220 (see Sridharan, FIG. 22) Sridharan further teaches parameter averaging through a central parameter server, whereby parameters from different training nodes are combined to form global parameters (see Sridharan pg. 20, ¶ [0190]). Thus, Sridharan expressly teaches parameter-server communication and synchronization between groups. When this known inter-group synchronization arrangement is used with Verma’s fusion-server averaging, the fusion occurs between model information originating from the respective groups. Accordingly, Sridharan teaches the “inter-group” relationship missing from Vera. “ r ∈ 1 , M ” – Sridharan renders obvious this portion. Sridharan teaches that, in combined model and data parallelism, each computational node includes multiple GPUs, each computational node may contain a complete model instance, and the separate GPUs within each node train different portions of that model: “separate GPUs within each node are used to train different portions of the model.” (Sridharan, pg. 20, ¶ [0191]; see also FIG. 19) Thus, when a plurality of GPUs/model-partition positions within each replicated computation node is designated as the claimed M device/model positions, there are M corresponding model portions capable of being synchronized across the replicated nodes. Sridharan’s parameter averaging then sets global model parameters based on the corresponding parameters from the respective nodes (see Sridharan, pg. 20, ¶ [0190]). Applying that averaging to the M respective model-partition positions predictably results in a respective synchronized/fused result for each of the M positions. It therefore would have been obvious to index a particular one of those M resulting fused model portions by r , where r ranges from 1 through M . Accordingly, Sridharan renders obvious “ r ∈ 1 , M ”. “ i traverses from 1 to M , and k traverses from 1 to K .” – Sridharan renders obvious this portion. Sridharan’s combined model/data-parallel architecture teaches multiple computational nodes, with each computational node including multiple GPUs, a complete instance of the model, and respective GPUs used to train different portions of the model (see Sridharan, pg. 20, ¶ [0191]). Figure 19 visually depicts this repeated arrangement of GPUs across multiple computational nodes. For purposes of the claimed arrangement, the Sridharan architecture may therefore be designated as follows: computational node k -> claimed k t h group of devices; GPU/model-partition position i within computational node k -> claimed i t h device/model position; and model information associated with GPU/model-partition position i in computation node k -> corresponding M i , n * M k model information. Sridharan further independently teaches that the number of distributed worker groups may vary based on the network topology and expressly establishes multiple groups of workers nodes for exchanging model-data parameters: “The specific number of groups that are created can vary based on the topology of the network.” (see Sridharan, pg. 23, ¶ [0223]) “The parameter servers can then be configured to enable efficient inter-group communication between the various groups” (see Sridharan, pg. 23, ¶ [0223]) Thus, where each computational-node group contains M respective device/model positions and the distributed architecture contains K such groups, ordinary indexing of those finite sets results in i traversing the M positions within each group and k traversing the K groups. The particular symbols i , M , k , and K are applicant’s nomenclature; Sridharan need not use those particular variable names where the underlying repeated group/device relationship is taught by the reference. Accordingly, Sridharan renders obvious “ i traverses from 1 to M , and k traverses from 1 to K .” It would have been obvious to one of ordinary skill in the art before the effective filing data to implement Verma’s distributed model training and fusion using Sridharan’s known combined model- and data-parallel architecture and parameter-synchronization techniques. Verma already teaches distributed training in which trained model information is provided to a fusion server, the fusion server averages the model weights to create a fused model, and the fused model information is returned to the participating system nodes for subsequent training. Verma pg. 4-5, ¶¶ [0047]-[0048]; FIG. 4). Sridharan teaches a known technique for scaling distributed neural-network training using combined model and data parallelism. In particular, Sridharan teaches complete model instances distributed across computational nodes while separate GPUs within each node train different portions of the model. Sridharan pg. 20, ¶ [0191]. Sridharan further teaches central parameter averaging across distributed nodes (see Sridharan, pg. 20, ¶ [0190]). One of ordinary skill therefore would have had reason to implement Verma’s distributed model training using Sridharan’s combined mode/data-parallel arrangement so that the computation and storage required for each model could be distributed among multiple devices while retaining Verma’s central fusion and redistribution for subsequent training. Applying Sridharan’s parameter-averaging technique to the data-parallel dimension of its combined architecture would predictably combine corresponding model parameters of model portions across the replicated computational-node groups, yielding a respective fused model portion for each corresponding model position. Sridharan further expressly teaches grouping distributed-training workers and using a parameter server for inter-group communication of model-data parameters. Sridharan, pg. 23, ¶ [0223]). Thus, using Sridharan’s grouping and synchronization technique with Verma’s existing server-side fusion would predictably cause Verma’s target/fusion server to fuse corresponding trained model information originating from different groups and return the respective fused information for subsequent training. The resulting arrangement would provide M corresponding device/model positions across K groups, with corresponding model information fused between the groups to obtain respective fused model results, as claimed. Regarding claim 12 Claim 12 recites substantially the same functional limitations already analyzed above with respect to corresponding method claim 6, but recast in communication-apparatus form. Accordingly, the teachings of the respective references identified above with respect to claim 6 are equally applicable to the corresponding functional limitations of claim 12. In particular, Verma teaches obtaining fused model information from the fusion server for subsequent training, while Sridharan teaches combined model and data parallelism in which complete model instances are replicated across computational nodes and respective model portions are distributed within each node, together with parameter averaging and parameter-server communication between groups. Accordingly, the same teachings relied upon for the r t h fused model, inter-group fusion, r ∈ 1 , M , and the traversal of the M model/device positions across the K groups apply equally to the apparatus of claim 12. Accordingly, the Verma-Lin-Sridharan combination applied to claim 6 likewise teaches or suggests the corresponding limitations of claim 12. Regarding claim 14, Verma in view of Lin and further in view of Sridharan, teach a communication apparatus comprising: “at least one processor; and one or more memories coupled to the at least one processor and storing programming instructions that, when executed by the at least one processor, cause the communication apparatus to:” – Verma teaches this limitation. Verma taches a computer system for distributed learning including: “a memory and a hardware processor system communicatively coupled to the memory.” (Verma, pg. 1, ¶ [0006]) “The processor system is configured to perform the computer implemented method.” (Verma, pg. 1, ¶ [0006]) Verma further depicts processing system 200 having CPUs 21a-21c coupled to RAM 34 and other system components. (See Verma, pg. 3, ¶¶ [0036]-[0037], Fig. 2) Thus, Verma teaches the claimed processor, memory, and stored-instruction apparatus structure. “receive a model M i , n * M k sent by an i t h device in a k t h group of devices in K groups of devices, wherein the model M i , n * M k is a model obtained by the i t h device in the k t h group of devices by completing n*M times of model training, a quantity of devices included in each group of devices is M, M is greater than or equal to 2, n is an integer, I traverses from 1 to M, and k traverses from 1 to K;” – Verma teaches this limitation in part. Verma teaches that each agent trains a local AI model on one or more mini-batches and may: “send the resulting AI model to the fusion server.” (Verma, pg. 4, ¶ [0045]) Verma further teaches that the fusion server can: “collect intelligence models trained on incremental mini-batches from each system node and average their weights to create a fused artificial intelligence model.” (Verma, pg. 6, ¶ [0048]) Thus, Verma teaches a fusion-server apparatus receiving a trained model sent by a training device after model training. “and perform inter-group fusion on K groups of models, wherein the K groups of models comprise K models M i , n * M k .” – Verma teaches this limitation in part. Verma teaches: “The fusion server creates a fused AI model” (Verma, pg. 4, ¶ [0047]) And further teaches that the fused model may be produced by: “averaging the weights of all of the AI models provided by the system node” (Verma, pg. 4, ¶ [0047]) Figure 4 likewise teaches collecting the trained AI models and averaging their weights to create the fused AI model: “The fusion server can collect intelligence models trained on incremental mini-batches from each system node and average their weights to create a fused artificial intelligence model.” (Verma, pg. 5, ¶ [0048]; FIG. 4) Thus, Verma teaches fusion, but does not teach that the fusion is the specifically claimed inter-group fusion or corresponding models across K groups. Verma does not teach these limitations and/or portions of: “in a k t h group of devices in K groups of devices” “the n * M times of” “a quantity of devices included in each group of devices is M , M is greater than or equal to 2, n is an integer, i traverses from 1 to M , and k traverses from 1 to K ” “inter-group” “on K groups of models, wherein the K groups of models comprise K models M i , n * M k ” Lin, however, teaches these limitations and /or portions of: “the n * M times of” – Lin teaches Local SGD in which each worker performs H sequential SGD updates before synchronization by averaging: “each worker k ∈ K evolves a local model by performing sequential SGD updates … before communication (synchronization by averaging) among the workers.” (Lin, pg. 1, §§ Introduction – Local SGD) Lin further expressly discloses the configuration: “A5: post-local SGD ( K = 16 , H = 16 )” (Lin, pg. 2, Figure 1) For purposes of the claimed relationship, Lin’s 16 workers correspond to the claimed quantity M = 16 devices, and Lin’s H = 16 local model-trainnig updates correspond to n * M training operations with n = 1 and M = 16 . Thus, Lin’s expressly disclosed embodiment satisfies M ≥ 2 , n being an integer, and completion of n * M model-training operations. This mapping relies on Lin’s specific K = 16 , H = 16 embodiment; it does not rely on a general proposition that Lin requires H = K . It would have been obvious to one of ordinary skill in the art before the effective filing date to employ Lin’s local-update training scheme in Verma’s distributed-learning system so that each device performs multiple local training updates before providing its trained model for fusion. Lin expressly motivates Local SGD as balancing computation and communication resources and reducing the frequency of communication required during distributed training. Neither Verma nor Lin teach these limitations and/or portions of: “in a k t h group of devices in K groups of devices” “a quantity of devices included in each group of devices is M ” “ i traverses from 1 to M , and k traverses from 1 to K ” “inter-group” “on K groups of models, wherein the K groups of models comprise K models M i , n * M k ” Sridharan, however, teaches these limitations and/or portions of: “in a k t h group of devices in K groups of devices” and “a quantity of devices included in each group of devices is M ,” – Sridharan teaches or renders obvious these portions. Sridharan teaches combined model and data parallelism in which: “each computational node includes multiple GPUs.” (Sridharan, pg. 19, ¶ [0191]) Each node may have: “a complete instance of the model with separate GPUs within each node are used to train different portions of the model.” (Sridharan, pg. 19, ¶ [0191]; FIG. 19) Thus, each computational node provides a plurality of training devices operating on respective model portions. Designating a computational node as a claimed k t h group and its multiple GPUs as the M devices within that group yields the claimed group/device arrangement. Sridharan separately teaches that its distributed-training communication framework can: “exchange model data parameters, such as gradient updates, between the worker nodes, the communications framework can establish multiple groups of worker nodes.” (Sridharan, pg. 23, ¶ [0223]) Thus, Sridharan expressly contemplates a distributed-training architecture containing multiple groups of training devices. “ i traverses from 1 to M , and k traverses from 1 to K ” – Sridharan renders obvious this portion. Sridharan’s combined model/data-parallel architecture provides multiple device/model-partition positions within each computational node, repeated across multiple computational nodes. Sridharan ¶ [0191], FIG. 19. Accordingly: computational node k -> claimed k t h group; GPU/model-partition position i -> claimed i t h device; and the trained model information associated with position i in group k -> claimed M i , n * M k . Where there are M such device/model positions within each of K groups, ordinary indexing of those finite sets results in i traversing from 1 to M and k traversing from 1 to K . Thus, Sridharan renders obvious “ i traverses from 1 to M , and k traverses from 1 to K ”. “inter-group” – Sridharan teaches this portion. Sridharan teaches that local nodes may be assembled into compute groups and that: “distant nodes are bridged via synchronization operations performed with a parameter server.” (Sridharan, pg. 23, ¶ [0222]) Sridharan then describes first and second groups 2210 and 2230 and expressly teaches a parameter server configured to provide inter-group communication between the groups: “The parameter servers can then be configured to enable efficient inter-group communication between the various groups 2210, 2230.” (Sridharan, pg. 23, ¶ [0223]; FIG. 22) Thus, Sridharan expressly teaches synchronization of model information between groups, corresponding to the claimed “inter-group” relationship. “on K groups of models, wherein the K groups of models comprise K models M i , n * M k ” – Sridharan renders obvious this portion. Sridharan teaches data-parallel training in which complete model instances exist on different computational nodes and: “the global parameters (e.g., weights, biases)” (Sridharan, pg. 19, ¶ [0190]) are set to averages of the parameters from the respective nodes. Parameter averaging uses a central parameter server maintaining the parameter data. Sridharan further teaches that, in combined model/data parallelism, the complete model at each computational node is divided among respective GPUs that train different model portions: “Combined model and data parallelism 1906 can be implemented, for example, in a distributed system in which each computational node includes multiple GPUs. Each node can have a complete instance of the model with separate GPUs within each node are used to train different portions of the model.” (Sridharan, pg. 19, ¶ [0191]) Applying Sridharan’s disclosed parameter averaging to this replicated, partitioned arrangement predictably combines corresponding model portions across the replicated groups. For a particular device/model position i , one corresponding trained model of model portion M i , n * M k is provided from each of the K groups. Those K corresponding models therefore constitute the claimed K models M i , n * M k on which inter-group fusion is performed. Accordingly, Sridharan renders obvious “on K groups of models, wherein the K groups of models comprise K models M i , n * M k ”. It would have been obvious to one of ordinary skill in the art before the effective filing data to implement the Verma-Lin distributed-learning and fusion arrangement using Sridharan’s known combined model/data-parallel architecture and topology-aware grouping. Verma already teaches a central fusion server receiving trained models form disturbed system nodes and averaging their model weights. Lin teaches permitting the training devices to perform multiple local updates before synchronization. Sridharan teaches scaling distributed neural-network training through replicated complete models across computational nodes, distributing portions of those models among multiple GPUs, and synchronizing model information across groups through a parameter server. One of ordinary skill therefore would have had reason to organize the training devices of Verma and Lin according to Sriharan’s multi-device, multi-group architecture to distribute model computation among multiple devices while retaining central parameter fusion. Applying Verma’s model averaging to Sridharan’s replicated model partitions would predictably fuse corresponding trained model information from the respective groups, thereby producing the claimed inter-group fusion of K corresponding models. Regarding claim 16, Verma in view of Lin and further in view of Sridharan, teach the communication apparatus according to claim 14, wherein the communication apparatus is further caused to: “perform inter-group fusion on a q t h model in each of the K groups of models according to a fusion algorithm, wherein q ∈ [ 1 , M ] ” – Verma does not teach this limitation. Sridharan, however, teaches or renders obvious this limitation. As discussed with respect to claim 14, Sridharan teaches combined model and data parallelism in which: “each computational node includes multiple GPUs” (Sridharan, pg. 19, ¶ [0191]) And: “Each node can have a complete instance of the model with separate GPUs within each node are used to train different portions of the model.” (Sridharan, pg. 19, ¶ [0191]) Thus, each replicated computational-node group contains corresponding model portions at respective device/model positions. For a particular position q , where q is one of the M model/device positions, a corresponding q t h model portion exists in each of the K replicated groups. Sridharan further teaches parameter averaging in which: “the global parameters (e.g., weights, biases)” (Sridharan, pg. 19, ¶ [0190]) are set to the average of the parameters from the respective nodes, using: “a central parameter server that maintains the parameter data.” (Sridharan, pg. 19, ¶ [0190]) Applying Sridharan’s disclosed parameter-averaging algorithm to the combined model/data-parallel arrangement predictably averages the corresponding parameters of the same model portion across the replicated groups. Thus, for a selected model position q , the corresponding q t h model in each of the K groups is fused according to the parameter-averaging algorithm. Because the combined architecture contains M respective model/device positions, indexing any one of those positions as q results in q ∈ [ 1 , M ] . The variable q is applicant’s nomenclature for selecting one of the corresponding model positions; Sridharan need not use that particular symbol. Sridharan expressly teaches parameter-server synchronization between groups, including parameter servers configured to enable: “inter-group communication between the various groups 2210, 2230” (Sridharan, pg. 23, ¶ [0223]; FIG. 22) Accordingly, Sridharan teaches or renders obvious performing inter-group fusion on a q t h model in each of the K groups of models according to a fusion algorithm, wherein q ∈ [ 1 , M ] . Claim 13 is rejected under 35 U.S.C. 103 as being unpatentable over Verma in view of Lin further in view of Sridharan and further in view of Keith Bonawitz et al. (Towards Federated Learning at Scale: System Design). Regarding claim 13, Verma in view of Lin further in view of Sridharan and further in view of Bonawitz, teach the communication apparatus according to claim 12, wherein the communication apparatus is further caused to: “receive a selection result sent by the target device;” – Verma does not teach this limitation. Bonawitz, however, teaches this limitation. Bonawitz teaches a federated-learning server that selects a subset of available devices to participate in a training round and communicates the selection to those devices: “The server selects a subset of connected devices” (Bonawitz, § 2.2) And: “If the device has been selected, the FL runtime receives the FL plan” (Bonawitz, § 2.3) Thus, Bonawitz teaches a device receiving from the server an indication that the device was selected to participate, corresponding to the claimed selection result. “and obtain the information about the model M r from the target device based on the selection result.” – Verma does not teach this limitation. Bonawitz, however, teaches this limitation. Bonawitz teaches that, after selection: “The server sends the FL plan and an FL checkpoint with the global model to each of the devices.” (Bonawitz, § 2.2) Figure 1 likewise states that the: “Model and configuration are sent to selected devices” (Bonawitz, Figure 1) Thus, model information is sent by the server to a device based on the device having been selected. Applying Bonawitz’s selection protocol to the apparatus of Verma, Lin, and Sridharan, the model information supplied to the selected device corresponds to the information about M r established with respect to claim 12. It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Bonawitz’s known device-selection protocol into the distributed-training apparatus of Verma, Lin, and Sridharan. As discussed above, the combination of Verma, Lin, and Sridharan teaches distributing fused model information to participating devices for subsequent training. Bonawitz teaches selecting which available devices will participate in a training round and thereafter sending the model and configuration to those selected devices. One of ordinary skill therefore would have had reason to use Bonawitz’s selection mechanism to determine which devices receive the fused model information and participate in subsequent training, thereby coordinating participation and avoiding distribution of model information to devices not selected for the training round. Claim 15 is rejected under 35 U.S.C. 103 as being unpatentable over Verma in view of Lin in further view of Sridharan and further in view of Rituparna Saha et al. (FogFL: Fog-Assisted Federated Learning for Resource-Constrained IoT Devices). Regarding claim 15, Verma in view of Lin in further view of Sridharan and further in view of Saha, teach the communication apparatus according to claim 14, wherein “the communication apparatus is ” “” “or the communication apparatus is a device specified by a device other than the K groups of devices.” – Verma does not teach this limitation. Saha, however, teaches this limitation. Saha teaches a federated-learning architecture in which a plurality of fog nodes perform local model aggregation and a separate cloud server selects one of the fog nodes to perform global model aggregation. Specifically, after the local aggregation operations: “the cloud server selects a fog node as a global aggregator” (Saha, pg. 8459, § III-A) Saha expressly states: “In each round, the cloud server selects a global aggregator” (Saha, pg. 8459, § III-B) Algorithm 1 likewise recites “Select global aggregator node G”, after which the global parameter is calculated (Saha, pg. 8459, Algorithm 1). The selected global aggregator then calculates the final model among the fog nodes. Thus, Saha teaches that the device performing the global model-aggregation function is specified by a separate cloud sever. In the proposed combination, the selected fog node corresponds to the claimed communication apparatus preforming the model-fusion function, while Saha’s cloud server corresponds to a device other than the K groups of devices that specifies which device is to perform that function. Accordingly, Saha teaches “the communication apparatus is a device specified by a device other than the K groups of devices”. Because claim 15 recites the three conditions in the alternative, Saha’s teaching of the third alternative is sufficient to satisfy the additional limitation of claim 15. It would have been obvious to one of ordinary skill in the art before the effective filing data to incorporate Saha’s global-aggregator selection technique into the distributed-training apparatus of Verma, Lin, and Sridharan. As discussed with respect to claim 14, Verma in view of Lin and Sridharan teaches a distributed-training architecture in which a communication apparatus receives trained model information from multiple groups and performs inter-group fusion. Saha similarly addresses distributed federated learning and teaches dynamically selecting a fog node to assume the global model-aggregation role rather than relying on a fixed global aggregator. Saha expressly identifies a variable global aggregator node as a mechanism for increasing system reliability and reducing dependence on a centralized global aggregator. One of ordinary skill therefore would have had reason to use Saha’s known selection technique to designate which participating device performs the fusion/aggregation function in the Verma-Lin-Sridharan architecture, thereby providing dynamic assignment of the aggregation role and reducing reliance on a fixed aggregation device. The predictable result is a communication apparatus performing the claimed fusion operations that is specified by a separate device outside the K groups, as claimed. Claim 17 is rejected under 35 U.S.C. 103 as being unpatentable over Verma in view of Lin further in view of Sridharan and further in view of Roy. Regarding claim 17, Verma in view of Lin further in view of Sridharan and further in view of Roy, teach the communication apparatus according to claim 14, wherein the communication apparatus is further caused to: “and perform inter-group fusion on a q t h model in each of the K groups of models i t h k t h q ∈ [ 1 , M ] .” – Verma does not teach this limitation. Sridharan, however, teaches or renders obvious this limitation in part. For substantially the same reasons stated with respect to claim 16, Sridharan’s combined model/data-parallel architecture provides corresponding model portions across replicated groups, and its parameter averaging/inter-group synchronization teaches or renders obvious fusion of the corresponding q t h model information across those groups. Sridharan teaches combined model and data parallelism in which: “each computational node includes multiple GPUs” (Sridharan, pg. 19, ¶ [0191]) And: “Each node can have a complete instance of the model with separate GPUs within each node are used to train different portions of the model.” (Sridharan, pg. 19, ¶ [0191]) Thus, each replicated computational-node group contains corresponding model portions at respective device/model positions. For a particular position q , where q is one of the M model/device positions, a corresponding q t h model portion exists in each of the K replicated groups. Sridharan further teaches parameter averaging in which: “the global parameters (e.g., weights, biases)” (Sridharan, pg. 19, ¶ [0190]) are set to the average of the parameters from the respective nodes, using: “a central parameter server that maintains the parameter data.” (Sridharan, pg. 19, ¶ [0190]) Applying Sridharan’s disclosed parameter-averaging algorithm to the combined model/data-parallel arrangement predictably averages the corresponding parameters of the same model portion across the replicated groups. Thus, for a selected model position q , the corresponding q t h model in each of the K groups is fused according to the parameter-averaging algorithm. Because the combined architecture contains M respective model/device positions, indexing any one of those positions as q results in q ∈ [ 1 , M ] . The variable q is applicant’s nomenclature for selecting one of the corresponding model positions; Sridharan need not use that particular symbol. Sridharan expressly teaches parameter-server synchronization between groups, including parameter servers configured to enable: “inter-group communication between the various groups 2210, 2230” (Sridharan, pg. 23, ¶ [0223]; FIG. 22) Accordingly, Sridharan teaches or renders obvious performing inter-group fusion on a q t h model in each of the K groups of models according to a fusion algorithm, wherein q ∈ [ 1 , M ] . Neither Verma nor Sridharan teach these limitations and/or portions of: “receive a sample quantity sent by the i t h device in the k t h group of devices, wherein the sample quantity comprises a sample quantity currently stored in the i t h device and a sample quantity obtained by exchanging with at least one other device in the k t h group of devices;” “based on the sample quantity sent by the i t h device in the k t h group of devices and according to a fusion algorithm” Roy, however, teaches these limitations and/or portions of: “receive a sample quantity sent by the i t h device in the k t h group of devices, wherein the sample quantity comprises a sample quantity currently stored in the i t h device and a sample quantity obtained by exchanging with at least one other device in the k t h group of devices;” – Roy teaches a peer-to-peer federated-learning system in which each client C i has a local dataset containing a i training samples. During an update, other clients having newer models send to the initiating client both: “their weights W j and the training sample size a j to C i .” (Roy, pg. 4, § 2.2) And the initiating client combines those received models with its own current model: “This subset of models is merged with C i ’s current model to a single model by weighted averaging.” (Roy, pg. 4, § 2.2) Thus, the initiating device possesses its own locally stored sample quantity a i and obtains additional sample quantities a j through exchange with other devices. Roy’s Algorithm 1 likewise initializes the aggregation using the initiating client’s own model and sample quantity and then, for each updated peer, receives W j and a j and incorporates the received contribution into the aggregate. Accordingly, Roy teaches sample-quantity information comprising both a locally stored sample quantity and sample-quantity information obtained from at least one other device. It would have been obvious to communicate that sample-quantity information together with the corresponding model information to the fusion apparatus so that the apparatus can perform the weighted aggregation taught by Roy. “based on the sample quantity sent by the i t h device in the k t h group of devices and according to a fusion algorithm” – Roy teaches that a contributing client sends its model weights W j together with its training sample size a j , and that the models are combined by weighted averaging based on those training sample sizes. In Roy’s server-based federated-learning embodiment, the server aggregates client models according to: “ W S = ∑ i a i a W i ” (Roy, pg. 3, § 2.1) where the weighting factor is based on the fraction of the total training data belonging to the respective client. Roy explains that its weighting emphasizes clients having more training data: “The rationale is to emphasize clients with more training data.” (Roy, pg. 3, § 2.1) Applying Roy’s weighting technique to Sridharan’s corresponding q t h models causes the inter-group fusion already supplied by Sridharan to be performed based on the sample quantity associated with the respective contributing model, as claimed. Claims 18-20 are rejected under 35 U.S.C. 103 as being unpatentable over Verma in view of Lin, further in view of Sridharan, and further in view of Takayuki Nishio (Client Selection for Federated Learning with Heterogeneous Resources in Mobile Edge). Regarding claim 18, Verma in view of Lin further in view of Sridharan and further in view of Nishio, teach the communication apparatus according to claim 14, wherein the communication apparatus is further caused to: “receive status information reported by N devices, ” – Verma teaches this limitation in part. Verma teaches that: “Each agent computes statistics of data distribution stored at its respective system node” (Verma, pg. 3, ¶ [0034]) “The agents share these statistics with the fusion server.” (Verma, pg. 3, ¶ [0034]) Thus, Verma teaches a coordinating/fusion apparatus receiving information concerning the respective participating training devices. Verma does not, however, teach the claimed selection arrangement in which status information is obtained from a larger population of N candidate devices from which the M devices included in each of the K groups are selected. Verma does not teach these limitations and/or portions of: “wherein the N devices comprise the M devices included in each group of devices in the K group of devices” “select, based on the status information, the M devices included in each group of devices in the K groups of devices form the N devices;” “and broadcast a selection result to the M devices included in each group of devices in the K groups of devices.” Sridharan, however, teaches these limitations and/or portions of: “wherein the N devices comprise the M devices included in each group of devices in the K group of devices” - As established with claim 14, Sridharan teaches a distributed-training architecture having multiple groups of worker devices: “local nodes can be assembled into compute groups based on network topology.” (Sridharan, pg. 23, ¶ [0222]) Sridharan further teaches that the communications framework can: “flexibly and transparently adjust node grouping” (Sridharan, pg. 23, ¶ [0222]) and, in the embodiment of FIG. 22: “the communications framework can establish multiple groups of worker nodes.” (Sridharan, pg. 23, ¶ [0223]) Sridharan specifically creates a first group 2210 and a second group 2230 and explains that: “The specific number of groups that are created can vary based on the topology of the network.” (Sridharan, pg. 23, ¶ [0223]) Thus, as established with respect to claim 14, Sridharan provides the claimed K-group architecture containing respective M device/model positions within each group. Designating the larger population of available worker devices as the claimed N devices, that population comprises the devices assigned to the respective groups. and broadcast ” – Sridharan teaches this limitation in part. Sridharan expressly teaches group-wide broadcast communication in its distributed machine-learning architecture. Table 1 defines ALLREDUCE as: “A reduce combined with a broadcast, where the outcome of the reduce operation is broadcast to all processes within a group.” (Sridharan, Table 1) Sridharan further teaches that: “The Allreduce operations are reduce operations for which the results are broadcast or transferred to the receive buffers of all processes in the communication group.” (Sridharan, pg. 21, ¶ [0206]) Thus, Sridharan expressly teaches broadcasting common information to all participating processes/devices within a communication group. Applied to the K-group architecture established with respect to claim 14, Sridharan teaches broadcasting information to the M devices included within the respective groups. Sridharan does not, however, expressly teach that the information being broadcast to those group members is “a selection result” resulting from the claimed selection of the M devices based on the status information reported by the N-device candidate population. Verma nor Sridharan teach these limitations and/or portions of “select, based on the status information, the M devices included in each group of devices in the K groups of devices form the N devices;” “a selection result” Nishio, however, teaches these limitations and/or portions of: select, based on the status information, the M devices included in each group of devices in the K groups of devices form the N devices;” – Nishio teaches a Resource Request step in which candidate population of clients reports resource/status information to a coordinating MEC operator. Nishio teaches that the requested information includes: “wireless channel states, computational capacities … and the size of data resources relevant to the current training task” (Nishio, pg. 3, § III-B) Nishio’s Protocol 2 expressly provides: “Clients who receive the request notify the operator of their resource information.” (Nishio, pg. 3, § III-B, Protocol 2) Nishio then expressly uses that reported information to select a subset of those reported clients: “Client Selection: Using the information, the MEC operator determines which of the clients go to the subsequent steps to complete the steps within a certain deadline.” (Nishio, § III-B, Protocol 2) Nishio further explains that the operator refers to this information in the subsequent Client Selection step and uses the reported resource information to determine which clients will participate in the subsequent distributed-training operations. Nishio formally defines K′ as the subset of candidate clients selected in the Resource Request step and S as: “a sequence of indices of the clients selected in Client Selection” where the selected clients k i are members of the candidate set K ' . Nishio further states that the client-wise parameters used in the selection procedure: “can be determined based on the resource information notified in the Resource Request step.” (Nishio, pg. 4, § III-C; Algorithm 3) Thus, Nishio expressly teaches receiving status information from a candidate population and selecting participating devices from that same candidate population based on the reported status information. Applying Nishio’s selection procedure to Sridharan’s K-group architecture predictably selects the M devices included in each of the K groups from the N reporting candidate devices. “a selection result” – Nishio teaches or renders obvious this remaining portion. Nishio teaches that immediately following its Client Selection operation, information resulting from that selection is communicated specifically to the selected devices. Protocol 2 provides: “Client Selection: Using the information, the MEC operator determines which of the clients go to the subsequent steps…” (Nishio, pg. 3, § III-B, Protocol 2) followed by: “Distribution: The server distributes the parameters of the global model to the selected clients.” (Nishio, pg. 3, § III-B, Protocol 2) Nishio further teaches that: “a global model is distributed to the selected clients via multicast from the BS because it is bandwidth effective for transmitting the same content (i.e., the global model) to client populations.” (Nishio, pg. 3, § III-B) Figure 2 further depicts, immediately following “3. Client Selection,” the communication: “Global model & schedule” (Nishio, pg. 4, § III-C) to the selected clients. The schedule is determined for those clients selected by the Client Selection procedure, and Nishio explains that the MEC operator: “schedules when the RBs for model uploads are allocated to the selected clients.” (Nishio, pg. 4, § III-C) Thus, Nishio teaches a post-selection communication directed specifically to the devices selected by the status-based selection operation and containing the model and schedule by which those selected devices participate in the training round. Under the broadest reasonable interpretation, this selection-conditioned model/schedule communication corresponds to a selection result, because it communicates and implements the outcome of the immediately preceding client-selection determination. Nishio further expressly recognizes both: “broadcast and multicast transmission by the BS” (Nishio, pg. 3, § III-A) as known one-to-many communication techniques in its disclosed MEC system. Accordingly, Sridharan’s group-wide broadcast teaching, in view of Nishio’s post-selection model/schedule communication to the selected client population and Nishio’s express recognition of broadcast and multicast transmission, teaches or renders obvious “broadcast a selection result to the M devices included in each group of devices in the K groups of devices.” It would have been obvious to one of ordinary skill in the art before the effective filing date to incorporate Nishio’s resource-aware client-selection technique into the grouped distributed-training architecture of Verma, Lin, and Sridharan. As established above, Verma teaches distributed training in which participating devices report information concerning their local training resources to a coordinating fusion server. Sridharan teaches grouping distributed-training devices based on communication conditions, including communication latency, and further teaches devices reporting communication metrics including throughput and latency. Nishio expressly addresses inefficiencies caused by heterogeneous device resources and wireless-channel conditions by obtaining resource information from a candidate population and selecting participating clients based on that reported information. One of ordinary skill therefore would have had reason to apply Nishio’s resource-aware client-selection procedure when determining which available devices participate in Sridharan’s distributed-training groups so that the selected devices are capable of completing the required training and communication operations within the applicable deadline. The modification would predictably result in status information being received from N candidate devices and the M devices included in each of the K groups being selected from those N devices based on that status information. Nishio further teaches communicating a common global model and schedule to the selected clients immediately following the selection operation and expressly contemplates broadcast and multicast transmission by the base station. Sridharan independently teaches broadcasting common information to all processes within a communication group. One of ordinary skill therefore would have had reason to use the known group-wide broadcast mechanism of Sridharan to communicate Nishio’s post-selection model/schedule information to the selected members of each group, thereby informing and configuring those devices for the training operation in which they were selected to participate. The combination would predictably result in receiving status information reported by N devices, selecting based on that status information the M devices included in each of K groups from the N devices, and broadcasting the resulting selection-conditioned information to the selected M devices in each group, as claimed. Regarding claim 19, Verma in view of Lin further in view of Sridharan and further in view of Nishio, teach the communication apparatus according to claim 18, wherein “the selection result comprises at least one of ” – Verma does not teach this limitation. Nishio, however, teaches this limitation. As discussed above with respect to claim 18, Nishio teaches a Client Selection operation in which the MEC operator determines, based on reported resource information, which candidate clients are selected to participate in the subsequent distributed-training operations: “Client Selection: Using the information, the MEC operator determines which of the clients go to the subsequent steps to complete the steps within a certain deadline.”(Nishio, pg. 3, § III-B, Protocol 2) Immediately following the Client Selection operation, Nishio teaches: “Distribution: The server distributes the parameters of the global model to the selected clients.” (Nishio, pg. 3, § III-B, Protocol 2) Nishio further explains: “In the Distribution step, a global model is distributed to the selected clients via multicast from the BS because it is bandwidth effective for transmitting the same content (i.e., the global model) to client populations.” (Nishio, pg. 3, § III-B) Thus, as established with respect to claim 18, the devices determined by the Client Selection operation are the recipients of the post-selection communication, and that communication conveys the information associated with their selected participation in the distributed-training operation. Nishio expressly teaches that the content of the communication includes the parameters of the global model, thereby teaching that the selection result comprises “information about a model”. Nishio's FIG. 2 further confirms this relationship. Following “3. Client Selection,” the communication transmitted to the selected clients in the Distribution step is expressly labeled: “Global model & schedule” (Nishio, pg. 4, FIG. 2). Accordingly, Nishio teaches a selection result communicated to the selected devices that comprises information about the model, including the global model and its parameters. Because claim 19 requires only that the selection result comprise “at least one of” a selected device, a grouping status, or information about a model, Nishio's disclosure of the information-about-a-model alternative is sufficient to satisfy the additional limitation of claim 19. Regarding claim 20, Verma in view of Lin further in view of Sridharan in further view of Nishio, teach the communication apparatus according to claim 19, wherein “the information about the model comprises a model structure of the model, a model parameter of the model, ”– Verma teaches this limitation in part. Verma teaches communicating information concerning both the structure and parameters of the model to the participating training devices. In particular, Verma teaches that: “The fusion server can transmit instructions to each agent to construct mini-batches according to an overall frequency distribution and the structure of the artificial intelligence model each system node should train” (Verma, pg. 5, ¶ [0048]) Thus, Verma expressly teaches transmitting information specifying the structure of the model that the respective system node is to train. Verma further teaches collecting the locally trained models and performing model fusion: “The fusion server can collect intelligence models trained on incremental mini-batches from each system node and average their weights to create a fused artificial intelligence model.” (Verma, pg. 5, ¶ [0048]) If additional training is required, Verma teaches that: “the fusion server can direct the fused intelligence model weights to each system node to create next mini-batch, update the model weights” (Verma, pg. 5, ¶ [0048]) Thus, Verma expressly teaches communicating the weights of the model to the participating system nodes, corresponding to “a model parameter of the model.” FIG. 4 likewise depicts the same relationship. Step 404 provides instructions according to the “structure of neural network each agent should train,” step 405 averages the model weights to create the fused model, and step 408 sends the “fused AI model weights to each system node” for the next training operation. Accordingly, Verma teaches that the information concerning the model provided to the participating training devices comprises: “a model structure of the model”; and “a model parameter of the model.” Verma does not, however, expressly teach that the information about the model further comprises a specified fusion round period and a total quantity of fusion rounds. Verma does not teach these limitations and/or portions of: “a fusion round period” “a total quantity of fusion rounds.” Lin, however, teaches these remaining limitations and/or portions of: “a fusion round period”– Lin teaches this portion. Lin teaches Local SGD in which a specified number of local model-training operations are performed between successive synchronization/fusion operations. Lin explains: “In local SGD, each worker k … evolves a local model by performing H sequential SGD updates … before communication (synchronization by averaging) among the workers” (Lin, pg. 1, Introduction – Local SGD) Thus, Lin expressly teaches that H determines the interval between successive model-fusion/synchronization operations. Lin's Algorithm 1 makes the relationship explicit by defining: “input: number of synchronization steps T;” (Lin, pg. 18, Algorithm 1) “input: number of local steps H;” (Lin, pg. 18, Algorithm 1) and, for each synchronization step, performing: “for h := 1, …, H do” (Lin, pg. 18, Algorithm 1) before performing: “all-reduce aggregation of the gradients” (Lin, pg. 18, Algorithm 1) and obtaining the: “new global (synchronized) model … for all K nodes.” (Lin, pg. 18, Algorithm 1) Accordingly, H specifies how many local model-training operations occur between successive fusion operations and therefore corresponds to the claimed “fusion round period.” Lin further teaches that the synchronization interval may be a training-dependent variable rather than merely a fixed constant. In the post-local-SGD embodiment, Lin changes from synchronized mini-batch SGD during the first training phase to Local SGD during the second phase, thereby using the Local-SGD synchronization interval during the later portion of training. Lin explains that the second phase switches to the communication-efficient effective batch H B l o c and expressly states that Local SGD is used in that second phase. Thus, Lin teaches a parameter controlling the period at which model fusion/synchronization is performed during model training. “a total quantity of fusion rounds” – Lin teaches this portion. Lin's Algorithm 1 expressly defines: “input: number of synchronization steps T;” (Lin, pg. 18, Algorithm 1) and thereafter executes: “for t := 1, …, T do” (Lin, pg. 18, Algorithm 1) with each iteration performing, after the H local model updates: “all-reduce aggregation of the gradients” (Lin, pg. 18, Algorithm 1) followed by generation of the: “new global (synchronized) model … for all K nodes.” (Lin, pg. 18, Algorithm 1) Thus, each of the T synchronization steps includes a model-fusion operation through the all-reduce aggregation. Lin's specified “number of synchronization steps T” therefore corresponds to the claimed “total quantity of fusion rounds.” Accordingly, Verma teaches model information including the model structure and model parameters, while Lin further teaches that the distributed model-training procedure is configured according to a fusion round period H and a total quantity T of fusion rounds. As discussed with respect to claim 19, Nishio teaches communicating the global model and associated schedule to the devices selected for participation in the distributed-training round. Incorporating Lin’s expressly disclosed synchronization parameters H and T into the model/training information supplied to those selected devices would predictably provide the participating devices with the information necessary to perform the distributed model-training and fusion procedure at the prescribed synchronization interval and for the prescribed number of synchronization rounds. The resulting combination therefore teaches or renders obvious that: “the information about the model comprises a model structure of the model, a model parameter of the model, a fusion round period, and a total quantity of fusion rounds”. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to Paul Coleman whose telephone number is (571)272-4687. The examiner can normally be reached Mon-Fri. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /PAUL COLEMAN/ Examiner, Art Unit 2126 /DAVID YI/ Supervisory Patent Examiner, Art Unit 2126
Read full office action

Prosecution Timeline

May 14, 2024
Application Filed
Jun 14, 2024
Response after Non-Final Action
Sep 11, 2026
Non-Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12748961
NEUROMORPHIC COMPUTING DEVICE WITH THREE-DIMENSIONAL MEMORY
4y 4m to grant Granted Sep 29, 2026
Patent 12731010
MULTIVARIABLE TIME-SERIES FEATURE EXTRACTION
3y 7m to grant Granted Sep 08, 2026
Patent 12711416
MODEL MODIFICATION AND DEPLOYMENT
5y 3m to grant Granted Aug 18, 2026
Patent 12688400
COMPUTATIONAL NEURAL NETWORK APPARATUS, CARD, METHOD, AND READABLE STORAGE MEDIUM
3y 6m to grant Granted Jul 21, 2026
Patent 12665745
MACHINE LEARNING/ARTIFICIAL INTELLIGENCE (ML/AI) SYSTEM WITH PROTECTED NEURAL NETWORKS
3y 5m to grant Granted Jun 23, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
65%
Grant Probability
99%
With Interview (+47.4%)
3y 8m (~1y 4m remaining)
Median Time to Grant
Low
PTA Risk
Based on 26 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month