Prosecution Insights
Last updated: August 18, 2026
Application No. 18/184,742

COLLABORATIVE INFERENCE METHOD AND COMMUNICATION APPARATUS

Non-Final OA §103
Filed
Mar 16, 2023
Priority
Sep 21, 2020 — CN 202010998618.7 +1 more
Examiner
STANLEY, JEREMY L
Art Unit
2127
Tech Center
2100 — Computer Architecture & Software
Assignee
Huawei Technologies Co., Ltd.
OA Round
2 (Non-Final)
49%
Grant Probability
Moderate
2-3
OA Rounds
0m
Est. Remaining
90%
With Interview

Examiner Intelligence

Grants 49% of resolved cases
49%
Career Allowance Rate
141 granted / 288 resolved
-6.0% vs TC avg
Strong +41% interview lift
Without
With
+40.9%
Interview Lift
resolved cases with interview
Typical timeline
3y 3m
Avg Prosecution
25 currently pending
Career history
313
Total Applications
across all art units

Statute-Specific Performance

§101
10.4%
-29.6% vs TC avg
§103
54.2%
+14.2% vs TC avg
§102
14.0%
-26.0% vs TC avg
§112
16.7%
-23.3% vs TC avg
Black line = Tech Center average estimate • Based on career data from 288 resolved cases

Office Action

§103
DETAILED ACTION Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This action is responsive to the Amendment filed on April 9, 2026. Claims 1, 2, 4, 6-9, 12, and 16 are amended. Claims 1-20 are pending in the case. Claims 1, 9, and 16 are the independent claims. This action is final. Applicant’s Response In the Amendment filed on April 9, 2026, Applicant amended the claims and provided arguments in response to the rejections of the claims under 35 USC 101, 102, and 103 in the previous office action. Response to Argument/Amendment Applicant’s amendments to the claims in response to the rejection of the claims under 35 USC 101 in the previous office action, and Applicant’s associated arguments have been fully considered. Applicant argues that the independent claims are now amended to recite “comprising a plurality of hidden layers and segmentation is performed between an input layer and a first hidden layer of the plurality of hidden layers, the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers, the second hidden layer of the plurality of hidden layers and a third hidden layer of the plurality of hidden layers, and the third hidden layer of the plurality of hidden layers and an output layer; send the first inference result, wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers,” and that this subject matter cannot be mentally performed as it is directed to hidden layers of a machine learning model. Applicant further notes that the limitation “send the first inference result, wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers) provides improvements to the functioning of a computer or another technology or technical field, because it reduces a risk of data privacy exposure and improves data security of the terminal device, and also argues that increasing data of hidden layers of the model can improve precision or capacity of the model. Applicant’s arguments are persuasive, and the rejection is withdrawn. Applicant’s amendments to the claims in response to the rejections of the claims under 35 USC 102 and 103 in the previous office action, and Applicant’s associated arguments have been fully considered. Applicant argues that the independent claims are now amended to recite “comprising a plurality of hidden layers and segmentation is performed between an input layer and a first hidden layer of the plurality of hidden layers, the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers, the second hidden layer of the plurality of hidden layers and a third hidden layer of the plurality of hidden layers, and the third hidden layer of the plurality of hidden layers and an output layer; send the first inference result, wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers,” and that Bloom and the other cited references fail to teach these limitations. However, Bloom does appear to teach an ML model comprising a plurality of hidden layers (e.g. Figs. 5 and 6, showing a plurality of at least three hidden layers in between input and output layers of the neural network; paragraph 0074, series of intervening layers of model, i.e. between initial/input layers and end/output layers, where these intermediate/intervening layers between input and output layers are hidden layers) and segmentation is performed between an input layer and a first hidden layer of the plurality of hidden layers (e.g. paragraph 0074, ML model may be split into at least three ML model components; ML model split between two adjacent layers of the neural network; ML model split such that the first ML model component comprises a series of initial layers at one end of the neural network and the second ML model component comprises a series of intervening layers of the neural network, the series of initial layers may receive and process input data to produce first intermediate neural network output data; Figs. 5 and 6, each showing the neural network having a plurality, at least three, intermediate layers (analogous to hidden layers) between the input and output layers; i.e. where the model includes a first series of initial/input layers, and is split between these and a series of intervening/hidden layers, this is analogous to a segmentation performed between an input layer (i.e. one of the series of initial layers responsible for processing inputs) and a first hidden layer of the plurality of hidden layers (i.e. a first layer of the intervening/hidden layers which is adjacent to the corresponding layer of the initial/input layers); Examiner notes that the claims do not appear to preclude the possibility of the neural network having multiple input layers/layers responsible for processing input data), the second hidden layer of the plurality of hidden layers and a third hidden layer of the plurality of hidden layers (e.g. as shown in Figs. 5 and 6, the neural network may also be split between layer 2 and layer 3; paragraph 0074, model split into at least three ML model components; ML model split between two adjacent layers; i.e. the model could be split into more than three components, such as four components, such that an additional split is introduced between intermediate/hidden layer components as shown in Figs. 5 and 6), and the third hidden layer of the plurality of hidden layers and an output layer (e.g. paragraph 0074, ML model may be split into at least three ML model components; ML model split between two adjacent layers of the neural network; ML model split such that the second ML model component comprises a series of intervening layers of the neural network, and the third ML model component comprises a series of end layers of the neural network; the series of intervening layers receive and process first intermediate neural network output data to produce second intermediate neural network output data; the series of end layers of the neural network may receive and process the second intermediate neural network output data to produce result data; Figs. 5 and 6, each showing the neural network having a plurality, at least three, intermediate layers (analogous to hidden layers) between the input and output layers; i.e. where the model includes a series of end/output layers, and is split between these and a series of intervening/hidden layers, this is analogous to a segmentation performed between an output/end layer (i.e. one of the series of end layers responsible for producing result data from intermediate data) and a third hidden layer of the plurality of hidden layers (i.e. a third layer of the intervening/hidden layers which is adjacent to the corresponding layer of the end/output layers); Examiner notes that the claims do not appear to preclude the possibility of the neural network having multiple output layers/layers responsible for producing output/result data); That is, Bloom appears to teach that the model may include at least an input layer, first, second, and third intervening (i.e. hidden) layers, and an output layer in the neural network, and further teaches that the model may be split in a variety of ways, including at least splitting the layers into component comprising input layers, intervening/hidden layers, and end/output layers (i.e. such that split/segmentation is provided between at least an input layer and a first layer of the hidden layers and between at least a third/final of the hidden layers and one of the output layers), and further separately teaches that splitting/segmentation can be performed between the intervening/hidden layers (i.e. such that splitting/segmentation is provided between at least a second hidden layer and a third hidden layer). However Examiner agrees that Bloom does not explicitly disclose segmentation is performed between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers; and wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers. Therefore, Applicant’s arguments are persuasive in part, and the rejections are withdrawn. New grounds of rejection are provided below. Claim Rejections – 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. The factual inquiries set forth in Graham v. John Deere Co., 383 U.S. 1, 148 USPQ 459 (1966), that are applied for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows: 1. Determining the scope and contents of the prior art. 2. Ascertaining the differences between the prior art and the claims at issue. 3. Resolving the level of ordinary skill in the pertinent art. 4. Considering objective evidence present in the application indicating obviousness or nonobviousness. This application currently names joint inventors. In considering patentability of the claims under pre-AIA 35 U.S.C. 103(a), the examiner presumes that the subject matter of the various claims was commonly owned at the time any inventions covered therein were made absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and invention dates of each claim that was not commonly owned at the time a later invention was made in order for the examiner to consider the applicability of pre-AIA 35 U.S.C. 103(c) and potential pre-AIA 35 U.S.C. 102€, (f) or (g) prior art under pre-AIA 35 U.S.C. 103(a). Claims 1-3, 7-10, and 13-18 are rejected under 35 U.S.C. 103 as being unpatentable over Bloom (US 20180336463 A1) in view of D’Ercoli et al. (US 20200210834 A1). With respect to claim 1, Bloom teaches a collaborative inference apparatus (e.g. paragraph 0086, Fig. 14, machine 1400 able to perform discussed methodologies), comprising: a transceiver (e.g. paragraph 0090 Fig. 14, machine 1400 including communication components 1440 operable to couple machine 1400 to communications network 1432 and devices 1424, and may include wired communication components, wireless communication components, cellular communication components, NFC components, Bluetooth components, Wi-Fi components, etc.; i.e. a device capable of both transmitting and receiving data, including wirelessly, such as via radio wave/signal); at least one processor; and one or more memories coupled to the at least one processor and storing programming instructions, when executed by the at least one processor (e.g. paragraph 0087, Fig. 14, machine 1400 including processors 1404, memory/storage 1406, etc.; storage unit 1416 and memory 1414 store instructions 1410 embodying described methodologies/functions, and processors execute the instructions), cause the apparatus to: determine a first inference result based on a first machine learning (ML) submodel, wherein the first ML submodel is a part of an ML model (e.g. paragraph 0022, performing remote inference using ML model which is split into components and distributed to multiple computing devices; paragraph 0023, trained ML model split into phase I and phase II components; sending computing device possessing the phase I component; sending computing device encoding data using phase I component; paragraph 0065, Fig. 8, step 804, split ML model into at least first and second ML model components; ML model comprising neural network, where first model component comprises first portion and second ML model component comprises second portion of the neural network; step 806, providing first ML model component to remote computing device; paragraph 0066, step 808 of Fig. 8, processing input data using second ML model component to generate intermediate neural network output data); comprising a plurality of hidden layers (e.g. Figs. 5 and 6, showing a plurality of at least three hidden layers in between input and output layers of the neural network; paragraph 0074, series of intervening layers of model, i.e. between initial/input layers and end/output layers, where these intermediate/intervening layers between input and output layers are hidden layers) and segmentation is performed between an input layer and a first hidden layer of the plurality of hidden layers (e.g. paragraph 0074, ML model may be split into at least three ML model components; ML model split between two adjacent layers of the neural network; ML model split such that the first ML model component comprises a series of initial layers at one end of the neural network and the second ML model component comprises a series of intervening layers of the neural network, the series of initial layers may receive and process input data to produce first intermediate neural network output data; Figs. 5 and 6, each showing the neural network having a plurality, at least three, intermediate layers (analogous to hidden layers) between the input and output layers; i.e. where the model includes a first series of initial/input layers, and is split between these and a series of intervening/hidden layers, this is analogous to a segmentation performed between an input layer (i.e. one of the series of initial layers responsible for processing inputs) and a first hidden layer of the plurality of hidden layers (i.e. a first layer of the intervening/hidden layers which is adjacent to the corresponding layer of the initial/input layers); Examiner notes that the claims do not appear to preclude the possibility of the neural network having multiple input layers/layers responsible for processing input data), the second hidden layer of the plurality of hidden layers and a third hidden layer of the plurality of hidden layers (e.g. as shown in Figs. 5 and 6, the neural network may also be split between layer 2 and layer 3; paragraph 0074, model split into at least three ML model components; ML model split between two adjacent layers; i.e. the model could be split into more than three components, such as four components, such that an additional split is introduced between intermediate/hidden layer components as shown in Figs. 5 and 6), and the third hidden layer of the plurality of hidden layers and an output layer (e.g. paragraph 0074, ML model may be split into at least three ML model components; ML model split between two adjacent layers of the neural network; ML model split such that the second ML model component comprises a series of intervening layers of the neural network, and the third ML model component comprises a series of end layers of the neural network; the series of intervening layers receive and process first intermediate neural network output data to produce second intermediate neural network output data; the series of end layers of the neural network may receive and process the second intermediate neural network output data to produce result data; Figs. 5 and 6, each showing the neural network having a plurality, at least three, intermediate layers (analogous to hidden layers) between the input and output layers; i.e. where the model includes a series of end/output layers, and is split between these and a series of intervening/hidden layers, this is analogous to a segmentation performed between an output/end layer (i.e. one of the series of end layers responsible for producing result data from intermediate data) and a third hidden layer of the plurality of hidden layers (i.e. a third layer of the intervening/hidden layers which is adjacent to the corresponding layer of the end/output layers); Examiner notes that the claims do not appear to preclude the possibility of the neural network having multiple output layers/layers responsible for producing output/result data); send the first inference result (e.g. paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided); and receive a target inference result, wherein the target inference result is an inference result of the ML model determined based on the first inference result (e.g. paragraph 0023, remote computing device further encoding received encoded data using phase II component to produce result data which can represent an inference made by the ML model; paragraph 0068, Fig. 8 step 812, receiving prediction data from the remote computing device, where the prediction data is based on result data generated at the remote computing device by the first ML model component processing the intermediate neural network output data; the result data comprises inference data produced by the first ML model component). Bloom does not explicitly disclose: segmentation is performed between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers; and wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers. However, D’Ercoli teaches: segmentation is performed between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers (e.g. paragraph 0027-0029, Fig. 2, DNN 200 includes input layers 204, multiple hidden layers 206, and an output layer 208; DNN 200 partitioned to form partitioned DNN 212 having multiple layers L1-L4; after performing the partitioning each of the layers L1-L4 includes one or more of the original DNN’s layers 204, 206, and 208; deploying partitioned DNN to CPs/computing devices; i.e. as described, the neural network may have input, output, and a plurality of hidden layers, where each of these may be individually partitioned as a separate layer of partitioned neural network (i.e. as noted, each layer of the partitioned DNN may include one or more, and therefore only one, of the layers of the original DNN, such that partitioning may be performed between a first hidden layer of the plurality of hidden layers 206 and a second hidden layer of the plurality of hidden layers 206 to form separate/individual partitioned layers, such as L2 and L3 of the partitioned DNN)), wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers (e.g. paragraphs 0054-0055, Fig. 7, DNN partitioned into layers L1-L4 on separate computation points CP1-CP6; L1 on CP1 performing processing of new input and transmitting available feature map to other available CP such as CP2; CP2 generating output O1; highlighted layers of computational points CP2-6 are layers available to process feature maps from CP1 or from CP2, etc.; i.e. where the neural network has been partitioned into individual layers (i.e. each layer of the partitioned NN having one of the layers of the original network) and individual layers are distributed to separate computing devices, a first inference result may take the form of an output/inference result of a first hidden layer of the plurality of hidden layers; e.g., where a first hidden layer of the original DNN is partitioned into L2 of the partitioned NN, the output O2 generated by CP2 using layer L2 as shown in Fig. 7 may be a first inference result of a first hidden layer of a plurality of hidden layers). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom and D’Ercoli in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing), to incorporate the teachings of D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points) to include the capability to perform the segmentation/partitioning of the neural network including the input, plurality of hidden, and output layers on a layer-by-layer basis such that each layer of the partitioned network corresponds to one layer of the original network (and the segmentation therefore includes segmentation between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers), and to individually process outputs of each of the partitioned layers (including hidden layers, such as by passing the intermediate output from a first hidden layer to a second hidden layer), such that the first inference result is an inference result of the first hidden layer of the plurality of hidden layers (as taught by D’Ercoli). One of ordinary skill would have been motivated to perform such a modification in order to overcome problems of latency and energy consumption in performing deep learning algorithms as described in D’Ercoli (paragraph 0013). With respect to claim 9, Bloom teaches a collaborative inference apparatus (e.g. paragraph 0086, Fig. 14, machine 1400 able to perform discussed methodologies), comprising: a transceiver (e.g. paragraph 0090 Fig. 14, machine 1400 including communication components 1440 operable to couple machine 1400 to communications network 1432 and devices 1424, and may include wired communication components, wireless communication components, cellular communication components, NFC components, Bluetooth components, Wi-Fi components, etc.; i.e. a device capable of both transmitting and receiving data, including wirelessly, such as via radio wave/signal); at least one processor; and one or more memories coupled to the at least one processor and storing programming instructions, when executed by the at least one processor (e.g. paragraph 0087, Fig. 14, machine 1400 including processors 1404, memory/storage 1406, etc.; storage unit 1416 and memory 1414 store instructions 1410 embodying described methodologies/functions, and processors execute the instructions), cause the apparatus to: receive first inference information from a terminal device, wherein the first inference information comprises all information or partial information of a first inference result, the first inference result is an inference result of a first machine learning (ML) submodel, and the first ML submodel is a part of an ML model (e.g. paragraph 0022, performing remote inference using ML model which is split into components and distributed to multiple computing devices; paragraph 0023, trained ML model split into phase I and phase II components; sending computing device possessing the phase I component; sending computing device encoding data using phase I component; paragraph 0070, Fig. 9, step 908 receiving intermediate neural network output data from remote computing device which is generated by processing input data at the remote computing device using first ML model component; paragraph 0076, Fig. 12, describing system in which model has been split into at least three components; intermediate neural network output data generated 1208 and provided to remote computing device); comprising a plurality of hidden layers (e.g. Figs. 5 and 6, showing a plurality of at least three hidden layers in between input and output layers of the neural network; paragraph 0074, series of intervening layers of model, i.e. between initial/input layers and end/output layers, where these intermediate/intervening layers between input and output layers are hidden layers) and segmentation is performed between an input layer and a first hidden layer of the plurality of hidden layers (e.g. paragraph 0074, ML model may be split into at least three ML model components; ML model split between two adjacent layers of the neural network; ML model split such that the first ML model component comprises a series of initial layers at one end of the neural network and the second ML model component comprises a series of intervening layers of the neural network, the series of initial layers may receive and process input data to produce first intermediate neural network output data; Figs. 5 and 6, each showing the neural network having a plurality, at least three, intermediate layers (analogous to hidden layers) between the input and output layers; i.e. where the model includes a first series of initial/input layers, and is split between these and a series of intervening/hidden layers, this is analogous to a segmentation performed between an input layer (i.e. one of the series of initial layers responsible for processing inputs) and a first hidden layer of the plurality of hidden layers (i.e. a first layer of the intervening/hidden layers which is adjacent to the corresponding layer of the initial/input layers); Examiner notes that the claims do not appear to preclude the possibility of the neural network having multiple input layers/layers responsible for processing input data), the second hidden layer of the plurality of hidden layers and a third hidden layer of the plurality of hidden layers (e.g. as shown in Figs. 5 and 6, the neural network may also be split between layer 2 and layer 3; paragraph 0074, model split into at least three ML model components; ML model split between two adjacent layers; i.e. the model could be split into more than three components, such as four components, such that an additional split is introduced between intermediate/hidden layer components as shown in Figs. 5 and 6), and the third hidden layer of the plurality of hidden layers and an output layer (e.g. paragraph 0074, ML model may be split into at least three ML model components; ML model split between two adjacent layers of the neural network; ML model split such that the second ML model component comprises a series of intervening layers of the neural network, and the third ML model component comprises a series of end layers of the neural network; the series of intervening layers receive and process first intermediate neural network output data to produce second intermediate neural network output data; the series of end layers of the neural network may receive and process the second intermediate neural network output data to produce result data; Figs. 5 and 6, each showing the neural network having a plurality, at least three, intermediate layers (analogous to hidden layers) between the input and output layers; i.e. where the model includes a series of end/output layers, and is split between these and a series of intervening/hidden layers, this is analogous to a segmentation performed between an output/end layer (i.e. one of the series of end layers responsible for producing result data from intermediate data) and a third hidden layer of the plurality of hidden layers (i.e. a third layer of the intervening/hidden layers which is adjacent to the corresponding layer of the end/output layers); Examiner notes that the claims do not appear to preclude the possibility of the neural network having multiple output layers/layers responsible for producing output/result data); and send second inference information to a second network device, wherein the second inference information is determined based on the first inference information, and the second inference information is for determining a target inference result of the ML model, or the second inference information is the target inference result (e.g. paragraph 0022, indicating that the trained ML model can be split into a plurality of components and individual components can be distributed to individual computing devices including at least a sending device, a remote computing device, and intervening computing devices; paragraph 0027, indicating that multiple remote devices may be present in the system; paragraph 0038, Fig. 2, indicating that multiple remote devices (each with their own ML model component) may be utilized in the system; paragraph 0070, Fig. 9, step 910 processing intermediate neural network output data using second ML model component to generate result data; result data used to provide prediction based on input data received by first ML model component; paragraph 0077, Fig. 12, second intermediate neural network output data generated at remote computing device using second ML model component received from remote computing device; second intermediate neural network output data is processed using third ML component to generate result data comprising inference data; i.e. first intermediate data, generated by a first device using a first ML model component and received at another device having a second ML model component, may be processed by the device having the second ML model component to generate second intermediate data/inference information, and this may then be sent to yet another device having a third ML model component, such as to an intervening device (in the case that the first device also includes the third ML model, where the data then is provided to the first device after it is sent through the intervening device), or directly to a third device, in the instance that there are individual devices each having an individual ML model component (as cited with respect to paragraph 0022), such that the third device has the third ML model component which is to process the second intermediate data in order to produce the resulting final inference of the overall ML model). Bloom does not explicitly disclose: segmentation is performed between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers; and wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers. However, D’Ercoli teaches: segmentation is performed between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers (e.g. paragraph 0027-0029, Fig. 2, DNN 200 includes input layers 204, multiple hidden layers 206, and an output layer 208; DNN 200 partitioned to form partitioned DNN 212 having multiple layers L1-L4; after performing the partitioning each of the layers L1-L4 includes one or more of the original DNN’s layers 204, 206, and 208; deploying partitioned DNN to CPs/computing devices; i.e. as described, the neural network may have input, output, and a plurality of hidden layers, where each of these may be individually partitioned as a separate layer of partitioned neural network (i.e. as noted, each layer of the partitioned DNN may include one or more, and therefore only one, of the layers of the original DNN, such that partitioning may be performed between a first hidden layer of the plurality of hidden layers 206 and a second hidden layer of the plurality of hidden layers 206 to form separate/individual partitioned layers, such as L2 and L3 of the partitioned DNN)), wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers (e.g. paragraphs 0054-0055, Fig. 7, DNN partitioned into layers L1-L4 on separate computation points CP1-CP6; L1 on CP1 performing processing of new input and transmitting available feature map to other available CP such as CP2; CP2 generating output O1; highlighted layers of computational points CP2-6 are layers available to process feature maps from CP1 or from CP2, etc.; i.e. where the neural network has been partitioned into individual layers (i.e. each layer of the partitioned NN having one of the layers of the original network) and individual layers are distributed to separate computing devices, a first inference result may take the form of an output/inference result of a first hidden layer of the plurality of hidden layers; e.g., where a first hidden layer of the original DNN is partitioned into L2 of the partitioned NN, the output O2 generated by CP2 using layer L2 as shown in Fig. 7 may be a first inference result of a first hidden layer of a plurality of hidden layers). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom and D’Ercoli in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing), to incorporate the teachings of D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points) to include the capability to perform the segmentation/partitioning of the neural network including the input, plurality of hidden, and output layers on a layer-by-layer basis such that each layer of the partitioned network corresponds to one layer of the original network (and the segmentation therefore includes segmentation between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers), and to individually process outputs of each of the partitioned layers (including hidden layers, such as by passing the intermediate output from a first hidden layer to a second hidden layer), such that the first inference result is an inference result of the first hidden layer of the plurality of hidden layers (as taught by D’Ercoli). One of ordinary skill would have been motivated to perform such a modification in order to overcome problems of latency and energy consumption in performing deep learning algorithms as described in D’Ercoli (paragraph 0013). With respect to claim 16, Bloom teaches a collaborative inference apparatus (e.g. paragraph 0086, Fig. 14, machine 1400 able to perform discussed methodologies), comprising: a transceiver (e.g. paragraph 0090 Fig. 14, machine 1400 including communication components 1440 operable to couple machine 1400 to communications network 1432 and devices 1424, and may include wired communication components, wireless communication components, cellular communication components, NFC components, Bluetooth components, Wi-Fi components, etc.; i.e. a device capable of both transmitting and receiving data, including wirelessly, such as via radio wave/signal); at least one processor; and one or more memories coupled to the at least one processor and storing programming instructions, when executed by the at least one processor (e.g. paragraph 0087, Fig. 14, machine 1400 including processors 1404, memory/storage 1406, etc.; storage unit 1416 and memory 1414 store instructions 1410 embodying described methodologies/functions, and processors execute the instructions), cause the apparatus to: obtain third inference information, wherein the third inference information is determined based on all information about a first inference result, the first inference result is an inference result obtained after an operation is performed based on a first machine learning (ML) submodel, and the first ML submodel is a part of an ML model (e.g. paragraph 0019, split ML model components placed at different computing devices; sending computing device with encoding component encoding original domain specific data using component and sending the encoded data to remote computing device, which ultimately uses the received data to generate prediction data; paragraph 0022, performing remote inference using ML model which is split into components and distributed to multiple computing devices; paragraph 0023, trained ML model split into phase I and phase II components; sending computing device possessing the phase I component; sending computing device encoding data using phase I component; paragraph 0045, Fig. 3, remote machine sending encoded data to inference machine, which uses decoding ML model component on the encoded data to make prediction/inference; paragraph 0067, Fig. 8, providing intermediate neural network output data to remote computing device; result data generated at remote computing device using ML model component to process the intermediate neural network output data; paragraph 0070, Fig. 9, step 908 receiving intermediate neural network output data from remote computing device which is generated by processing input data at the remote computing device using first ML model component; step 910 processing intermediate neural network output data using second ML model component to generate result data; result data used to provide prediction based on input data received by first ML model component); comprising a plurality of hidden layers (e.g. Figs. 5 and 6, showing a plurality of at least three hidden layers in between input and output layers of the neural network; paragraph 0074, series of intervening layers of model, i.e. between initial/input layers and end/output layers, where these intermediate/intervening layers between input and output layers are hidden layers) and segmentation is performed between an input layer and a first hidden layer of the plurality of hidden layers (e.g. paragraph 0074, ML model may be split into at least three ML model components; ML model split between two adjacent layers of the neural network; ML model split such that the first ML model component comprises a series of initial layers at one end of the neural network and the second ML model component comprises a series of intervening layers of the neural network, the series of initial layers may receive and process input data to produce first intermediate neural network output data; Figs. 5 and 6, each showing the neural network having a plurality, at least three, intermediate layers (analogous to hidden layers) between the input and output layers; i.e. where the model includes a first series of initial/input layers, and is split between these and a series of intervening/hidden layers, this is analogous to a segmentation performed between an input layer (i.e. one of the series of initial layers responsible for processing inputs) and a first hidden layer of the plurality of hidden layers (i.e. a first layer of the intervening/hidden layers which is adjacent to the corresponding layer of the initial/input layers); Examiner notes that the claims do not appear to preclude the possibility of the neural network having multiple input layers/layers responsible for processing input data), the second hidden layer of the plurality of hidden layers and a third hidden layer of the plurality of hidden layers (e.g. as shown in Figs. 5 and 6, the neural network may also be split between layer 2 and layer 3; paragraph 0074, model split into at least three ML model components; ML model split between two adjacent layers; i.e. the model could be split into more than three components, such as four components, such that an additional split is introduced between intermediate/hidden layer components as shown in Figs. 5 and 6), and the third hidden layer of the plurality of hidden layers and an output layer (e.g. paragraph 0074, ML model may be split into at least three ML model components; ML model split between two adjacent layers of the neural network; ML model split such that the second ML model component comprises a series of intervening layers of the neural network, and the third ML model component comprises a series of end layers of the neural network; the series of intervening layers receive and process first intermediate neural network output data to produce second intermediate neural network output data; the series of end layers of the neural network may receive and process the second intermediate neural network output data to produce result data; Figs. 5 and 6, each showing the neural network having a plurality, at least three, intermediate layers (analogous to hidden layers) between the input and output layers; i.e. where the model includes a series of end/output layers, and is split between these and a series of intervening/hidden layers, this is analogous to a segmentation performed between an output/end layer (i.e. one of the series of end layers responsible for producing result data from intermediate data) and a third hidden layer of the plurality of hidden layers (i.e. a third layer of the intervening/hidden layers which is adjacent to the corresponding layer of the end/output layers); Examiner notes that the claims do not appear to preclude the possibility of the neural network having multiple output layers/layers responsible for producing output/result data); and send a target inference result to a terminal device, wherein the target inference result is an inference result that is of the ML model and that is determined based on the third inference information (e.g. paragraph 0019, computing device returning the generated prediction data to the sending computing device; paragraph 0045, Fig. 3, the inference machine provides/sends the prediction/inference back to the remote machine; paragraph 0068, prediction data is received (i.e. at a device, different from the remote computing device) from the remote computing device, where the prediction data is based on result data generated by first ML model component processing intermediate NN output data; result data may comprise inference data produced by the first ML model component). Bloom does not explicitly disclose: segmentation is performed between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers; and wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers. However, D’Ercoli teaches: segmentation is performed between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers (e.g. paragraph 0027-0029, Fig. 2, DNN 200 includes input layers 204, multiple hidden layers 206, and an output layer 208; DNN 200 partitioned to form partitioned DNN 212 having multiple layers L1-L4; after performing the partitioning each of the layers L1-L4 includes one or more of the original DNN’s layers 204, 206, and 208; deploying partitioned DNN to CPs/computing devices; i.e. as described, the neural network may have input, output, and a plurality of hidden layers, where each of these may be individually partitioned as a separate layer of partitioned neural network (i.e. as noted, each layer of the partitioned DNN may include one or more, and therefore only one, of the layers of the original DNN, such that partitioning may be performed between a first hidden layer of the plurality of hidden layers 206 and a second hidden layer of the plurality of hidden layers 206 to form separate/individual partitioned layers, such as L2 and L3 of the partitioned DNN)), wherein the first inference result is an inference result of the first hidden layer of the plurality of hidden layers (e.g. paragraphs 0054-0055, Fig. 7, DNN partitioned into layers L1-L4 on separate computation points CP1-CP6; L1 on CP1 performing processing of new input and transmitting available feature map to other available CP such as CP2; CP2 generating output O1; highlighted layers of computational points CP2-6 are layers available to process feature maps from CP1 or from CP2, etc.; i.e. where the neural network has been partitioned into individual layers (i.e. each layer of the partitioned NN having one of the layers of the original network) and individual layers are distributed to separate computing devices, a first inference result may take the form of an output/inference result of a first hidden layer of the plurality of hidden layers; e.g., where a first hidden layer of the original DNN is partitioned into L2 of the partitioned NN, the output O2 generated by CP2 using layer L2 as shown in Fig. 7 may be a first inference result of a first hidden layer of a plurality of hidden layers). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom and D’Ercoli in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing), to incorporate the teachings of D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points) to include the capability to perform the segmentation/partitioning of the neural network including the input, plurality of hidden, and output layers on a layer-by-layer basis such that each layer of the partitioned network corresponds to one layer of the original network (and the segmentation therefore includes segmentation between the first hidden layer of the plurality of hidden layers and a second hidden layer of the plurality of hidden layers), and to individually process outputs of each of the partitioned layers (including hidden layers, such as by passing the intermediate output from a first hidden layer to a second hidden layer), such that the first inference result is an inference result of the first hidden layer of the plurality of hidden layers (as taught by D’Ercoli). One of ordinary skill would have been motivated to perform such a modification in order to overcome problems of latency and energy consumption in performing deep learning algorithms as described in D’Ercoli (paragraph 0013). With respect to claim 13, Bloom in view of D'Ercoli teaches all of the limitations of claim 9 as previously discussed, and Bloom further teaches wherein the first inference information comprises all information about the first inference result (e.g. paragraph 0022, performing remote inference using ML model which is split into components and distributed to multiple computing devices; paragraph 0023, trained ML model split into phase I and phase II components; sending computing device possessing the phase I component; sending computing device encoding data using phase I component; paragraph 0029, encoded/intermediate data also provided with related metadata such as information regarding the ML model component that generated the data including versioning information, architecture of the model component, etc.; paragraph 0052, metadata including instructions to transform unprocessed raw data, links to code/containers that transform unprocessed data, encoding component of the model, governance/provenance data such as details regarding appropriate domain data, lifetime of validity, origin of the model, relevant URLs/links to send data to, ID of the model, creators of the model, etc.; paragraphs 0055 and 0067, intermediate neural network output data sent with metadata; paragraph 0070, Fig. 9, step 908 receiving intermediate neural network output data from remote computing device which is generated by processing input data at the remote computing device using first ML model component; paragraph 0076, Fig. 12, describing system in which model has been split into at least three components; intermediate neural network output data generated 1208 and provided to remote computing device); and the programming instructions, when executed by the at least one processor, further cause the apparatus to: determine the target inference result based on all information about the first inference result and a target ML submodel, wherein the second inference information is the target inference result, and input data of the target ML submodel corresponds to output data of the first ML submodel (e.g. paragraph 0070, Fig. 9, step 910 processing intermediate neural network output data using second ML model component to generate result data; result data used to provide prediction based on input data received by first ML model component; paragraph 0077, Fig. 12, second intermediate neural network output data generated at remote computing device using second ML model component received from remote computing device; second intermediate neural network output data is processed using third ML component to generate result data comprising inference data). With respect to claim 14, Bloom in view of D'Ercoli teaches all of the limitations of claim 9 as previously discussed, and Bloom further teaches wherein the first inference information comprises all information about the first inference result (e.g. paragraph 0022, performing remote inference using ML model which is split into components and distributed to multiple computing devices; paragraph 0023, trained ML model split into phase I and phase II components; sending computing device possessing the phase I component; sending computing device encoding data using phase I component; paragraph 0029, encoded/intermediate data also provided with related metadata such as information regarding the ML model component that generated the data including versioning information, architecture of the model component, etc.; paragraph 0052, metadata including instructions to transform unprocessed raw data, links to code/containers that transform unprocessed data, encoding component of the model, governance/provenance data such as details regarding appropriate domain data, lifetime of validity, origin of the model, relevant URLs/links to send data to, ID of the model, creators of the model, etc.; paragraphs 0055 and 0067, intermediate neural network output data sent with metadata; paragraph 0070, Fig. 9, step 908 receiving intermediate neural network output data from remote computing device which is generated by processing input data at the remote computing device using first ML model component; paragraph 0076, Fig. 12, describing system in which model has been split into at least three components; intermediate neural network output data generated 1208 and provided to remote computing device), and the programming instructions, when executed by the at least one processor, further cause the apparatus to: determine a second inference result based on all information about the first inference result and a second ML submodel, wherein the second inference information is the second inference result, and input data of the second ML submodel corresponds to output data of the first ML submodel (e.g. paragraph 0070, Fig. 9, step 910 processing intermediate neural network output data (from first ML model component) using second ML model component to generate result data; result data used to provide prediction based on input data received by first ML model component; paragraph 0077, Fig. 12, second intermediate neural network output data generated at remote computing device using second ML model component received from remote computing device; second intermediate neural network output data is processed using third ML component to generate result data comprising inference data). With respect to claim 2, Bloom in view of D'Ercoli teaches all of the limitations of claim 1 as previously discussed, and Bloom further teaches wherein when the apparatus accesses a first network device before determining the first inference result (e.g. as shown in Fig. 8, prior to generating the intermediate NN output data at step 808, the ML model is split into components at step 804 and the first ML model component is provided to the remote computing device at step 806; therefore the device (executing the method of Fig. 8) accesses the remote computing device (in order to provide the first ML model component) before determining the first inference result), the programming instructions, when executed by the at least one processor, further cause the apparatus to: send all information about the first inference result to the first network device (e.g. paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided); and receive the target inference result from the first network device, wherein the target inference result is an inference result that is of the ML model and that is determined based on all the information about the first inference result (e.g. paragraph 0023, remote computing device further encoding received encoded data using phase II component to produce result data which can represent an inference made by the ML model; paragraph 0068, Fig. 8 step 812, receiving prediction data from the remote computing device, where the prediction data is based on result data generated at the remote computing device by the first ML model component processing the intermediate neural network output data; the result data comprises inference data produced by the first ML model component). With respect to claim 17, Bloom in view of D'Ercoli teaches all of the limitations of claim 16 as previously discussed, and Bloom further teaches wherein when the terminal device accesses the apparatus before the apparatus obtains the third inference information, the third inference information is all information about the first inference result (e.g. as shown in Figs. 8/9, prior to generating the intermediate NN output data at step 808, the ML model is split into components at step 804/904 and the first ML model component is provided to the remote computing device at step 806/906; therefore the devices (executing the method of Fig. 8/9) access one another (in order to provide the first ML model component) before determining the inference information); and the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive all information about the first inference result from the terminal device (e.g. paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided); and determine the target inference result based on all information about the first inference result and a target ML submodel, wherein input data of the target ML submodel corresponds to output data of the first ML submodel (e.g. paragraph 0023, remote computing device further encoding received encoded data using phase II component to produce result data which can represent an inference made by the ML model; paragraph 0068, Fig. 8 step 812, receiving prediction data from the remote computing device, where the prediction data is based on result data generated at the remote computing device by the first ML model component processing the intermediate neural network output data; the result data comprises inference data produced by the first ML model component; paragraph 0070, Fig. 9, step 910 processing intermediate neural network output data using second ML model component to generate result data; result data used to provide prediction based on input data received by first ML model component; paragraph 0077, Fig. 12, second intermediate neural network output data generated at remote computing device using second ML model component received from remote computing device; second intermediate neural network output data is processed using third ML component to generate result data comprising inference data). With respect to claim 3, Bloom in view of D'Ercoli teaches all of the limitations of claim 2 as previously discussed, and Bloom further teaches wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive information about the first ML submodel from the first network device (e.g. paragraph 0023-0024, remote computing device can further encode received encoded data and send result/intervening encoded data to the sending computing device; paragraph 0029, encoded data/intermediate neural network output data provided with related metadata, such as information regarding the remote ML mode component that generated the data, including versioning information regarding the architecture of the ML model component; paragraph 0055, indicating that metadata includes additional information regarding the model, including origin of the model, ID of the model, creators of the model, etc.). With respect to claim 10, Bloom in view of D'Ercoli teaches all of the limitations of claim 9 as previously discussed, and Bloom further teaches wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: determine information about the first ML submodel; and send the information about the first ML submodel to the terminal device (e.g. paragraph 0023-0024, remote computing device can further encode received encoded data and send result/intervening encoded data to the sending computing device; paragraph 0029, encoded data/intermediate neural network output data provided with related metadata, such as information regarding the remote ML mode component that generated the data, including versioning information regarding the architecture of the ML model component). With respect to claim 18, Bloom in view of D'Ercoli teaches all of the limitations of claim 17 as previously discussed, and Bloom further teaches wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: send information about the first ML submodel to the terminal device (e.g. paragraph 0023-0024, remote computing device can further encode received encoded data and send result/intervening encoded data to the sending computing device; paragraph 0029, encoded data/intermediate neural network output data provided with related metadata, such as information regarding the remote ML mode component that generated the data, including versioning information regarding the architecture of the ML model component; paragraph 0055, indicating that metadata includes additional information regarding the model, including origin of the model, ID of the model, creators of the model, etc.). With respect to claim 7, Bloom in view of D'Ercoli teaches all of the limitations of claim 1 as previously discussed, and Bloom further teaches wherein when the apparatus accesses a first network device before sending the first inference result (e.g. as shown in Fig. 8, prior to generating the intermediate NN output data at step 808, the ML model is split into components at step 804 and the first ML model component is provided to the remote computing device at step 806; therefore the device (executing the method of Fig. 8) accesses the remote computing device (in order to provide the first ML model component) before determining the first inference result), and the apparatus accesses a second network device after sending the first inference result and before receiving the target inference result (e.g. paragraph 0019, computing device returning the generated prediction data to the sending computing device; paragraph 0022, indicating that model components are distributed to computing devices sending devices, remote devices, and intervening devices; i.e. where an intervening device exists between the sending computing device and the computing device returning the generated prediction/inference, the sending device will access this intervening device after sending the intermediate data/first inference result and before receiving the result, such as in order to receive the final result from the other computing device, since the sending device would not attempt to access the intervening device for the purpose of receiving the final result until after the first/intermediate result has been sent, and must first access the intervening device in order to be able to receive the result (generated by the other computing device) via the intervening device; Examiner notes that Fig. 1, for example, shows a remote device 130 with a respective model component and an application server 140 with a respective model component, and additionally shows at least one intervening device, such as API server 120 or Web Server 122, on the communication pathway between the remote device and the application server, such that inference results generated at the application server would be received directly from one of these intervening devices), the programming instructions, when executed by the at least one processor, further cause the apparatus to: send all information about the first inference result to the first network device (e.g. paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided); and receive the target inference result from the second network device, wherein the target inference result is an inference result that is of the ML model and that is determined based on all the information about the first inference result (e.g. paragraph 0019, computing device returning the generated prediction data to the sending computing device; paragraph 0022, indicating that model components are distributed to computing devices sending devices, remote devices, and intervening devices; paragraph 0023, remote computing device further encoding received encoded data using phase II component to produce result data which can represent an inference made by the ML model; paragraph 0045, Fig. 3, the inference machine provides/sends the prediction/inference back to the remote machine; paragraph 0068, Fig. 8 step 812, receiving prediction data from the remote computing device, where the prediction data is based on result data generated at the remote computing device by the first ML model component processing the intermediate neural network output data; the result data comprises inference data produced by the first ML model component; i.e. where an intervening device is present between the sending device and the other computing device, the inference result would be sent by the other computing device, via the intervening device, and ultimately received by the sending device (where this receiving by the sending device would be both receiving directly from the intervening device and receiving ultimately from the other computing device (via the intervening device); Examiner notes that Fig. 1, for example, shows a remote device 130 with a respective model component and an application server 140 with a respective model component, and additionally shows at least one intervening device, such as API server 120 or Web Server 122, on the communication pathway between the remote device and the application server, such that inference results generated at the application server would be received directly from one of these intervening devices). With respect to claim 8, Bloom in view of D'Ercoli teaches all of the limitations of claim 1 as previously discussed, and Bloom further teaches wherein when the apparatus accesses a second network device before sending the first inference result (e.g. paragraph 0022, indicating that model components are distributed to computing devices sending devices, remote devices, and intervening devices; Fig. 1, for example, shows a remote device 130 with a respective model component and an application server 140 with a respective model component, and additionally shows at least one intervening device, such as API server 120 or Web Server 122, on the communication pathway between the remote device and the application server, such that data sent between these two devices would also include accessing the intervening device on the communication pathway; as shown in Fig. 8, prior to generating the intermediate NN output data at step 808, the ML model is split into components at step 804 and the first ML model component is provided to the remote computing device at step 806; therefore the device (executing the method of Fig. 8) accesses the remote computing device (in order to provide the first ML model component) before determining the first inference result and, in a case where an intervening/second device is present in the communication pathway, also accesses this intervening/second computing device (such as during the process of providing the ML model component) before determining/sending the first inference result), the programming instructions, when executed by the at least one processor, further cause the apparatus to: send all information about the first inference result to the second network device (e.g. paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided; as noted above with respect to Fig. 1 and paragraph 0022, where an intervening/second device is present in the communication pathway between sending and remote devices, the sending of the information about the first inference result would include sending the information to the intervening/second device as well as to its ultimate destination); and receive the target inference result from the second network device, wherein the target inference result is an inference result that is of the ML model and that is determined based on all the information about the first inference result (e.g. paragraph 0023, remote computing device further encoding received encoded data using phase II component to produce result data which can represent an inference made by the ML model; paragraph 0068, Fig. 8 step 812, receiving prediction data from the remote computing device, where the prediction data is based on result data generated at the remote computing device by the first ML model component processing the intermediate neural network output data; the result data comprises inference data produced by the first ML model component; as noted above with respect to Fig. 1 and paragraph 0022, where an intervening/second device is present in the communication pathway between sending and remote devices, the receiving of the target inference result would include receiving this information from/via the intervening/second device). With respect to claim 15, Bloom in view of D'Ercoli teaches all of the limitations of claim 14 as previously discussed, and Bloom further teaches wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: send information about a target ML submodel to the second network device, wherein input data of the target ML submodel corresponds to output data of the second ML submodel; and the target ML submodel is used by the second network device to determine the target inference result (e.g. paragraph 0022, indicating that the trained ML model can be split into a plurality of components and individual components can be distributed to individual computing devices including at least a sending device, a remote computing device, and intervening computing devices; paragraph 0027, indicating that multiple remote devices may be present in the system; paragraph 0038, Fig. 2, indicating that multiple remote devices (each with their own ML model component) may be utilized in the system; paragraph 0070, Fig. 9, step 910 processing intermediate neural network output data using second ML model component to generate result data; result data used to provide prediction based on input data received by first ML model component; paragraph 0077, Fig. 12, second intermediate neural network output data generated at remote computing device using second ML model component received from remote computing device; second intermediate neural network output data is processed using third ML component to generate result data comprising inference data; i.e. where, as described in paragraph 0022, each individual model component is provided to a corresponding individual device and, as described in paragraph 0077, there are at least three model components, this third/target component/submodel is sent to the corresponding device (including information about the third/target component/submodel), where the second intermediate data, generated by a second device using a second ML model component and received as input at a third/different device having a third ML model component, may be processed by the device using the third ML model component to generate resulting inference information such as a final inference of the overall ML model). With respect to claim 6, Bloom in view of D'Ercoli teaches all of the limitations of claim 1 as previously discussed, and Bloom further teaches wherein when the apparatus accesses a first network device before sending the first inference result (e.g. as shown in Fig. 8, prior to generating the intermediate NN output data at step 808, the ML model is split into components at step 804 and the first ML model component is provided to the remote computing device at step 806; therefore the device (executing the method of Fig. 8) accesses the remote computing device (in order to provide the first ML model component) before determining the first inference result), and accesses a second network device in a process of sending the first inference result by the apparatus (e.g. paragraph 0019, computing device returning the generated prediction data to the sending computing device; paragraph 0022, indicating that model components are distributed to computing devices sending devices, remote devices, and intervening devices; i.e. where an intervening device exists between the sending computing device and the computing device returning the generated prediction/inference, the sending device will access this intervening device after sending the intermediate data/first inference result and before receiving the result, such as in order to receive the final result from the other computing device, since the sending device would not attempt to access the intervening device for the purpose of receiving the final result until after the first/intermediate result has been sent, and must first access the intervening device in order to be able to receive the result (generated by the other computing device) via the intervening device; Examiner notes that Fig. 1, for example, shows a remote device 130 with a respective model component and an application server 140 with a respective model component, and additionally shows at least one intervening device, such as API server 120 or Web Server 122, on the communication pathway between the remote device and the application server, such that inference results generated at the application server would be received directly from one of these intervening devices), and the programming instructions, when executed by the at least one processor, further cause the apparatus to: send first partial information about the first inference result to the first network device (e.g. paragraph 0022, indicating that the trained ML model can be split into a plurality of components and individual components can be distributed to individual computing devices including at least a sending device, a remote computing device, and intervening computing devices; paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided; i.e. where the system includes sending device/apparatus, a receiving/first device, and an intervening/second device, and inference information made up of at least two different components (actual result, and metadata, where these may be considered to be first and second partial information about an inference result) is sent to the receiving/first device via/through the intervening/second device, both the first partial information and the second partial information may be sent to the receiving/first device (i.e. after being sent from the sending device to the intervening device, and then from the intervening device to the receiving device)); send second partial information about the first inference result to the second network device (e.g. paragraph 0022, indicating that the trained ML model can be split into a plurality of components and individual components can be distributed to individual computing devices including at least a sending device, a remote computing device, and intervening computing devices; paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided; i.e. where the system includes sending device/apparatus, a receiving/first device, and an intervening/second device, and inference information made up of at least two different components (actual result, and metadata, where these may be considered to be first and second partial information about an inference result) is sent to the receiving/first device via/through the intervening/second device, both the first partial information and the second partial information may be sent to the intervening/second device (i.e. after being sent from the sending device to the intervening device, where this information is subsequently sent from the intervening device to the receiving device)); and receive the target inference result from the second network device, wherein the target inference result is an inference result that is of the ML model and that is determined based on the first partial information and the second partial information (e.g. paragraph 0022, indicating that the trained ML model can be split into a plurality of components and individual components can be distributed to individual computing devices including at least a sending device, a remote computing device, and intervening computing devices; paragraph 0023, remote computing device further encoding received encoded data using phase II component to produce result data which can represent an inference made by the ML model; paragraph 0068, Fig. 8 step 812, receiving prediction data from the remote computing device, where the prediction data is based on result data generated at the remote computing device by the first ML model component processing the intermediate neural network output data; the result data comprises inference data produced by the first ML model component; i.e. where the system includes sending device/apparatus, a receiving/first device, and an intervening/second device, and inference information made up of at least two different components is received at the receiving/first device, and the receiving/first device generates the target inference result based on both of these information components, this receiving/first device will then send the result back to the sending device/apparatus via/through the intervening device, such that the result is also received from the intervening device). With respect to claim 20, Bloom in view of D'Ercoli teaches all of the limitations of claim 16 as previously discussed, and Bloom further teaches wherein when the terminal device accesses the apparatus in a process of obtaining the third inference information by the apparatus, the third inference information is all information about the first inference result (e.g. as shown in Figs. 8/9, prior to generating the intermediate NN output data at step 808, the ML model is split into components at step 804/904 and the first ML model component is provided to the remote computing device at step 806/906; therefore the devices (executing the method of Fig. 8/9) access one another (in order to provide the first ML model component) before determining the inference information); and the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive first partial information about the first inference result from the terminal device (e.g. paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided; i.e. where the received information about the inference result includes at least two parts, including the intermediate neural network output data and the metadata relating to it, this is analogous to receiving at least first partial information about the inference result (such as the output data itself)); and receive second partial information about the first inference result from a first network device (e.g. paragraph 0023, sending device sending the encoded data to the remote computing device; paragraph 0067, Fig. 8 step 810, intermediate neural network output data (representing data transferred between layers of neural network) is provided to the remote computing device; metadata relating to the intermediate neural network output data, including versioning information regarding architecture, etc., also provided; i.e. where the received information about the inference result includes at least two parts, including the intermediate neural network output data and the metadata relating to it, this is analogous to receiving at least second partial information about the inference result (such as the corresponding metadata relating to the output data)); and determine the target inference result based on the first partial information, the second partial information, and a target ML submodel, wherein input data of the target ML submodel corresponds to output data of the first ML submodel (e.g. paragraph 0023, remote computing device further encoding received encoded data using phase II component to produce result data which can represent an inference made by the ML model; paragraph 0068, Fig. 8 step 812, receiving prediction data from the remote computing device, where the prediction data is based on result data generated at the remote computing device by the first ML model component processing the intermediate neural network output data; the result data comprises inference data produced by the first ML model component; paragraph 0070, Fig. 9, step 910 processing intermediate neural network output data using second ML model component to generate result data; result data used to provide prediction based on input data received by first ML model component; paragraph 0077, Fig. 12, second intermediate neural network output data generated at remote computing device using second ML model component received from remote computing device; second intermediate neural network output data is processed using third ML component to generate result data comprising inference data; i.e. generating the target/prediction/inference by the corresponding submodel, the received intermediate output data, and the received metadata (such as when the metadata includes information required to transform raw/unprocessed data, details for what domain data is appropriate for the process, links to send encoded data to, etc., as discussed in paragraph 0052)). Claims 4 and 12 are rejected under 35 U.S.C. 103 as being unpatentable over Bloom in view of D'Ercoli, further in view of Pogorelik et al. (US 20210319098 A1). With respect to claim 4, Bloom in view of D'Ercoli teaches all of the limitations of claim 3 as previously discussed. Bloom does not explicitly disclose wherein the information about the first ML submodel comprises first target indication information, and the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive first model information from the first network device, wherein the first model information comprises a correspondence between first candidate indication information and a first segmentation location; at least one piece of first candidate indication information and at least one first segmentation location are provided; and one piece of first candidate indication information indicates to segment the ML model, and a location at which the ML model is segmented is a first segmentation location that has a correspondence with the one piece of first candidate indication information; and determine the first ML submodel based on the first target indication information and the correspondence between the first candidate indication information and the first segmentation location. However, Pogorelik teaches wherein the information about the first ML submodel comprises first target indication information, and the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive first model information from the first network device, wherein the first model information comprises a correspondence between first candidate indication information and a first segmentation location; at least one piece of first candidate indication information and at least one first segmentation location are provided; and one piece of first candidate indication information indicates to segment the ML model, and a location at which the ML model is segmented is a first segmentation location that has a correspondence with the one piece of first candidate indication information (e.g. paragraph 0362, interrogating distributed computing platforms to determine security capabilities; paragraphs 0365-0366, client devices and servers communicating/exchanging data over network; client device memory storing model security requirements and server capabilities; paragraph 0369, model security requirements include indications of security requirements for respective layers, portions, or parts of the inference model; inference model split into slices based on the model security requirements and server capabilities; i.e. receiving model security requirements and network device capabilities, where this information includes information indicating different locations (i.e. layers, parts, or portions) of the model along with indicating corresponding security requirements associated with the model, where this information collectively provides indications of locations at which the model may be split/segmented such that the resulting model portions may be distributed to different devices for distributed execution); and determine the first ML submodel based on the first target indication information and the correspondence between the first candidate indication information and the first segmentation location (e.g. paragraph 0366, splitting inference model into model slices based on model security requirements and server capabilities so that model can be executed in distributed manner by servers while providing for security protection of portions of the model based on server capabilities and security requirements of the model; paragraph 0374, splitting model into model slices based on model security requirements and received capabilities). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom, D’Ercoli, and Pogorelik in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing) and D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points), to incorporate the teachings of Pogorelik (directed to securing systems employing AI, including via model splitting and distributed inferencing) to include the capability to receive, from a network device, information indicating different parts, portions, layers, locations, etc. of the model which may be split, including model security requirements corresponding to each of these candidate parts, portions, layers, or locations, along with corresponding security capability information of distributed network computing devices which will potentially receive and execute the corresponding parts, portions, layers, or locations of the model after the splitting is performed, and to determine various different split model portions (i.e. submodels) based on this information (as taught by Pogorelik). One of ordinary skill would have been motivated to perform such a modification in order to mitigate risk associated with attack vectors/vulnerabilities of AI systems as described in Pogorelik (paragraphs 0087-0088). D’Ercoli further teaches that the segmentation location corresponds to a respective segmentation option of a plurality of different segmentation options directed to segmenting by layers (e.g. paragraph 0027-0029, Fig. 2, DNN 200 includes input layers 204, multiple hidden layers 206, and an output layer 208; DNN 200 partitioned to form partitioned DNN 212 having multiple layers L1-L4; after performing the partitioning each of the layers L1-L4 includes one or more of the original DNN’s layers 204, 206, and 208; deploying partitioned DNN to CPs/computing devices; i.e. as described, the neural network may have input, output, and a plurality of hidden layers, where each of these may be individually partitioned as a separate layer of partitioned neural network (i.e. as noted, each layer of the partitioned DNN may include one or more, and therefore only one, of the layers of the original DNN)). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom, Pogorelik, and D’Ercoli in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing) and Pogorelik (directed to securing systems employing AI, including via model splitting and distributed inferencing), to incorporate the teachings of D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points) to include the capability to perform the segmentation/partitioning of the neural network according to a plurality of different segmentation options, including segmenting by layers (i.e. partitioning the neural network such that each layer of the partitioned network corresponds to one layer of the original network). One of ordinary skill would have been motivated to perform such a modification in order to overcome problems of latency and energy consumption in performing deep learning algorithms as described in D’Ercoli (paragraph 0013). With respect to claim 12, Bloom in view of D'Ercoli teaches all of the limitations of claim 10 as previously discussed. Bloom does not explicitly disclose wherein the information about the first ML submodel comprises first target indication information; and the programming instructions, when executed by the at least one processor, further cause the apparatus to: send first model information to the terminal device, wherein the first model information comprises a correspondence between first candidate indication information and a first segmentation location; at least one piece of first candidate indication information and at least one first segmentation location are provided; one piece of first candidate indication information indicates to segment the ML model, and a location at which the ML model is segmented is a first segmentation location that has a correspondence with the one piece of first candidate indication information (e.g. paragraph 0362, interrogating distributed computing platforms to determine security capabilities; paragraphs 0365-0366, client devices and servers communicating/exchanging data over network; client device memory storing model security requirements and server capabilities; paragraph 0369, model security requirements include indications of security requirements for respective layers, portions, or parts of the inference model; inference model split into slices based on the model security requirements and server capabilities; i.e. receiving model security requirements and network device capabilities, where this information includes information indicating different locations (i.e. layers, parts, or portions) of the model along with indicating corresponding security requirements associated with the model, where this information collectively provides indications of locations at which the model may be split/segmented such that the resulting model portions may be distributed to different devices for distributed execution); and the first model information and the first target indication information are used by the terminal device to determine the first ML submodel (e.g. paragraph 0366, splitting inference model into model slices based on model security requirements and server capabilities so that model can be executed in distributed manner by servers while providing for security protection of portions of the model based on server capabilities and security requirements of the model; paragraph 0374, splitting model into model slices based on model security requirements and received capabilities). However, Pogorelik teaches wherein the information about the first ML submodel comprises first target indication information; and the programming instructions, when executed by the at least one processor, further cause the apparatus to: send first model information to the terminal device, wherein the first model information comprises a correspondence between first candidate indication information and a first segmentation location; at least one piece of first candidate indication information and at least one first segmentation location are provided; one piece of first candidate indication information indicates to segment the ML model, and a location at which the ML model is segmented is a first segmentation location that has a correspondence with the one piece of first candidate indication information; and the first model information and the first target indication information are used by the terminal device to determine the first ML submodel. Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom, D’Ercoli, and Pogorelik in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing) and D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points), to incorporate the teachings of Pogorelik (directed to securing systems employing AI, including via model splitting and distributed inferencing) to include the capability to receive, from a network device, information indicating different parts, portions, layers, locations, etc. of the model which may be split, including model security requirements corresponding to each of these candidate parts, portions, layers, or locations, along with corresponding security capability information of distributed network computing devices which will potentially receive and execute the corresponding parts, portions, layers, or locations of the model after the splitting is performed, and to determine various different split model portions (i.e. submodels) based on this information (as taught by Pogorelik). One of ordinary skill would have been motivated to perform such a modification in order to mitigate risk associated with attack vectors/vulnerabilities of AI systems as described in Pogorelik (paragraphs 0087-0088). D’Ercoli further teaches that the segmentation location corresponds to a respective segmentation option of a plurality of different segmentation options directed to segmenting by layers (e.g. paragraph 0027-0029, Fig. 2, DNN 200 includes input layers 204, multiple hidden layers 206, and an output layer 208; DNN 200 partitioned to form partitioned DNN 212 having multiple layers L1-L4; after performing the partitioning each of the layers L1-L4 includes one or more of the original DNN’s layers 204, 206, and 208; deploying partitioned DNN to CPs/computing devices; i.e. as described, the neural network may have input, output, and a plurality of hidden layers, where each of these may be individually partitioned as a separate layer of partitioned neural network (i.e. as noted, each layer of the partitioned DNN may include one or more, and therefore only one, of the layers of the original DNN)). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom, Pogorelik, and D’Ercoli in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing) and Pogorelik (directed to securing systems employing AI, including via model splitting and distributed inferencing), to incorporate the teachings of D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points) to include the capability to perform the segmentation/partitioning of the neural network according to a plurality of different segmentation options, including segmenting by layers (i.e. partitioning the neural network such that each layer of the partitioned network corresponds to one layer of the original network). One of ordinary skill would have been motivated to perform such a modification in order to overcome problems of latency and energy consumption in performing deep learning algorithms as described in D’Ercoli (paragraph 0013). Claims 11 and 19 are rejected under 35 U.S.C. 103 as being unpatentable over Bloom in view of D'Ercoli, further in view of Kim et al. (US 20210287085 A1). With respect to claim 11, Bloom in view of D'Ercoli teaches all of the limitations of claim 10 as previously disclosed. Bloom does not explicitly disclose wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive inference requirement information from the terminal device, wherein the inference requirement information comprises information about a time at which the terminal device obtains the target inference result; and determine the information about the first ML submodel based on the inference requirement information. However, Kim teaches wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive inference requirement information from the terminal device, wherein the inference requirement information comprises information about a time at which the terminal device obtains the target inference result; and determine the information about the first ML submodel based on the inference requirement information (e.g. paragraph 0045, parallelization schemes/policies used for parallelization strategy, including intra-layer parallelism, inter-layer parallelism, partition dimensions indicating direction in which model, layer, etc. is divided, a division number indicating number of models or number of layers to be divided, etc.; paragraph 0046, generating parallelization strategy for target model; paragraph 0053, executing target model based on parallelization strategy of each target layer of the target model; outputting execution time of the target model or each target layer; the execution time used to evaluate performance of the parallelization strategy; paragraph 0054, reference layer information associated with reference layers, including metadata and reference parallelization strategy corresponding to layers, along with performance (execution time) of the parallelization strategy; paragraph 0058, comparing metadata of each target layer of target model and reference metadata of each reference layer in reference DM, measuring similarity; paragraph 0060, selecting layer corresponding to target layer based on similarity and generating parallelization strategy for the target layer based on matching; i.e. the system obtains information regarding time (execution time) in which the various results (such as first results, second results, target results, etc., corresponding to different model layers) are obtained by the device executing the corresponding model portions/layers/submodels (analogous to inference requirement information as defined in the claim), and determines corresponding information, such as metadata information, similarity measures, parallelization strategy, etc., for the model portions/layers/submodels based on this information). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom, D’Ercoli, and Kim in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing) and D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points), to incorporate the teachings of Kim (directed to parallel processing methods for neural network models) to include the capability to receive, from a corresponding device, information about a time (i.e. inference requirement information), such as an execution time, at which the device obtains corresponding inference results using a corresponding model layer/portion/submodel, and determine various information about one or more model layers/portions/submodels based on this information, including metadata information, similarity measures (between a target model portion/layer and a reference portion/layer), corresponding parallelization strategies, etc. (as taught by Kim). One of ordinary skill would have been motivated to perform such a modification in order to allow for more quickly converging to results in neural network model training and inference as described in Kim (paragraph 0003). With respect to claim 19, Bloom in view of D'Ercoli teaches all of the limitations of claim 18 as previously discussed. Bloom does not explicitly disclose wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive inference requirement information from the terminal device, wherein the inference requirement information comprises information about a time at which the terminal device obtains the target inference result; and determine the information about the first ML submodel based on the inference requirement information. However, Kim teaches wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: receive inference requirement information from the terminal device, wherein the inference requirement information comprises information about a time at which the terminal device obtains the target inference result; and determine the information about the first ML submodel based on the inference requirement information (e.g. paragraph 0045, parallelization schemes/policies used for parallelization strategy, including intra-layer parallelism, inter-layer parallelism, partition dimensions indicating direction in which model, layer, etc. is divided, a division number indicating number of models or number of layers to be divided, etc.; paragraph 0046, generating parallelization strategy for target model; paragraph 0053, executing target model based on parallelization strategy of each target layer of the target model; outputting execution time of the target model or each target layer; the execution time used to evaluate performance of the parallelization strategy; paragraph 0054, reference layer information associated with reference layers, including metadata and reference parallelization strategy corresponding to layers, along with performance (execution time) of the parallelization strategy; paragraph 0058, comparing metadata of each target layer of target model and reference metadata of each reference layer in reference DM, measuring similarity; paragraph 0060, selecting layer corresponding to target layer based on similarity and generating parallelization strategy for the target layer based on matching; i.e. the system obtains information regarding time (execution time) in which the various results (such as first results, second results, target results, etc., corresponding to different model layers) are obtained by the device executing the corresponding model portions/layers/submodels (analogous to inference requirement information as defined in the claim), and determines corresponding information, such as metadata information, similarity measures, parallelization strategy, etc., for the model portions/layers/submodels based on this information). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom, D’Ercoli, and Kim in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing) and D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points), to incorporate the teachings of Kim (directed to parallel processing methods for neural network models) to include the capability to receive, from a corresponding device, information about a time (i.e. inference requirement information), such as an execution time, at which the device obtains corresponding inference results using a corresponding model layer/portion/submodel, and determine various information about one or more model layers/portions/submodels based on this information, including metadata information, similarity measures (between a target model portion/layer and a reference portion/layer), corresponding parallelization strategies, etc. (as taught by Kim). One of ordinary skill would have been motivated to perform such a modification in order to allow for more quickly converging to results in neural network model training and inference as described in Kim (paragraph 0003). Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Bloom in view of D'Ercoli, further in view of Pogorelik, further in view of Kim. With respect to claim 5, Bloom in view of D'Ercoli, further in view of Pogorelik teaches all of the limitations of claim 4, as previously discussed. Bloom and Pogorelik do not explicitly disclose wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: send inference requirement information to the first network device, wherein the inference requirement information comprises information about a time at which the apparatus obtains the target inference result; and the inference requirement information is for determining the information about the first ML submodel. However, Kim teaches wherein the programming instructions, when executed by the at least one processor, further cause the apparatus to: send inference requirement information to the first network device, wherein the inference requirement information comprises information about a time at which the apparatus obtains the target inference result; and the inference requirement information is for determining the information about the first ML submodel (e.g. paragraph 0045, parallelization schemes/policies used for parallelization strategy, including intra-layer parallelism, inter-layer parallelism, partition dimensions indicating direction in which model, layer, etc. is divided, a division number indicating number of models or number of layers to be divided, etc.; paragraph 0046, generating parallelization strategy for target model; paragraph 0053, executing target model based on parallelization strategy of each target layer of the target model; outputting execution time of the target model or each target layer; the execution time used to evaluate performance of the parallelization strategy; paragraph 0054, reference layer information associated with reference layers, including metadata and reference parallelization strategy corresponding to layers, along with performance (execution time) of the parallelization strategy; paragraph 0058, comparing metadata of each target layer of target model and reference metadata of each reference layer in reference DM, measuring similarity; paragraph 0060, selecting layer corresponding to target layer based on similarity and generating parallelization strategy for the target layer based on matching; i.e. the system obtains information regarding time (execution time) in which the various results (such as first results, second results, target results, etc., corresponding to different model layers) are obtained by the device executing the corresponding model portions/layers/submodels (analogous to inference requirement information as defined in the claim), and determines corresponding information, such as metadata information, similarity measures, parallelization strategy, etc., for the model portions/layers/submodels based on this information). Accordingly, it would have been obvious to one of ordinary skill in the art before the effective filing date of the invention having the teachings of Bloom, D’Ercoli, and Kim in front of him to have modified the teachings of Bloom (directed to ML model splitting and distributed inferencing) and D’Ercoli (directed to deployment of deep neural networks, including by partitioning layers of the neural network and distributing them to different computational points), to incorporate the teachings of Kim (directed to parallel processing methods for neural network models) to include the capability to receive, from a corresponding device, information about a time (i.e. inference requirement information), such as an execution time, at which the device obtains corresponding inference results using a corresponding model layer/portion/submodel, and determine various information about one or more model layers/portions/submodels based on this information, including metadata information, similarity measures (between a target model portion/layer and a reference portion/layer), corresponding parallelization strategies, etc. (as taught by Kim). One of ordinary skill would have been motivated to perform such a modification in order to allow for more quickly converging to results in neural network model training and inference as described in Kim (paragraph 0003). It is noted that any citation to specific pages, columns, lines, or figures in the prior art references and any interpretation of the references should not be considered to be limiting in any way. “The use of patents as references is not limited to what the patentees describe as their own inventions or to the problems with which they are concerned. They are part of the literature of the art, relevant for all they contain,” In re Heck, 699 F.2d 1331, 1332-33, 216 USPQ 1038, 1039 (Fed. Cir. 1983) (quoting in re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (GCPA 1968)). Further, a reference may be relied upon for all that it would have reasonably suggested to one having ordinary skill the art, including nonpreferred embodiments. Merck & Co, v. Biocraft Laboratories, 874 F.2d 804, 10 USPQ2d 1843 (Fed. Cir.), cert, denied, 493 U.S. 975 (1989). See also Upsher-Smith Labs. v. Pamlab, LLC, 412 F,3d 1319, 1323, 75 USPQ2d 1213, 1215 (Fed. Cir, 2005): Celeritas Technologies Ltd. v. Rockwell International Corp., 150 F.3d 1354, 1361, 47 USPQ2d 1516, 1522-23 (Fed. Cir. 1998). Conclusion The prior art made of record and not relied upon is considered pertinent to applicant’s disclosure. Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. Any inquiry concerning this communication or earlier communications from the examiner should be directed to JEREMY L STANLEY whose telephone number is (469)295-9105. The examiner can normally be reached on Monday-Friday from 9:00 AM to 5:00 PM CST. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Abdullah Al Kawsar, can be reached at telephone number (571) 270-3169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from Patent Center and the Private Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from Patent Center or Private PAIR. Status information for unpublished applications is available through Patent Center and Private PAIR for authorized users only. Should you have questions about access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) Form at https://www.uspto.gov/patents/uspto-automated- interview-request-air-form. /JEREMY L STANLEY/ Primary Examiner, Art Unit 2127
Read full office action

Prosecution Timeline

Mar 16, 2023
Application Filed
Jan 27, 2026
Non-Final Rejection mailed — §103
Apr 09, 2026
Response Filed
Jun 17, 2026
Final Rejection mailed — §103
Aug 11, 2026
Response after Non-Final Action

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12688456
SECURE AND FAIR COMPETITIVE BIDDING
4y 5m to grant Granted Jul 21, 2026
Patent 12688443
INFORMATION PROCESSING METHOD AND APPARATUS, AND COMPUTER-READABLE STORAGE MEDIUM
4y 3m to grant Granted Jul 21, 2026
Patent 12670362
VIDEO SYNTHESIS WITHIN A MESSAGING SYSTEM
4y 9m to grant Granted Jun 30, 2026
Patent 12670198
Textual Summaries In Information Systems Based On Personalized Prior Knowledge
3y 4m to grant Granted Jun 30, 2026
Patent 12660205
TEMPORAL KERNEL DEVICES, TEMPORAL KERNEL COMPUTING SYSTEMS, AND METHODS OF THEIR OPERATION
3y 9m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

2-3
Expected OA Rounds
49%
Grant Probability
90%
With Interview (+40.9%)
3y 3m (~0m remaining)
Median Time to Grant
Moderate
PTA Risk
Based on 288 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month