DETAILED ACTION
Claims 1 and 12-20 are presented for examination.
This office action is in response to submission of application on 04-FEBRUARY-2026.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01-DECEMBER-2022 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted on 08-AUGUST-2023 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted on 10-JANUARY-2024 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
The information disclosure statement (IDS) submitted on 17-JUNE-2026 is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Response to Amendment
The amendment filed 04-FEBRUARY-2026 in response to the non-final office action mailed DATE has been entered. Claims 1 and 12-20 remain pending in the application.
With regards to the non-final office action’s rejection under 112, the amendments to the claims have overcome the original rejection with regards to the claims being directed towards an abstract idea.
With regards to the non-final office action’s rejection under 101, the amendments to the claims have overcome the original rejection with regards to the claims being directed towards an abstract idea.
With regards to the non-final office action’s rejections under 103, the amendments to the claims necessitated a new consideration of the art. After this consideration, the examiner respectfully disagrees with the applicant’s arguments that the art referenced in the previous office action does not teach the amendment claim limitations. A new 103 rejection over the prior art has been provided.
Claim Rejections - 35 USC § 103
Claims 1-14 and 19-20 are rejected under 35 U.S.C. 103 as being unpatentable over Yin et al. (Pub. No. US 20220351020 A1, filed April 30th 2021, hereinafter Yin) in view of Zhao et al. (Pub. No. US 20190312772 A1, filed April 4th 2018, hereinafter Zhao).
Regarding claim 1:
Claim 1 recites:
A distributed training method based on end-to-end adaption, executed by a model training platform, comprising: obtaining slicing results by slicing a model to be trained; obtaining an attribute of computing resources allocated to the model for training by parsing the computing resources, wherein the computing resources are determined based on a number of clients that initiate a model training request and a minimum computing power and a maximum computing power of required computing resources parsed from a model training request of the model to be trained in response to receiving the model training request of the model to be trained submitted by a client, and the attribute of the computing resources comprises a communication topology relation and a hardware topology relation, the hardware topology relation comprises a connection relation of the computing resources, bandwidth information, and a task processing capability; determining a distribution strategy of each of the slicing results in the computing resources based on the attributes of the computing resources; and performing distributed training on the model using the computing resources based on the distribution strategy wherein obtaining the slicing results by slicing the model to be trained. comprises: obtaining N slices by slicing the operators and the tensors in the model based on the slicing strategy, the N being a positive integer; for each of the N slices, loading distributed attribute information of the slice, wherein the distributed attribute information comprises at least one of process topology information of the slice in the model, slicing mapping information of the slice and slice size information of the slice; and taking the slice loaded with the distributed attribute information as the slicing result; wherein the method further comprises: determining placement information of each of the N slices based on the distributed attribute information of each of the N slices, wherein the placement information is configured to represent a physical mapping relation between the N slices and the computing resources; wherein, when the slices are located at adjacent network layers of the model and have different placement information, the method comprises: determining a communication auxiliary operator based on the placement information, wherein the communication auxiliary operator is configured to represent a logical operation relation between the slices; and transmitting, by using the communication auxiliary operator, a computing result obtained by a computing resource corresponding to an upstream network layer in the adjacent network layers to a computing resource corresponding to a downstream network layer in the adjacent network layers; wherein, when the slices are located at a same network layer of the model, the method comprises: determining a recombination transformation operator, wherein the recombination transformation operator is configured to represent a network layer consistency relation between the slices; and merging, by using the recombination transformation operator, computing results obtained by computing resources corresponding to the slices located at the same network layer; wherein, obtaining the attribute of the computing resources allocated to the model for training by parsing the computing resources, comprises: determining a minimum component in the computing resources, wherein the minimum component comprises a processor or a memory; determining a machine comprising at least one minimum component, wherein the minimum component in each machine is not repeated; determining a cluster comprising at least one machine, wherein the machine in each cluster is not repeated; determining an affinity list of each minimum component, wherein the affinity list comprises at least one of a connection relation between a source minimum component and a target minimum component, bandwidth information and latency information; and taking the minimum component, the machine, the cluster and the affinity list as the hardware topology relation of the computing resources; acquiring a communication path of the computing resources, wherein the communication path of the computing resources is configured to represent a communication connection state, a communication connection mode and a communication speed of a source communication resource and a target communication resource; and constructing a communication topology relation between the computing resources based on the communication path of the computing resources.
Yin discloses distributed training method based on end-to-end adaption, executed by a model training platform, comprising: obtaining slicing results by slicing a model to be trained;
Yin teaches that a model is split into different slices, wherein a machine learning model would inherently include training (Paragraph 23). Furthermore, Yin teaches a model deployment program which is used to train the model, which would make it a model training platform (Paragraph 54).
Yin discloses obtaining an attribute of computing resources allocated to the model for training by parsing the computing resources, wherein the computing resources are determined [based on a number of clients that initiate a model training request] and a minimum computing power and a maximum computing power of required computing resources parsed from a model training request of the model to be trained in response to receiving the model training request of the model to be trained submitted by a client:
Yin teaches performing a search to match slices with computing resource, which would be obtaining an attribute of computing resources allocated to the model for training, including the capability of an edge device which would be a minimum computing power and a maximum computing power of a required computing resource since the capability would describe the full range of ability of the edge device (Paragraph 23).
Furthermore, Yin teaches a virtual model cache that selects a candidate model and model slices, wherein the selection would be a model training request of the model to be trained where the virtual model cache is a client (Paragraph 21).
However, Yin does not teach wherein the computing resources are determined based on a number of clients that initiate a model training request. This limitation is taught by Zhao further below.
Yin discloses determining a distribution strategy of each of the slicing results in the computing resources based on the attributes of the computing resources;
Yin teaches distributing model slices based on the resources of the matching device (Paragraph 23) wherein the resources of the matching device would be attributes of the computing resource and the particular distribution of slices to matchings devices would be a distribution strategy as it may vary based on the available edge devices and hence the available best matches.
Yin discloses wherein obtaining the slicing results by slicing the model to be trained. comprises: obtaining N slices by slicing the operators and the tensors in the model based on the slicing strategy, the N being a positive integer:
Yin teaches splitting a model into slices by a slicing generator, wherein the slices may be divided by layers (Paragraph 23), wherein tensors would be contained with layers and hence determined and sliced with the rest of the model. The operators that connect the tensors would therefore also be within the layers and determined and sliced when the layers go through that same process.
Furthermore, Yin teaches determining a different slicing strategy based on the complexity of the network (Paragraph 23) which would in turn impact the slicing of the operators and tensors.
Finally, Yin teaches splitting a model into a plurality of slices based on a set of predetermined rules, wherein predetermined rules would describe a slicing strategy (Paragraph 6) and the number of slices in the plurality would be the positive integer N.
Yin teaches for each of the N slices, loading distributed attribute information of the slice, wherein the distributed attribute information comprises at least one of process topology information of the slice in the model, slicing mapping information of the slice and slice size information of the slice; and taking the slice loaded with the distributed attribute information as the slicing result:
Yin teaches that in order to match model slices with resources, a best match is determined based on the size and capability of the device (Paragraph 23). If this information is used to pair the device with a model slice, then the slice size information of the slice (wherein slice size information is at least one of the list of distributed attribute information) must also be loaded during the matching process and taken in order to match it with a device as a result of the slice size attribute. As each slice is distributed, (Paragraph 23) this process happens for each of the N slices.
Yin discloses wherein the method further comprises: determining placement information of each of the N slices based on the distributed attribute information of each of the N slices, wherein the placement information is configured to represent a physical mapping relation between the N slices and the computing resources:
Yin teaches that each of the N slices are distributed in accordance with the above limitation, wherein the matchings of slices to resources would be determining placement information of the slices based on the attribute information i.e. the slices size of each slice (Paragraph 23) wherein the placement information is a physical mapping relation between the N slices and the computing resources as the mapping determines the distribution of the slices (Paragraph 23), which is a physical relation.
Yin discloses wherein, when the slices are located at adjacent network layers of the model and have different placement information, the method comprises: determining a communication auxiliary operator based on the placement information, wherein the communication auxiliary operator is configured to represent a logical operation relation between the slices:
Yin teaches that the slices of the model may communication with each other (Paragraph 6) and that more specifically, since the slices are based on the layers, the information exchange may be performed through related layers (Paragraph 22). The slices would have different placement information as they comprise different layers of the network and would have an adjacent relationship based on adjacent layers being within them. Furthermore, communication auxiliary operator would be the communication between adjacent layers, as the related layers would be a logical operation relation between the slices as it represents part of the original operation of the base machine learning model that was sliced.
Yin discloses transmitting, by using the communication auxiliary operator, a computing result obtained by a computing resource corresponding to an upstream network layer in the adjacent network layers to a computing resource corresponding to a downstream network layer in the adjacent network layers:
Yin teaches cutting the model into different network layers with data exchange between related layers (Paragraph 22). Related layers would include layers downstream and upstream of each other as well as adjacent network layers wherein the data exchange between layers would be transmitting a computing result obtained by a computing resources corresponding to an upstream layer to a downstream layer.
Yin discloses wherein, when the slices are located at a same network layer of the model, the method comprises: determining a recombination transformation operator, wherein the recombination transformation operator is configured to represent a network layer consistency relation between the slices:
Yin teaches that for some model which are hard to parallelize, the model may be split into the smallest parallelizable layers, wherein the determination of smallest parallelizable layers would be a recombination transformation operator as it represents a network layer consistency relation between the slices as the sliced layers must also form a relationship with each other in order to keep the state of the network intact (Paragraph 39). Furthermore, in some cases slices may be located at the same network layer as layers may need to record part of the layer before and ahead of it in order to link layers together, wherein the layers containing part of another layer would be a slice and therefore occupying the same layer as the slice before it and after it (Paragraph 39).
Yin discloses merging, by using the recombination transformation operator, computing results obtained by computing resources corresponding to the slices located at the same network layer:
Yin teaches above a recombination transformation operator. Furthermore, Yin teaches selected slices to create new virtual models from the slices (Paragraph 9). This would be a form of merging the computing results obtained by computing resources corresponding to the slices located at the same network layer.
Yin does not disclose wherein the computing resources are determined based on a number of clients that initiate a model training request. Instead, this limitation is disclosed by Zhao in the same field of endeavor of computing resource allocation discloses:
Zhao teaches load balancing, wherein computing resources are determined based of a number of client devices in order to appropriately distribute resources (Paragraph 82).
Zhao and the present application are analogous art because they are both in the same field of computing resource allocation.
Yin does not disclose and the attribute of the computing resources comprises a communication topology relation and a hardware topology relation, the hardware topology relation comprises a connection relation of the computing resources, bandwidth information, and a task processing capability. Instead, this limitation is disclosed by Zhao:
Zhao teaches a service platform that provides a topology-aware provisioning of computer resources (Paragraph 6) wherein the providing of a topology-aware provisioning would demonstrate that a hardware topology relation of the computing resources has been determined.
Zhao teaches that the topology database includes information on the computer resources such as connections between nodes as well as performance metrics which are analogous to task processing capabilities (Paragraph 27) as it is given for the computing resources which would be the minimum components. Furthermore, Zhao teaches a database that regarding current bandwidth usage (Paragraph 30).
However, Yin does not teach performing distributed training on the model using the computing resources based on the distribution strategy. Instead, Zhao discloses:
Zhao teaches performed distributed training (Paragraph 35), which in combination with the particular distribution strategy of Yin would be performing distributed training of the model using the computing resource based on the distribution strategy.
Yin does not disclose wherein, obtaining the attribute of the computing resources allocated to the model for training by parsing the computing resources comprises: determining a minimum component in the computing resources, wherein the minimum component comprises a processor or a memory;
Zhao teaches determining for a server node the number of hardware processor devices (Paragraph 25) which would be as described in the claim language a minimum component in the computing resources.
determining a machine comprising at least one minimum component, wherein the minimum component in each machine is not repeated; determining a cluster comprising at least one machine, wherein the machine in each cluster is not repeated:
Zhao teaches a given server node which contains hardware processor resources, which would mean that it comprises a minimum component which is a process (Paragraph 25). Furthermore, the processor resources may include GPU resources (Paragraph 25) which would include one or more GPU devices (Paragraph 84). In the case where there in one GPU resource, the minimum component is not repeated. Furthermore, a server cluster may contain one or more server nodes (Paragraph 21) wherein is one server node is present than the server node is not repeated.
determining an affinity list of each minimum component, wherein the affinity list comprises at least one of a connection relation between a source minimum component and a target minimum component, bandwidth information and latency information:
Zhao teaches that the topology database includes information on the resources such as connections between nodes as well as performance metrics (Paragraph 27) which would be an affinity list of each minimum component as it is given for the computing resources which would be the minimum components. Furthermore, as the topology database includes the connection topologies, it would include a connection relation between a source minimum component and a target minimum component which would be at least one of the above list.
taking the minimum component, the machine, the cluster and the affinity list as the hardware topology relation of the computing resources;
Zhao teaches that the topology database include information on the topology of each active server in a server cluster (Paragraph 25) which demonstrates that the topology database contains both the cluster, the machine (e.g. the server node) and the minimum component which would be the server topology.
Yin does not disclose acquiring a communication path of the computing resources, wherein the communication path of the computing resources is configured to represent a communication connection state, a communication connection mode and a communication speed of a source communication resource and a target communication resource; and constructing a communication topology relation between the computing resources based on the communication path of the computing resources:
Zhao teaches that for provisioning accelerator devices in a logical ring, which would be a form of communication path, the slowest communication path in the logical ring determines the overall performance, and therefore a provisioning model avoids provisioning low-performance and high-performance connection topologies (Paragraph 28). This would demonstrate representation of a communication speed of a source communication resource and a target communication resource through the communication path, as well as a communication connection state of either low or high performance and a communication connection mode of provisioned or not. The provisioning model may therefore acquire a communication path of the computing resources and construct a topology relation between the resources based on the computing path when determining the logical ring of the devices, and may take the performance of each path between devices as an attribute for provisioning.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that used the teachings of Yin and the teachings of Zhao. This would have provided the advantage of efficiently provisioning a group of computing resources (Zhao, Paragraph 24).
Regarding claim 12, which depends upon claim 1:
Claim 12 recites:
The method of claim 1, wherein, determining the distribution strategy of each of the slicing results in the computing resources based on the attribute of the computing resources, comprises: acquiring candidate distribution strategies of respective slicing results in the computing resources; determining an efficiency of each of the candidate distribution strategies; and determining a target distribution strategy in the candidate distribution strategies based on the efficiency of each of the candidate distribution strategies
Yin in view of Zhao teach the method of claim 1 upon which claim 12 depends. Furthermore, regarding the limitation of claim 12:
Yin teaches performing a search in order to determine the resources of best match for a model slice, wherein a search would comprise determining an efficiency of the match in question based on the size and capability of the computing resource (Paragraph 23). The match would be a candidate distribution strategy as each resource the search looks through would be a candidate for the distribution of the slice. Therefore, the search would also be acquiring candidate distribution strategies as it must determine the resources to search through. When a best match is found through the search, the target distribution strategy would be determined based on the efficiency of each of the strategies as the match is based on capability of the resource, which would lead to efficiency.
Regarding claim 13, which depends upon claim 1:
Claim 13 recites:
The method of claim 12, wherein, determining the target distribution strategy in the candidate distribution strategies based on the efficiency of each of the candidate distribution strategies, comprises: sorting the candidate distribution strategies based on a predetermined rule; and determining the target distribution strategy in the candidate distribution strategies based on a sorting result
Yin in view of Zhao teach the method of claim 12 upon which claim 13 depends. Furthermore, regarding the limitation of claim 13:
Yin teaches that in order to determine the most optimal model, the top n models may be recorded and displayed based on an analysis of the model, wherein n is a predetermined value (Paragraph 43). This would comprise sorting based on a predetermined rule and determining a target based on a sort result as the sorting is based on the suitability of the model and the chosen n, wherein the top n models are determined as possible target models. This could be used in conjunction with Yin’s distribution strategies in order to sort the candidate distribution strategies based on a predetermined rule and determine a target distribution strategy based on a sort result in the same manner.
Regarding claim 14, which depends upon claim 1:
Claim 14 recites:
The method of claim 1, wherein, performing distributed training on the model using the computing resources based on the distribution strategy, comprises: periodically detecting availability of the computing resources; and performing a remedial measure in response to a detection result indicating that the computing resources are in an unavailable condition, the unavailable condition comprising computing resource failure or shrinkage in a number of the computing resources
Yin in view of Zhao teach the method of claim 1 upon which claim 14 depends. However, Yin does not teach the limitation performing distributed training on the model using the computing resources based on the distribution strategy, comprises: periodically detecting availability of the computing resources:
Zhao teaches that whenever a service request is received, determine available devices in order to process server nodes (Paragraph 21). The periodic detection of availability in Zhao is therefore tied to a service request, which may happen periodically.
Regarding the limitation performing a remedial measure in response to a detection result indicating that the computing resources are in an unavailable condition, the unavailable condition comprising computing resource failure or shrinkage in a number of the computing resources:
Yin teaches that for each device when running the model, it may be determined that a particular device is at high risk or failing (Paragraph 48). This would be an indication that the compute resource is in an unavailable condition, which would be compute resource failure. In response to this, Yin replaced the failure device with a healthy one (Paragraph 48) which would be a remedial measure.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that used the teachings of Yin and the teachings of Zhao. This would have provided the advantage of efficiently provisioning a group of computing resources (Zhao, Paragraph 24).
Claim 19 recites a device that parallels the method of claim 1. Therefore, the analysis discussed above with respect to claim 1 also applies to claim 19. Accordingly, claim 19 is rejected based on substantially the same rationale as set forth above with respect to claim 1.
Claim 20 recites a non-transitory computer readable medium that parallels the method of claim 1. Therefore, the analysis discussed above with respect to claim 1 also applies to claim 20. Accordingly, claim 20 is rejected based on substantially the same rationale as set forth above with respect to claim 1.
Claims 15-16 are rejected under 35 U.S.C. 103 as being unpatentable over Yin in view of Zhao further in view of Qiao et al. (Pub. No. CN 112000473 A, published August 12th 2020, hereinafter Qiao).
Regarding claim 15, which depends upon claim 14:
Claim 15 recites:
The method of claim 14, wherein, performing the remedial measure in response to the unavailable condition being the computing resource failure, comprises: acquiring a training mode comprised in a model training request initiated by a client; waiting for failure recovery of the computing resources in response to the training mode being a fault-tolerant training mode; and determining that performing ends in response to the computing resource failure is not recovered within a predetermined time
Yin in view of Zhao teach the method of claim 14 upon which claim 15 depends. However, Yin in view of Zhao does not teach the limitation of claim 15:
“In one embodiment, further comprising: a data storage module for controlling the main node to store the training state data to the database; a fault-tolerant recovery module, used for when the training node executes the training task failure, restarting the training node, and loading the training state data in the database, to recover the execution training task” Qiao.
Qiao in the same field of endeavor of machine learning teaches a fault-tolerant recovery module, which would be the training mode. In combination with the request of Yin as described above in claim 14, this fault-tolerant recovery module would be acquired when there is a service request and a fault is detected (Yin, Paragraph 48).
Qiao and the present application are analogous art because they are in the same field of endeavor of machine learning
Furthermore, Qiao teaches that the recovery module is used for the training node executes the training task failure upon which the training task is recovered, which would be waiting for the failure recovery of the computing resource by the module.
Regarding the limitation determining that performing ends in response to the computing resource failure is not recovered within a predetermined time:
“the main node also will store the index list of the data partition and the different training nodes into the persistent storage (ETCD) together with the use progress of the data, so as to recover the training task in time when the training fails” Qiao.
Finally, Qiao teaches that the training task must be recovered in time, which would be the predetermined time to recover the computing resource failure. While Qiao does not explicitly say that performing the task ends in response to failure to recover the task, the task could be continued if not recovered and would therefore be ended.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that used the teachings of Yin in view of Zhao and the teachings of Qiao. This would have provided the advantage of improving the utilization rate of various computing resources (Qiao, “One embodiment of the above application has the following advantages or beneficial effects […] improves the GPU or CPU resource utilization rate”).
Regarding claim 16, which depends upon claim 15:
Claim 16 recites:
The method of claim 15, wherein, performing the remedial measure in response to the unavailable condition being computing resource failure, further comprises: determining candidate computing resources in response to the training mode being an elastic training mode; and retrying training in the candidate computing resources
Yin in view of Zhao further in view of Qiao teach the method of claim 15 upon which claim 16 depends. Furthermore, regarding the limitation performing the remedial measure in response to the unavailable condition being computing resource failure, further comprises: determining candidate computing resources in response to the training mode being an elastic training mode:
Yin teaches that when a compute resource fails, an healthy device is identified to replace to failure device (Paragraph 48), which would be a remedial measure that determine a candidate computing resource in response to an unavailable condition of the original.
Furthermore, Qiao teaches “generating a first elastic telescopic strategy according to the cluster resource requirement sent by the user; the first elastic telescopic strategy comprises increasing or reducing the number of the training node” which would be an elastic training mode.
However, Yin in view of Zhao does not teach the limitation retrying training in the candidate computing resources:
“In one embodiment, further comprising: a data storage module for controlling the main node to store the training state data to the database; a fault-tolerant recovery module, used for when the training node executes the training task failure, restarting the training node, and loading the training state data in the database, to recover the execution training task” Qiao.
Qiao teaches recovering the training task, which would be a form of retrying training in the candidate compute resource when combined with the above-described methodology of Yin.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that used the teachings of Yin in view of Zhao and the teachings of Qiao. This would have provided the advantage of improving the utilization rate of various computing resources (Qiao, “One embodiment of the above application has the following advantages or beneficial effects […] improves the GPU or CPU resource utilization rate”).
Claims 17-18 are rejected under 35 U.S.C. 103 as being unpatentable over Yin in view of Zhao further in view of Tasinga et al. (Pub. No. US 20220180178 A1, filed December 8th 2020, hereinafter Tasinga).
Regarding claim 17, which depends upon claim 14:
Claim 17 recites:
The method of claim 14, wherein, performing the remedial measure in response to the unavailable condition being the shrinkage in the number of the computing resources, comprises: determining a first number of remaining computing resources after the shrinkage; obtaining first re-slicing results by re-slicing the model based on the first number; determining a first distribution strategy of each of the first re-slicing results in the remaining computing resources based on the attribute of the remaining computing resources; and performing distributed training on the model using the remaining computing resources based on the first distribution strategy
Yin in view of Zhao teach the method of claim 14 upon which claim 17 depends. However, Yin in view of Zhao does not teach the limitation performing the remedial measure in response to the unavailable condition being the shrinkage in the number of the computing resources, comprises: determining a first number of remaining computing resources after the shrinkage:
Tasinga in the same field of endeavor of machine learning teaches a cloud computing environment wherein an AI-assisted load-balancer employs load-balancing techniques to evenly distribute resources utilization across available servers, i.e. computing resources (Paragraph 63). Load-balancing among available servers would include responding to shrinkage of the number of servers and determining a number of remaining servers.
Tasinga and the present application are analogous art because they are in the same field of endeavor of machine learning
Regarding the limitation obtaining first re-slicing results by re-slicing the model based on the first number; determining a first distribution strategy of each of the first re-slicing results in the remaining computing resources based on the attribute of the remaining computing resources; and performing distributed training on the model using the remaining computing resources based on the first distribution strategy:
Tasinga in the same field of endeavor of machine learning teaches a cloud computing environment wherein an AI-assisted load-balancer employs load-balancing techniques to evenly distribute resources utilization across available servers, i.e. computing resources (Paragraph 63). In combination with the model slicing as described in claim 1 by Yin in view of Zhao, the load-balancing of Tasinga is meant to evenly distribute tasks, wherein the model slices would be a task. Therefore, a redistribution by load-balancing of Tasinga would be obtaining re-slicing results based on the number of compute resource and determining a distribution strategy for the re-slicing results in the remaining compute resources.
Furthermore, Yin in view of Zhao has previously taught performing distributed training on the model, wherein the distributed training would take place in the load-balanced environment of Tasinga.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that used the teachings of Yin in view of Zhao and the teachings of Tasinga. This would have provided the advantage of improving performance of the load-balanced machine learning models (Tasinga, Paragraph 63).
Regarding claim 18, which depends upon claim 14:
Claim 18 recites:
The method of claim 14, in response to the detection result indicating that there are available additional computing resources, comprising: determining a second number of the available additional computing resources; obtaining second re-slicing results by re-slicing the model based on the second number; determining a second distribution strategy of each of the second re-slicing results in computing resources after expansion using an attribute of the additional computing resources; and performing distributed training on the model using the computing resources after the expansion based on the second distribution strategy
Yin in view of Zhao teach the method of claim 14 upon which claim 18 depends. Furthermore, either Yin nor Zhao teach the limitation in response to the detection result indicating that there are available additional computing resources, comprising: determining a second number of the available additional computing resources:
Tasinga teaches a cloud computing environment wherein an AI-assisted load-balancer employs load-balancing techniques to evenly distribute resources utilization across available servers, i.e. computing resources (Paragraph 63). Load-balancing among available servers would include responding to additional availability of the number of servers and determining a number of additional servers.
However, Yin in view of Zhao does not teach the limitation obtaining second re-slicing results by re-slicing the model based on the second number; determining a second distribution strategy of each of the second re-slicing results in computing resources after expansion using an attribute of the additional computing resources; and performing distributed training on the model using the computing resources after the expansion based on the second distribution strategy:
Tasinga in the same field of endeavor of machine learning teaches a cloud computing environment wherein an AI-assisted load-balancer employs load-balancing techniques to evenly distribute resources utilization across available servers, i.e. computing resources (Paragraph 63). In combination with the model slicing as described in claim 1 by Yin in view of Zhao, the load-balancing of Tasinga is meant to evenly distribute tasks, wherein the model slices would be a task. Therefore, a redistribution by load-balancing of Tasinga would be obtaining re-slicing results based on the number of compute resource and determining a distribution strategy for the re-slicing results in the remaining compute resources.
Furthermore, Yin in view of Zhao has previously taught performing distributed training on the model, wherein the distributed training would take place in the load-balanced environment of Tasinga.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the present application to implement a methodology that used the teachings of Yin in view of Zhao and the teachings of Tasinga. This would have provided the advantage of improving performance of the load-balanced machine learning models (Tasinga, Paragraph 63).
Response to Arguments
Applicant’s arguments filed 04-FEBRUARY-2026 have been fully considered, but the examiner believes that not all are fully persuasive.
Regarding the applicant’s remarks on the non-final office action’s 103 rejection of the claims, the applicant argues that Yin in view of Zhao does not teach the amended limitations of these claims. As such, the applicant argues that all claims dependent on the above would additionally not be obvious under 103. However, the examiner believes that Yin in view of Zhao does teach the amended limitations and respectfully requests applicant’s consideration of the following:
Regarding the applicant’s argument Yin does not teach or suggest determining the computer resources allocated to the model for training based on the number of clients that initiate the model training request and the minimum and maximum computer power of required computer resources parsed from the model training request to be trained:
Yin teaches performing a search to match slices with computing resource, which would be obtaining an attribute of computing resources allocated to the model for training, including the capability of an edge device which would be a minimum computing power and a maximum computing power of a required computing resource since the capability would describe the full range of ability of the edge device (Paragraph 23).
Furthermore, Yin teaches a virtual model cache that selects a candidate model and model slices, wherein the selection would be a model training request of the model to be trained where the virtual model cache is a client (Paragraph 21).
However, Yin does not teach wherein the computing resources are determined based on a number of clients that initiate a model training request. This limitation is taught by Zhao.
Zhao teaches load balancing, wherein computing resources are determined based of a number of client devices in order to appropriately distribute resources (Paragraph 82).
In combination, the consideration of the number of client devices seen in Zhao may be incorporated into the determining of the computer resources. This would have provided the advantage of efficiently provisioning a group of computing resources (Zhao, Paragraph 24).
Yin discloses wherein the method further comprises: determining placement information of each of the N slices based on the distributed attribute information of each of the N slices, wherein the placement information is configured to represent a physical mapping relation between the N slices and the computing resources:
Yin teaches that each of the N slices are distributed in accordance with the above limitation, wherein the matchings of slices to resources would be determining placement information of the slices based on the attribute information i.e. the slices size of each slice (Paragraph 23) wherein the placement information is a physical mapping relation between the N slices and the computing resources as the mapping determines the distribution of the slices (Paragraph 23), which is a physical relation.
Furthermore, the applicant argues that Yin does not disclose slices that are located at adjacent network layers corresponding to different computing resources, and hence that the placement information of the slices located at adjacent network layers is different:
Yin teaches that the slices of the model may communicate with each other (Paragraph 6) and that more specifically, since the slices are based on the layers, the information exchange may be performed through related layers (Paragraph 22). The slices would have different placement information as they comprise different layers of the network and would have an adjacent relationship based on adjacent layers being within them. Furthermore, communication auxiliary operator would be the communication between adjacent layers, as the related layers would be a logical operation relation between the slices as it represents part of the original operation of the base machine learning model that was sliced.
Furthermore, the applicant argues that Yin does not disclose how to handle the computing results obtained by the computing resources corresponding to slices that are located at the same network layer of the model to be trained and likewise does not teach a recombination transformation operator:
Yin teaches that for some model which are hard to parallelize, the model may be split into the smallest parallelizable layers, wherein the determination of smallest parallelizable layers would be a recombination transformation operator as it represents a network layer consistency relation between the slices as the sliced layers must also form a relationship with each other in order to keep the state of the network intact (Paragraph 39). Furthermore, in some cases slices may be located at the same network layer as the smallest parallelizable layers are mentioned in opposition to each layer being a separate slice, and the split layers exist as the same network layer as other slices between there is overlap between the layer information recorded in that slice and the layer information of other slices (Paragraph 39).
Furthermore, regarding the limitation merging, by using the recombination transformation operator, computing results obtained by computing resources corresponding to the slices located at the same network layer:
Yin teaches above a recombination transformation operator. Furthermore, Yin teaches selected slices to create new virtual models from the slices (Paragraph 9). This would be a form of merging the computing results obtained by computing resources corresponding to the slices located at the same network layer.
For these reasons, the examiner believes that the amended claims are rejectable under 103 over Yin in view of Zhao. Claims 19 and 20, which contain analogous limitations, are rejected based on substantially the same rationale as set forth above with respect to claim 1.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ALEXANDRIA JOSEPHINE MILLER whose telephone number is (703)756-5684. The examiner can normally be reached Monday-Thursday: 7:30 - 5:00 pm, every other Friday 7:30 - 4:00.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Mariela Reyes can be reached at (571) 270-1006. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/A.J.M./Examiner, Art Unit 2142 /Mariela Reyes/Supervisory Patent Examiner, Art Unit 2142