Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Information Disclosure Statement
The information disclosure statement (IDS) submitted on 01-17-2024 and 06-25-2024
is in compliance with the provisions of 37 CFR 1.97. Accordingly, the information disclosure statement is being considered by the examiner.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-6, 13-14 are rejected under U.S.C 101 for containing an abstract idea without significantly more.
Regarding claim 1:
Step 1 – Is the claim to a process, machine, manufacture or composition of matter?
Yes, the claim is a process.
Step 2A – Prong 1 – Does the claim recite an abstract idea, law of nature, or natural phenomenon?
Yes, the claim recites an abstract idea.
determining a target model to be trained and dividing the target model into sub- models; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
in response to monitoring that the model training task for the target model is abnormal during execution, determining a faulty node from the device nodes, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
determining an execution progress when the model training task for the target model is abnormal as a first progress; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
determining a backup node for the faulty node, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
monitoring whether the faulty node returns to normal within a set period; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
in response to determining that the faulty node returns to normal within the set period, determining an execution progress of the backup node executing the model training task for the sub-model deployed in the faulty node when the faulty node returns to normal as a second progress, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
in response to determining that the faulty node does not return to normal within the set period, re-dividing the target model according to the number of normal device nodes to obtain re-divided sub-models, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
Step 2A – Prong 2 – Does the claim recite additional elements that integrate the judicial exception into a practical application?
No, there are no additional elements that integrate the judicial exception into a practical application. The additional elements:
deploying the sub-models respectively in device nodes to perform a model training task for the target model through the device nodes; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
continuing to execute, by the backup node, a model training task for a sub-model deployed in the faulty node from the first progress; and Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
continuing to execute, by the faulty node, the model training task for the sub-model deployed in the faulty node from the second progress; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
deploying the re-divided sub-models respectively in the normal device nodes, to perform the model training task for the target model. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Step 2B – Does the claim recite additional elements that amount to significantly more than the judicial exception?
No, there are no additional elements that amount to significantly more than the judicial exception. The additional elements are:
deploying the sub-models respectively in device nodes to perform a model training task for the target model through the device nodes; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
continuing to execute, by the backup node, a model training task for a sub-model deployed in the faulty node from the first progress; and Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
continuing to execute, by the faulty node, the model training task for the sub-model deployed in the faulty node from the second progress; Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
deploying the re-divided sub-models respectively in the normal device nodes, to perform the model training task for the target model. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 2,
Claim 2 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
monitoring whether heartbeat signals from the device nodes are received at default time intervals; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
in response to determining that a heartbeat signal transmitted by at least one of the device nodes is not received within a designated period, determining that the model training task for the target model is abnormal during execution, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
determining a device node that does not transmit a heartbeat signal within the designated period as a faulty node. This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
Regarding claim 3,
Claim 3 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 2 which includes an abstract idea (see rejection for claim 2). The additional limitations:
transmitting a start signal to the backup node corresponding to the faulty node, such that after receiving the start signal, the backup node corresponding to the faulty node reads out the sub-model deployed in the faulty node that is locally pre-stored in the backup node, and Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
continues to execute the model training task for the sub-model deployed in the faulty node from the first progress. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 4,
Claim 4 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein determining the execution progress of the backup node executing the model training task for the sub-model deployed in the faulty node when the faulty node returns to normal as the second progress, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
continuing to execute, by the faulty node, the model training task for the sub-model deployed in the faulty node from the second progress comprises: Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
in response to determining that the faulty node returns to normal, according to execution progress information of the model training task for the target model carried by the heartbeat signal transmitted by the backup node, determining the execution progress of the backup node executing the model training task for the sub-model deployed in the faulty node, as the second progress; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
transmitting model data of the sub-model deployed in the backup node to the faulty node, such that the faulty node updates a sub-model deployed in the faulty node according to the model data; and Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
transmitting a restart signal to the faulty node, such that after receiving the restart signal, the faulty node continues to execute the model training task for an updated sub- model deployed in the faulty node from the second progress. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 5,
Claim 5 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein re-dividing the target model according to the number of the normal device nodes to obtain the re-divided sub-models, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
deploying the re-divided sub-models respectively in the normal device nodes comprises: Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
re-dividing the target model according to the number of the normal device nodes, to obtain a dividing result; This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
for each of the normal device nodes, based on the dividing result, determining a network layer in the target model that needs to be migrated to the device node as a supplementary network layer corresponding to the device node, and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
determining a current device node where the supplementary network layer corresponding to the device node is currently located as a network layer source node corresponding to the device node; and This limitation is directed to the abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
based on the supplementary network layer corresponding to each of the normal device nodes and the network layer source node corresponding to each of the normal device nodes, adjusting a network layer currently contained in each of the normal device nodes, to deploy the re-divided sub-models respectively to the normal device nodes. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 6,
Claim 6 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
wherein the backup node is a predecessor node for the faulty node, and This claim merely recites a further limitation on the determining a backup node for the faulty node from Claim 1 which was directed to abstract idea of a mental process (including an observation, evaluation, judgement, opinion) which can be performed in the human mind, or by a human using pen and paper (see MPEP 2106.04(a)(2) Ill. C.)
the predecessor node is configured to – This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
transmit a result of a forward calculation to the faulty node after completing the forward calculation of the sub-model deployed to the predecessor node. Adding the words "apply it" (or an equivalent) with the judicial exception, or mere instructions to implement an abstract idea on a computer, or merely uses a computer as a tool to perform an abstract idea [see MPEP 2106.05(f)] and therefore fails to integrate the exception into a practical application.
Regarding claim 13,
Claim 13 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
A computer readable storage medium, wherein the storage medium stores a computer program, and the computer program when executed by a processor achieves a method according to claim 1. – This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
Regarding claim 14,
Claim 14 is rejected under 35 U.S.C 101 because the claimed invention is directed to an abstract idea without significantly more. The claim is dependent on claim 1 which includes an abstract idea (see rejection for claim 1). The additional limitations:
An electronic device comprising a memory, a processor and a computer program stored on the memory and runnable on the processor, wherein the processor, when the program is executed by the processor, achieves a method according to claim 1 . – This limitation is directed to a computer merely used as a tool to perform an existing process (see MPEP 2106.05(f) (2)).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
Claim(s) 1, 3-6 and 13-14 are rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (US 2023/0169351 A1) in view of YU (US 2023/0229905 A1)
Regarding claim 1, Wang explicitly discloses:
A method for training a distributed model based on node fault perception, comprising: (Wang, ¶[0031]: “the present disclosure relates to a distributed training method based on end-to-end adaption, and the method may include the followings.”, ¶[0153]: “As mentioned above, the training mode may include a fault-tolerant training mode and an elastic training mode”)
determining a target model to be trained and dividing the target model into sub- models; (Wang, ¶[0078]: “At block S502, a target category from a plurality of candidate categories is determined as the category of the distributed attribute information”)
deploying the sub-models respectively in device nodes to perform a model training task for the target model through the device nodes; (Wang, ¶[0066]: “By parsing the model to be trained, the operators and the tensors in the model to be trained may be determined.”, ¶[0060]: “When a model and a resource change, the whole system may automatically regenerate a unified computing view and a unified resource view, and automatically trigger subsequent sub-modules without interrupting a task to perform reslicing, resources mapping again, allocate new distributed tasks to each device, and start a process to perform task allocation and execution.”)
monitoring whether the faulty node returns to normal within a set period; (Wang, ¶[0205]: “For a task in the fault-tolerant mode, after a failure occurs in one or more computing resources during the training process (a node is unavailable or a GPU card fails), the entire training task may not exit, and a current computing resource is not released either, and training may be continued if a failure node is recovered within a timeout period (continue training from a failure moment), otherwise the task fails and exits.”)
in response to determining that the faulty node returns to normal within the set period, determining an execution progress of the backup node executing the model training task for the sub-model deployed in the faulty node when the faulty node returns to normal as a second progress, and (Wang, ¶[0205]: “For a task in the fault-tolerant mode, after a failure occurs in one or more computing resources during the training process (a node is unavailable or a GPU card fails), the entire training task may not exit, and a current computing resource is not released either, and training may be continued if a failure node is recovered within a timeout period (continue training from a failure moment), otherwise the task fails and exits.”)
in response to determining that the faulty node does not return to normal within the set period, re-dividing the target model according to the number of normal device nodes to obtain re-divided sub-models, and (Wang, ¶[0170]: “At block S1402, first re-slicing results are obtained by re-slicing the model to be trained based on the first number.”, ¶[0171]: “At block S1403, a first distribution strategy of each of the first re-slicing results in the remaining computing resources; is determined using the attribute of the redetermined remaining computing resources.”)
deploying the re-divided sub-models respectively in the normal device nodes, to perform the model training task for the target model. (Wang, ¶[0168-0172]: “…when the unavailable condition is the shrinkage in the number of computing resources, a remedial measure is performed as follows. At block S1401, a first number of remaining computing resources after the shrinkage is determined. At block S1402, first re-slicing results are obtained by re-slicing the model to be trained based on the first number. At block S1403, a first distribution strategy of each of the first re-slicing results in the remaining computing resources; is determined using the attribute of the redetermined remaining computing resources. At block S1404, distributed training is performed on the model to be trained using the remaining computing resources based on the first distribution strategy.”)
Wang fails to disclose:
in response to monitoring that the model training task for the target model is abnormal during execution, determining a faulty node from the device nodes, and
determining an execution progress when the model training task for the target model is abnormal as a first progress;
determining a backup node for the faulty node, and
continuing to execute, by the backup node, a model training task for a sub-model deployed in the faulty node from the first progress; and
continuing to execute, by the faulty node, the model training task for the sub-model deployed in the faulty node from the second progress;
Yu explicitly discloses:
in response to monitoring that the model training task for the target model is abnormal during execution, determining a faulty node from the device nodes, and (Yu, ¶[0060]: “All agents may be monitored for potential faults during and/or between training iterations. Failed agents may be detected at a node, within a pod, by the scheduler, and/or by processes running on the system. For example, an agent that reports a failed worker to the scheduler may cause the scheduler to initiate the restart process. If that agent fails to reply to the restart request within a threshold period of time the scheduler may indicate that the agent has failed.”)
determining an execution progress when the model training task for the target model is abnormal as a first progress; (Yu, ¶[0043]: “For globally distributed training, a coherent, consistent checkpointing of the state is needed to ensure the correctness of the entire training in result of a failure. Each checkpoint may be considered roughly equal to the current state of the training, indicating the current values of the parameters of the model. Such a checkpointing step typically occurs at the end of a training iteration.”)
determining a backup node for the faulty node, and (Yu, ¶[0049]: “In other words, a global backup may be generated infrequently, e.g., daily, following a threshold number of training iterations, following a threshold number of faults.”)
continuing to execute, by the backup node, a model training task for a sub-model deployed in the faulty node from the first progress; and (Yu, ¶0029]: “To guard against failure, the system may be periodically checkpointed, storing a checkpoint state (e.g., current parameters) so that any failure does not require the training to start at the very beginning, only at the most recent checkpoint.”, ¶[0030]: “Typically, the checkpoint state is saved to a global storage system. Recovery from a failure necessitates loading the checkpoint state to all processors from global storage and then restarting every processor from the checkpoint state.”)
continuing to execute, by the faulty node, the model training task for the sub-model deployed in the faulty node from the second progress; (Yu, ¶0029]: “To guard against failure, the system may be periodically checkpointed, storing a checkpoint state (e.g., current parameters) so that any failure does not require the training to start at the very beginning, only at the most recent checkpoint.”, ¶[0030]: “Typically, the checkpoint state is saved to a global storage system. Recovery from a failure necessitates loading the checkpoint state to all processors from global storage and then restarting every processor from the checkpoint state.”)
The combination of Wang and Yu are analogous art because they are in the same field of training time series data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, having the teachings of Wang and Yu before them, to modify the teachings of Wang to include the teachings of Yu to improve the reliability of computing devices by limiting progression of the worker processor units through training iterations such that all worker processor units are maintained within a threshold number of training iterations
Regarding to claim 3, the combination of Wang and Yu discloses all the limitations of claim 1 (as shown in the rejection above).
Wang in view of Yu further discloses:
transmitting a start signal to the backup node corresponding to the faulty node, (Yu, ¶[0060]: “an agent that reports a failed worker to the scheduler may cause the scheduler to initiate the restart process.”)
such that after receiving the start signal, the backup node corresponding to the faulty node reads out the sub-model deployed in the faulty node that is locally pre-stored in the backup node, and (Yu, ¶[0032]: “During training, each worker processor unit stores its share of the state locally in memory associated with an agent processor unit. If a worker processor unit fails, it can thus be recovered quickly based on the locally stored checkpoint state, as long as the respective agent is healthy”, ¶[0056]: “When checkpoint states are distributed and maintained locally in parallel, an entire agent, node, and/or pod can be rescued in the result of a failure by retrieving the checkpoint states from peers and either re-initializing or re-allocating the failed system components.”, ¶[0103]: “In any of the preceding examples, or any other example, each agent is additionally or alternatively configured to upload checkpoint states from local memory to a global backup at a lower frequency than the checkpoint states are recorded at the respective agent”)
continues to execute the model training task for the sub-model deployed in the faulty node from the first progress. (Yu, ¶0029]: “To guard against failure, the system may be periodically checkpointed, storing a checkpoint state (e.g., current parameters) so that any failure does not require the training to start at the very beginning, only at the most recent checkpoint.”, ¶[0030]: “Typically, the checkpoint state is saved to a global storage system. Recovery from a failure necessitates loading the checkpoint state to all processors from global storage and then restarting every processor from the checkpoint state.”)
Regarding to claim 4, the combination of Wang and Yu discloses all the limitations of claim 1 (as shown in the rejection above).
Wang in view of Yu further discloses:
wherein determining the execution progress of the backup node executing the model training task for the sub-model deployed in the faulty node when the faulty node returns to normal as the second progress, and (Wang, ¶[0205]: “For a task in the fault-tolerant mode, after a failure occurs in one or more computing resources during the training process (a node is unavailable or a GPU card fails), the entire training task may not exit, and a current computing resource is not released either, and training may be continued if a failure node is recovered within a timeout period (continue training from a failure moment), otherwise the task fails and exits.”)
continuing to execute, by the faulty node, the model training task for the sub-model deployed in the faulty node from the second progress comprises: (Yu, ¶0029]: “To guard against failure, the system may be periodically checkpointed, storing a checkpoint state (e.g., current parameters) so that any failure does not require the training to start at the very beginning, only at the most recent checkpoint.”, ¶[0030]: “Typically, the checkpoint state is saved to a global storage system. Recovery from a failure necessitates loading the checkpoint state to all processors from global storage and then restarting every processor from the checkpoint state.”)
in response to determining that the faulty node returns to normal, according to execution progress information of the model training task for the target model carried by the heartbeat signal transmitted by the backup node, determining the execution progress of the backup node executing the model training task for the sub-model deployed in the faulty node, as the second progress; (Wang, ¶[0205]: “For a task in the fault-tolerant mode, after a failure occurs in one or more computing resources during the training process (a node is unavailable or a GPU card fails), the entire training task may not exit, and a current computing resource is not released either, and training may be continued if a failure node is recovered within a timeout period (continue training from a failure moment), otherwise the task fails and exits.”)
transmitting model data of the sub-model deployed in the backup node to the faulty node, such that the faulty node updates a sub-model deployed in the faulty node according to the model data; and ¶[YU, ¶[0060]: “At 530, method 500 includes based on an indication of an agent failing, reassigning and initializing the agent and associated worker processing units based at least on one or more checkpoint states stored at an agent associated with a different worker processing unit in a respective peer group. All agents may be monitored for potential faults during and/or between training iterations. Failed agents may be detected at a node, within a pod, by the scheduler, and/or by processes running on the system… A most recent checkpoint state may be selected to initialize the reassigned agent and worker, though in some examples, the scheduler may indicate an older stored checkpoint state for initialization.”]
transmitting a restart signal to the faulty node, (Yu, ¶[0060]: “an agent that reports a failed worker to the scheduler may cause the scheduler to initiate the restart process.”)
such that after receiving the restart signal, the faulty node continues to execute the model training task for an updated sub- model deployed in the faulty node from the second progress. (Yu, ¶0029]: “To guard against failure, the system may be periodically checkpointed, storing a checkpoint state (e.g., current parameters) so that any failure does not require the training to start at the very beginning, only at the most recent checkpoint.”, ¶[0030]: “Typically, the checkpoint state is saved to a global storage system. Recovery from a failure necessitates loading the checkpoint state to all processors from global storage and then restarting every processor from the checkpoint state.”)
Regarding to claim 5, the combination of Wang and Yu discloses all the limitations of claim 1 (as shown in the rejection above).
Wang in view of Yu further discloses:
wherein re-dividing the target model according to the number of the normal device nodes to obtain the re-divided sub-models, and (Wang, ¶[0170]: “At block S1402, first re-slicing results are obtained by re-slicing the model to be trained based on the first number.”, ¶[0171]: “At block S1403, a first distribution strategy of each of the first re-slicing results in the remaining computing resources; is determined using the attribute of the redetermined remaining computing resources.”)
deploying the re-divided sub-models respectively in the normal device nodes comprises: (Wang, ¶[0168-0172]: “…when the unavailable condition is the shrinkage in the number of computing resources, a remedial measure is performed as follows. At block S1401, a first number of remaining computing resources after the shrinkage is determined. At block S1402, first re-slicing results are obtained by re-slicing the model to be trained based on the first number. At block S1403, a first distribution strategy of each of the first re-slicing results in the remaining computing resources; is determined using the attribute of the redetermined remaining computing resources. At block S1404, distributed training is performed on the model to be trained using the remaining computing resources based on the first distribution strategy.”)
re-dividing the target model according to the number of the normal device nodes, to obtain a dividing result; (Wang, ¶[0168-0172]: “…when the unavailable condition is the shrinkage in the number of computing resources, a remedial measure is performed as follows. At block S1401, a first number of remaining computing resources after the shrinkage is determined. At block S1402, first re-slicing results are obtained by re-slicing the model to be trained based on the first number. At block S1403, a first distribution strategy of each of the first re-slicing results in the remaining computing resources; is determined using the attribute of the redetermined remaining computing resources. At block S1404, distributed training is performed on the model to be trained using the remaining computing resources based on the first distribution strategy.”)
for each of the normal device nodes, based on the dividing result, determining a network layer in the target model that needs to be migrated to the device node as a supplementary network layer corresponding to the device node, and (Wang, ¶[0161]: “In this case, when the computing resource failure occurs, other available computing resources may be selected as candidate computing resources. Training is retried in the candidate computing resources by migrating model training data.”, ¶[0162]: “The computing resources include DO to D4. The example as illustrated on the left side of FIG. 13 is that a failure occurs in the computing resource D2 in a model training process using the computing resources DO to D3. On the basis of this, the example as illustrated on the right side of FIG. 13 is that the computing resource D4 is taken as a candidate computing resource, and the data originally trained in the computing resource D2 is migrated to the computing source D4 for retrying.”)
determining a current device node where the supplementary network layer corresponding to the device node is currently located as a network layer source node corresponding to the device node; and (Wang, ¶[0083]: “determining placement information of each of the slices using the distributed attribute of each of the slices, the placement information is configured to represent a physical mapping relation between the slices and the computing resources.”, ¶0084]: “The placement information (device_placement) of the slice may represent the computing resource required by the slice.”, ¶[0104]: “when the current computing resource is taken as a source computing resource, the connection relation of the computing resources may include a connection relation between the source computing resource and a target computing resource”)
based on the supplementary network layer corresponding to each of the normal device nodes and the network layer source node corresponding to each of the normal device nodes, adjusting a network layer currently contained in each of the normal device nodes, to deploy the re-divided sub-models respectively to the normal device nodes. (Wang, ¶0160]: “the number of requested computing resources has been declared within one range in the model training request, and further the number of computing resources may be adjusted based on the range”, ¶[0170]: “At block S1402, first re-slicing results are obtained by re-slicing the model to be trained based on the first number.”, ¶[0171]: “At block S1403, a first distribution strategy of each of the first re-slicing results in the remaining computing resources; is determined using the attribute of the redetermined remaining computing resources “, ¶[0175]: “the computing resource DO is allocated to the new slice rank0 and the computing resource D1 is allocated to the new slice rank1 based on the first distribution strategy.”)
Regarding to claim 6, the combination of Wang and Yu discloses all the limitations of claim 1 (as shown in the rejection above).
Wang in view of Yu further discloses:
wherein the backup node is a predecessor node for the faulty node, and (Wang, ¶[0088]: “For the adjacent network layers, output data of an upstream network layer may be taken as input data of a downstream network layer. The computing resource corresponding to the first hidden layer is XPU0 in FIG. 3, and the computing resource corresponding to the second hidden layer is GPU0 in FIG. 3.”)
the predecessor node is configured to transmit a result of a forward calculation to the faulty node after completing the forward calculation of the sub-model deployed to the predecessor node. (WANG, ¶[0089]: “In this case, a communication auxiliary operator may be determined, to inform two slices and corresponding computing resources, and for example, informed information may indicate to transmit a computing result to GPU0 using the communication auxiliary operator when XPU0 in FIG. 3 completes corresponding computation and obtains the computing result, to achieve continuity of computing and ensure correctness of cross-device slicing.”, FIG3:
PNG
media_image1.png
366
1020
media_image1.png
Greyscale
)
Regarding to claim 13, the combination of Wang and Yu discloses all the limitations of claim 1 (as shown in the rejection above).
Wang in view of Yu further discloses:
A computer readable storage medium, wherein the storage medium stores a computer program, and the computer program when executed by a processor achieves a method according to claim 1. (Wang, ¶[0005-0006]: “The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor;… According to another aspect of the present disclosure, a non-transitory computer-readable storage medium stored with computer instructions is provided, the computer instructions are configured to cause a computer to perform the method in any one embodiment of the present disclosure”)
Regarding to claim 14, the combination of Wang and Yu discloses all the limitations of claim 1 (as shown in the rejection above).
Wang in view of Yu further discloses:
An electronic device comprising a memory, a processor and a computer program stored on the memory and runnable on the processor, wherein the processor, when the program is executed by the processor, achieves a method according to claim 1 . (Wang, ¶[0005-0006]: “The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor;… According to another aspect of the present disclosure, a non-transitory computer-readable storage medium stored with computer instructions is provided, the computer instructions are configured to cause a computer to perform the method in any one embodiment of the present disclosure”)
Claim(s) 2 is rejected under 35 U.S.C. 103 as being unpatentable over Wang et al. (US 2023/0169351 A1) in view of YU (US 2023/0229905 A1) and further in view of Fleming et al. (US 2002/0152431 A1)
Regarding claim 2, the combination of Wang and Yu discloses all the limitations of claim 1 (as shown in the rejection above).
Wang in view of Yu fails to disclose:
monitoring whether heartbeat signals from the device nodes are received at default time intervals; and
in response to determining that a heartbeat signal transmitted by at least one of the device nodes is not received within a designated period, determining that the model training task for the target model is abnormal during execution, and
determining a device node that does not transmit a heartbeat signal within the designated period as a faulty node.
However, Fleming explicitly discloses:
monitoring whether heartbeat signals from the device nodes are received at default time intervals; and (Fleming, ¶[0004]: “According to the heartbeat technique, if a process does not receive a heartbeat from a remote process prior to the expiration of a predetermined length of time, i.e., the heartbeat timeout, the remote process is suspected to have failed. Corrective action, such as eliminating the suspected process, may thus be taken.”)
in response to determining that a heartbeat signal transmitted by at least one of the device nodes is not received within a designated period, determining that the model training task for the target model is abnormal during execution, and (Fleming, ¶[0008]: “process in the distributed system monitors a plurality of other processes in the distributed system via a network. If the process fails to receive a heartbeat from any one of the plurality of processes within a network failure time limit, the network in the distributed system is suspected of failing.”)
determining a device node that does not transmit a heartbeat signal within the designated period as a faulty node. (Fleming, ¶[0017]: “The distributed system 100 includes host 1, host 2 and host 3 executing process A, process B and process C, respectively. Processes A-C function to provide a service to a plurality of users via distributed system 100”, ¶[0018]: “process A monitors processes B and C by monitoring heartbeats transmitted on communication paths 110 and 130 from processes B and C, respectively. If the difference between a period of time to receive a heartbeat from process B and a period of time to receive a heartbeat from process C exceeds a process failure threshold, process B is suspected of failing. The process failure threshold may be a predetermined threshold or a threshold that can automatically adapt to varying network conditions.”)
The combination of Wang, Yu and Fleming are analogous art because they are in the same field of training time series data. Therefore, it would have been obvious to one of ordinary skill in the art before the effective filling date of the claimed invention, having the teachings of Wang, Yu and Fleming before them, to modify the teachings of Wang and Yu include the teachings of Fleming to provide low cost simplistic techniques for detecting network and process failures in a distributed system by taking corrective action immediately when failures are detected, which will help to minimize down-time for a service provided by the processes in the distributed system.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to AMY TRAN whose telephone number is (571)270-0693. The examiner can normally be reached Monday - Friday 7:30 am - 5:00 pm EST.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, David Yi can be reached at (571) 270-7519. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/AMY TRAN/Examiner, Art Unit 2126
/DAVID YI/Supervisory Patent Examiner, Art Unit 2126