DETAILED ACTION
This office action is responsive to the above identified application filed 1/15/2024. The application contains claims 1-10, 24, 26-32, all examined and rejected.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Priority
Acknowledgment is made of applicant's claim for foreign priority based on an application filed in United Kingdom on 7/15/2021. It is noted, however, that applicant has not filed a certified copy of the PCT/IB2021/056390 application as required by 37 CFR 1.55.
Information Disclosure Statement
The Information Disclosure Statement with references submitted 1/15/2024, has been considered and entered into the file.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-10, 24, 26-32 are rejected under 35 U.S.C. 101 because the claimed invention is directed to non-statutory subject matter. Claim 1 is rejected under 35 USC 101 because the claimed inventions are directed to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more.
While independent claims 1, and 24 are each directed to a statutory category, they recites a series of steps, which appears to be directed to an abstract idea (mental process, mathematical concept).
Claims 1-10, 24, 26-32 are rejected under 35 U.S.C. § 101 because the instant application is directed to non-patentable subject matter. Specifically, the claims are directed toward at least one judicial exception without reciting additional elements that amount to significantly more than the judicial exception. The rationale for this determination is in accordance with the guidelines of USPTO, applies to all statutory categories, and is explained in detail below.
When considering subject matter eligibility under 35 U.S.C. 101, (1) it must be determined whether the claim is directed to one of the four statutory categories of invention, i.e., process, machine, manufacture, or composition of matter. If the claim does fall within one of the statutory categories, (2a) it must then be determined whether the claim is directed to a judicial exception (i.e., law of nature, natural phenomenon, and abstract idea), and if so (2b), it must additionally be determined whether the claim is a patent-eligible application of the exception. If an abstract idea is present in the claim, any element or combination of elements in the claim must be sufficient to ensure that the claim amounts to significantly more than the abstract idea itself. Examples of abstract ideas include certain methods of organizing human activities; a mental processes; and mathematical concepts, (2019 PEG)
STEP 1.
Per Step 1, the claims are determined to include process, and machine as in independent Claim 1, and 24, and in the therefrom dependent claims. Therefore, the claims are directed to a statutory eligibility category.
At step 2A, prong 1, The invention is directed to identifying features within received data that could be an indication of the probability of occurrence of a machine failure based on analyzed historic data which is akin to Mental Process (see Alice), As such, the claims include an abstract idea. When considering the limitations individually and as a whole the limitations directed to the abstract idea are:
“generating a placement map for the ML model, wherein the placement map specifies, for each of the functional model parts, a mapping between the functional model part and at least one resource node of the system that is to execute the functional model part; identifying, from the placement map, a functional model part that is to be executed by the resource node” (Mental process, observation, evaluation and judgment).
The claim recites additional elements as
“A computer implemented method for facilitating execution of a Machine Learning, ML, model by a system of resource nodes, wherein the ML model comprises a plurality of functional model parts, the method, performed by a resource node of the system”, “executing the identified functional model part” (“Using a computer as a tool to perform a mental process”, MPEP 2106.04(a)(2)(III)(C)).
This judicial exception is not integrated into a practical application. The elements are recited at a high level of generality, i.e. a generic computing system performing generic functions including generic processing of data. Accordingly the additional elements do not integrate the abstract into a practical application because it does not impose any meaningful limits on practicing the abstract idea. Therefore the claims are directed to an abstract idea. (2019 Revised Patent Subject Matter Eligibility Guidance ("2019 PEG"). Thus, under Step 2A of the Mayo framework, the Examiner holds that the claims are directed to concepts identified as abstract.
STEP 2B.
Because the claims include one or more abstract ideas, the examiner now proceeds to Step 2B of the analysis, in which the examiner considers if the claims include individually or as an ordered combination limitations that are "significantly more" than the abstract idea itself. This includes analysis as to whether there is an improvement to either the "computer itself," "another technology," the "technical field," or significantly more than what is "well-understood, routine, or conventional" (WURC) in the related arts.
The instant application includes in Claim 1 additional steps to those deemed to be abstract idea(s).
When taken the steps individually, these steps are:
“A computer implemented method for facilitating execution of a Machine Learning, ML, model by a system of resource nodes, wherein the ML model comprises a plurality of functional model parts, the method, performed by a resource node of the system”, “executing the identified functional model part” (“Using a computer as a tool to perform a mental process”, MPEP 2106.05(f)(2)).
In the instant case, Claim 1 is directed to above mentioned abstract idea. Technical functions such as receiving, and extracting are common and basic functions in computer technology. The individual limitations are recited at a high level and do not provide any specific technology or techniques to perform the functions claimed.
In addition, when the claims are taken as a whole, as an ordered combination, the combination of steps does not add "significantly more" by virtue of considering the steps as a whole, as an ordered combination. The instant application, therefore, still appears only to implement the abstract idea to the particular technological environments using what is well-understood, routine, and conventional in the related arts. The steps are still a combination made to the abstract idea. The additional steps only add to those abstract ideas using well understood and conventional functions, and the claims do not show improved ways of, for example, an unconventional non-routine functions for analyzing model operations or updating the model that could then be pointed to as being "significantly more" than the abstract ideas themselves.
Moreover, Examiner was not able to identify any "unconventional" steps, which, when considered in the ordered combination with the other steps, could have transformed the nature of the abstract idea previously identified. The instant application, therefore, still appears to only implement the abstract ideas to the particular technological environments using what is well-understood, routine, and conventional (WURC) in the related arts.
Further, note that the limitations, in the instant claims, are done by the generically
recited computing devices. The limitations are merely instructions to implement the abstract idea on a computing device that is recited in an abstract level and require no more than a generic computing devices to perform generic functions.
. Claim 24 recites a system comprising “A resource node of a system of resource nodes, wherein the resource node is for facilitating execution of a Machine Learning, ML, model by the system, and wherein the ML model comprises a plurality of functional model parts, the resource node comprising processing circuitry” configured to perform the same method as set forth in claim 1, the added element of “A resource node of a system of resource nodes, wherein the resource node is for facilitating execution of a Machine Learning, ML, model by the system, and wherein the ML model comprises a plurality of functional model parts, the resource node comprising processing circuitry configured” do not transform the judicial exception into a practical application because they are tantamount to a mere instruction to apply the judicial exception to a generic computer. The additional elements are also not sufficient to amount to significantly more than the judicial exception because the action of implementing the method on a general purpose computer with at least one processor and at least one memory is tantamount to a mere instruction to apply the judicial exception to a computer.
Claim 24 is therefore rejected according to the same findings and rationale as provided above.
Independent claim 24 is the same analogy and rejected using similar analysis as claim 1.
CONCLUSION
It is therefore determined that the instant application not only represents an abstract idea identified as such based on criteria defined by the Courts and on USPTO examination guidelines, but also lacks the capability to bring about "Improvements to another technology or technical field" (Alice), bring about "Improvements to the functioning of the computer itself" (Alice), "Apply the judicial exception with, or by use of, a particular machine" (Bilski), "Effect a transformation or reduction of a particular article to a different state or thing" (Diehr), "Add a specific limitation other than what is well-understood, routine and conventional in the field" (Mayo), "Add unconventional steps that confine the claim to a particular useful application" (Mayo), or contain "Other meaningful limitations beyond generally linking the use of the judicial exception to a particular technological environment" (Alice), transformed a traditionally subjective process performed by humans into a mathematically automated process executed on computers (McRO), or limitations directed to improvements in computer related technology, including claims directed to software (Enfish).
The dependent claims, when considered individually and as a whole, likewise do not provide "significantly more" than the abstract idea for similar reasons as the independent claim.
claims 2 disclose “the resource node comprises a managing agent and at least one unit of resource, the resource comprising at least one of: storage resource; computational resource; networking resource” (data description, which is directed to generally linking the use of a judicial exception to a particular technological environment or type or source of data or field of use MPEP 2106.05(h)), It does not integrate the abstract idea into a practical application and did not add significantly more to the abstract idea, claims 3 disclose “configuring resources of the resource node for execution of the identified functional model part” (generally linking the use of a judicial exception to a particular technological environment or type or source of data or field of use MPEP 2106.05(h)), It does not integrate the abstract idea into a practical application and did not add significantly more to the abstract idea, claims 4 disclose “ updating the placement map” (mental process); identifying, from the updated placement map, a new functional model part that is to be executed by the resource node (mental process); ceasing to execute the previously identified functional model part; and executing the functional model part identified from the updated placement map” (data description, which is directed to generally linking the use of a judicial exception to a particular technological environment or type or source of data or field of use MPEP 2106.05(h)), It does not integrate the abstract idea into a practical application and did not add significantly more to the abstract idea, claims 5 disclose “reconfiguring resource of the resource node for execution of the functional model part identified from the updated placement map) (data description, which is directed to generally linking the use of a judicial exception to a particular technological environment or type or source of data or field of use MPEP 2106.05(h)), It does not integrate the abstract idea into a practical application and did not add significantly more to the abstract idea, claims 6 disclose “the method is for facilitating execution of a plurality of ML models, and wherein the functional model part identified from the updated placement map and the previously identified functional model part are functional parts of different ML models”, It does not integrate the abstract idea into a practical application and did not add significantly more to the abstract idea, claims 7 disclose “identifying, from the placement map, a functional model part that is to be executed by the resource node (mental process) comprises identifying a functional model part that orchestrates execution of the ML model (mental process); and wherein: executing the identified functional model part comprises writing values to and reading values from other resource nodes in the system in accordance with the placement map” (data description, which is directed to generally linking the use of a judicial exception to a particular technological environment or type or source of data or field of use MPEP 2106.05(h)), It does not integrate the abstract idea into a practical application and did not add significantly more to the abstract idea, claims 8 disclose “identifying, from the placement map, a functional model part that is to be executed by the resource node comprises identifying a functional model part that orchestrates execution of a deployment instance of the ML model (mental process); wherein: executing the identified functional model part comprises reading values for trainable parameters of the ML model from a shared memory; and wherein: the resource node has read only access to the shared memory, and a resource node of the system that is orchestrating execution of a training instance of the ML model has read and write access to the shared memory” (data description, which is directed to generally linking the use of a judicial exception to a particular technological environment or type or source of data or field of use MPEP 2106.05(h)), It does not integrate the abstract idea into a practical application and did not add significantly more to the abstract idea, claims 9 disclose “updating the placement map (mental process); and continuing to execute the identified functional model part by writing values to and reading values from other resource nodes in the system in accordance with the updated placement map” (data description, which is directed to generally linking the use of a judicial exception to a particular technological environment or type or source of data or field of use MPEP 2106.05(h)), It does not integrate the abstract idea into a practical application and did not add significantly more to the abstract idea, claims 10 disclose “the updated placement map specifies addition or removal of a functional part of the model; the method further comprising: continuing to execute the identified functional model part by reading values for trainable parameters of the ML model from the shared memory (data description, which is directed to generally linking the use of a judicial exception to a particular technological environment or type or source of data or field of use MPEP 2106.05(h)).
The dependent claims which impose additional limitations also fail to claim patent eligible subject matter because the limitations cannot be considered statutory. The dependent claim(s) have been examined individually and in combination with the preceding claims, however they do not cure the deficiencies of claim 1 ; where all claims are directed to the same abstract idea, "addressing each claim of the asserted patents [is] unnecessary." Content Extraction &. Transmission LLC v, Wells Fargo Bank, Natl Ass'n, 776 F.3d 1343, 1348 (Fed. Cir. 2014). If applicant believes the dependent claims are directed towards patent eligible subject matter, they are invited to point out the specific limitations in the claim that are directed towards patent eligible subject matter. Claims for the other statutory classes are similarly analyzed.
For at least these reasons, the claimed inventions of each of dependent claims 2-10, and claims 26-32 that are similar in scope to claims 2-10 are directed or indirect to a judicial exception (i.e., a law of nature, a natural phenomenon, or an abstract idea) without significantly more and are rejected under 35 USC 101.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(1) the claimed invention was patented, described in a printed publication, or in public use, on sale, or otherwise available to the public before the effective filing date of the claimed invention.
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 1-5, 7, 9-10 are rejected under 35 U.S.C. 102(a)(1) and 35 U.S.C. 102(a)(2) as being anticipated by John et al. [US 2019/0378016 A1, hereinafter John].
With regard to Claim 1,
John teach a computer implemented method for facilitating execution of a Machine Learning, ML, model by a system of resource nodes (¶2, “Aspects of the present disclosure are directed toward a computer-implemented method comprising generating a model mapping table (MMT) storing information regarding respective portions of a deep learning model distributed amongst a plurality of interconnected host nodes”, ¶¶29-30, “ Network architecture 100 can comprise a large model manager (LMM) 102 communicatively coupled to a large model pooler (LMP) 104 and a model mapping table (MMT) 120. The LMM 102 can manage training the deep learning model based on information stored in MMT 120 and allocations made to hosts 106, CPU memories 108, CPUs 110, GPU memories 112, and/or GPUs 114 by the LMP 104”), wherein the ML model comprises a plurality of functional model parts (¶40, “a model portion (e.g., model portion X 116A, 116B, and/or 116C) can comprise individual layers, error functions (e.g., gradients), parameters (e.g., variables, weights, biases, etc.), and/or datasets associated with a deep learning model. In some embodiments, a model portion can comprise a single layer of the deep learning model, a portion of a single layer of the deep learning model, data associated with an operation of the deep learning model, or data associated with a portion of an operation of the deep learning model”), the method, performed by a resource node of the system (¶82, “ LMM 500 performs any of the methods described in FIGS. 2-4”, ¶83, “The LMM 500 includes a memory 525, storage 530, an interconnect (e.g., BUS) 520, one or more CPUs 505 (also referred to as processors 505 herein), an I/O device interface 510, I/O devices 512, and a network interface 515”, ¶89), comprising:
generating a placement map for the ML model (¶58, “In operation 208, the LMM can generate a deep learning model mapping table (MMT)”, “The LMM can populate the MMT with information regarding the LMM, the host nodes, the LMP, and/or the deep learning model”), wherein the placement map specifies, for each of the functional model parts, a mapping between the functional model part and at least one resource node of the system that is to execute the functional model part (¶42, “MMT 120 can be used to store information regarding model portions (e.g., model portion X 116A, 116B, and 116C), CPU memories 108, CPUs 110, GPU memories 112, GPUs 114, and/or hosts 106. MMT 120 can store pointers 122, layer identifiers 124, ranks 126, memory handles 128, memory offsets 130, metadata 132, and/or flags 134”, ¶58, “the MMT stores pointers, layer identifiers, process ranks, memory handles, memory offsets, metadata, and/or flags for respective portions of the deep learning model distributed amongst the host nodes”, ¶43, “Pointers 122 can comprise pointers indicating a host 106, CPU memory 108, CPU 110, GPU memory 112, and/or GPU 114 associated with a respective portion of the deep learning model”, ¶45, “Ranks 126 can comprise respective process ranks associated with a process to be implemented by a requesting GPU 114 for a portion of the deep learning model”, ¶62, “In operation 304, the LMM can allocate the required size from the LMP … The LMM can create an entry in the MMT (e.g., MMT 120 of FIG. 1) having a data pointer, a layer identifier, a rank of the process requesting the allocation, a remote memory handle, a remote memory offset, metadata, and/or flags for each respective portion of the deep learning model”);
identifying, from the placement map, a functional model part that is to be executed by the resource node (¶64, “In operation 308, the LMM can query the MMT to identify a host node where the requested data is located”, ¶51, “LMP 104 uses MMT 120 to identify that model portion X 116A resides in CPU 1 memory 108A”, ¶23, “In some embodiments of the present disclosure, a large model manager (LMM) manages an interconnected cluster of host nodes using a large model pooler (LMP) and a model mapping table (MMT) to transparently train a large deep learning model using model parallelism”); and
executing the identified functional model part (¶31, “ LMM 102 stores MMT 120 and contains functionality equivalent to LMP 104”, ¶87, “The instructions 560 are processor executable instructions for executing any portion of, any combination of, or all of the methods previously discussed in FIGS. 2-4”, ¶84, Each CPU 505 retrieves and executes programming instructions stored in the memory 525 or storage 530”, ¶51, ¶82, “In some embodiments, LMM 500 provides instructions for one or more of the methods described in FIGS. 2-4 to a client machine such that the client machine executes the method, or a portion of the method, based on the instructions provided by the LMM 500”).
With regard to Claim 2,
John teach the method of claim 1, wherein the resource node comprises a managing agent (Fig. 5, ¶30, “LMP 104 can comprise pooling functionality capable of organizing and deploying a set of computational resources”, ¶31, “ LMM 102 stores MMT 120 and contains functionality equivalent to LMP 104”) and at least one unit of resource, the resource comprising at least one of: storage resource; computational resource; networking resource (Fig. 5, 505, 515, 525, 530, ¶83).
With regard to Claim 3,
John teach the method of claim 1, further comprising: configuring resources of the resource node for execution of the identified functional model part (¶56, “In operation 204, the LMM can establish MPI communication across the list of host nodes”, ¶57, “In operation 206, the LMM can initialize a large model pooler (LMP) by registering with a handle of a memory region (e.g., a window region) on all host nodes in the list of host nodes”, ¶62, “In operation 304, the LMM can allocate the required size from the LMP (e.g., LMP 104 of FIG. 1) for respective portions of the deep learning model”, ¶30, “LMP 104 can comprise pooling functionality capable of organizing and deploying a set of computational resources”, ¶31, “ LMM 102 stores MMT 120 and contains functionality equivalent to LMP 104”, ¶66, “In operation 312, the LMM can copy the requested data from the requesting host node to the memory associated with the requesting GPU (e.g., GPU memory 112). Operation 312 can comprise creating a working copy of the requested data”).
With regard to Claim 4,
John teach the method of claim 1,further comprising:
updating the placement map (¶68, “In operation 316, the LMM can copy updates from the LMP to the MMT in response to performing processing at the requesting GPU”, ¶23, “The MMT can be updated once any allocation is made”);
identifying, from the updated placement map, a new functional model part that is to be executed by the resource node (Fig. 3, ¶63, “In operation 306, the LMM can receive a request for data relevant to the deep learning model by a requesting GPU (e.g., GPU 114 of FIG. 1) of a requesting host node (e.g., host 106 of FIG. 1) for forward propagation and/or backpropagation of a portion of the deep learning model”, ¶64, “In operation 308, the LMM can query the MMT to identify a host node where the requested data is located”, ¶70, “Operations 306-318 can occur any number of times for any number of portions of the deep learning model until the deep learning model is fully trained”);
ceasing to execute the previously identified functional model part (¶69, “In operation 318, the LMM can relinquish the data pointer for the requested data of the deep learning model once the forward propagation and/or backpropagation for the requested data is complete”); and
executing the functional model part identified from the updated placement map (¶52, “The aforementioned example process can occur any number of times for any number of model portions of a deep learning model until the deep learning model is fully trained”, ¶¶69-70, “Operations 306-318 can occur any number of times for any number of portions of the deep learning model until the deep learning model is fully trained”).
With regard to Claim 5,
John teach the method of claim 4, further comprising:
reconfiguring resource of the resource node for execution of the functional model part identified from the updated placement map (¶66, “In operation 312, the LMM can copy the requested data from the requesting host node to the memory associated with the requesting GPU (e.g., GPU memory 112). Operation 312 can comprise creating a working copy of the requested data (e.g., copy model portion X 116C of FIG. 1)”, ¶68, “In some embodiments, the LMP identifies a beneficial location for the updated data to be stored in the distributed network architecture”, ¶51, “ In some embodiments, LMP 104 transfers the synchronized model portion X 116B to a different host 106 for efficient storage (and subsequently updates MMT 120)”, ¶70).
With regard to Claim 7,
John teach the method of claim 1,wherein:
identifying, from the placement map, a functional model part that is to be executed by the resource node comprises identifying a functional model part that orchestrates execution of the ML model (¶23, “In some embodiments of the present disclosure, a large model manager (LMM) manages an interconnected cluster of host nodes using a large model pooler (LMP) and a model mapping table (MMT) to transparently train a large deep learning model using model parallelism”, “The LMM can manage the deep learning model distribution using the LMP and the MMT. The LMP can allocate portions of the deep learning model (e.g., layer, gradients, parameters, datasets, etc.) from a CPU memory on one host node to an available GPU memory on the same or a different host node for processing. Such allocations can be based on information in the MMT”, ¶29, “he LMM 102 can manage training the deep learning model based on information stored in MMT 120 and allocations made to hosts 106, CPU memories 108, CPUs 110, GPU memories 112, and/or GPUs 114 by the LMP 104”, ¶51, “LMM 102 instructs LMP 104 to train the deep learning model, including model portion X 116A”); and
wherein: executing the identified functional model part comprises writing values to and reading values from other resource nodes in the system (¶51, “ LMP 104 uses an MPI RMA communication protocol to transfer 118A model portion X 116B into CPU 2 memory 108B and to then generate and store 118B copy model portion X 116C in GPU 2 memory“, ¶56, “In some embodiments, MPI communication comprises a one-way messaging protocol that can read from and/or write to selected portions (e.g., window regions) of different host nodes without the involvement of the other host nodes”, ¶65, “In operation 310, the LMM can transfer (e.g., copy, transmit, replicate, etc.) the requested data from the identified host node to the requesting host node (e.g., using MPI RMA)”, ¶66, “In operation 312, the LMM can copy the requested data from the requesting host node to the memory associated with the requesting GPU “, ¶2, “The training can further comprise synchronizing the first copy of the first portion of the deep learning model with the first portion of the deep learning model in response to performing processing”) in accordance with the placement map (¶64, “In operation 308, the LMM can query the MMT to identify a host node where the requested data is located. The identified host node can be the requesting host node or a different host node”, ¶51, “ LMP 104 uses MMT 120 to identify that model portion X 116A resides in CPU 1 memory 108A”, ¶46, “Memory handles 128 can comprise a reference to a resource associated with a portion of the deep learning model. In some embodiments, memory handles 128 indicate a window of available memory configured for MPI RMA communication in a CPU memory 108, GPU memory 112, or a different memory associated with a host”, ¶47, “Memory offsets 130 can be used to indicate locations of portions of the deep learning model. Memory offsets 130 can indicate an offset relative to a window of accessible memory in any CPU memory 108, GPU memory 112, or other memory associated with a host”).
With regard to Claim 9,
John teach the method of claim 7, further comprising:
updating the placement map (¶23, “The MMT can be updated once any allocation is made”, ¶68, “In operation 316, the LMM can copy updates from the LMP to the MMT in response to performing processing at the requesting GPU”, ¶51, “ LMP 104 updates MMT 120 with the copy model portion X 116C on GPU 2 memory”); and
continuing to execute the identified functional model part by writing values to and reading values from other resource nodes in the system (¶70, “Operations 306-318 can occur any number of times for any number of portions of the deep learning model until the deep learning model is fully trained”, ¶52, “The aforementioned example process can occur any number of times for any number of model portions of a deep learning model until the deep learning model is fully trained”, ¶56, “In some embodiments, MPI communication comprises a one-way messaging protocol that can read from and/or write to selected portions (e.g., window regions) of different host nodes without the involvement of the other host nodes”, ¶65, “In operation 310, the LMM can transfer (e.g., copy, transmit, replicate, etc.) the requested data from the identified host node to the requesting host node (e.g., using MPI RMA)”, ¶66, “In operation 312, the LMM can copy the requested data from the requesting host node to the memory associated with the requesting GPU”) in accordance with the updated placement map (¶64, “In operation 308, the LMM can query the MMT to identify a host node where the requested data is located. The identified host node can be the requesting host node or a different host node”, ¶68, “ LMP identifies a beneficial location for the updated data to be stored in the distributed network architecture”, ¶51, “ LMP 104 transfers the synchronized model portion X 116B to a different host 106 for efficient storage (and subsequently updates MMT 120)”).
With regard to Claim 10,
John teach the method of claim 9,
wherein the updated placement map specifies addition or removal of a functional part of the model (¶40, “In various embodiments, a model portion (e.g., model portion X 116A, 116B, and/or 116C) can comprise individual layers, error functions (e.g., gradients), parameters (e.g., variables, weights, biases, etc.), and/or datasets associated with a deep learning model”, ¶43, “Pointers 122 can comprise pointers indicating a host 106, CPU memory 108, CPU 110, GPU memory 112, and/or GPU 114 associated with a respective portion of the deep learning model”, ¶62, “The LMM can create an entry in the MMT (e.g., MMT 120 of FIG. 1) having a data pointer, a layer identifier, a rank of the process requesting the allocation, a remote memory handle, a remote memory offset, metadata, and/or flags for each respective portion of the deep learning model”, ¶69, “In operation 318, the LMM can relinquish the data pointer for the requested data of the deep learning model once the forward propagation and/or backpropagation for the requested data is complete”, ¶68, “In operation 316, the LMM can copy updates from the LMP to the MMT in response to performing processing at the requesting GPU”, ¶49, “Flags 134 can indicate functions associated with portions of the deep learning model, such as, but not limited to reuse data functions, recompute functions, and/or other functions”);
the method further comprising:
continuing to execute the identified functional model part (¶70, “Operations 306-318 can occur any number of times for any number of portions of the deep learning model until the deep learning model is fully trained “, ¶52) by reading values for trainable parameters of the ML model from the shared memory (¶40, “In various embodiments, a model portion (e.g., model portion X 116A, 116B, and/or 116C) can comprise individual layers, error functions (e.g., gradients), parameters (e.g., variables, weights, biases, etc.), and/or datasets associated with a deep learning model “, ¶56, “In some embodiments, MPI communication comprises a one-way messaging protocol that can read from and/or write to selected portions (e.g., window regions) of different host nodes without the involvement of the other host nodes”, ¶57, “In operation 206, the LMM can initialize a large model pooler (LMP) by registering with a handle of a memory region (e.g., a window region) on all host nodes in the list of host nodes”, ¶46, “Memory handles 128 can comprise a reference to a resource associated with a portion of the deep learning model. In some embodiments, memory handles 128 indicate a window of available memory configured for MPI RMA communication in a CPU memory 108, GPU memory 112, or a different memory associated with a host 106”).
With regard to Claim 24,
Claim 24 is similar in scope to claim 1. Therefore it is rejected under similar rationale. John further teach resource node of a system of resource nodes, wherein the resource node is for facilitating execution of a Machine Learning, ML, model by the system, and wherein the ML model comprises a plurality of functional model parts, the resource node comprising processing circuitry configured (¶¶3-4, “a system comprising a processor and a computer-readable storage medium storing program instructions for deep learning model training which, when executed by the processor, are configured to cause the processor to perform a method comprising generating a model mapping table (MMT) storing information regarding respective portions of a deep learning model distributed amongst a plurality of interconnected host nodes. Respective host nodes can comprise at least one central processing unit (CPU), at least one CPU memory, at least one graphics processing unit (GPU), and at least one GPU memory …”, ¶¶29-30, “ Network architecture 100 can comprise a large model manager (LMM) 102 communicatively coupled to a large model pooler (LMP) 104 and a model mapping table (MMT) 120. The LMM 102 can manage training the deep learning model based on information stored in MMT 120 and allocations made to hosts 106, CPU memories 108, CPUs 110, GPU memories 112, and/or GPUs 114 by the LMP 104”, ¶¶32-35, ¶¶83-84, ¶87)
With regard to Claim 26,
Claim 26 is similar in scope to claim 3. Therefore it is rejected under similar rationale.
With regard to Claim 27,
Claim 27 is similar in scope to claim 4. Therefore it is rejected under similar rationale.
With regard to Claim 28,
Claim 28 is similar in scope to claim 5. Therefore it is rejected under similar rationale.
With regard to Claim 29,
Claim 29 is similar in scope to claim 7. Therefore it is rejected under similar rationale.
With regard to Claim 31,
Claim 31 is similar in scope to claim 9. Therefore it is rejected under similar rationale.
With regard to Claim 32,
Claim 32 is similar in scope to claim 10. Therefore it is rejected under similar rationale.
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim 6 is rejected under 35 U.S.C. 103 as being unpatentable over John et al. [US 2019/0378016 A1, hereinafter John] in view of Hughes et al. [US 2021/0149680 A1, hereinafter Hughes].
With regard to Claim 6,
John teach the method of claim 4, wherein the method is for facilitating execution of a plurality of ML models (¶79, “the deep learning model can be associated with cybersecurity (e.g., operation 404). The input data can comprise log data, network data, firewall data, or other data from one or more computing devices (e.g., operation 406). The output data can be a malware notification based on the deep learning model identifying malware in the input data (e.g., operation 408).”, ¶80, “As another example, the deep learning model can be associated with quality control for a manufacturing and assembly line (e.g., operation 404). The input data can be a series of measurements from a series of parts”, ¶86, “The deep learning model 536 can be any deep learning model (e.g., ANN, DNN, CNN, etc.), or portion thereof”, ¶17, ¶96, “Resource pooling: the provider's computing resources are pooled to serve multiple consumers using a multi-tenant model, with different physical and virtual resources dynamically assigned and reassigned according to demand”, ¶31, “hosts 106 comprise virtual resources provisioned in a cloud computing environment”), and wherein the functional model part identified from the updated placement map and the previously identified functional model part are functional parts of different ML models (¶23, “The MMT can be updated once any allocation is made”, ¶68, “In operation 316, the LMM can copy updates from the LMP to the MMT in response to performing processing at the requesting GPU”, ¶42, “MMT 120 can be used to store information regarding model portions (e.g., model portion X 116A, 116B, and 116C), CPU memories 108, CPUs 110, GPU memories 112, GPUs 114, and/or hosts 106. MMT 120 can store pointers 122, layer identifiers 124, ranks 126, memory handles 128, memory offsets 130, metadata 132, and/or flags”, ¶48, “Metadata 132 can comprise data types (e.g., parameter, gradient, temperature data, etc.) and/or data characteristics (e.g., times, origins, etc.)”, ¶69, “In operation 318, the LMM can relinquish the data pointer for the requested data of the deep learning model once the forward propagation and/or backpropagation for the requested data is complete”, ¶70, “Operations 306-318 can occur any number of times for any number of portions of the deep learning model until the deep learning model is fully trained”).
John does not explicitly teach
Hughes teach previously identified functional model part are functional parts of different ML models (¶121, “A set of registers 445 store context data for threads executed by the graphics processing engines 431-432, N and a context management circuit 448 manages the thread contexts. For example, the context management circuit 448 may perform save and restore operations to save and restore contexts of the various threads during contexts switches (e.g., where a first thread is saved and a second thread is restored so that the second thread can be execute by a graphics processing engine) …”, ¶122, “The graphics accelerator module 446 may be dedicated to a single application executed on the processor 407 or may be shared between multiple applications. Optionally, a virtualized graphics execution environment is provided in which the resources of the graphics processing engines 431-432, N are shared with multiple applications, virtual machines (VMs), or containers. The resources may be subdivided into “slices” which are allocated to different VMs and/or applications based on the processing requirements and priorities associated with the VMs and/or applications”, ¶142, “There are two programming models where the graphics acceleration module 446 is shared by multiple processes and partitions: time-sliced shared and graphics directed shared”, ¶¶201-202)
John and Hughes are analogous art to the claimed invention because they are from a similar field of endeavor of GPU accelerated distributed training across interconnected nodes. Thus, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify John resulting in resolutions as disclosed by Hughes with a reasonable expectation of success.
One of ordinary skill in the art would be motivated to modify John as described above to increase processing efficiency by allowing subdividing resources into “slices” which are allocated to different VMs and/or applications based on the processing requirements and priorities associated with the VMs and/or applications (Hughes, ¶122). This is simply Combining prior art elements according to known methods to yield predictable results; use of known technique to improve similar devices (methods, or products) in the same way, and applying a known technique to a known device (method, or product) ready for improvement to yield predictable results (MPEP 2143).
Claims 8 and 30 are rejected under 35 U.S.C. 103 as being unpatentable over John et al. [US 2019/0378016 A1, hereinafter John] in view of Hassanzadeh et al. [US 2022/0414661 A1, hereinafter Hassanzadeh].
With regard to Claim 8,
John teach the method of claim 1,wherein:
identifying, from the placement map, a functional model part that is to be executed by the resource node comprises identifying a functional model part that orchestrates execution of a deployment instance of the ML model (¶23, “In some embodiments of the present disclosure, a large model manager (LMM) manages an interconnected cluster of host nodes using a large model pooler (LMP) and a model mapping table (MMT) to transparently train a large deep learning model using model parallelism”, ¶29, “The LMM 102 can manage training the deep learning model based on information stored in MMT 120 and allocations made to hosts 106, CPU memories 108, CPUs 110, GPU memories 112, and/or GPUs 114 by the LMP 104”, ¶76, “In operation 406, the LMM can input data into the trained deep learning model”, ¶77, “In operation 408, the LMM can receive output based on the input data provided to the trained deep learning model”, ¶71, “outputting a trained deep learning model comprises utilizing the trained deep learning model by inputting new data into the trained learning model and receiving output data as a result of inputting the new data”);
wherein: executing the identified functional model part comprises reading values for trainable parameters of the ML model from a shared memory (¶40, “In various embodiments, a model portion (e.g., model portion X 116A, 116B, and/or 116C) can comprise individual layers, error functions (e.g., gradients), parameters (e.g., variables, weights, biases, etc.), and/or datasets associated with a deep learning model”, ¶46, “memory handles 128 indicate a window of available memory configured for MPI RMA communication in a CPU memory 108, GPU memory 112, or a different memory associated with a host”, ¶48, “Metadata 132 can comprise data types (e.g., parameter, gradient, temperature data, etc.) and/or data characteristics”, ¶56, “In some embodiments, MPI communication comprises a one-way messaging protocol that can read from and/or write to selected portions (e.g., window regions) of different host nodes without the involvement of the other host nodes”, ¶57, “In operation 206, the LMM can initialize a large model pooler (LMP) by registering with a handle of a memory region (e.g., a window region) on all host nodes in the list of host nodes”).
John does not explicitly teach the resource node has read only access to the shared memory, and a resource node of the system that is orchestrating execution of a training instance of the ML model has read and write access to the shared memory.
Hassanzadeh teach orchestrates execution of a deployment instance of the ML model (¶64, “ In some implementations, the server 340 includes or supports an endpoint node 344 that can be configured to provide ML services based on a trained ML model”, ¶71, “ uploading the fraud predictor parameter set may configure the endpoint node 344 to receive a fraud prediction request from one of the clients 310, 320, or 330, or a user device (e.g., a client that subscribes to the ML service but that does not participate in the training process) and to transmit a prediction to the requester”, ¶73, “ the server 440 includes or supports an endpoint node 446 that can be configured to provide ML services based on a trained ML model”)
wherein: executing the identified functional model part comprises reading values for trainable parameters of the ML model from a shared memory (¶37, “the structural parameters may include a number of layers, a number of hidden layers, a number of nodes per layer or per type of layer, a number of input nodes, a number of output nodes, a number of hidden nodes, a number of connections per node, weights of connections, activation functions associated with nodes, or the like”, ¶73, “The server 440 may include an aggregation storage location 442 that is configured to store parameter sets corresponding to local ML models trained by the clients and a distribution storage location 444 that is configured to store parameter sets corresponding to global ML models to be distributed to the clients or to other entities”, ¶64, “The server 340 may include a storage location 342 that is configured to store ML models for use by the clients. Because the ML models are the same for each client in federated learning, the ML models may be referred to as global ML models. The storage location 342 may be accessible to the clients 310, 320, and 330”)
wherein: the resource node has read only access to the shared memory (¶74, “ the clients 410, 420, and 430 may be granted write-only access to the aggregation storage location 442 and read-only access to the distribution storage location 444, the server-side executable file package 404 may be granted read-only access to the aggregation storage location 442 and write-only access to the distribution storage location 444, and the endpoint node 446 may be granted read-only access to the distribution storage location 444”) and a resource node of the system that is orchestrating execution of a training instance of the ML model has read and write access to the shared memory (¶65, “the server 340 initializes a global ML model (e.g., an initial ML model) and shares copies of the model with participating clients … The trained local ML models are collected and aggregated by the server 340 to construct a new global ML model (e.g., an aggregated ML model) for deployment”, “The server-side executable file package 304 may be configured with read and write access to the storage location 342 and may include one or more ML libraries and any other files needed to support initialization, aggregation, and validation of ML models using federated learning”, ¶75, “ the server 440 write a fraud predictor parameter set corresponding to the fraud prediction model to the distribution storage location 444 for deployment to the endpoint node 446, one or more of the clients 410, 420, and 430, other clients or users, or a combination thereof”).
John and Hassanzadeh are analogous art to the claimed invention because they are from a similar field of endeavor of distributed machine learning across networked nodes with server side storage of model parameter sets. Thus, it would have been obvious to a person of ordinary skill in the art before the effective filing date of the claimed invention to modify John resulting in resolutions as disclosed by Hughes with a reasonable expectation of success.
One of ordinary skill in the art would be motivated to modify John as described above to use fewer storage locations at the clients and an additional storage location at the server by leveraging access controls (Hassanzadeh, ¶74).
This is simply Combining prior art elements according to known methods to yield predictable results; use of known technique to improve similar devices (methods, or products) in the same way, and applying a known technique to a known device (method, or product) ready for improvement to yield predictable results (MPEP 2143).
With regard to Claim 30,
Claim 30 is similar in scope to claim 8. Therefore it is rejected under similar rationale.
Conclusion
The prior art made of record and not relied upon is considered pertinent to the applicant’s disclosure.
US Patent Application Publication No. 20220292303 filed by CAO et al. that disclose a method involves determining multi-computing resource configurations used to perform machine learning model training jobs. The computing resource configuration is provided with a first tuple including numbers of worker nodes and parameter server nodes. A second tuple is provided with resource allocations for the worker nodes and the parameter server nodes. A machine learning training job is executed by using the first computing resource configuration. A first set of values associated with the first tuple during executing the machine learning training job. Resource usage of the worker nodes and parameter server nodes caused by a second set of values associated with the second tuple is monitored.
Examiner has pointed out particular references contained in the prior arts of record in the body of this action for the convenience of the applicant. Although the specified citations are representative of the teachings in the art and are applied to the specific limitations within the individual claim, other passages and Figures may apply as well. It is respectfully requested from the applicant, in preparing the response, to consider fully the entire references as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior arts or disclosed by the examiner. It is noted that any citation to specific pages, columns, figures, or lines in the prior art references any interpretation of the references should not be considered to be limiting in any way. A reference is relevant for all it contains and may be relied upon for all that it would have reasonably suggested to one having ordinary skill in the art. In re Heck, 699 F.2d 1331-33, 216 USPQ 1038-39 (Fed. Cir. 1983) (quoting In re Lemelson, 397 F.2d 1006, 1009, 158 USPQ 275, 277 (CCPA 1968)).
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MOHAMED ABOU EL SEOUD whose telephone number is (303)297-4285. The examiner can normally be reached Monday-Thursday 9:00am-6:00pm MT.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Michelle Bechtold can be reached at (571) 431-0762. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MOHAMED ABOU EL SEOUD/Primary Examiner, Art Unit 2148