Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Objections
Claim 4 is objected to because of the following informalities: “The method of claim 3” should read “The method of claim 1”. Appropriate correction is required.
Claim 21 is objected to because of the following informalities: “and the initialization of the second ML model includes allocating a second portion of the memory and selecting a first processing core on which to run the portion of the first ML model” should read “and the initialization of the second ML model includes allocating a second portion of the memory and selecting a second processing core on which to run the portion of the second ML model”. Appropriate correction is required.
Claim Rejections - 35 USC § 101
35 U.S.C. 101 reads as follows:
Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title.
Claims 1-2, 4-19, and 21-22 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Regarding Claim 1,
Claim 1 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 1 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“allocating a first portion of a memory and selecting a first processing core on which to run the portion of the first ML model”
“allocating a second portion of the memory and selecting a second processing core on which to run the portion of the second ML model”
“determining, based on the synchronization information, a time to run the first ML model”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)) and insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“running the first ML model at the time such that the first ML model and the second ML model run concurrently, wherein the running of the first ML model at the time comprises inserting a delay between the initialization of the first ML model and the running of the first ML model”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception for the abstract ideas. See MPEP 2106.05(f).
The limitations:
“receiving an indication to run a first machine learning (ML) model”
“receiving synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: a minimum duration between initialization of the first ML model and running of a portion of the first ML model”
“receiving synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: … a minimum duration between initialization of the second ML model and running of a portion of the second ML model”
“receiving synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: …a minimum duration between running of the portion of the second ML model and the portion of the first ML model”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply” and “insignificant extra-solution activity”. Further, the receiving an indication limitation recites the well-understood, routine, and conventional activity of receiving or transmitting data over a network. MPEP 2106.05(d)(II); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network). Additionally, the receiving synchronization information limitations recite the well-understood, routine, and conventional activity of storing and retrieving information in memory. MPEP 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Mere instructions to apply an exception and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 2,
Claim 2 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 2 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 1.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the synchronization information includes: an indication of the minimum duration between running of the portion of the second ML model and the portion of the first ML model”
“wherein the synchronization information includes: ...an indication of the first ML model”
“wherein the synchronization information includes: ...an indication of a core to run the first ML model”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 4,
Claim 4 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 4 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 3.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the delay is based on a callback function or a parallel thread”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 5,
Claim 5 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 5 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“wherein the determining of the time to run the first ML model comprises determining whether to insert a delay between layers of the first ML model”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 1.
Step 2B Analysis: See corresponding analysis of claim 1.
Regarding Claim 6,
Claim 6 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 6 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“determining, based on the synchronization information a time to run the first ML model, wherein the determining of the time to run the first ML model comprises determining a difference between an expected time to complete running the first ML model and a current time”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)) and insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“running the first ML model at the time such that the first ML model and the second ML model run concurrently, wherein beginning the run of the first ML model is based on the difference”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception for the abstract ideas. See MPEP 2106.05(f).
The limitations:
“receiving an indication to run a first machine learning (ML) model”
“receiving synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: a minimum duration between initialization of the first ML model and running of a portion of the first ML model”
“receiving synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: …a minimum duration between initialization of the second ML model and running of a portion of the second ML model”
“receiving synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: …a minimum duration between running of the portion of the second ML model and the portion of the first ML model”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply” and “insignificant extra-solution activity”. Further, the receiving an indication limitation recites the well-understood, routine, and conventional activity of receiving or transmitting data over a network. MPEP 2106.05(d)(II); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network). Additionally, the receiving synchronization information limitations recite the well-understood, routine, and conventional activity of storing and retrieving information in memory. MPEP 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Mere instructions to apply an exception and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 7,
Claim 7 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 7 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“adjusting a next expected time to run the first ML model based on the difference”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 6.
Step 2B Analysis: See corresponding analysis of claim 6.
Regarding Claim 8,
Claim 8 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 8 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 6.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)).
The limitations:
“removing a delay of the running of the first ML model based on the difference”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception. See MPEP 2106.05(f).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply”. The claim is not patent eligible.
Regarding Claim 9,
Claim 9 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 9 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“simulate running the set of ML models …to determine resources utilized by running the set of ML models and to determine timing information”
“determine, based on the simulation: a minimum delay between initialization of a first ML model of the set of ML models and running a portion of the first ML model”
“determine, based on the simulation: …a minimum delay between initialization of a second ML model of the set of ML models and running a portion of the second ML model”
“determine, based on the simulation: …a minimum delay between running the portion of the first ML model and the portion of the second ML model”
“generate synchronization information to run the first ML model and the second ML model in parallel based on the determining”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)) and insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“A non-transitory program storage device comprising instructions stored thereon to cause one or more processors to”
“on a target hardware”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception for the abstract ideas. See MPEP 2106.05(f).
The limitations:
“receive a set of machine learning (ML) models”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply” and “insignificant extra-solution activity”. Additionally, the receiving limitation recites the well-understood, routine, and conventional activity of storing and retrieving information in memory. MPEP 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Mere instructions to apply an exception and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 10,
Claim 10 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 10 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 9.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the target hardware includes at least two cores for executing the set of ML models and wherein the synchronization information includes timing information for coordinating execution of the set of ML models across the at least two cores”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 11,
Claim 11 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 11 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 10.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the synchronization information includes: an indication of the minimum delay between running the portion of the first ML model and the portion of the second ML model”
“wherein the synchronization information includes: …an indication of the second ML model”
“wherein the synchronization information includes: …an indication of a core of the target hardware to run the second ML model”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 12,
Claim 12 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 12 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 10.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein the synchronization information is organized in a lookup table”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 13,
Claim 13 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 13 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 9.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)).
The limitations:
“wherein meeting the minimum delay between running the portion of the first ML model and the portion of the second ML model includes delaying running of the second ML model by inserting a delay before beginning to run the second ML model”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception. See MPEP 2106.05(f).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply”. The claim is not patent eligible.
Regarding Claim 14,
Claim 14 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 14 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 9.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)) and mere instructions to apply an exception (See MPEP 2106.05(f)).
The limitations:
“the portion of the first ML model is an intermediate portion”
“the portion of the second ML model is an intermediate portion”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
The limitations:
“meeting the minimum delay between running the portion of the first ML model and the portion of the second ML model includes inserting a delay between layers of the second ML model”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception. See MPEP 2106.05(f).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way or are “mere instructions to apply”. The claim is not patent eligible.
Regarding Claim 15,
Claim 15 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 15 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 9.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein determining the minimum delay between running the portion of the first ML model and the portion of the second ML model is based on one or more cost functions”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 16,
Claim 16 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 16 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 15.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein a cost function of the one or more cost functions is based on at least one of a memory bandwidth, an amount of power consumed, and a size of available memory”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 17,
Claim 17 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 17 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: See corresponding analysis of claim 15.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are additional details that do not apply the exception in a meaningful way (See MPEP 2106.05(e)).
The limitations:
“wherein a cost function of the one or more cost functions is based on an amount of delays added to the set of ML models”
As drafted, are additional elements that do not apply an exception for the abstract ideas in a meaningful way. See MPEP 2106.05(e).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements do not apply the exception in a meaningful way. The claim is not patent eligible.
Regarding Claim 18,
Claim 18 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 18 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“determining at least an amount of memory bandwidth, power, and memory size used when executing the set of ML models on the target hardware”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 9.
Step 2B Analysis: See corresponding analysis of claim 9.
Regarding Claim 19,
Claim 19 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 19 is directed to an electronic device, comprising: a memory; and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“allocation of a first portion of the memory and selection of a first processor of the one or more processor on which to run the portion of the first ML model”
“allocation of a second portion of the memory and selection of a second processor of the one or more processors on which to run the portion of the second ML model”
“determining, based on the synchronization information, a time to run the first ML model”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: The judicial exceptions are not integrated into a practical application. In particular, the claim recited additional elements that are mere instructions to apply an exception (See MPEP 2106.05(f)) and insignificant extra-solution activity (See MPEP 2106.05(g)).
The limitations:
“An electronic device, comprising: a memory; and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions”
“run the first ML model at the time such that the first ML model and the second ML model are run in parallel, wherein the running of the first ML model at the time comprises inserting a delay after the initialization of the first ML model and before the running of the first ML model”
As drafted, are additional elements that amount to no more than mere instructions to apply an exception for the abstract ideas. See MPEP 2106.05(f).
The limitations:
“receive an indication to run a first machine learning (ML) model”
“receive synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: a minimum interval between initialization of the first ML model and running a portion of the first ML model”
“receive synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: …a minimum interval between running a portion of the second ML model and running the portion of the first ML model”
As drafted, are additional elements that amount to no more than insignificant extra-solution activity. See MPEP 2106.05(g).
Therefore, the additional elements do not integrate the abstract ideas into a practical application.
Step 2B Analysis: The claim does not include additional elements that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to the integration of the abstract ideas into a practical application, all of the additional elements are “mere instructions to apply” and “insignificant extra-solution activity”. Further, the receiving an indication limitation recites the well-understood, routine, and conventional activity of receiving or transmitting data over a network. MPEP 2106.05(d)(II); OIP Techs., Inc., v. Amazon.com, Inc., 788 F.3d 1359, 1363, 115 USPQ2d 1090, 1093 (Fed. Cir. 2015) (sending messages over a network). Additionally, the receiving synchronization information limitations recite the well-understood, routine, and conventional activity of storing and retrieving information in memory. MPEP 2106.05(d)(II); Versata Dev. Group, Inc. v. SAP Am., Inc., 793 F.3d 1306, 1334, 115 USPQ2d 1681, 1701 (Fed. Cir. 2015). Mere instructions to apply an exception and insignificant extra-solution activity cannot provide an inventive concept. The claim is not patent eligible.
Regarding Claim 21,
Claim 21 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 21 is directed to a method, which is directed to a process, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“allocating a first portion of a memory and selecting a first processing core on which to run the portion of the first ML model”
“allocating a second portion of the memory and selecting a first processing core on which to run the portion of the first ML model”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 6.
Step 2B Analysis: See corresponding analysis of claim 6.
Regarding Claim 22,
Claim 22 is rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more.
Step 1 Analysis: Claim 22 is directed to a non-transitory program storage device comprising instructions stored thereon, which is directed to a machine, one of the statutory categories.
Step 2A Prong One Analysis: The limitations:
“allocating a first portion of a memory and selecting a first processing core on which to run the portion of the first ML model”
“allocating a second portion of the memory and selecting a second processing core on which to run the portion of the second ML model”
As drafted, under their broadest reasonable interpretations, cover mental processes, i.e., concepts performed in the human mind (including an observation, evaluation, judgement, opinion). The above limitations in the context of this claim correspond to mental processes, e.g., evaluation and judgement with assistance of pen and paper.
Step 2A Prong Two Analysis: See corresponding analysis of claim 9.
Step 2B Analysis: See corresponding analysis of claim 9.
Claim Rejections - 35 USC § 102
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of the appropriate paragraphs of 35 U.S.C. 102 that form the basis for the rejections under this section made in this Office action:
A person shall be entitled to a patent unless –
(a)(2) the claimed invention was described in a patent issued under section 151, or in an application for patent published or deemed published under section 122(b), in which the patent or application, as the case may be, names another inventor and was effectively filed before the effective filing date of the claimed invention.
Claims 6-13 and 15-18 are rejected under 35 U.S.C. 102(a)(2) as being unpatentable over Chen et al. (U.S. Patent Publication No. 2022/0374274) (“Chen”).
Regarding claim 6, Chen teaches a method, comprising: receiving an indication to run a first machine learning (ML) model (Chen [0072] “For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402, which provides an indication to run a machine learning model corresponding to receiving an indication to run a first machine learning model.); receiving synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: a minimum duration between initialization of the first ML model and running of a portion of the first ML model (Chen [0072] “In some embodiments, timer 402 is configured to track an amount of time that has elapsed since a machine learning model has executed, an amount of time that has elapsed since production data has been retrieved, or other time periods. As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402, which determines an amount of time that has elapsed since production data has been retrieved and when the machine learning model is to run corresponding to a minimum duration between initialization of the first ML model and running of a portion of the first ML model.); a minimum duration between initialization of the second ML model and running of a portion of the second ML model (Chen [0072] “In some embodiments, timer 402 is configured to track an amount of time that has elapsed since a machine learning model has executed, an amount of time that has elapsed since production data has been retrieved, or other time periods …For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run” Chen provides deploying at least two machine learning models, wherein timer 402 tracks an amount of time elapsed since production data has been received and when to run for both machine learning models, corresponding to a minimum duration between initialization of the second ML model and running of a portion of the second ML model.); and a minimum duration between running of the portion of the second ML model and the portion of the first ML model (Chen [0072] “As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402 which determines the frequency which multiple machine learning models are to run, which includes a minimum during between the running of each one, corresponding to a minimum duration between running of the portion of the second ML model and the portion of the first ML model.); determining, based on the synchronization information a time to run the first ML model, wherein the determining of the time to run the first ML model comprises determining a difference between an expected time to complete running the first ML model and a current time (Chen [0072] “When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.”; [0103] “Model database 134 may include multiple sets of machine learning models, each of which may have a different execution frequency, purpose. For example, model database 134 includes a first set of machine learning models 702a-702n, each having a first execution frequency, F1, and may also include a second set of machine learning models 704a-704m, each having a second execution frequency, F2. For example, first execution frequency F1 may be an hourly, daily, weekly, or other frequencies with which machine learning models 702a-702n execute. Second execution frequency F2 may be weekly, monthly, bi-monthly, quarterly, yearly, or other frequencies with which machine learning models 704a-704m execute.” Timer 402 provides determining a time to run a machine learning model including a first machine learning model.); and running the first ML model at the time such that the first ML model and the second ML model run concurrently, wherein beginning the run of the first ML model is based on the difference (Chen [0072] “When timer 402 determines that a predetermined amount of time has elapsed corresponding to an execution frequency of a machine learning model, timer 402 may be configured to generate a trigger to cause one or more actions facilitating a machine learning model's execution.”; [0076] “In some embodiments, upon timer 402 determining that a first amount of time associated with a first execution frequency of machine learning models 410a and 410b has elapsed, model execution subsystem 114 may cause machine learning models 410a and 410b to execute on production data 202.” Chen provides running models at a time determined by timer 402, where multiple models may be executed simultaneously, such as models 410a and 410b, corresponding to running the first ML model at the time such that the first ML model and the second ML model run concurrently.).
Regarding claim 7, Chen teaches the method of claim 6, further comprising adjusting a next expected time to run the first ML model based on the difference (Chen [0072] “For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides adjusting the frequency of when to run machine learning models though timer 402 corresponding to adjusting a next expected time to run the first ML model based on the difference.).
Regarding claim 8, Chen teaches the method of claim 6, wherein the determining of the time to run the first ML model further comprises removing a delay of the running of the first ML model based on the difference (Chen [0076] “In some embodiments, upon timer 402 determining that a first amount of time associated with a first execution frequency of machine learning models 410a and 410b has elapsed, model execution subsystem 114 may cause machine learning models 410a and 410b to execute on production data 202. However, machine learning model 410n (and/or other machine learning models) may not execute because machine learning model 410n has a different execution frequency than machine learning models 410a and 410b.” Chen provides executing the models based on a difference in frequency corresponding to removing the delay of the running of the ML model based on the difference.).
Regarding claim 9, Chen teaches a non-transitory program storage device (Chen [0128] “The electronic storages may include non-transitory storage media that electronically stores information.” Chen provides a non-transitory program storage device.) comprising instructions stored thereon to cause one or more processors (Chen [0129] “The processors may be programmed to provide information processing capabilities in the computing devices.” Chen provides a processor executing stored instructions) to: receive a set of machine learning (ML) models (Chen [0004] “In some embodiments, production data to be provided to a plurality of machine learning models may be obtained via a data feed, which may be configured to receive updated application data from one or more real-time applications. The plurality of machine learning models may include, for example, a first machine learning model and a second machine learning model, which each have a first execution frequency.” Chen provides receiving a set of machine learning models.); simulate running the set of ML models on a target hardware to determine resources utilized by running the set of ML models (Chen [0018] “When it is determined that executing some machine learning models on the data will cause issues (e.g., running on one or more processing cores), based on the results of other machine learning models executing on that data (e.g., running on different processing cores), preventative actions may be initiated to conserve computing resources and ensure that those models are not executed.” Chen provides simulating (producing a computer model of) resources utilized on a target hardware by running machine learning models.) and to determine timing information (Chen [0017] “The outputs from the machine learning models can then be incorrect, inconsistent, or invalid, creating technical problems such as valuable computational resources processing the data with the machine learning model being wasted as the model will likely need to be re-run at a later time once the data has been updated or cleaned.” Chen provides determinations models need to be re-run at a later time corresponding to timing information.); determine, based on the simulation: a minimum delay between initialization of a first ML model of the set of ML models and running a portion of the first ML model (Chen [0072] “In some embodiments, timer 402 is configured to track an amount of time that has elapsed since a machine learning model has executed, an amount of time that has elapsed since production data has been retrieved, or other time periods. As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402, which determines an amount of time that has elapsed since production data has been retrieved and when the machine learning model is to run corresponding to a minimum delay between initialization of the first ML model and running of a portion of the first ML model.); a minimum delay between initialization of a second ML model of the set of ML models and running a portion of the second ML model (Chen [0072] “In some embodiments, timer 402 is configured to track an amount of time that has elapsed since a machine learning model has executed, an amount of time that has elapsed since production data has been retrieved, or other time periods …For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run” Chen provides deploying at least two machine learning models, wherein timer 402 tracks an amount of time elapsed since production data has been received and when to run for both machine learning models, corresponding to a minimum delay between initialization of the second ML model and running of a portion of the second ML model); and a minimum delay between running the portion of the first ML model and the portion of the second ML model (Chen [0072] “As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402 which determines the frequency which multiple machine learning models are to run, which includes a minimum during between the running of each one, corresponding to a minimum delay between running the portion of the first ML model and the portion of the second ML model.); and generate synchronization information to run the first ML model and the second ML model in parallel based on the determining (Chen [0018] “In particular, in multi-thread environments, multiple machine learning models may be executed in parallel or substantially in parallel.”; [0072] “When timer 402 determines that a predetermined amount of time has elapsed corresponding to an execution frequency of a machine learning model, timer 402 may be configured to generate a trigger to cause one or more actions facilitating a machine learning model's execution.”; [0076] “In some embodiments, upon timer 402 determining that a first amount of time associated with a first execution frequency of machine learning models 410a and 410b has elapsed, model execution subsystem 114 may cause machine learning models 410a and 410b to execute on production data 202.” Chen provides parallelism and running models at a time determined by timer 402, where multiple models may be executed simultaneously, such as models 410a and 410b, corresponding to generate synchronization information to run the first ML model and the second ML model in parallel based on the determining.).
Regarding claim 10, Chen teaches the non-transitory program storage device of claim 9, wherein the target hardware includes at least two cores for executing the set of ML models (Chen [0019] “For example, computer system 102 may include a plurality of computing devices (e.g., multiple processing cores), to implement the disclosed techniques. As a result, latency in obtaining results can be reduced from thirty hours to as few as thirty minutes.” Chen provides multiple processing cores corresponding to at least two cores for executing machine learning models.) and wherein the synchronization information includes timing information for coordinating execution of the set of ML models across the at least two cores (Chen [0018] “For instance, while one processing core is used to execute one machine learning model, a different processing core can be used to execute another machine learning model. When it is determined that executing some machine learning models on the data will cause issues (e.g., running on one or more processing cores), based on the results of other machine learning models executing on that data (e.g., running on different processing cores), preventative actions may be initiated to conserve computing resources and ensure that those models are not executed.” Chen provides timing information for coordinating execution of the models between at least two cores.).
Regarding claim 11, Chen teaches the non-transitory program storage device of claim 10, wherein the synchronization information includes: an indication of the minimum delay between running the portion of the first ML model and the portion of the second ML model (Chen [0072] “As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.”; [0076] “In response to timer 402 determining that a second amount of time associated with a second execution frequency of machine learning model 410n has elapsed, model execution subsystem 114 may cause machine learning model 410n to execute on production data 202 to cause datasets 412n to be generated.” Chen provides an indication of the minimum delay between running the portion of the first ML model and the portion of the second ML model.); an indication of the second ML model (Chen [0072] “In particular, in multi-thread environments, multiple machine learning models may be executed in parallel or substantially in parallel. For instance, while one processing core is used to execute one machine learning model, a different processing core can be used to execute another machine learning model.” Chen provides an indication of a second machine learning model.) and an indication of a core of the target hardware to run the second ML model (Chen [0072] “In particular, in multi-thread environments, multiple machine learning models may be executed in parallel or substantially in parallel. For instance, while one processing core is used to execute one machine learning model, a different processing core can be used to execute another machine learning model.” Chen provides an indication of a second machine learning model and a corresponding core to execute said model.).
Regarding claim 12, Chen teaches the non-transitory program storage device of claim 10, wherein the synchronization information is organized in a lookup table (Chen [0012] “FIG. 7 shows a database storing machine learning models having various execution frequencies, in accordance with one or more embodiments.” Chen provides storing data in a database corresponding to a lookup table, as shown in FIG. 7, which is merely an array of data that maps input values to output values.).
Regarding claim 13, Chen teaches the non-transitory program storage device of claim 9, wherein meeting the minimum delay between running the portion of the first ML model and the portion of the second ML model includes delaying running of the second ML model by inserting a delay before beginning to run the second ML model (Chen [0072] “As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides delaying running of the second ML model by inserting a delay before beginning to run the second ML model by scheduling set frequencies of model execution.).
Regarding claim 15, Chen teaches the non-transitory program storage device of claim 9, wherein determining the minimum delay between running the portion of the first ML model and the portion of the second ML model is based on one or more cost functions (Chen [0017] “In some instances, the data can cause errors with some of the machine learning models, such as inaccurate predictions, null result sets, or other issues. However, in many cases, these issues are not recognized until run time when the machine learning models execute on the data. The outputs from the machine learning models can then be incorrect, inconsistent, or invalid, creating technical problems such as valuable computational resources processing the data with the machine learning model being wasted as the model will likely need to be re-run at a later time once the data has been updated or cleaned. In some cases, to address the issues, the model may need to be rebuilt, re-trained, or replaced with another model.” Chen provides error determinations corresponding to cost functions causing models to be rebuilt, retrained, or replaced corresponding to determining delays between running the portion of the first ML model and the portion of the second ML model based on cost functions (errors).).
Regarding claim 16, Chen teaches the non-transitory program storage device of claim 15, wherein a cost function of the one or more cost functions is based on at least one of a memory bandwidth, an amount of power consumed, and a size of available memory (Chen [0017] “In addition to wasting computational resources, the aforementioned scenarios are time consuming, particularly when a model needs to be rebuilt or re-trained. In real-world applications, latency in obtaining results from a machine learning model can be tremendously impactful.” Chen provides time consuming scenarios where models need to be rebuilt and retrained which wastes computational resources and corresponds to the one or more cost functions is based on at least one of a memory bandwidth, an amount of power consumed, and a size of available memory since rebuilding and retraining models requires additional memory bandwidth, power consumption and available memory.).
Regarding claim 17, Chen teaches the non-transitory program storage device of claim 15, wherein a cost function of the one or more cost functions is based on an amount of delays added to the set of ML models (Chen [0018] “For instance, while one processing core is used to execute one machine learning model, a different processing core can be used to execute another machine learning model. When it is determined that executing some machine learning models on the data will cause issues (e.g., running on one or more processing cores), based on the results of other machine learning models executing on that data (e.g., running on different processing cores), preventative actions may be initiated to conserve computing resources and ensure that those models are not executed. For example, the models may be replaced with other existing models, rebuilt, or re-trained” Chen provides preventative actions for the ML models delaying their execution corresponding to a number of delays added to the ML models.).
Regarding claim 18, Chen teaches the non-transitory program storage device of claim 9, wherein simulating running the set of ML models on the target hardware comprises determining at least an amount of memory bandwidth, power, and memory size used when executing the set of ML models on the target hardware (Chen [0001] “However, these issues are typically detected after running the machine learning model, wasting valuable processing resources, memory, and time.”, [0019] “Additionally, the technical solutions described herein reduce latency in obtaining valid machine learning results by minimizing an amount of time wasted on machine learning models whose results will not be used, as well as having a model ready to execute at the desired execution frequency that will not cause invalid results to be produced. In some embodiments, the technical solutions may be implemented using a distributed computing environment. For example, computer system 102 may include a plurality of computing devices (e.g., multiple processing cores), to implement the disclosed techniques. As a result, latency in obtaining results can be reduced from thirty hours to as few as thirty minutes.” Chen provides latency reduction and reducing time/resources on machine learning models corresponding to an amount of memory bandwidth, power, and memory size used when executing the set of ML models on the target hardware.).
Claim Rejections - 35 USC § 103
In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status.
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
The factual inquiries for establishing a background for determining obviousness under 35 U.S.C. 103 are summarized as follows:
1. Determining the scope and contents of the prior art.
2. Ascertaining the differences between the prior art and the claims at issue.
3. Resolving the level of ordinary skill in the pertinent art.
4. Considering objective evidence present in the application indicating obviousness or nonobviousness.
This application currently names joint inventors. In considering patentability of the claims the examiner presumes that the subject matter of the various claims was commonly owned as of the effective filing date of the claimed invention(s) absent any evidence to the contrary. Applicant is advised of the obligation under 37 CFR 1.56 to point out the inventor and effective filing dates of each claim that was not commonly owned as of the effective filing date of the later invention in order for the examiner to consider the applicability of 35 U.S.C. 102(b)(2)(C) for any potential 35 U.S.C. 102(a)(2) prior art against the later invention.
Claim 1-2, 4, 19, and 21-22 are rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (U.S. Patent Publication No. 2022/0374274) (“Chen”) in view of Zhang et al. (U.S. Patent Publication No. 2019/0325305) (“Zhang”).
Regarding claim 1, Chen teaches a method, comprising: receiving an indication to run a first machine learning (ML) model (Chen [0072] “For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402, which provides an indication to run a machine learning model corresponding to receiving an indication to run a first machine learning model.); receiving synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: a minimum duration between initialization of the first ML model and running of a portion of the first ML model (Chen [0072] “In some embodiments, timer 402 is configured to track an amount of time that has elapsed since a machine learning model has executed, an amount of time that has elapsed since production data has been retrieved, or other time periods. As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402, which determines an amount of time that has elapsed since production data has been retrieved and when the machine learning model is to run corresponding to a minimum duration between initialization of the first ML model and running of a portion of the first ML model.), …a minimum duration between initialization of the second ML model and running of a portion of the second ML model (Chen [0072] “In some embodiments, timer 402 is configured to track an amount of time that has elapsed since a machine learning model has executed, an amount of time that has elapsed since production data has been retrieved, or other time periods …For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run” Chen provides deploying at least two machine learning models, wherein timer 402 tracks an amount of time elapsed since production data has been received and when to run for both machine learning models, corresponding to a minimum duration between initialization of the second ML model and running of a portion of the second ML model.), …and a minimum duration between running of the portion of the second ML model and the portion of the first ML model (Chen [0072] “As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402 which determines the frequency which multiple machine learning models are to run, which includes a minimum during between the running of each one, corresponding to a minimum duration between running of the portion of the second ML model and the portion of the first ML model.); determining, based on the synchronization information, a time to run the first ML model (Chen [0072] “When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Timer 402 provides determining a time to run a machine learning model including a first machine learning model.); and running the first ML model at the time such that the first ML model and the second ML model run concurrently (Chen [0072] “When timer 402 determines that a predetermined amount of time has elapsed corresponding to an execution frequency of a machine learning model, timer 402 may be configured to generate a trigger to cause one or more actions facilitating a machine learning model's execution.”; [0076] “In some embodiments, upon timer 402 determining that a first amount of time associated with a first execution frequency of machine learning models 410a and 410b has elapsed, model execution subsystem 114 may cause machine learning models 410a and 410b to execute on production data 202.” Chen provides running models at a time determined by timer 402, where multiple models may be executed simultaneously, such as models 410a and 410b, corresponding to running the first ML model at the time such that the first ML model and the second ML model run concurrently.), wherein the running of the first ML model at the time comprises inserting a delay between the initialization of the first ML model and the running of the first ML model (Chen [0072] “When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run. In some embodiments, timer 402 may be a physical timer having hardware components configured to monitor an amount of time that has elapsed since a particular event (e.g., a spring-based timing mechanism, a quartz clock, etc.), computer software (e.g., an electronic oscillator), or another timing mechanism. When timer 402 determines that a predetermined amount of time has elapsed corresponding to an execution frequency of a machine learning model, timer 402 may be configured to generate a trigger to cause one or more actions facilitating a machine learning model's execution.” Chen provides timer 402, which determines when to run machine learning models after they are deployed to a production environment, corresponding to inserting a delay between the initialization of the first ML model and the running of the first ML model.).
Chen fails to teach …wherein the initialization of the first ML model includes allocating a first portion of a memory and selecting a first processing core on which to run the portion of the first ML model; …wherein the initialization of the second ML model includes allocating a second portion of the memory and selecting a second processing core on which to run the portion of the second ML model.
However, Zhang teaches …wherein the initialization of the first ML model includes allocating a first portion of a memory (Zhang [0029] “The identification of which processing elements are in the first portion and which processing elements are in the second portion is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing. Also, the identification of the first and second communication networks from communication networks 430A-N is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing”; [0034] “In various implementations, different types of techniques for determining an optimal or preferred allocation of a machine learning model on a multi-core inference accelerator engine based on constrained memory bandwidth can be accommodated by control unit 605 and table 610.” Zhang provides memory bandwidth utilization and/or load balancing for machine learning memory allocation running on a multi-score system, corresponding to allocating a first portion of a memory to initialize a first machine learning model.) and selecting a first processing core on which to run the portion of the first ML model (Zhang [0023] “While input channel data is broadcast to inference cores 202-216 from inference core 201, each inference core 201-216 fetches its own coefficients for a corresponding set of filters. After receiving the input data and fetching the coefficients, inference cores 201-216 perform the calculations to implement a given layer of a machine learning model (e.g., convolutional neural network).”; [0030] “Referring now to FIG. 5, a block diagram of one implementation of a convolutional neural network executing on a single inference core 500. Depending on the implementation, inference core 500 is implemented as one of the inference cores of a multi-core inference accelerator engine (e.g., multi-core inference accelerator engine 105 of system 100 of FIG. 1). Inference core 500 includes a plurality of channel processing engines 502A-N.” Zhang provides a multi-core processing system for machine learning models, including single core execution of a machine learning model, corresponding to selecting a first processing core on which to run the portion of the first ML model.) …wherein the initialization of the second ML model includes allocating a second portion of the memory (Zhang [0022] “In one implementation, an “inference core” is defined as a collection of computing elements that are supervised as a unit and execute a collection of instructions to support various machine learning models. In one implementation, a “multi-core inference accelerator engine” is defined as a combination of multiple inference cores working together to implement any of various machine learning models.”; [0029] “The identification of which processing elements are in the first portion and which processing elements are in the second portion is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing. Also, the identification of the first and second communication networks from communication networks 430A-N is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing”; [0034] “In various implementations, different types of techniques for determining an optimal or preferred allocation of a machine learning model on a multi-core inference accelerator engine based on constrained memory bandwidth can be accommodated by control unit 605 and table 610.” Zhang provides memory bandwidth utilization and/or load balancing for machine learning memory allocation running on a multi-score system for various machine learning models, corresponding to allocating a second portion of the memory to initialize a second machine learning model.) and selecting a second processing core on which to run the portion of the second ML model (Zhang [0022] “In one implementation, an “inference core” is defined as a collection of computing elements that are supervised as a unit and execute a collection of instructions to support various machine learning models. In one implementation, a “multi-core inference accelerator engine” is defined as a combination of multiple inference cores working together to implement any of various machine learning models.”; [0023] “While input channel data is broadcast to inference cores 202-216 from inference core 201, each inference core 201-216 fetches its own coefficients for a corresponding set of filters. After receiving the input data and fetching the coefficients, inference cores 201-216 perform the calculations to implement a given layer of a machine learning model (e.g., convolutional neural network).”; [0030] “Referring now to FIG. 5, a block diagram of one implementation of a convolutional neural network executing on a single inference core 500. Depending on the implementation, inference core 500 is implemented as one of the inference cores of a multi-core inference accelerator engine (e.g., multi-core inference accelerator engine 105 of system 100 of FIG. 1). Inference core 500 includes a plurality of channel processing engines 502A-N.” Zhang provides a multi-core processing system for various machine learning models, including a plurality of cores and single core execution of a machine learning model, corresponding to selecting a second processing core on which to run the portion of the second ML model.);
Chen and Zhang are both considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to multi-core machine learning systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen with the above teachings of Zhang. Doing so would allow for achieving the desired level of memory bandwidth utilization and/or load balancing for the execution of machine learning models (Zhang [0026] “Control unit 410 programs logic 415, communication networks 430A-N, and processing elements 435A-N to implement a given machine learning model while achieving the desired level of memory bandwidth utilization and/or load balancing.”).
Regarding claim 2, Chen in view of Zhang teaches the method of claim 1, as discussed above in the rejection of claim 1, wherein the synchronization information includes an indication of the minimum duration between running of the portion of the second ML model and the portion of the first ML model (Chen [0072] “As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402 which determines the frequency which multiple machine learning models are to run, which includes a minimum during between the running of each one, corresponding to a minimum duration between running of the portion of the second ML model and the portion of the first ML model.); an indication of the first ML model (Chen [0018] “For instance, while one processing core is used to execute one machine learning model, a different processing core can be used to execute another machine learning model.” Chen provides an association indication of a first model.); and an indication of a core to run the first ML model (Chen [0018] “For instance, while one processing core is used to execute one machine learning model, a different processing core can be used to execute another machine learning model.” Chen provides an association indication of a first model and a core to run said model.).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen in view of Zhang for the same reasons disclosed above in the rejection of claim 1.
Regarding claim 4, Chen in view of Zhang teaches the method of claim 1, as discussed above in the rejection of claim 1, wherein the delay is based on a callback function or a parallel thread (Chen [0018] In particular, in multi-thread environments, multiple machine learning models may be executed in parallel or substantially in parallel. For instance, while one processing core is used to execute one machine learning model, a different processing core can be used to execute another machine learning model. When it is determined that executing some machine learning models on the data will cause issues (e.g., running on one or more processing cores), based on the results of other machine learning models executing on that data (e.g., running on different processing cores), preventative actions may be initiated to conserve computing resources and ensure that those models are not executed.” Chen provides preventative actions for executing models in parallel including not executing them corresponding to a delay based on a parallel thread.).
It would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen in view of Zhang for the same reasons disclosed above in the rejection of claim 1.
Regarding claim 19, Chen teaches an electronic device, comprising: a memory (Chen [0071] “Each module of model execution subsystem 114 may be implemented by one or more processors executing computer program instructions stored in memory of computer system 102.” Chen provides a memory.); and one or more processors operatively coupled to the memory, wherein the one or more processors are configured to execute instructions (Chen [0071] “Each module of model execution subsystem 114 may be implemented by one or more processors executing computer program instructions stored in memory of computer system 102.” Chen provides processors coupled to memory configured to execute instructions.) causing the one or more processors to: receive an indication to run a first machine learning (ML) model (Chen [0072] “For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides time 402 which provides an indication to run a machine learning model corresponding to receiving an indication to run a machine learning model.); receive synchronization information for organizing the running of the first ML model with respect to a second ML model that specifies: a minimum interval between initialization of the first ML model and running a portion of the first ML model (Chen [0072] “In some embodiments, timer 402 is configured to track an amount of time that has elapsed since a machine learning model has executed, an amount of time that has elapsed since production data has been retrieved, or other time periods. As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402, which determines an amount of time that has elapsed since production data has been retrieved and when the machine learning model is to run corresponding to a minimum interval between initialization of the first ML model and running of a portion of the first ML model.); …and a minimum interval between running a portion of the second ML model and running the portion of the first ML model (Chen [0072] “As mentioned previously, machine learning models may have various execution frequencies with which each runs. For example, one machine learning model may execute weekly, while another machine learning model may execute monthly. When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402 which determines the frequency which multiple machine learning models are to run, which includes a minimum during between the running of each one, corresponding to a minimum interval between running of the portion of the second ML model and the portion of the first ML model.) …determine, based on the synchronization information, a time to run the first ML model (Chen [0072] “When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run.” Chen provides timer 402 that provides determining a time to run a machine learning model including a first machine learning model.); and run the first ML model at the time such that the first ML model and the second ML model are run in parallel (Chen [0018] “In particular, in multi-thread environments, multiple machine learning models may be executed in parallel or substantially in parallel”; [0072] “When timer 402 determines that a predetermined amount of time has elapsed corresponding to an execution frequency of a machine learning model, timer 402 may be configured to generate a trigger to cause one or more actions facilitating a machine learning model's execution.”; [0076] “In some embodiments, upon timer 402 determining that a first amount of time associated with a first execution frequency of machine learning models 410a and 410b has elapsed, model execution subsystem 114 may cause machine learning models 410a and 410b to execute on production data 202.” Chen provides running models at a time determined by timer 402, where multiple models may be executed simultaneously, including parallelism, such as from models 410a and 410b, corresponding run the first ML model at the time such that the first ML model and the second ML model are run in parallel.), wherein the running of the first ML model at the time comprises inserting a delay after the initialization of the first ML model and before the running of the first ML model (Chen [0072] “When both of these machine learning models are deployed to a production environment, timer 402 can determine whether a machine learning model is to run and/or when the machine learning model is to run. In some embodiments, timer 402 may be a physical timer having hardware components configured to monitor an amount of time that has elapsed since a particular event (e.g., a spring-based timing mechanism, a quartz clock, etc.), computer software (e.g., an electronic oscillator), or another timing mechanism. When timer 402 determines that a predetermined amount of time has elapsed corresponding to an execution frequency of a machine learning model, timer 402 may be configured to generate a trigger to cause one or more actions facilitating a machine learning model's execution.” Chen provides timer 402, which determines when to run machine learning models after they are deployed to a production environment, corresponding to inserting a delay between the initialization of the first ML model and the running of the first ML model.).
Chen fails to teach …wherein the initialization of the first ML model includes allocation of a first portion of the memory and selection of a first processor of the one or more processor on which to run the portion of the first ML model ...wherein the initialization of the second ML model includes allocation of a second portion of the memory and selection of a second processor of the one or more processors on which to run the portion of the second ML model.
However, Zhang teaches …wherein the initialization of the first ML model includes allocation of a first portion of the memory (Zhang [0029] “The identification of which processing elements are in the first portion and which processing elements are in the second portion is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing. Also, the identification of the first and second communication networks from communication networks 430A-N is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing”; [0034] “In various implementations, different types of techniques for determining an optimal or preferred allocation of a machine learning model on a multi-core inference accelerator engine based on constrained memory bandwidth can be accommodated by control unit 605 and table 610.” Zhang provides memory bandwidth utilization and/or load balancing for machine learning memory allocation running on a multi-score system, corresponding to allocating a first portion of a memory to initialize a first machine learning model.) and selection of a first processor of the one or more processor on which to run the portion of the first ML model (Zhang [0023] “While input channel data is broadcast to inference cores 202-216 from inference core 201, each inference core 201-216 fetches its own coefficients for a corresponding set of filters. After receiving the input data and fetching the coefficients, inference cores 201-216 perform the calculations to implement a given layer of a machine learning model (e.g., convolutional neural network).”; [0030] “Referring now to FIG. 5, a block diagram of one implementation of a convolutional neural network executing on a single inference core 500. Depending on the implementation, inference core 500 is implemented as one of the inference cores of a multi-core inference accelerator engine (e.g., multi-core inference accelerator engine 105 of system 100 of FIG. 1). Inference core 500 includes a plurality of channel processing engines 502A-N.” Zhang provides a multi-core processing system for machine learning models, including single core execution of a machine learning model, corresponding to selecting a first processing core on which to run the portion of the first ML model.) ...wherein the initialization of the second ML model includes allocation of a second portion of the memory (Zhang [0022] “In one implementation, an “inference core” is defined as a collection of computing elements that are supervised as a unit and execute a collection of instructions to support various machine learning models. In one implementation, a “multi-core inference accelerator engine” is defined as a combination of multiple inference cores working together to implement any of various machine learning models.”; [0029] “The identification of which processing elements are in the first portion and which processing elements are in the second portion is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing. Also, the identification of the first and second communication networks from communication networks 430A-N is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing”; [0034] “In various implementations, different types of techniques for determining an optimal or preferred allocation of a machine learning model on a multi-core inference accelerator engine based on constrained memory bandwidth can be accommodated by control unit 605 and table 610.” Zhang provides memory bandwidth utilization and/or load balancing for machine learning memory allocation running on a multi-score system for various machine learning models, corresponding to allocating a second portion of the memory to initialize a second machine learning model.) and selection of a second processor of the one or more processors on which to run the portion of the second ML model (Zhang [0022] “In one implementation, an “inference core” is defined as a collection of computing elements that are supervised as a unit and execute a collection of instructions to support various machine learning models. In one implementation, a “multi-core inference accelerator engine” is defined as a combination of multiple inference cores working together to implement any of various machine learning models.”; [0023] “While input channel data is broadcast to inference cores 202-216 from inference core 201, each inference core 201-216 fetches its own coefficients for a corresponding set of filters. After receiving the input data and fetching the coefficients, inference cores 201-216 perform the calculations to implement a given layer of a machine learning model (e.g., convolutional neural network).”; [0030] “Referring now to FIG. 5, a block diagram of one implementation of a convolutional neural network executing on a single inference core 500. Depending on the implementation, inference core 500 is implemented as one of the inference cores of a multi-core inference accelerator engine (e.g., multi-core inference accelerator engine 105 of system 100 of FIG. 1). Inference core 500 includes a plurality of channel processing engines 502A-N.” Zhang provides a multi-core processing system for various machine learning models, including a plurality of cores and single core execution of a machine learning model, corresponding to selecting a second processing core on which to run the portion of the second ML model.).
Chen and Zhang are both considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to multi-core machine learning systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen with the above teachings of Zhang. Doing so would allow for achieving the desired level of memory bandwidth utilization and/or load balancing for the execution of machine learning models (Zhang [0026] “Control unit 410 programs logic 415, communication networks 430A-N, and processing elements 435A-N to implement a given machine learning model while achieving the desired level of memory bandwidth utilization and/or load balancing.”).
Regarding claim 21, Chen teaches the method of claim 6 as discussed above in the rejection of claim 6, but fails to teach wherein: the initialization of the first ML model includes allocating a first portion of a memory and selecting a first processing core on which to run the portion of the first ML model; and the initialization of the second ML model includes allocating a second portion of the memory and selecting a first processing core on which to run the portion of the first ML model.
However, Zhang teaches wherein: the initialization of the first ML model includes allocating a first portion of a memory (Zhang [0029] “The identification of which processing elements are in the first portion and which processing elements are in the second portion is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing. Also, the identification of the first and second communication networks from communication networks 430A-N is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing”; [0034] “In various implementations, different types of techniques for determining an optimal or preferred allocation of a machine learning model on a multi-core inference accelerator engine based on constrained memory bandwidth can be accommodated by control unit 605 and table 610.” Zhang provides memory bandwidth utilization and/or load balancing for machine learning memory allocation running on a multi-score system, corresponding to allocating a first portion of a memory to initialize a first machine learning model.) and selecting a first processing core on which to run the portion of the first ML model (Zhang [0023] “While input channel data is broadcast to inference cores 202-216 from inference core 201, each inference core 201-216 fetches its own coefficients for a corresponding set of filters. After receiving the input data and fetching the coefficients, inference cores 201-216 perform the calculations to implement a given layer of a machine learning model (e.g., convolutional neural network).”; [0030] “Referring now to FIG. 5, a block diagram of one implementation of a convolutional neural network executing on a single inference core 500. Depending on the implementation, inference core 500 is implemented as one of the inference cores of a multi-core inference accelerator engine (e.g., multi-core inference accelerator engine 105 of system 100 of FIG. 1). Inference core 500 includes a plurality of channel processing engines 502A-N.” Zhang provides a multi-core processing system for machine learning models, including single core execution of a machine learning model, corresponding to selecting a first processing core on which to run the portion of the first ML model.); and the initialization of the second ML model includes allocating a second portion of the memory (Zhang [0022] “In one implementation, an “inference core” is defined as a collection of computing elements that are supervised as a unit and execute a collection of instructions to support various machine learning models. In one implementation, a “multi-core inference accelerator engine” is defined as a combination of multiple inference cores working together to implement any of various machine learning models.”; [0029] “The identification of which processing elements are in the first portion and which processing elements are in the second portion is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing. Also, the identification of the first and second communication networks from communication networks 430A-N is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing”; [0034] “In various implementations, different types of techniques for determining an optimal or preferred allocation of a machine learning model on a multi-core inference accelerator engine based on constrained memory bandwidth can be accommodated by control unit 605 and table 610.” Zhang provides memory bandwidth utilization and/or load balancing for machine learning memory allocation running on a multi-score system for various machine learning models, corresponding to allocating a second portion of the memory to initialize a second machine learning model.) and selecting a first processing core on which to run the portion of the first ML model (Zhang [0022] “In one implementation, an “inference core” is defined as a collection of computing elements that are supervised as a unit and execute a collection of instructions to support various machine learning models. In one implementation, a “multi-core inference accelerator engine” is defined as a combination of multiple inference cores working together to implement any of various machine learning models.”; [0023] “While input channel data is broadcast to inference cores 202-216 from inference core 201, each inference core 201-216 fetches its own coefficients for a corresponding set of filters. After receiving the input data and fetching the coefficients, inference cores 201-216 perform the calculations to implement a given layer of a machine learning model (e.g., convolutional neural network).”; [0030] “Referring now to FIG. 5, a block diagram of one implementation of a convolutional neural network executing on a single inference core 500. Depending on the implementation, inference core 500 is implemented as one of the inference cores of a multi-core inference accelerator engine (e.g., multi-core inference accelerator engine 105 of system 100 of FIG. 1). Inference core 500 includes a plurality of channel processing engines 502A-N.” Zhang provides a multi-core processing system for various machine learning models, including a plurality of cores and single core execution of a machine learning model, corresponding to selecting a first or second processing core on which to run the portion of the first or second ML model.).
Chen and Zhang are both considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to multi-core machine learning systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen with the above teachings of Zhang. Doing so would allow for achieving the desired level of memory bandwidth utilization and/or load balancing for the execution of machine learning models (Zhang [0026] “Control unit 410 programs logic 415, communication networks 430A-N, and processing elements 435A-N to implement a given machine learning model while achieving the desired level of memory bandwidth utilization and/or load balancing.”).
Regarding claim 22, Chen teaches the non-transitory program storage device of claim 9 as discussed above in the rejection of claim 9, but fails to teach wherein: the initialization of the first ML model includes allocating a first portion of a memory and selecting a first processing core on which to run the portion of the first ML model; and the initialization of the second ML model includes allocating a second portion of the memory and selecting a second processing core on which to run the portion of the second ML model.
However, Zhang teaches wherein: the initialization of the first ML model includes allocating a first portion of a memory (Zhang [0029] “The identification of which processing elements are in the first portion and which processing elements are in the second portion is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing. Also, the identification of the first and second communication networks from communication networks 430A-N is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing”; [0034] “In various implementations, different types of techniques for determining an optimal or preferred allocation of a machine learning model on a multi-core inference accelerator engine based on constrained memory bandwidth can be accommodated by control unit 605 and table 610.” Zhang provides memory bandwidth utilization and/or load balancing for machine learning memory allocation running on a multi-score system, corresponding to allocating a first portion of a memory to initialize a first machine learning model.) and selecting a first processing core on which to run the portion of the first ML model (Zhang [0023] “While input channel data is broadcast to inference cores 202-216 from inference core 201, each inference core 201-216 fetches its own coefficients for a corresponding set of filters. After receiving the input data and fetching the coefficients, inference cores 201-216 perform the calculations to implement a given layer of a machine learning model (e.g., convolutional neural network).”; [0030] “Referring now to FIG. 5, a block diagram of one implementation of a convolutional neural network executing on a single inference core 500. Depending on the implementation, inference core 500 is implemented as one of the inference cores of a multi-core inference accelerator engine (e.g., multi-core inference accelerator engine 105 of system 100 of FIG. 1). Inference core 500 includes a plurality of channel processing engines 502A-N.” Zhang provides a multi-core processing system for machine learning models, including single core execution of a machine learning model, corresponding to selecting a first processing core on which to run the portion of the first ML model.); and the initialization of the second ML model includes allocating a second portion of the memory (Zhang [0022] “In one implementation, an “inference core” is defined as a collection of computing elements that are supervised as a unit and execute a collection of instructions to support various machine learning models. In one implementation, a “multi-core inference accelerator engine” is defined as a combination of multiple inference cores working together to implement any of various machine learning models.”; [0029] “The identification of which processing elements are in the first portion and which processing elements are in the second portion is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing. Also, the identification of the first and second communication networks from communication networks 430A-N is determined by control unit 410 based on the type of machine learning model being implemented and based on the desired memory bandwidth utilization and/or load balancing”; [0034] “In various implementations, different types of techniques for determining an optimal or preferred allocation of a machine learning model on a multi-core inference accelerator engine based on constrained memory bandwidth can be accommodated by control unit 605 and table 610.” Zhang provides memory bandwidth utilization and/or load balancing for machine learning memory allocation running on a multi-score system for various machine learning models, corresponding to allocating a second portion of the memory to initialize a second machine learning model.) and selecting a second processing core on which to run the portion of the second ML model (Zhang [0022] “In one implementation, an “inference core” is defined as a collection of computing elements that are supervised as a unit and execute a collection of instructions to support various machine learning models. In one implementation, a “multi-core inference accelerator engine” is defined as a combination of multiple inference cores working together to implement any of various machine learning models.”; [0023] “While input channel data is broadcast to inference cores 202-216 from inference core 201, each inference core 201-216 fetches its own coefficients for a corresponding set of filters. After receiving the input data and fetching the coefficients, inference cores 201-216 perform the calculations to implement a given layer of a machine learning model (e.g., convolutional neural network).”; [0030] “Referring now to FIG. 5, a block diagram of one implementation of a convolutional neural network executing on a single inference core 500. Depending on the implementation, inference core 500 is implemented as one of the inference cores of a multi-core inference accelerator engine (e.g., multi-core inference accelerator engine 105 of system 100 of FIG. 1). Inference core 500 includes a plurality of channel processing engines 502A-N.” Zhang provides a multi-core processing system for various machine learning models, including a plurality of cores and single core execution of a machine learning model, corresponding to selecting a second processing core on which to run the portion of the second ML model.).
Chen and Zhang are both considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to multi-core machine learning systems. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen with the above teachings of Zhang. Doing so would allow for achieving the desired level of memory bandwidth utilization and/or load balancing for the execution of machine learning models (Zhang [0026] “Control unit 410 programs logic 415, communication networks 430A-N, and processing elements 435A-N to implement a given machine learning model while achieving the desired level of memory bandwidth utilization and/or load balancing.”).
Claim 5 is rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (U.S. Patent Publication No. 2022/0374274) (“Chen”) in view of Zhang et al. (U.S. Patent Publication No. 2019/0325305) (“Zhang”) in further view of Lee et al. (U.S. Patent Publication No. 2022/0114015) (“Lee”).
Regarding claim 5, Chen in view of Zhang teaches the method of claim 1 as discussed above in the rejection of claim 1, but fails to teach wherein the determining of the time to run the first ML model comprises determining whether to insert a delay between layers of the first ML model.
However, Lee teaches wherein the determining of the time to run the first ML model comprises determining whether to insert a delay between layers of the first ML model (Lee [0111] “In some case, due to a great idle time of the accelerator for a certain candidate layer, the candidate layer may not be selected by the scheduler. In such a case, a latency of a model in which the candidate layer is included may increase greatly. To prevent this, when there is a layer for which execution is delayed a preset number of times or more among candidate layers of a plurality of models, the scheduler may perform scheduling on the layer and allow the layer to be forced to be executed.” Lee provides a scheduler for layers of a neural network, including determinations of a delay, corresponding to determining of the time to run the first ML model comprises determining whether to insert a delay between layers of the first ML model.).
Chen, Zhang and Lee are all considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to machine learning model scheduling. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen in view of Zhang with the above teachings of Lee. Doing so would allow for effectively managing latency of an accelerator applied to layers of machine learning models (Lee [0111] “In such a case, a latency of a model in which the candidate layer is included may increase greatly. To prevent this, when there is a layer for which execution is delayed a preset number of times or more among candidate layers of a plurality of models, the scheduler may perform scheduling on the layer and allow the layer to be forced to be executed. Through this, it is possible to effectively manage a latency of the accelerator.”).
Claim 14 is rejected under 35 U.S.C. 103 as being unpatentable over Chen et al. (U.S. Patent Publication No. 2022/0374274) (“Chen”) in view of Lee et al. (U.S. Patent Publication No. 2022/0114015) (“Lee”).
Regarding claim 14, Chen teaches the non-transitory program storage device of claim 9 as discussed above in the rejection of claim 9, wherein: the portion of the first ML model is an intermediate portion (Chen [0076] “In some embodiments, upon timer 402 determining that a first amount of time associated with a first execution frequency of machine learning models 410a and 410b has elapsed, model execution subsystem 114 may cause machine learning models 410a and 410b to execute on production data 202.” Chen provides executing a portion of a plurality of models corresponding to the portion of the first ML model is an intermediate portion.); the portion of the second ML model is an intermediate portion (Chen [0076] “In some embodiments, upon timer 402 determining that a first amount of time associated with a first execution frequency of machine learning models 410a and 410b has elapsed, model execution subsystem 114 may cause machine learning models 410a and 410b to execute on production data 202.” Chen provides executing a portion of a plurality of models corresponding to the portion of the second ML model is an intermediate portion.), but fails to teach and meeting the minimum delay between running the portion of the first ML model and the portion of the second ML model includes inserting a delay between layers of the second ML model.
However, Lee teaches and meeting the minimum delay between running the portion of the first ML model and the portion of the second ML model includes inserting a delay between layers of the second ML model (Lee [0110] “When the scheduler is called after the first layer of the model A is executed in an accelerator, the scheduler may perform scheduling on a first layer of the model B that has a minimum idle time among candidate layers 620 (e.g., a second layer of the model A and the first layer of the model B) which are a target for the scheduling in the models A and B.”; [0111] “In some case, due to a great idle time of the accelerator for a certain candidate layer, the candidate layer may not be selected by the scheduler. In such a case, a latency of a model in which the candidate layer is included may increase greatly. To prevent this, when there is a layer for which execution is delayed a preset number of times or more among candidate layers of a plurality of models, the scheduler may perform scheduling on the layer and allow the layer to be forced to be executed.” Lee provides a scheduler for layers of a neural network including models A and B and including determinations of a delay, wherein model B corresponds to the second ML model, corresponding to meeting the minimum delay between running the portion of the first ML model and the portion of the second ML model includes inserting a delay between layers of the second ML model.).
Chen and Lee are both considered to be analogous to the claimed invention because they are in the same field of artificial intelligence and more specifically applied to machine learning model scheduling. Therefore, it would have been obvious to someone of ordinary skill in the art before the effective filing date of the claimed invention to have modified Chen with the above teachings of Lee. Doing so would allow for effectively managing latency of an accelerator applied to layers of machine learning models (Lee [0111] “In such a case, a latency of a model in which the candidate layer is included may increase greatly. To prevent this, when there is a layer for which execution is delayed a preset number of times or more among candidate layers of a plurality of models, the scheduler may perform scheduling on the layer and allow the layer to be forced to be executed. Through this, it is possible to effectively manage a latency of the accelerator.”).
Response to Arguments
Regarding the rejection applied under 35 U.S.C. 101, Applicant firstly asserts that the limitation “determining, based on the synchronization information, a time to run the first ML model” is not practically performable in the human mind (“Remarks”, Page 13). Applicant further asserts that claim 1 is more akin to the claim of Example 39, where none of the limitations are practically performed in the human mind (“Remarks”, Page 13). Applicant further asserts the claim 1 does not recite any judicial exceptions (“Remarks”, Page 13).
However, “determining, based on the synchronization information, a time to run the first ML model”, is practically performable in the human mind. For example, the various “synchronization information”, as defined by the claims, are periods of time. Therefore, making a determination on a time to run a machine learning model, based on a plurality of other time periods, is a mentally performable task. For example, one can determine, mentally, a time to run a machine learning model, based on a variety of calculated time periods. Therefore, the limitation is a mental process. Example 39 merely provides an example claim which does not recite any abstract ideas, wherein the Present Application recites at least one abstract idea, as discussed above.
Applicant further asserts that even if claim 1 recites a judicial exception, the claim is still eligible because it improves the functioning of computer running a first and second machine learning model using synchronization points to balance the utilization of computing hardware resources (“Remarks”, Page 14). Applicant further asserts that the Present Application teaches a technique to control a duration between synchronization points of various ML models to balance computing hardware resource utilization (“Remarks”, Page 15). Applicant further asserts that the “running the first ML model at the time such that the first ML model and second ML model run concurrently” limits the claim to the use case where the problem to be solved occurs (“Remarks”, Page 16). Applicant therefore asserts that the claims are patent eligible (“Remarks”, Pages 16-19).
However, even assuming the claims did recite an improvement, it would in the abstract idea of determining times to run machine learning models. The MPEP notes that it is important to keep in mind that an improvement in the abstract idea itself is not an improvement in technology. MPEP 2106.05(a)(II). Further, the “running” limitation recites a generic running/execution of machine learning models, and therefore corresponds to mere instructions to apply (See MPEP 2106.05(f)). Therefore, the claims remain rejected under 35 U.S.C. 101.
Regarding the rejection applied under 35 U.S.C. 102, Applicant firstly asserts that Chen fails to teach “inserting a delay between the initialization of the first ML model and the running of the first ML model”. Applicant further asserts that timer 402 in Chen does not provide the delay, and simply schedules the running of machine learning models based on a defined frequency (e.g., daily, weekly, monthly, etc.) (“Remarks”, Pages 21-22). Applicant further asserts that Chen fails to teach the amended limitations of “allocating memory” and “selecting a processor” recited in claim 1 for the initialization of the machine learning models (“Remarks”, Page 22).
However, as discussed above in the 35 U.S.C. 103 rejection of claim 1 above, Zhang teaches the amended limitations of an initialization comprising allocating memory and selecting a processing core. Further, Chen does teach “inserting a delay between the initialization of the first ML model and the running of the first ML model”. For example, as discussed in paragraph [0072] of Chen, timer 402 determines whether a machine learning model is to run and/or when the machine learning model is to run after being deployed to a production environment. In this instance, Examiner is interpretating the deployment of the machine learning model to a production environment as an “initialization” of a machine learning model. Further, the models deployed to the production environment are not immediately executed, but rather, a period of time elapses after being deployed to the production environment, as determined by timer 402. Since a “delay” is a period of time by which something is postponed, and the machine learning models deployed to the production environment in Chen are not immediately run/executed, there exists a “delay” in the running of the machine learning models after their initialization, as determined by timer 402. Therefore, Chen teaches the limitations, and the claims remain rejected under 35 U.S.C. 102/103.
Conclusion
Applicant's amendment necessitated the new ground(s) of rejection presented in this Office action. Accordingly, THIS ACTION IS MADE FINAL. See MPEP § 706.07(a). Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a).
A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any nonprovisional extension fee (37 CFR 1.17(a)) pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to KURT NICHOLAS PRESSLY whose telephone number is (703)756-4639. The examiner can normally be reached M-F 8-4.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Kamran Afshar can be reached at (571) 272-7796. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/KURT NICHOLAS PRESSLY/Examiner, Art Unit 2125
/KAMRAN AFSHAR/Supervisory Patent Examiner, Art Unit 2125