Prosecution Insights
Last updated: October 02, 2026
Application No. 18/732,715

POLLER WITH SHARED RECEIVE QUEUES AND CPU GROUPS FOR STORAGE CLUSTER

Non-Final OA §101§103§112
Filed
Jun 04, 2024
Examiner
ALAM, SHIHAB
Art Unit
Tech Center
Assignee
Dell Products L.P.
OA Round
1 (Non-Final)
Grant Probability
Favorable
1-2
OA Rounds

Examiner Intelligence

Grants only 0% of cases
0%
Career Allowance Rate
0 granted / 0 resolved
-60.0% vs TC avg
Minimal +0% lift
Without
With
+0.0%
Interview Lift
resolved cases with interview
Typical timeline
Avg Prosecution
15 currently pending
Career history
13
Total Applications
across all art units
This examiner has no resolved cases yet (career too new); statute-level performance unavailable. The Grant Probability card shows Tech Center averages instead.

Office Action

§101 §103 §112
DETAILED ACTION The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . This Office Action is in response to claims filed 07/31/2023. Claims 1-20 are pending. Claim Rejections - 35 USC § 112 The following is a quotation of 35 U.S.C. 112(b): (b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention. The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph: The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention. Claims 2-8 and 12-17 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention. Claims 2 and 12 recite the limitation "the CPU cores". There is insufficient antecedent basis for this limitation in the claim. For the purposes of compact prosecution, Examiner will interpret “the CPU cores” as referring to the “multiple CPU cores” established in independent Claims 1 and 11. Claim 7 recites the limitation "the same group". There is insufficient antecedent basis for this limitation in the claim. For the purposes of compact prosecution, Examiner will interpret the first “the same group” as “a same group” and the subsequent recitations of “the same group” as referring to that first interpreted “a same group”. Claims 8 and 17 recite the limitation "the other storage node". There is insufficient antecedent basis for this limitation in the claim. For the purposes of compact prosecution, Examiner will interpret “the other storage node” as referring to the “another storage node” recited earlier in the Claims 8 and 17. Claims 3-6 and 13-16 are further rejected based on their dependency to the aforementioned rejected Claims 2 and 12. Claim Rejections - 35 USC § 101 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. Claims 1-20 are rejected under 35 U.S.C. 101 because the claimed invention recites a judicial exception, is directed to that judicial exception, an abstract idea, as it has not been integrated into practical application and the claims further do not recite significantly more than the judicial exception. Examiner has evaluated the claims under the framework provided in the 2019 Patent Eligibility Guidance published in the Federal Register 01/07/2019 and has provided such analysis below. Step 1: Claims 1-10 are directed to methods and fall within the statutory category of processes; Claims 11-17 are directed to a system and falls within the statutory category of machines; Claims 18-20 a computer program product and falls within the statutory category of manufacture. Therefore, “Are the claims to a process, machine, manufacture or composition of matter?” Yes. In order to evaluate the Step 2A inquiry “Is the claim directed to a law of nature, a natural phenomenon or an abstract idea?” we must determine, at Step 2A Prong 1, whether the claim recites a law of nature, a natural phenomenon or an abstract idea and further whether the claim recites additional elements that integrate the judicial exception into a practical application. Step 2A Prong 1: Claims 1, 11 and 18: The limitations “monitoring a system load of a storage node”, “detecting a level of the system load relative to a threshold value”, “dynamically changing a number of groups of CPU cores”, and “dynamically changing a number of CPU cores in each group” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, a person can mentally. Further, but for the recitation of generic computing components being used as a tool to perform the functionality, a person mentally evaluate whether a threshold is being approached/exceeded. Further, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can change the grouping of resources based on that evaluation. Therefore, yes, Claims 1, 11, and 18 recite judicial exceptions. The claims have been identified to recite judicial exceptions, Step 2A Prong 2 will evaluate whether the claims are directed to the judicial exception. Step 2A Prong 2: Claims 1, 11, and 18: The judicial exceptions are not integrated into practical applications. In particular, the claims recite the following additional elements – “the storage node including a multi core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages”, is a recitation of generic computing components and functions merely being used as a tool to apply the abstract idea (see MPEP § 2106.05(f)). Therefore, “Do the claims recite additional elements that integrate the judicial exception into a practical application? No, these additional elements do not integrate the abstract idea into a practical application and they do not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea. After having evaluating the inquires set forth in Steps 2A Prong 1 and 2, it has been concluded that the Claims 1, 11 and 18 not only recite a judicial exception but that the claims are directed to a judicial exception as a judicial exception has not been integrated into a practical application. Step 2B: Claims 1, 11, and 18: The claims do not include additional elements, alone or in combination, that are sufficient to amount to significantly more than the judicial exception. As discussed above with respect to integration of the abstract idea into a practical application, the additional elements amount to no more than field of use/technological environment which do not amount to significantly more than the abstract idea. Therefore, “Do the claims recite additional elements that amount to significantly more than the judicial exception? No, these additional elements, alone or in combination, do not amount to significantly more than the judicial exception. Having concluded analysis within the provided framework, Claims 1, 11 and 18 do not recite patent eligible subject matter under 35 U.S.C. § 101. Regarding Claims 2 and 12: “upon initialization of the storage node, assigning all the CPU cores to a single group; allocating a single shared queue for the single group;” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate when to add to a resource to a group. Further, “poll, by the CPU cores assigned to the single group, the single shared queue for received messages”, reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). With regard to integration into practical application and whether additional elements amount to significantly more, Claims 2 and 12 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. iv. Storing and retrieving information in memory”. See MPEP § 2106.05(d)(II). Therefore, Claims 2 and 12 do not recite patent eligible subject matter under 35 U.S.C. § 101. Regarding Claims 3 and 13: “detecting the level of the system load includes detecting that the system load has increased relative to the threshold value, wherein dynamically changing the number of groups includes increasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes decreasing the number of CPU cores in each group by the predetermined factor”, as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate when a threshold is being reached and adjust resources groups based on that evaluation. With regard to integration into practical application and whether additional elements amount to significantly more, Claims 3 and 13 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claims 3 and 13 do not recite patent eligible subject matter under 35 U.S.C. § 101. Claims 4 and 14: “assigning the decreased number of CPU cores to each of the increased number of groups; allocating a plurality of shared queues for the increased number of groups, respectively;” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate when to add to a resource to a group. Further, “polling the plurality of shared queues for received messages by the decreased number of CPU cores assigned to the increased number of groups, respectively”, reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). With regard to integration into practical application and whether additional elements amount to significantly more, Claims 4 and 14 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. iv. Storing and retrieving information in memory”. See MPEP § 2106.05(d)(II). Therefore, Claims 4 and 14 do not recite patent eligible subject matter under 35 U.S.C. § 101. Claims 5 and 15: “detecting the level of the system load includes detecting that the system load has decreased relative to the threshold value, wherein dynamically changing the number of groups includes decreasing the number of groups by the predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes increasing the number of CPU cores in each group by the predetermined factor” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate when a threshold is being reached and adjust resources groups based on that evaluation. With regard to integration into practical application and whether additional elements amount to significantly more, Claims 5 and 15 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claims 5 and 15 do not recite patent eligible subject matter under 35 U.S.C. § 101. Claims 6 and 16: “assigning the increased number of CPU cores to each of the decreased number of groups; allocating a decreased plurality of shared queues for the decreased number of groups, respectively;” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate when to add to a resource to a group. Further, “polling the decreased plurality of shared queues for received messages by the increased number of CPU cores assigned to the decreased number of groups, respectively”, reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). With regard to integration into practical application and whether additional elements amount to significantly more, Claims 6 and 16 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. iv. Storing and retrieving information in memory”. See MPEP § 2106.05(d)(II). Therefore, Claims 6 and 16 do not recite patent eligible subject matter under 35 U.S.C. § 101. Claim 7: “assigning CPU cores that share a cache level to the same group; assigning CPU cores that execute similar application threads to the same group; assigning CPU cores that utilize resources local to a NUMA (non-uniform memory access) node to the same group; assigning a CPU core with a high average queue polling frequency to each group; and assigning a low-stressed CPU core to each group including a high-stressed CPU core” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate what group a resource belongs to based on similar traits. With regard to integration into practical application and whether additional elements amount to significantly more, Claim 7 fails both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claim 7 does not recite patent eligible subject matter under 35 U.S.C. § 101. Claims 8 and 17: “the multiple CPU cores are grouped for polling the at least one shared queue for received remote procedure call (RPC) messages”, “generating an RPC reply message;” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate what group a resource belongs to and generate a response to a message. Further, “processing the RPC request message;” is a recitation of generic computing components and functions merely being used as a tool to apply the abstract idea (see MPEP § 2106.05(f)). Lastly “polling, by the number of groups of CPU cores, the at least one shared queue for an RPC request message from another storage node;” and “sending the RPC reply message to the other storage node”, reciting insignificant extra-solution data gathering activity, MPEP § 2106.05(g). With regard to integration into practical application and whether additional elements amount to significantly more, Claims 8 and 17 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more, performing a well understood, routine, and conventional task of data gathering. “The courts have recognized the following computer functions as well‐understood, routine, and conventional functions when they are claimed in a merely generic manner (e.g., at a high level of generality) or as insignificant extra-solution activity. iv. Storing and retrieving information in memory”. See MPEP § 2106.05(d)(II). Therefore, Claims 8 and 17 do not recite patent eligible subject matter under 35 U.S.C. § 101. Claims 9-10 and 19-20: “detecting the level of the system load includes detecting that the system load has increased relative to the threshold value, wherein dynamically changing the number of groups includes increasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes decreasing the number of CPU cores in each group by the predetermined factor”, and “detecting the level of the system load includes detecting that the system load has decreased relative to the threshold value, wherein dynamically changing the number of groups includes decreasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes increasing the number of CPU cores in each group by the predetermined factor” as drafted, is a process that, but for the recitation of generic computing components, under its broadest reasonable interpretation, covers performance of the limitation in the mind. For example, but for the recitation of generic computing components being used as a tool to perform the functionality, a person can mentally evaluate when a threshold is being reached and adjust resources groups based on that evaluation. With regard to integration into practical application and whether additional elements amount to significantly more, Claims 9-10 and 19-20 fail both prongs of Step 2A, thus the claims are directed to the judicial exception as it has not been integrated into practical application, and fails Step 2B as not amounting to significantly more. Therefore, Claims 9-10 and 19-20 do not recite patent eligible subject matter under 35 U.S.C. § 101. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. Claims 1, 7, 11 and 18 are rejected under 35 U.S.C. 103(a) as being unpatentable over Yang et al. (US 20220210097 A1) (hereinafter Yang), in view of Kondapuram et al. (US 20210149736 A1) (hereinafter Kondapuram). Regarding Claim 1, Yang teaches: A method comprising: monitoring a system load of a storage node, the storage node including a multi core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; “Target 150 can access NVMe subsystems to read or write data. Target 150 can provide responses to NVMe commands in accordance with applicable protocols”, (Yang: ¶14), “Storage subsystem can include NVMe subsystems 180-0 to 180-n, where n is an integer. One or more storage and/or memory devices can be accessed by target in NVMe subsystems 180-0 to 180-n”, (Yang: ¶23), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair.”, (Yang: ¶25), “Network interface device 160 can utilize queues 162 to store packets received from an initiator or sender or packets to be transmitted by target 170. Queues 162 can store packets received through one or multiple connections for processing by a central processing unit (CPU) core by grouping connections together under the same identifier and avoiding locking or stalling from contention for queue accesses”, (Yang: ¶18), “packets and/or commands in the packets stored in queues 162 can be processed by a polling group executed in a thread which executes on a CPU”, (Yang: ¶19), " An NVMe-oF target application can be executed by cores 300-0 and 300-1. NVMe over QUIC transport on the target side can start-up with a fixed number of CPU cores allocated to monitor for received NVMe commands.”, (Yang: ¶37). Examiner notes: Queues 162 is being interpreted as the shared queue and the “threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands” is being interpreted as the group of cores. detecting a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically changing a number of groups of CPU cores, and dynamically changing a number of CPU cores in each group. “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 1, Yang fails to explicitly teach: A method comprising: monitoring a system load of a storage node, the storage node including a multi core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; However, Kondapuram teaches: “monitors CPU core utilization as an indicator of an amount of stress to which the CPU cores are subjected and makes cores switches based on the monitoring”, (Kondapuram: ¶62), “a policy is based on degradation of the cores which can be caused by a temperature in which the cores are operating and/or an amount of usage (workload) the cores are experiencing”, (Kondapuram: ¶33), “CPU core utilization is monitored as an indicator of an amount of stress to which the CPU cores are subjected.”, (Kondapuram: ¶41). detecting a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically changing a number of groups of CPU cores, and dynamically changing a number of CPU cores in each group. However, Kondapuram teaches: “The number of groups/subsets can vary based on the configuration of the cores and/or the electronic device in which the cores are disposed, a manner in which monitoring data is collected, etc. In some examples, the determination as to the number of groups/subsets of cores to be created is informed, in part, by the example policy selector 412” … “a policy is based on degradation of the cores which can be caused by a temperature in which the cores are operating and/or an amount of usage (workload) the cores are experiencing”, (Kondapuram: ¶33), “the example subset selector 414 operates to select one or more groups/subsets of cores to be switched from a first state (e.g., inactive) to a second state (e.g., active) and vice versa.”, (Kondapuram: ¶34), “the cores switcher algorithm causes switching groups of round robin (RR) cores formed/created by the cores partition 604 (616). In some examples, a different RR group is deactivated during each cores switch and the remaining ones of the RR group are activated (or, if already active, are not affected)”, (Kondapuram: ¶62). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine monitoring a system load of a storage node, the storage node including a multi core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; detecting a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically changing a number of groups of CPU cores, and dynamically changing a number of CPU cores in each group of Kondapuram with the methods and systems of Yang resulting in a system being able to resize resource groups based on monitored usage. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “extended the lifespan of CPU cores through strategic switching of core usage” … “improve the efficiency of using a computing device by extending the lifespan of a product that incorporates the computing device” … “cause cores to experience less degradation from high temperatures and heavy workloads and thereby can permit the usage of commercial grade cores in industrial applications”, (Kondapuram: ¶73). Regarding Claim 7, Yang teaches: assigning a CPU core with a high average queue polling frequency to each group; “a polling group executed in a thread which executes on a CPU”, (Yang: ¶19), “more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair.”, (Yang: ¶25). Further regarding Claim 7, Yang fails to teach: and assigning a low-stressed CPU core to each group including a high-stressed CPU core. However, Kondapuram teaches: “CPU core utilization is monitored as an indicator of an amount of stress to which the CPU cores are subjected” … “round-robin load balancing algorithm to distribute workloads across the cores based on the CPU core utilization data. In some examples, the workloads are distributed with a goal of achieving an extended CPU core lifespan of twice the non-extended lifespan.”, (Kondapuram: ¶41). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine assigning a low-stressed CPU core to each group including a high-stressed CPU core of Kondapuram with the methods and systems of Yang resulting in a system being able to groups cores based on core stress. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “extended the lifespan of CPU cores through strategic switching of core usage” … “improve the efficiency of using a computing device by extending the lifespan of a product that incorporates the computing device” … “cause cores to experience less degradation from high temperatures and heavy workloads and thereby can permit the usage of commercial grade cores in industrial applications”, (Kondapuram: ¶73). Regarding Claim 11, Yang teaches: A system comprising: a memory; and processing circuitry configured to execute program instructions out of the memory to: monitor a system load of a storage node, the storage node including a multi-core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; “Target 150 can access NVMe subsystems to read or write data. Target 150 can provide responses to NVMe commands in accordance with applicable protocols”, (Yang: ¶14), “Storage subsystem can include NVMe subsystems 180-0 to 180-n, where n is an integer. One or more storage and/or memory devices can be accessed by target in NVMe subsystems 180-0 to 180-n”, (Yang: ¶23), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair.”, (Yang: ¶25), “Network interface device 160 can utilize queues 162 to store packets received from an initiator or sender or packets to be transmitted by target 170. Queues 162 can store packets received through one or multiple connections for processing by a central processing unit (CPU) core by grouping connections together under the same identifier and avoiding locking or stalling from contention for queue accesses”, (Yang: ¶18), “packets and/or commands in the packets stored in queues 162 can be processed by a polling group executed in a thread which executes on a CPU”, (Yang: ¶19), " An NVMe-oF target application can be executed by cores 300-0 and 300-1. NVMe over QUIC transport on the target side can start-up with a fixed number of CPU cores allocated to monitor for received NVMe commands.”, (Yang: ¶37). Examiner notes: Queues 162 is being interpreted as the shared queue and the “threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands” is being interpreted as the group of cores. detect a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically change a number of groups of CPU cores, and dynamically change a number of CPU cores in each group. “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 11, Yang fails to explicitly teach: monitor a system load of a storage node, the storage node including a multi-core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; However, Kondapuram teaches: “monitors CPU core utilization as an indicator of an amount of stress to which the CPU cores are subjected and makes cores switches based on the monitoring”, (Kondapuram: ¶62), “a policy is based on degradation of the cores which can be caused by a temperature in which the cores are operating and/or an amount of usage (workload) the cores are experiencing”, (Kondapuram: ¶33), “CPU core utilization is monitored as an indicator of an amount of stress to which the CPU cores are subjected.”, (Kondapuram: ¶41). detect a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically change a number of groups of CPU cores, and dynamically change a number of CPU cores in each group. However, Kondapuram teaches: “The number of groups/subsets can vary based on the configuration of the cores and/or the electronic device in which the cores are disposed, a manner in which monitoring data is collected, etc. In some examples, the determination as to the number of groups/subsets of cores to be created is informed, in part, by the example policy selector 412” … “a policy is based on degradation of the cores which can be caused by a temperature in which the cores are operating and/or an amount of usage (workload) the cores are experiencing”, (Kondapuram: ¶33), “the example subset selector 414 operates to select one or more groups/subsets of cores to be switched from a first state (e.g., inactive) to a second state (e.g., active) and vice versa.”, (Kondapuram: ¶34), “the cores switcher algorithm causes switching groups of round robin (RR) cores formed/created by the cores partition 604 (616). In some examples, a different RR group is deactivated during each cores switch and the remaining ones of the RR group are activated (or, if already active, are not affected)”, (Kondapuram: ¶62). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine monitor a system load of a storage node, the storage node including a multi-core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; detect a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically change a number of groups of CPU cores, and dynamically change a number of CPU cores in each group of Kondapuram with the methods and systems of Yang resulting in a system being able to resize resource groups based on monitored usage. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “extended the lifespan of CPU cores through strategic switching of core usage” … “improve the efficiency of using a computing device by extending the lifespan of a product that incorporates the computing device” … “cause cores to experience less degradation from high temperatures and heavy workloads and thereby can permit the usage of commercial grade cores in industrial applications”, (Kondapuram: ¶73). Regarding Claim 18, Yang teaches: A computer program product including a set of non-transitory, computer-readable media having instructions that, when executed by processing circuitry, cause the processing circuitry to perform a method comprising: monitoring a system load of a storage node, the storage node including a multi core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; “Target 150 can access NVMe subsystems to read or write data. Target 150 can provide responses to NVMe commands in accordance with applicable protocols”, (Yang: ¶14), “Storage subsystem can include NVMe subsystems 180-0 to 180-n, where n is an integer. One or more storage and/or memory devices can be accessed by target in NVMe subsystems 180-0 to 180-n”, (Yang: ¶23), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair.”, (Yang: ¶25), “Network interface device 160 can utilize queues 162 to store packets received from an initiator or sender or packets to be transmitted by target 170. Queues 162 can store packets received through one or multiple connections for processing by a central processing unit (CPU) core by grouping connections together under the same identifier and avoiding locking or stalling from contention for queue accesses”, (Yang: ¶18), “packets and/or commands in the packets stored in queues 162 can be processed by a polling group executed in a thread which executes on a CPU”, (Yang: ¶19), " An NVMe-oF target application can be executed by cores 300-0 and 300-1. NVMe over QUIC transport on the target side can start-up with a fixed number of CPU cores allocated to monitor for received NVMe commands.”, (Yang: ¶37). Examiner notes: Queues 162 is being interpreted as the shared queue and the “threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands” is being interpreted as the group of cores. detecting a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically changing a number of groups of CPU cores, and dynamically changing a number of CPU cores in each group. “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 18, Yang fails to explicitly teach: monitoring a system load of a storage node, the storage node including a multi core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; However, Kondapuram teaches: “monitors CPU core utilization as an indicator of an amount of stress to which the CPU cores are subjected and makes cores switches based on the monitoring”, (Kondapuram: ¶62), “a policy is based on degradation of the cores which can be caused by a temperature in which the cores are operating and/or an amount of usage (workload) the cores are experiencing”, (Kondapuram: ¶33), “CPU core utilization is monitored as an indicator of an amount of stress to which the CPU cores are subjected.”, (Kondapuram: ¶41). detecting a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically changing a number of groups of CPU cores, and dynamically changing a number of CPU cores in each group. However, Kondapuram teaches: “The number of groups/subsets can vary based on the configuration of the cores and/or the electronic device in which the cores are disposed, a manner in which monitoring data is collected, etc. In some examples, the determination as to the number of groups/subsets of cores to be created is informed, in part, by the example policy selector 412” … “a policy is based on degradation of the cores which can be caused by a temperature in which the cores are operating and/or an amount of usage (workload) the cores are experiencing”, (Kondapuram: ¶33), “the example subset selector 414 operates to select one or more groups/subsets of cores to be switched from a first state (e.g., inactive) to a second state (e.g., active) and vice versa.”, (Kondapuram: ¶34), “the cores switcher algorithm causes switching groups of round robin (RR) cores formed/created by the cores partition 604 (616). In some examples, a different RR group is deactivated during each cores switch and the remaining ones of the RR group are activated (or, if already active, are not affected)”, (Kondapuram: ¶62). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine monitoring a system load of a storage node, the storage node including a multi core central processing unit (CPU) having multiple CPU cores grouped for polling at least one shared queue for received messages; detecting a level of the system load relative to a threshold value; and in response to the detected level of the system load: dynamically changing a number of groups of CPU cores, and dynamically changing a number of CPU cores in each group of Kondapuram with the methods and systems of Yang resulting in a system being able to resize resource groups based on monitored usage. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “extended the lifespan of CPU cores through strategic switching of core usage” … “improve the efficiency of using a computing device by extending the lifespan of a product that incorporates the computing device” … “cause cores to experience less degradation from high temperatures and heavy workloads and thereby can permit the usage of commercial grade cores in industrial applications”, (Kondapuram: ¶73). Claims 2 and 12 are rejected under 35 U.S.C. 103(a) as being unpatentable over Yang in view of Kondapuram, in further view of Nichols et al. (US 20180113738 A1) (hereinafter Nichols). Regarding Claim 2, Yang teaches: upon initialization of the storage node “socket initialization (init) module 404 can create a socket descriptor, set a socket priority based on the configuration of queues, bind a socket to a particular port, and monitor the socket for received NVMe commands.”, (Yang: ¶44), “NVMe over QUIC transport on the target side can start-up with a fixed number of CPU cores allocated to monitor for received NVMe commands”, (Yang: ¶37), “At 602, a new connection for NVMe-over-QUIC commands can be detected” … “At 606, processor resources can be allocated to monitor for received NVMe commands on the QUIC connection.”, (Yang: ¶49). Examiner notes: when a new connection is detected, processor resources are allocated to monitor the connection. Further regarding Claim 2, Yang in view of Kondapuram fails to teach: upon initialization of the storage node, assigning all the CPU cores to a single group; allocating a single shared queue for the single group; and polling, by the CPU cores assigned to the single group, the single shared queue for received messages. However, Nichols teaches: “a global queue that is accessible by all of the processing cores.”, (Nichols: Abstract), “These work node software objects are placed in a global queue, which each of the cores 204.a-204.d have access to”, (Nichols: ¶41), “any available CPU core may access the global queue and pull the next available work node software object from the global queue to process (or multiple work node software objects, where the global queue is a collection of queue heads)”, (Nichols: ¶19), “the global queue 402 includes a plurality of slots 404.a, 404.b, 404.c, 404.d, 404.e, 404.f, and 404.g” … “The slots 404.f, 404.g, and 404.a (i.e., in a circular queue as queue 402) are available to receive work node software objects 302”, (Nichols: ¶51). Examiner notes: the single group of cores that are able to access the global queue are all the cores of the system. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine upon initialization of the storage node, assigning all the CPU cores to a single group; allocating a single shared queue for the single group; and polling, by the CPU cores assigned to the single group, the single shared queue for received messages of Nichols with the methods and systems of Yang in view of Kondapuram resulting in a system having a global queue that may be accessed by all cores. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “the storage system 102 to more efficiently process I/O by dynamically balancing CPU workload to under-utilized CPU cores” … “efficiently execute non-SMP applications in a multicore environment, even where the tasks do not enjoy multicore protection” … “allows the use of multiple CPU cores in order to provide higher storage system performance while utilizing CPU resources in a more balanced fashion.”, (Nichols: ¶107). Regarding Claim 12, Yang teaches: upon initialization of the storage node “socket initialization (init) module 404 can create a socket descriptor, set a socket priority based on the configuration of queues, bind a socket to a particular port, and monitor the socket for received NVMe commands.”, (Yang: ¶44), “NVMe over QUIC transport on the target side can start-up with a fixed number of CPU cores allocated to monitor for received NVMe commands”, (Yang: ¶37), “At 602, a new connection for NVMe-over-QUIC commands can be detected” … “At 606, processor resources can be allocated to monitor for received NVMe commands on the QUIC connection.”, (Yang: ¶49). Examiner notes: when a new connection is detected, processor resources are allocated to monitor the connection. Further regarding Claim 12, Yang in view of Kondapuram fails to teach: upon initialization of the storage node, assign all the CPU cores to a single group; allocate a single shared queue for the single group; and poll, by the CPU cores assigned to the single group, the single shared queue for received messages. However, Nichols teaches: “a global queue that is accessible by all of the processing cores.”, (Nichols: Abstract), “These work node software objects are placed in a global queue, which each of the cores 204.a-204.d have access to”, (Nichols: ¶41), “any available CPU core may access the global queue and pull the next available work node software object from the global queue to process (or multiple work node software objects, where the global queue is a collection of queue heads)”, (Nichols: ¶19), “the global queue 402 includes a plurality of slots 404.a, 404.b, 404.c, 404.d, 404.e, 404.f, and 404.g” … “The slots 404.f, 404.g, and 404.a (i.e., in a circular queue as queue 402) are available to receive work node software objects 302”, (Nichols: ¶51). Examiner notes: the single group of cores that are able to access the global queue are all the cores of the system. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine upon initialization of the storage node, assign all the CPU cores to a single group; allocate a single shared queue for the single group; and poll, by the CPU cores assigned to the single group, the single shared queue for received messages of Nichols with the methods and systems of Yang in view of Kondapuram resulting in a system having a global queue that may be accessed by all cores. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “the storage system 102 to more efficiently process I/O by dynamically balancing CPU workload to under-utilized CPU cores” … “efficiently execute non-SMP applications in a multicore environment, even where the tasks do not enjoy multicore protection” … “allows the use of multiple CPU cores in order to provide higher storage system performance while utilizing CPU resources in a more balanced fashion.”, (Nichols: ¶107). Claims 3-7 and 13-16 are rejected under 35 U.S.C. 103(a) as being unpatentable over Yang in view of Kondapuram and Nichols, in further view of Ibrahim et al. (US 20200293445 A1) (hereinafter Ibrahim). Regarding Claim 3, Yang teaches: detecting the level of the system load includes detecting that the system load has increased relative to the threshold value, “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 3, Yang in view of Kondapuram and Nichols fails to teach: wherein dynamically changing the number of groups includes increasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes decreasing the number of CPU cores in each group by the predetermined factor. However, Ibrahim teaches: “the processing system 100 balances competing factors of the L1 cache 116 miss rate and L1 cache 116 access latency to fit a target application profile by dynamically changing the number of clusters”, (Ibrahim: ¶24), “the GPU 204 change from a single, shared L1 cache configuration (i.e., all four CUs grouped into one single CU cluster) to the configuration illustrated for GPU 214 in which two CU clusters 120(2) and 120(3) each include two CUs”, (Ibrahim: ¶46), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23), “the available clustering options group the CUs into two (as shown in FIG. 5), four (e.g., CU1 and CU2 belonging to one cluster, CU3 and CU4 to another, and the like), and eight (i.e., the default private L1 cache model) clusters.”, (Ibrahim: ¶40). Examiner notes: a predetermined factor of 2 is used when scaling clusters which are being interpreted as groups and CU’s which are being interpreted as cores. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine wherein dynamically changing the number of groups includes increasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes decreasing the number of CPU cores in each group by the predetermined factor of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “a fine-grained interleaving at the cache line granularity to ensure better distribution or dynamically increasing the number of clusters (decreasing the CUs per CU cluster) to better distribute the processing load”, (Ibrahim: ¶67), “the CU clustering discussed herein reduces pressure on LLC and increases compute performance by improving L1 hit rates”, (Ibrahim: ¶73). Regarding Claim 4, Yang teaches: allocating a plurality of shared queues for the increased number of groups, respectively; and polling the plurality of shared queues for received messages by the decreased number of CPU cores assigned to the increased number of groups, respectively. “One or more queues can be allocated to store packets for NVMe-over-QUIC”, (Yang: ¶49), “allocation of exclusive consumer queues for NVMe over QUIC packets (or packets transmitted using other transport protocols) 154 can provide for queues allocated in a memory of a network interface device” … “exclusively accessed by one or more processors in connection with packet receipt or transmission.”, (Yang: ¶16), “Group-based connection handling 410 can perform polling for newly received NVMe commands received using a QUIC connection”, (Yang: ¶46). Further regarding Claim 4, Yang in view of Kondapuram and Nichols fails to teach: assigning the decreased number of CPU cores to each of the increased number of groups; allocating a plurality of shared queues for the increased number of groups, respectively; and polling the plurality of shared queues for received messages by the decreased number of CPU cores assigned to the increased number of groups, respectively. However, Ibrahim teaches: “the processing system 100 balances competing factors of the L1 cache 116 miss rate and L1 cache 116 access latency to fit a target application profile by dynamically changing the number of clusters”, (Ibrahim: ¶24), “the GPU 204 change from a single, shared L1 cache configuration (i.e., all four CUs grouped into one single CU cluster) to the configuration illustrated for GPU 214 in which two CU clusters 120(2) and 120(3) each include two CUs”, (Ibrahim: ¶46), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23), “the available clustering options group the CUs into two (as shown in FIG. 5), four (e.g., CU1 and CU2 belonging to one cluster, CU3 and CU4 to another, and the like), and eight (i.e., the default private L1 cache model) clusters.”, (Ibrahim: ¶40). Examiner notes: a predetermined factor of 2 is used when scaling clusters which are being interpreted as groups and CU’s which are being interpreted as cores. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine assigning the decreased number of CPU cores to each of the increased number of groups; allocating a plurality of shared queues for the increased number of groups, respectively and polling the plurality of shared queues for received messages by the decreased number of CPU cores assigned to the increased number of groups, respectively of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “a fine-grained interleaving at the cache line granularity to ensure better distribution or dynamically increasing the number of clusters (decreasing the CUs per CU cluster) to better distribute the processing load”, (Ibrahim: ¶67), “the CU clustering discussed herein reduces pressure on LLC and increases compute performance by improving L1 hit rates”, (Ibrahim: ¶73). Regarding Claim 5, Yang teaches: detecting the level of the system load includes detecting that the system load has decreased relative to the threshold value, “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 5, Yang in view of Kondapuram and Nichols fails to teach: wherein dynamically changing the number of groups includes decreasing the number of groups by the predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes increasing the number of CPU cores in each group by the predetermined factor. However, Ibrahim teaches: “resulting CU/L1 configuration, with fewer clusters and more CUs per cluster”, (Ibrahim: ¶15), “The L1 miss rate determines whether the GPU should keep using the current clustering configuration or reconfigure to a more fine-grained address interleaving (i.e., fewer clusters and more CUs per cluster)”, (Ibrahim: ¶58), “the GPU increases the number of CUs per CU cluster from a first number (e.g., two CUs 112 per CU cluster 120(2)/120(3) in FIG. 2) to a second number greater than the first number (e.g., four CUs 112 in the CU cluster 120(1) of FIG. 2)”, (Ibrahim: ¶60), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23). Examiner notes: the number of CU’s per cluster increases by a factor of 2 and the number of clusters decreases by 2 It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine wherein dynamically changing the number of groups includes decreasing the number of groups by the predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes increasing the number of CPU cores in each group by the predetermined factor of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “with fewer clusters and more CUs per cluster, provides for higher hit rates and reduces pressure on LLC caches”, (Ibrahim: ¶15), “decreasing the number of CU clusters 120 results in a decrease in the number of cache line replicas at the GPU 104 and a larger effective L1 cache capacity within each CU cluster that decreases miss rates to the L1 cache at the computational expense of longer L1 access latency”, (Ibrahim: ¶24). Regarding Claim 6, Yang teaches: allocating a decreased plurality of shared queues for the decreased number of groups, respectively; and polling the decreased plurality of shared queues for received messages by the increased number of CPU cores assigned to the decreased number of groups, respectively. “One or more queues can be allocated to store packets for NVMe-over-QUIC”, (Yang: ¶49), “allocation of exclusive consumer queues for NVMe over QUIC packets (or packets transmitted using other transport protocols) 154 can provide for queues allocated in a memory of a network interface device” … “exclusively accessed by one or more processors in connection with packet receipt or transmission.”, (Yang: ¶16), “Group-based connection handling 410 can perform polling for newly received NVMe commands received using a QUIC connection”, (Yang: ¶46), “ A one-to-one mapping between queues and processors can be made, so that with x queues and x threads, x cores are utilized”, (Yang: ¶18). Examiner notes: yang teaches a direct relationship between queues and cores so if the number of cores utilized are decreased then the number of queues are decreased as well. Further regarding Claim 6, Yang in view of Kondapuram and Nichols fails to teach: assigning the increased number of CPU cores to each of the decreased number of groups; allocating a decreased plurality of shared queues for the decreased number of groups, respectively; and polling the decreased plurality of shared queues for received messages by the increased number of CPU cores assigned to the decreased number of groups, respectively. However, Ibrahim teaches: “resulting CU/L1 configuration, with fewer clusters and more CUs per cluster”, (Ibrahim: ¶15), “The L1 miss rate determines whether the GPU should keep using the current clustering configuration or reconfigure to a more fine-grained address interleaving (i.e., fewer clusters and more CUs per cluster)”, (Ibrahim: ¶58), “the GPU increases the number of CUs per CU cluster from a first number (e.g., two CUs 112 per CU cluster 120(2)/120(3) in FIG. 2) to a second number greater than the first number (e.g., four CUs 112 in the CU cluster 120(1) of FIG. 2)”, (Ibrahim: ¶60), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23). Examiner notes: the number of CU’s per cluster increases by a factor of 2 and the number of clusters decreases by 2 It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine assigning the increased number of CPU cores to each of the decreased number of groups; allocating a decreased plurality of shared queues for the decreased number of groups, respectively; and polling the decreased plurality of shared queues for received messages by the increased number of CPU cores assigned to the decreased number of groups, respectively of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “with fewer clusters and more CUs per cluster, provides for higher hit rates and reduces pressure on LLC caches”, (Ibrahim: ¶15), “decreasing the number of CU clusters 120 results in a decrease in the number of cache line replicas at the GPU 104 and a larger effective L1 cache capacity within each CU cluster that decreases miss rates to the L1 cache at the computational expense of longer L1 access latency”, (Ibrahim: ¶24). Regarding Claim 7, Yang teaches: assigning a CPU core with a high average queue polling frequency to each group; “a polling group executed in a thread which executes on a CPU”, (Yang: ¶19), “more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair.”, (Yang: ¶25). Further regarding Claim 7, Yang fails to teach: and assigning a low-stressed CPU core to each group including a high-stressed CPU core. However, Kondapuram teaches: “CPU core utilization is monitored as an indicator of an amount of stress to which the CPU cores are subjected” … “round-robin load balancing algorithm to distribute workloads across the cores based on the CPU core utilization data. In some examples, the workloads are distributed with a goal of achieving an extended CPU core lifespan of twice the non-extended lifespan.”, (Kondapuram: ¶41). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine assigning a low-stressed CPU core to each group including a high-stressed CPU core of Kondapuram with the methods and systems of Yang resulting in a system being able to groups cores based on core stress. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “extended the lifespan of CPU cores through strategic switching of core usage” … “improve the efficiency of using a computing device by extending the lifespan of a product that incorporates the computing device” … “cause cores to experience less degradation from high temperatures and heavy workloads and thereby can permit the usage of commercial grade cores in industrial applications”, (Kondapuram: ¶73). Further regarding Claim 7, Yang in view of Kondapuram fails to teach: assigning CPU cores that execute similar application threads to the same group; However, Nichols teaches: “The application 214 and the operating system 216 may each be broken down into a discrete amount of tasks. For example, an application 214 may be broken down into multiple tasks based on their similarity to each other”, (Nichols: ¶39), “may be scheduled on any available CPU core together”, (Nichols: ¶46), “These multiple CPU cores are configured to execute one or more application tasks that have been combined into different work node software objects and placed on a queue for the next available CPU core”, (Nichols: ¶25), “Multi-core processor 202 executes instructions of the application 214 and/or operating system 216 to perform operations of the storage controller 108”, (Nichols: ¶39). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine assigning CPU cores that execute similar application threads to the same group of Nichols with the methods and systems of Yang in view of Kondapuram resulting in a system grouping cores based on what application they are processing. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “the storage system 102 to more efficiently process I/O by dynamically balancing CPU workload to under-utilized CPU cores” … “efficiently execute non-SMP applications in a multicore environment, even where the tasks do not enjoy multicore protection” … “allows the use of multiple CPU cores in order to provide higher storage system performance while utilizing CPU resources in a more balanced fashion.”, (Nichols: ¶107). Further regarding Claim 7, Yang in view of Kondapuram and Nichols fails to teach: assigning CPU cores that share a cache level to the same group; However, Ibrahim teaches: “the CU cluster 120(2) shares L1 caches 116(1) through 116(K) amongst the CUs 112 of CU cluster 120(2) by interleaving the memory address range for operating the shared L1 caches 116(1)-116(K) as one logical cache. The shared L1 caches 116(1) through 116(K) in CU cluster 120(2) (which are private to the CU cluster 120(2) but available for sharing to the CUs 112(N-2)-112(N)) operates as a shared resource and allows for a larger effective L1 cache capacity without increasing the actual L1 cache size of each individual L1 cache 116”, (Ibrahim: ¶22), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure”, (Ibrahim: ¶23). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine assigning CPU cores that share a cache level to the same group of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of grouping cores based on their shared cache levels. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “with fewer clusters and more CUs per cluster, provides for higher hit rates and reduces pressure on LLC caches”, (Ibrahim: ¶15), “decreasing the number of CU clusters 120 results in a decrease in the number of cache line replicas at the GPU 104 and a larger effective L1 cache capacity within each CU cluster that decreases miss rates to the L1 cache at the computational expense of longer L1 access latency”, (Ibrahim: ¶24). Regarding Claim 13, Yang teaches: detect that the system load has increased relative to the threshold value; “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 13, yang in view of Kondapuram and Nichols fails to teach: increase the number of groups by a predetermined factor; and decrease the number of CPU cores in each group by the predetermined factor. However, Ibrahim teaches: “the processing system 100 balances competing factors of the L1 cache 116 miss rate and L1 cache 116 access latency to fit a target application profile by dynamically changing the number of clusters”, (Ibrahim: ¶24), “the GPU 204 change from a single, shared L1 cache configuration (i.e., all four CUs grouped into one single CU cluster) to the configuration illustrated for GPU 214 in which two CU clusters 120(2) and 120(3) each include two CUs”, (Ibrahim: ¶46), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23), “the available clustering options group the CUs into two (as shown in FIG. 5), four (e.g., CU1 and CU2 belonging to one cluster, CU3 and CU4 to another, and the like), and eight (i.e., the default private L1 cache model) clusters.”, (Ibrahim: ¶40). Examiner notes: a predetermined factor of 2 is used when scaling clusters which are being interpreted as groups and CU’s which are being interpreted as cores. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine increase the number of groups by a predetermined factor; and decrease the number of CPU cores in each group by the predetermined factor of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “a fine-grained interleaving at the cache line granularity to ensure better distribution or dynamically increasing the number of clusters (decreasing the CUs per CU cluster) to better distribute the processing load”, (Ibrahim: ¶67), “the CU clustering discussed herein reduces pressure on LLC and increases compute performance by improving L1 hit rates”, (Ibrahim: ¶73). Regarding Claim 14, Yang teaches: allocate a plurality of shared queues for the increased number of groups, respectively; and poll the plurality of shared queues for received messages by the decreased number of CPU cores assigned to the increased number of groups, respectively. “One or more queues can be allocated to store packets for NVMe-over-QUIC”, (Yang: ¶49), “allocation of exclusive consumer queues for NVMe over QUIC packets (or packets transmitted using other transport protocols) 154 can provide for queues allocated in a memory of a network interface device” … “exclusively accessed by one or more processors in connection with packet receipt or transmission.”, (Yang: ¶16), “Group-based connection handling 410 can perform polling for newly received NVMe commands received using a QUIC connection”, (Yang: ¶46). Further regarding Claim 14, Yang in view of Kondapuram and Nichols fails to teach: assign the decreased number of CPU cores to each of the increased number of groups; allocate a plurality of shared queues for the increased number of groups, respectively; and poll the plurality of shared queues for received messages by the decreased number of CPU cores assigned to the increased number of groups, respectively. However, Ibrahim teaches: “the processing system 100 balances competing factors of the L1 cache 116 miss rate and L1 cache 116 access latency to fit a target application profile by dynamically changing the number of clusters”, (Ibrahim: ¶24), “the GPU 204 change from a single, shared L1 cache configuration (i.e., all four CUs grouped into one single CU cluster) to the configuration illustrated for GPU 214 in which two CU clusters 120(2) and 120(3) each include two CUs”, (Ibrahim: ¶46), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23), “the available clustering options group the CUs into two (as shown in FIG. 5), four (e.g., CU1 and CU2 belonging to one cluster, CU3 and CU4 to another, and the like), and eight (i.e., the default private L1 cache model) clusters.”, (Ibrahim: ¶40). Examiner notes: a predetermined factor of 2 is used when scaling clusters which are being interpreted as groups and CU’s which are being interpreted as cores. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine assign the decreased number of CPU cores to each of the increased number of groups; allocating a plurality of shared queues for the increased number of groups, respectively and polling the plurality of shared queues for received messages by the decreased number of CPU cores assigned to the increased number of groups, respectively of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “a fine-grained interleaving at the cache line granularity to ensure better distribution or dynamically increasing the number of clusters (decreasing the CUs per CU cluster) to better distribute the processing load”, (Ibrahim: ¶67), “the CU clustering discussed herein reduces pressure on LLC and increases compute performance by improving L1 hit rates”, (Ibrahim: ¶73). Regarding Claim 15, Yang teaches: detect that the system load has decreased relative to the threshold value; “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 15, Yang in view of Kondapuram and Nichols fails to teach: decrease the number of groups by the predetermined factor; and increase the number of CPU cores in each group by the predetermined factor. However, Ibrahim teaches: “resulting CU/L1 configuration, with fewer clusters and more CUs per cluster”, (Ibrahim: ¶15), “The L1 miss rate determines whether the GPU should keep using the current clustering configuration or reconfigure to a more fine-grained address interleaving (i.e., fewer clusters and more CUs per cluster)”, (Ibrahim: ¶58), “the GPU increases the number of CUs per CU cluster from a first number (e.g., two CUs 112 per CU cluster 120(2)/120(3) in FIG. 2) to a second number greater than the first number (e.g., four CUs 112 in the CU cluster 120(1) of FIG. 2)”, (Ibrahim: ¶60), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23). Examiner notes: the number of CU’s per cluster increases by a factor of 2 and the number of clusters decreases by 2 It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to decrease the number of groups by the predetermined factor; and increase the number of CPU cores in each group by the predetermined factor of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “with fewer clusters and more CUs per cluster, provides for higher hit rates and reduces pressure on LLC caches”, (Ibrahim: ¶15), “decreasing the number of CU clusters 120 results in a decrease in the number of cache line replicas at the GPU 104 and a larger effective L1 cache capacity within each CU cluster that decreases miss rates to the L1 cache at the computational expense of longer L1 access latency”, (Ibrahim: ¶24). Regarding Claim 16, Yang teaches: allocate a decreased plurality of shared queues for the decreased number of groups, respectively; and poll the decreased plurality of shared queues for received messages by the increased number of CPU cores assigned to the decreased number of groups, respectively. “One or more queues can be allocated to store packets for NVMe-over-QUIC”, (Yang: ¶49), “allocation of exclusive consumer queues for NVMe over QUIC packets (or packets transmitted using other transport protocols) 154 can provide for queues allocated in a memory of a network interface device” … “exclusively accessed by one or more processors in connection with packet receipt or transmission.”, (Yang: ¶16), “Group-based connection handling 410 can perform polling for newly received NVMe commands received using a QUIC connection”, (Yang: ¶46), “ A one-to-one mapping between queues and processors can be made, so that with x queues and x threads, x cores are utilized”, (Yang: ¶18). Examiner notes: yang teaches a direct relationship between queues and cores so if the number of cores utilized are decreased then the number of queues are decreased as well. Further regarding Claim 16, Yang in view of Kondapuram and Nichols fails to teach: assign the increased number of CPU cores to each of the decreased number of groups; allocate a decreased plurality of shared queues for the decreased number of groups, respectively; and poll the decreased plurality of shared queues for received messages by the increased number of CPU cores assigned to the decreased number of groups, respectively. However, Ibrahim teaches: “resulting CU/L1 configuration, with fewer clusters and more CUs per cluster”, (Ibrahim: ¶15), “The L1 miss rate determines whether the GPU should keep using the current clustering configuration or reconfigure to a more fine-grained address interleaving (i.e., fewer clusters and more CUs per cluster)”, (Ibrahim: ¶58), “the GPU increases the number of CUs per CU cluster from a first number (e.g., two CUs 112 per CU cluster 120(2)/120(3) in FIG. 2) to a second number greater than the first number (e.g., four CUs 112 in the CU cluster 120(1) of FIG. 2)”, (Ibrahim: ¶60), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23). Examiner notes: the number of CU’s per cluster increases by a factor of 2 and the number of clusters decreases by 2 It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine assign the increased number of CPU cores to each of the decreased number of groups; allocating a decreased plurality of shared queues for the decreased number of groups, respectively; and polling the decreased plurality of shared queues for received messages by the increased number of CPU cores assigned to the decreased number of groups, respectively of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “with fewer clusters and more CUs per cluster, provides for higher hit rates and reduces pressure on LLC caches”, (Ibrahim: ¶15), “decreasing the number of CU clusters 120 results in a decrease in the number of cache line replicas at the GPU 104 and a larger effective L1 cache capacity within each CU cluster that decreases miss rates to the L1 cache at the computational expense of longer L1 access latency”, (Ibrahim: ¶24). Claims 8 and 17 are rejected under 35 U.S.C. 103(a) as being unpatentable over Yang in view of Kondapuram, in further view of Hayes et al. (US 20160004613 A1) (hereinafter Hayes). Regarding Claim 8, Yang teaches: the multiple CPU cores are grouped for polling the at least one shared queue for received remote procedure call (RPC) messages, and wherein the method comprises: polling, by the number of groups of CPU cores, the at least one shared queue for an RPC request message from another storage node; “One or more queues can be allocated to store packets for NVMe-over-QUIC”, (Yang: ¶49), “allocation of exclusive consumer queues for NVMe over QUIC packets (or packets transmitted using other transport protocols) 154 can provide for queues allocated in a memory of a network interface device” … “exclusively accessed by one or more processors in connection with packet receipt or transmission.”, (Yang: ¶16), “Group-based connection handling 410 can perform polling for newly received NVMe commands received using a QUIC connection”, (Yang: ¶46). Further regarding Claim 8, Yang in view of Kondapuram fails to teach: the multiple CPU cores are grouped for polling the at least one shared queue for received remote procedure call (RPC) messages, and wherein the method comprises: polling, by the number of groups of CPU cores, the at least one shared queue for an RPC request message from another storage node; processing the RPC request message; generating an RPC reply message; and sending the RPC reply message to the other storage node. However, Hayes teaches: “When a remote procedure call arrives for servicing, the storage node 150 or the non-volatile solid-state storage 152 determines whether the remote procedure call has already been serviced.”, (Hayes: ¶46), “the remote procedure call is serviced and responded to with the result of that service”, (Hayes: ¶47), “The remote procedure call cache belonging to this (destination) storage node is checked to see if there is already a result from executing the remote procedure call”, (Hayes: ¶52), “. If the answer is no, the mirrored remote procedure call cache is reachable, flow branches to the action 618, in which the remote procedure call is serviced”, (Hayes: ¶54), “If a result has been posted, the result is returned as a response to the remote procedure call”, (Hayes: ¶46), “each storage node configured to forward a remote procedure call to one of the plurality of storage nodes”, (Hayes: Claim 11). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine the multiple CPU cores are grouped for polling the at least one shared queue for received remote procedure call (RPC) messages, and wherein the method comprises: polling, by the number of groups of CPU cores, the at least one shared queue for an RPC request message from another storage node; processing the RPC request message; generating an RPC reply message; and sending the RPC reply message to the other storage node of Hayes with the methods and systems of Yang in view on Kondapuram resulting in a system with storage nodes that are able to us RPC messages to transmit data to one another. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “prevents corruption that can occur when a failure disrupts a remote procedure call in progress” … “a storage cluster is able to continue operating despite a failure causing disruption of a remote procedure call service, loss of contents of or access to a remote procedure call cache, or loss of other resource(s) involved in servicing the remote procedure call”, (Hayes: ¶16). Regarding Claim 17, Yang teaches: the multiple CPU cores are grouped for polling the at least one shared queue for received remote procedure call (RPC) messages, poll, by the number of groups of CPU cores, the at least one shared queue for an RPC request message from another storage node; “One or more queues can be allocated to store packets for NVMe-over-QUIC”, (Yang: ¶49), “allocation of exclusive consumer queues for NVMe over QUIC packets (or packets transmitted using other transport protocols) 154 can provide for queues allocated in a memory of a network interface device” … “exclusively accessed by one or more processors in connection with packet receipt or transmission.”, (Yang: ¶16), “Group-based connection handling 410 can perform polling for newly received NVMe commands received using a QUIC connection”, (Yang: ¶46). Further regarding Claim 17, Yang in view of Kondapuram fails to teach: the multiple CPU cores are grouped for polling the at least one shared queue for received remote procedure call (RPC) messages, poll, by the number of groups of CPU cores, the at least one shared queue for an RPC request message from another storage node; process the RPC request message; generate an RPC reply message; and send the RPC reply message to the other storage node. However, Hayes teaches: “When a remote procedure call arrives for servicing, the storage node 150 or the non-volatile solid-state storage 152 determines whether the remote procedure call has already been serviced.”, (Hayes: ¶46), “the remote procedure call is serviced and responded to with the result of that service”, (Hayes: ¶47), “The remote procedure call cache belonging to this (destination) storage node is checked to see if there is already a result from executing the remote procedure call”, (Hayes: ¶52), “. If the answer is no, the mirrored remote procedure call cache is reachable, flow branches to the action 618, in which the remote procedure call is serviced”, (Hayes: ¶54), “If a result has been posted, the result is returned as a response to the remote procedure call”, (Hayes: ¶46), “each storage node configured to forward a remote procedure call to one of the plurality of storage nodes”, (Hayes: Claim 11). It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine the multiple CPU cores are grouped for polling the at least one shared queue for received remote procedure call (RPC) messages, poll, by the number of groups of CPU cores, the at least one shared queue for an RPC request message from another storage node; process the RPC request message; generate an RPC reply message; and send the RPC reply message to the other storage node of Hayes with the methods and systems of Yang in view on Kondapuram resulting in a system with storage nodes that are able to us RPC messages to transmit data to one another. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “prevents corruption that can occur when a failure disrupts a remote procedure call in progress” … “a storage cluster is able to continue operating despite a failure causing disruption of a remote procedure call service, loss of contents of or access to a remote procedure call cache, or loss of other resource(s) involved in servicing the remote procedure call”, (Hayes: ¶16). Claims 9-10 and 19-20 are rejected under 35 U.S.C. 103(a) as being unpatentable over Yang in view of Kondapuram, in further view of Ibrahim. Regarding Claim 9, Yang teaches: detecting the level of the system load includes detecting that the system load has increased relative to the threshold value, “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 9, Yang in view of Kondapuram fails to teach: wherein dynamically changing the number of groups includes increasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes decreasing the number of CPU cores in each group by the predetermined factor. However, Ibrahim teaches: “the processing system 100 balances competing factors of the L1 cache 116 miss rate and L1 cache 116 access latency to fit a target application profile by dynamically changing the number of clusters”, (Ibrahim: ¶24), “the GPU 204 change from a single, shared L1 cache configuration (i.e., all four CUs grouped into one single CU cluster) to the configuration illustrated for GPU 214 in which two CU clusters 120(2) and 120(3) each include two CUs”, (Ibrahim: ¶46), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23), “the available clustering options group the CUs into two (as shown in FIG. 5), four (e.g., CU1 and CU2 belonging to one cluster, CU3 and CU4 to another, and the like), and eight (i.e., the default private L1 cache model) clusters.”, (Ibrahim: ¶40). Examiner notes: a predetermined factor of 2 is used when scaling clusters which are being interpreted as groups and CU’s which are being interpreted as cores. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine wherein dynamically changing the number of groups includes increasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes decreasing the number of CPU cores in each group by the predetermined factor of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “a fine-grained interleaving at the cache line granularity to ensure better distribution or dynamically increasing the number of clusters (decreasing the CUs per CU cluster) to better distribute the processing load”, (Ibrahim: ¶67), “the CU clustering discussed herein reduces pressure on LLC and increases compute performance by improving L1 hit rates”, (Ibrahim: ¶73). Regarding Claim 10, Yang teaches: detecting the level of the system load includes detecting that the system load has decreased relative to the threshold value, “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 10, Yang in view of Kondapuram fails to teach: wherein dynamically changing the number of groups includes decreasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes increasing the number of CPU cores in each group by the predetermined factor. However, Ibrahim teaches: “resulting CU/L1 configuration, with fewer clusters and more CUs per cluster”, (Ibrahim: ¶15), “The L1 miss rate determines whether the GPU should keep using the current clustering configuration or reconfigure to a more fine-grained address interleaving (i.e., fewer clusters and more CUs per cluster)”, (Ibrahim: ¶58), “the GPU increases the number of CUs per CU cluster from a first number (e.g., two CUs 112 per CU cluster 120(2)/120(3) in FIG. 2) to a second number greater than the first number (e.g., four CUs 112 in the CU cluster 120(1) of FIG. 2)”, (Ibrahim: ¶60), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23). Examiner notes: the number of CU’s per cluster increases by a factor of 2 and the number of clusters decreases by 2 It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine wherein dynamically changing the number of groups includes decreasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes increasing the number of CPU cores in each group by the predetermined factor of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “with fewer clusters and more CUs per cluster, provides for higher hit rates and reduces pressure on LLC caches”, (Ibrahim: ¶15), “decreasing the number of CU clusters 120 results in a decrease in the number of cache line replicas at the GPU 104 and a larger effective L1 cache capacity within each CU cluster that decreases miss rates to the L1 cache at the computational expense of longer L1 access latency”, (Ibrahim: ¶24). Regarding Claim 19, Yang teaches: detecting the level of the system load includes detecting that the system load has increased relative to the threshold value, “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 19, Yang in view of Kondapuram fails to teach: wherein dynamically changing the number of groups includes increasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes decreasing the number of CPU cores in each group by the predetermined factor. However, Ibrahim teaches: “the processing system 100 balances competing factors of the L1 cache 116 miss rate and L1 cache 116 access latency to fit a target application profile by dynamically changing the number of clusters”, (Ibrahim: ¶24), “the GPU 204 change from a single, shared L1 cache configuration (i.e., all four CUs grouped into one single CU cluster) to the configuration illustrated for GPU 214 in which two CU clusters 120(2) and 120(3) each include two CUs”, (Ibrahim: ¶46), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23), “the available clustering options group the CUs into two (as shown in FIG. 5), four (e.g., CU1 and CU2 belonging to one cluster, CU3 and CU4 to another, and the like), and eight (i.e., the default private L1 cache model) clusters.”, (Ibrahim: ¶40). Examiner notes: a predetermined factor of 2 is used when scaling clusters which are being interpreted as groups and CU’s which are being interpreted as cores. It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine wherein dynamically changing the number of groups includes increasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes decreasing the number of CPU cores in each group by the predetermined factor of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “a fine-grained interleaving at the cache line granularity to ensure better distribution or dynamically increasing the number of clusters (decreasing the CUs per CU cluster) to better distribute the processing load”, (Ibrahim: ¶67), “the CU clustering discussed herein reduces pressure on LLC and increases compute performance by improving L1 hit rates”, (Ibrahim: ¶73). Regarding Claim 20, Yang teaches: detecting the level of the system load includes detecting that the system load has decreased relative to the threshold value, “Some examples can scale a number of threads to perform an epoll group based on rate of NVMe over QUIC command receipt or NVMe over QUIC packet receipt so that more threads are used for a receive rate higher than a first threshold and fewer threads are used for a receive rate less than a second threshold”, (Yang: ¶21), “load balancing of detection received commands 152 can increase or decrease a number of threads or processors (e.g., cores) allocated to poll or monitor for received NVMe commands from initiator 100”, (Yang: ¶15), “a number of epoll group threads can be scaled up (increased) when more traffic load is present on a queue pair. For example, a number of epoll group threads can be scaled down (decreased) when less traffic load is present on queue pair”, (Yang: ¶25). Further regarding Claim 20, Yang in view of Kondapuram fails to teach: wherein dynamically changing the number of groups includes decreasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes increasing the number of CPU cores in each group by the predetermined factor. However, Ibrahim teaches: “resulting CU/L1 configuration, with fewer clusters and more CUs per cluster”, (Ibrahim: ¶15), “The L1 miss rate determines whether the GPU should keep using the current clustering configuration or reconfigure to a more fine-grained address interleaving (i.e., fewer clusters and more CUs per cluster)”, (Ibrahim: ¶58), “the GPU increases the number of CUs per CU cluster from a first number (e.g., two CUs 112 per CU cluster 120(2)/120(3) in FIG. 2) to a second number greater than the first number (e.g., four CUs 112 in the CU cluster 120(1) of FIG. 2)”, (Ibrahim: ¶60), “the CU clustering may be performed with CPU cores and the like without departing from the scope of this disclosure.”, (Ibrahim: ¶23). Examiner notes: the number of CU’s per cluster increases by a factor of 2 and the number of clusters decreases by 2 It would have been obvious to a person having ordinary skill in the art prior to the effective filing date of the claimed invention to combine wherein dynamically changing the number of groups includes decreasing the number of groups by a predetermined factor, and wherein dynamically changing the number of CPU cores in each group includes increasing the number of CPU cores in each group by the predetermined factor of Ibrahim with the methods and systems of Yang in view of Kondapuram and Nichols resulting in a system capable of scaling groups and cores based on utilization. A person having ordinary skill in the art would have been motivated to make this combination, with a reasonable expectation of success, for the purpose of “with fewer clusters and more CUs per cluster, provides for higher hit rates and reduces pressure on LLC caches”, (Ibrahim: ¶15), “decreasing the number of CU clusters 120 results in a decrease in the number of cache line replicas at the GPU 104 and a larger effective L1 cache capacity within each CU cluster that decreases miss rates to the L1 cache at the computational expense of longer L1 access latency”, (Ibrahim: ¶24). Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHIHAB ALAM whose telephone number is (571)272-8705. The examiner can normally be reached Mon - Fri 7:30am-5pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bradley Teets can be reached at (571) 272-3338. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /S.A./Examiner, Art Unit 2197 /BRADLEY A TEETS/Supervisory Patent Examiner, Art Unit 2197
Read full office action

Prosecution Timeline

Jun 04, 2024
Application Filed
Aug 21, 2026
Non-Final Rejection mailed — §101, §103, §112 (current)

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
Grant Probability
Low
PTA Risk
Based on 0 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month