DETAILED ACTION
Claims 1-20 are presented for examination.
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Continued Examination Under 37 CFR 1.114
A request for continued examination under 37 CFR 1.114, including the fee set forth in 37 CFR 1.17(e), was filed in this application after final rejection. Since this application is eligible for continued examination under 37 CFR 1.114, and the fee set forth in 37 CFR 1.17(e) has been timely paid, the finality of the previous Office action has been withdrawn pursuant to 37 CFR 1.114. Applicant's submission filed on 06/29/2026 has been entered.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1-3, 8-10 and 15-17 is/are rejected under 35 U.S.C. 103 as being unpatentable over Van De Groenendaal et al. (US PG Pub No. 2018/0026904 A1, hereinafter Groenendaal) in view of Gray (US Pat No. 10,048,871), further in view of Blocksome et al. (US Pat No. 8,930,956).
Regarding claim 1, Groenendaal teaches a processor, comprising: one or more circuits to:
receive,
generate a binding policy, based at least in part, on the one
cause performance of a software workload using
Groenendaal does not teach receive, via an interface, a selection of one non-uniform memory access (NUMA) binding option of a plurality of different NUMA binding options.
Gray teaches an interface to select a subset of one or more processors of a non-uniform memory access (NUMA) node to perform a software workload based, at least in part, on one or more user-specified parameters provided to the interface (col 2 line 35 to col 3 line 35, wherein a user-level tool (i.e. interface) is provided to allow users to assign new and/or pre-existing or executing processes to NUMA resources and can provide for achieving desirable NUMA affinity and the performance benefits). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to include an interface for selecting NUMA resources to be assigned. One would be motivated by the desire to predefine sets of NUMA resources to be allocated as taught by Gray (Abstract).
Groenendaal and Gray do not teach causing performance of a software workload using different selected subset of one or more processing resources of a NUMA node unique to different rank processes of the software workload.
Blocksome teaches well-known rank-based process identification wherein “rank” is defined as identifying a process wherein rank actually identifies a task or process that is executing a parallel operation (col 10 line 61 to col 11 line 4, wherein “Using the rank to identify a node assumes that only one such task is executing on each node. To the extent that more than one participating task executes on a single node, the rank identifies the task as such rather than the node”). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to assign different rank processes to different selected subsets of the one or more processing resources of the NUMA node. One would be motivated by the desire to utilize well known standardized indexing scheme used widely in parallel computing system as taught by Blocksome.
Regarding claim 2, Gray teaches wherein the user plurality of different NUMA binding options indicate one or more criteria for binding a process of the software workload to one or more of the one or more processing resources (col 3 lines 11-61).
Regarding claim 3, Gray does not teach wherein said software workload is a machine learning workload.
It is old and known to perform machine learning workload using NUMA systems. For example, Radke teaches using NUMA processing nodes for performing machine learning tasks ([0015]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention that the software workload is a machine learning workload. One would be motivated by the desire to utilize Gray for performing common tasks including machine learning workloads.
Regarding claims 8-10, and 15-17, they are the system and medium claims of claims 1-3 above.
Claim(s) 4, 11 and 18 is/are rejected under 35 U.S.C. 103 as being unpatentable over Van De Groenendaal et al. (US PG Pub No. 2018/0026904 A1, hereinafter Groenendaal) in view of Gray (US Pat No. 10,048,871), in view of Blocksome et al. (US Pat No. 8,930,956), further in view of Sato (US Pat No. 8,873,076).
Regarding claim 4, Van De Groenendaal, Gray, and Blocksome do not teach wherein the different, selected subsets of the one or more processing resources comprises unique to the different rank processes ensures that one of the different rank processes is bound to one or more central processing unit (CPUs) that are not shared with another one of the different rank processes.
Sato teaches assigning cores to different processing threads/processes via affinity masks such that core sets do not overlap (col 5 lines 15-16). By using this affinity mask, the control process may limit the core to be used for the thread to one or more cores designated by the affinity mask and if a plurality of threads is running, each core used by each of the threads is controlled by the affinity mask so that it is not used by a different thread (col 4 lines 56-61). Sato further teaches assigning a core to a thread is defined as the control process setting an affinity mask so that one thread runs on a particular core and other threads do not run on that core and assigning cores to each process is defined as assigning cores to one or more threads included in the process (col 4 line 62 to col 5 line 5; col 7 lines 51-62). In other words, each of the cores is assigned to each thread by the affinity mask so that the use of the cores does not overlap.
It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention to teach the different, selected subsets of the one or more processing resources comprises unique to the different rank processes ensures that one of the different rank processes is bound to one or more central processing unit (CPUs) that are not shared with another one of the different rank processes. One would be motivated by the desire to ensure to use affinity masks so that the use of the cores does not overlap between different concurrently running processes as taught by Sato.
Regarding claims 11 and 18, it is the system and medium claims of claim 4 above.
Claim(s) 5, 12, and 19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Van De Groenendaal et al. (US PG Pub No. 2018/0026904 A1, hereinafter Groenendaal) in view of Gray (US Pat No. 10,048,871), in view of Blocksome et al. (US Pat No. 8,930,956), further in view of Zhao et al. (US Pat No. 10,325,343).
Regarding claim 5, Gray does not teach wherein the subset of one or more processing resources is selected based, at least in part, on nearness of one or more of the subset of one or more central processing units (CPUs) to a graphics processing unit (GPU).
Zhao teaches the use of interconnects between CPUs and GPUs in NUMA nodes (col 11 lines 47 to col 12 line 8). Zhao teaches that static hardware factors, i.e. topology, can impact performance GPU services provided by GPU server node, such as the types of GPUs implemented in the GPU server node, the manner in which the GPUs are connected to CPUs and other GPUs, the distance of the communication path between a GPU and a network adapter (col 11 lines 50-56). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention that the subset of processors is selected based, at least in part, on nearness of one or more of the subset of one or more processors to a GPU. One would be motivated by the desire to account for performance factors based on GPU distance as taught by Zhao.
Regarding claims 12 and 19, they are the system and medium claims of claim 5 above.
Claim(s) 6-7, 13-14, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over Van De Groenendaal et al. (US PG Pub No. 2018/0026904 A1, hereinafter Groenendaal) in view of Gray (US Pat No. 10,048,871), in view of Blocksome et al. (US Pat No. 8,930,956), further in view of Ganguly et al. (US PG Pub No. 2022/0214825 A1).
Regarding claim 6, Gray does not teach wherein the subset of the one or more processing resources is selected based, at least in part, on a cache shared by one or more of the subset of the one or more processing resources.
Ganguly teaches defining multiple NUMA zones wherein each NUMA zone comprises of a socket, the processors within it, and the attached physical memory module, i.e. cache ([0006]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention that the subset of processors is selected based, at least in part, on a cache shared by one or more of the subset of one or more processors. One would be motivated by the desire to reduce the cost of access latency as taught by Ganguly.
Regarding claim 7, Gray does not teach wherein the subset of the one or more processing resources is selected based, at least in part, on a socket shared by one or more of the subset of the one or more processing resources.
Ganguly teaches defining multiple NUMA zones wherein each NUMA zone comprises of a socket, the processors within it, and the attached physical memory module ([0006]). It would have been obvious to one of ordinary skill in the art before the effective filing date of the invention that the subset of processors is selected based, at least in part, on a cache shared by one or more of the subset of one or more processors. One would be motivated by the desire to reduce the cost of access latency as taught by Ganguly.
Regarding claims 13-14 and 20, they are the system and medium claims of claims 6-7 above.
Response to Arguments
Applicant’s arguments with respect to claim(s) 1-20 have been considered but are moot because the new ground of rejection does not rely on any reference applied in the prior rejection of record for any teaching or matter specifically challenged in the argument.
Conclusion
Any inquiry concerning this communication or earlier communications from the examiner should be directed to ERIC C WAI whose telephone number is (571)270-1012. The examiner can normally be reached Monday - Friday 9-5.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Aimee Li can be reached at (571) 272-4169. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/Eric C Wai/Primary Examiner, Art Unit 2195