Non-Final Office Action
Notice of Pre-AIA or AIA Status
The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA .
Claim Rejections - 35 USC § 112
The following is a quotation of 35 U.S.C. 112(b):
(b) CONCLUSION.—The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the inventor or a joint inventor regards as the invention.
The following is a quotation of 35 U.S.C. 112 (pre-AIA ), second paragraph:
The specification shall conclude with one or more claims particularly pointing out and distinctly claiming the subject matter which the applicant regards as his invention.
Claims 12, 16, and 20 are rejected under 35 U.S.C. 112(b) or 35 U.S.C. 112 (pre-AIA ), second paragraph, as being indefinite for failing to particularly point out and distinctly claim the subject matter which the inventor or a joint inventor (or for applications subject to pre-AIA 35 U.S.C. 112, the applicant), regards as the invention.
Claim 12 recites “a voltage threshold draft sensor”, “a reference critical delay draft sensor”, and “a replication path timing/voltage draft sensor”. These terms do not have an ordinary and customary meaning to those of ordinary skill in the art. It is unclear as to what a “draft sensor” is. The specification does not provide a definition for these terms. Additionally, a search of the prior art did not provide a definition for the terms. The meaning of every term used in a claim should be apparent from the prior art or from the specification and drawings at the time the application is filed. Claim language may not be "ambiguous, vague, incoherent, opaque, or otherwise unclear in describing and defining the claimed invention." In re Packard, 751 F.3d 1307, 1311, 110 USPQ2d 1785, 1787 (Fed. Cir. 2014). Applicants need not confine themselves to the terminology used in the prior art, but are required to make clear and precise the terms that are used to define the invention whereby the metes and bounds of the claimed invention can be ascertained (see MPEP 2173.05(a)).
Claims 16 and 20 are rejected for similar reasons as claim 12.
Claim Rejections - 35 USC § 103
The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action:
A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made.
Claim(s) 1, 2, 4, 6, 8, 9, 11-15, and 17-19 is/are rejected under 35 U.S.C. 103 as being unpatentable over Thapa et al., US 2022/0253337 A1, and further in view of Kulkarni et al., US 2016/0140001 A1.
Referring to claim 1:
In para. 0010, Thapa et al. disclose a computer processing system comprising: a plurality of computational units (computing clusters) to execute tasks for one or more applications.
In para. 0016, Thapa et al. disclose a plurality of sensors, coupled to the plurality of computational units, to collect measurement data of the plurality of computational units.
In para. 0019 and 0023, Thapa et al. disclose the processing logic can flag or indicate the CPU as having an impending failure (wherein the hardware health statuses are based on the measurement data). However, Thapa et al. do not explicitly disclose a storing hardware health statuses corresponding to each of the plurality of computational units in a storage. In para. 0016, Kulkarni et al. disclose determining the health of a cluster (health status based on metrics). A controller does not consider a failed compute node when scheduling a job. In para. 0023-0024 and 0026, Kulkarni et al. disclose removing a node from an active list (see para. 0026) and placing a node with a major error on a list of nodes not to be used (see para. 0024). The list of Kulkarni et al. is based on health of the node (see para. 0016). Given the broadest, reasonable interpretation, the list of active nodes and major error nodes is a storage storing hardware health statuses (available or have major errors) corresponding to each of the plurality of computational units (nodes). It would have been obvious to one of ordinary skill at the time of filing of the invention to include the available node list of Kulkarni et al. into the system of Thapa et al. A person of ordinary skill in the art would have been motivated to make the modification because similar to Thapa et al., Kulkarni et al. is concerned with excluding failed nodes or nodes with impending failures from a task. Thapa et al. disclose flagging or indicating the CPU as having an impending failure and the list of Kulkarni et al. provides a means to flag or indicate. Using the list of Kulkarni et al. in the system of Thapa et al. would yield predictable results since both remove unhealthy nodes.
In para. 0010 and 0019, Thapa et al. disclose the plurality of computational units to be scheduled to perform task execution on the computer processing system for the one or more applications (para. 0010: plurality of computing clusters can be networked as a supercomputer to perform various complex computing tasks) based on the hardware health statuses of the plurality of computational units, wherein a first computational unit of the plurality of computational units is to be excluded from the task execution responsive to its corresponding hardware health status indicating an impending hardware failure (para. 0019: the management logic may preemptively allocate computing tasks assigned to the node processing module to other processing modules in response to an impending failure).
Referring to claims 2, 14, and 18, in para. 0024, Thapa et al. disclose wherein error screening content is to be scheduled to be executed along with the tasks for the one or more applications to further determine the hardware health statuses of the plurality of computational units (the telemetry monitoring logic can continuously monitor a node processing module while the processing module is executing computing tasks).
Referring to claim 4, in para. 0017, Thapa et al. disclose wherein the error screening content is to be executed as an application (processing logic is programmed).
Referring to claims 6, 15, and 19, in para. 0024, Thapa et al. disclose wherein the statuses of the plurality of computational units are to be updated based on the measurement data periodically obtained from the plurality of sensors (continuously monitor telemetries and predict impending faults).
Referring to claim 8, in para 0016, Thapa et al. disclose wherein the plurality of computational units includes one or more of following entities in the computer processing system: a set of central processing unit (CPU) cores, a set of graphics processing unit (GPU) cores, and a set of memory units.
Referring to claim 9, in para. 0016, Thapa et al. disclose wherein the plurality of computational units includes one or more components of the sets of CPU cores, GPU cores, and memory units.
Referring to claim 11, in para. 0010, 0019, and 0023, Thapa et al. disclose preemptively allocating computing tasks assigned to a node processing module to other processing module to mitigate an impending failure (wherein responsive to the hardware health status of a first set of computational units of no hardware health concerns (only nodes with an impending failure are excluded) and a hardware health status of a second set of computational units of no failure but an impending hardware failure, the second set of computational units are to be excluded from the task execution. Further, in para. 0016, Kulkarni et al. disclose determining the health of a cluster (health status based on metrics). A controller does not consider a failed compute node when scheduling a job. In para. 0023-0024 and 0026, Kulkarni et al. disclose removing a node from an active list (see para. 0026) and placing a node with a major error on a list of nodes not to be used (see para. 0024).
Referring to claim 12, in para. 0010, 0019, and 0023, Thapa et al. disclose preemptively allocating computing tasks assigned to a node processing module to other processing module to mitigate an impending failure (wherein tasks of a second computational unit of the plurality of computational units are to be moved (preemptively allocated) to a third computational unit of the plurality of computational units responsive to a corresponding second hardware health status of the second computational unit of an impending hardware failure and a corresponding third hardware health status of the third computational unit of no hardware health concerns (only nodes with an impending failure are excluded)).
Referring to claim 13:
In para. 0010, Thapa et al. disclose a method comprising: executing, by a plurality of computational units (computing clusters) of a computer processor system, tasks for one or more applications.
In para. 0016, Thapa et al. disclose collecting, by a plurality of sensors coupled to the plurality of computational units, measurement data of the plurality of computational units.
In para. 0019, Thapa et al. disclose the processing logic can flag or indicate the CPU as having an impending failure (wherein the hardware health statuses are based on the measurement data). However, Thapa et al. do not explicitly disclose storing hardware health statuses corresponding to each of the plurality of computational units in a storage. In para. 0016, Kulkarni et al. disclose determining the health of a cluster (health status based on metrics). A controller does not consider a failed compute node when scheduling a job. In para. 0023-0024 and 0026, Kulkarni et al. disclose removing a node from an active list (see para. 0026) and placing a node with a major error on a list of nodes not to be used (see para. 0024). The list of Kulkarni et al. is based on health of the node (see para. 0016). Given the broadest, reasonable interpretation, the list of active nodes and major error nodes is a storage storing hardware health statuses (available or have major errors) corresponding to each of the plurality of computational units (nodes). It would have been obvious to one of ordinary skill at the time of filing of the invention to include the available node list of Kulkarni et al. into the system of Thapa et al. A person of ordinary skill in the art would have been motivated to make the modification because similar to Thapa et al., Kulkarni et al. is concerned with excluding failed nodes or nodes with impending failures from a task. Thapa et al. disclose flagging or indicating the CPU as having an impending failure and the list of Kulkarni et al. provides a means to flag or indicate. Using the list of Kulkarni et al. in the system of Thapa et al. would yield predictable results since both remove unhealthy nodes.
In para. 0010 and 0019, Thapa et al. disclose scheduling the plurality of computational units to perform task execution on the computer processing system for the one or more applications (para. 0010: plurality of computing clusters can be networked as a supercomputer to perform various complex computing tasks) based on the hardware health statuses of the plurality of computational units, wherein a first computational unit of the plurality of computational units is to be excluded from the task execution responsive to its corresponding hardware health status indicating an impending hardware failure (para. 0019: the management logic may preemptively allocate computing tasks assigned to the node processing module to other processing modules in response to an impending failure).
Referring to claim 17:
In para. 0040-0041, Thapa et al. disclose a non-transitory computer-readable storage medium storing instructions that when executed by a processor of a computing system, are capable of causing the computing system to perform.
In para. 0010, Thapa et al. disclose executing, by a plurality of computational units (computing clusters) of a computer processor system, tasks for one or more applications.
In para. 0016, Thapa et al. disclose collecting, by a plurality of sensors coupled to the plurality of computational units, measurement data of the plurality of computational units.
In para. 0019, Thapa et al. disclose the processing logic can flag or indicate the CPU as having an impending failure (wherein the hardware health statuses are based on the measurement data). However, Thapa et al. do not explicitly disclose storing hardware health statuses corresponding to each of the plurality of computational units in a storage. In para. 0016, Kulkarni et al. disclose determining the health of a cluster (health status based on metrics). A controller does not consider a failed compute node when scheduling a job. In para. 0023-0024 and 0026, Kulkarni et al. disclose removing a node from an active list (see para. 0026) and placing a node with a major error on a list of nodes not to be used (see para. 0024). The list of Kulkarni et al. is based on health of the node (see para. 0016). Given the broadest, reasonable interpretation, the list of active nodes and major error nodes is a storage storing hardware health statuses (available or have major errors) corresponding to each of the plurality of computational units (nodes). It would have been obvious to one of ordinary skill at the time of filing of the invention to include the available node list of Kulkarni et al. into the system of Thapa et al. A person of ordinary skill in the art would have been motivated to make the modification because similar to Thapa et al., Kulkarni et al. is concerned with excluding failed nodes or nodes with impending failures from a task. Thapa et al. disclose flagging or indicating the CPU as having an impending failure and the list of Kulkarni et al. provides a means to flag or indicate. Using the list of Kulkarni et al. in the system of Thapa et al. would yield predictable results since both remove unhealthy nodes.
In para. 0010 and 0019, Thapa et al. disclose scheduling the plurality of computational units to perform task execution on the computer processing system for the one or more applications (para. 0010: plurality of computing clusters can be networked as a supercomputer to perform various complex computing tasks) based on the hardware health statuses of the plurality of computational units, wherein a first computational unit of the plurality of computational units is to be excluded from the task execution responsive to its corresponding hardware health status indicating an impending hardware failure (para. 0019: the management logic may preemptively allocate computing tasks assigned to the node processing module to other processing modules in response to an impending failure).
Claim(s) 3 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Thapa et al., US 2022/0253337 A1 and Kulkarni et al., US 2016/0140001 A1 as applied to claim 1 above, and further in view of “Detecting silent data corruptions in the wild” by Dixit et al. (on IDS).
Regarding claim 3, in para. 0017, Thapa et al. disclose monitoring telemetries such as temperature, voltages, and currents. However, neither Thapa et al. nor Kulkarni et al. explicitly disclose wherein the error screening content includes a test package for silent data error (SDE). In section 5, Dixit et al. disclose testing for silent data errors at the fleet level. It would have been obvious to one of ordinary skill at the time of filing of the invention to include the silent data error testing of Dixit et al. into the combined system of Thapa et al. and Kulkarni et al. A person of ordinary skill in the art would have been motivated to make the modification because silent data corruptors in hardware impact computational integrity for large scale applications (see Dixit et al.: Abstract). Testing for silent data errors mitigates the hardware impact and improves the combined system of Thapa et al. and Kulkarni et al.
Claim(s) 7, 10, 16, and 20 is/are rejected under 35 U.S.C. 103 as being unpatentable over the combination of Thapa et al., US 2022/0253337 A1 and Kulkarni et al., US 2016/0140001 A1 as applied to claims 1, 13, and 17 above, and further in view of Venkatesh et al., US 12,298,841 B1.
Regarding claims 7, 16, and 20, in para. 0017, Thapa et al. disclose monitoring telemetries such as temperature, voltages, and currents. However, neither Thapa et al. nor Kulkarni et al. explicitly disclose where the plurality of sensors includes one or more of a reference critical path delay draft (drift) sensor. In col. 6, lines 42-46, Venkatesh et al. disclose the embedded monitors include a process monitor, a gate delay monitor, a workload monitor, a signal monitor, a thermal sensor, and a current sensor. In col. 11, lines 40-45, Venkatesh et al. disclose monitoring and predicting delay along a critical path. It would have been obvious to one of ordinary skill at the time of filing of the invention to include the critical path delay drift sensor of Venkatesh et al. into the combined system of Thapa et al. and Kulkarni et al. A person of ordinary skill in the art would have been motivated to make the modification because age-related chip failure mechanisms adversely affect the gate and path delays of the circuit. These adversely impact operation within a chip. Monitoring these aspects is essential to ensure proper operation of the chip (see Venkatesh et al.: col. 6, lines 31-37).
Regarding claim 10, in para. 0019, Thapa et al. disclose determining an impending CPU failure. In para. 0016, Kulkarni et al. disclose determining the health of a cluster. However, neither Thapa et al. nor Kulkarni et al. explicitly disclose wherein machine learning is to be used to determine the statuses of the plurality of computational units. In col. 1, lines 64-67 continued in col. 2, lines 1-11, Venkatesh et al. disclose training a machine learning model and using the trained model to predict the failure of an IC chip. It would have been obvious to one of ordinary skill at the time of filing of the invention to include the machine learning model of Venkatesh et al. into the combined system of Thapa et al. and Kulkarni et al. A person of ordinary skill in the art would have been motivated to make the modification because smart telematic systems, equipped with artificial intelligence and machine learning models help mitigate or prevent failure by enabling timely and accurate intervention (see Venkatesh et al.: col. 3, lines 50-53).
Allowable Subject Matter
Claims 5 is objected to as being dependent upon a rejected base claim, but would be allowable if rewritten in independent form including all of the limitations of the base claim and any intervening claims.
The following is a statement of reasons for the indication of allowable subject matter.
With respect to claim 5, US 2020/0019434 A1 discloses a task queue and prioritizing tasks in the queue. However, the prior art does not teach or reasonably suggest, in combination with the remaining limitations, wherein the error screening content is to be executed on a computational unit responsive to the computational unit running below capacity during the task execution for the one or more applications.
Conclusion
The prior art made of record and not relied upon is considered pertinent to applicant's disclosure.
“A Unified Statistical Model for Inter-Die and Intra-Die Process Variation” by Doh et al. disclose analyzing circuit performance by modeling inter-die and intra-die process variations.
US 11,680,983 B1 discloses using a voltage sensor to detect an impending circuit failure.
US 2023/0315599 A1 discloses evaluating memory device health with a voltage drift sensor.
Any inquiry concerning this communication or earlier communications from the examiner should be directed to MICHAEL C MASKULINSKI whose telephone number is (571)272-3649. The examiner can normally be reached Monday-Friday 8:00 am-5:00 pm.
Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice.
If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Bryce Bonzo can be reached at (571) 272-3655. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300.
Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000.
/MICHAEL MASKULINSKI/Primary Examiner, Art Unit 2113