Prosecution Insights
Last updated: August 17, 2026
Application No. 18/773,363

TECHNIQUES FOR PARALLEL EXECUTION

Non-Final OA §101§103
Filed
Jul 15, 2024
Priority
Apr 23, 2021 — continuation of 12/056,494
Examiner
JEON, JAE UK
Art Unit
Tech Center
Assignee
NVIDIA Corporation
OA Round
1 (Non-Final)
75%
Grant Probability
Favorable
1-2
OA Rounds
1y 0m
Est. Remaining
99%
With Interview

Examiner Intelligence

Grants 75% — above average
75%
Career Allowance Rate
309 granted / 412 resolved
+15.0% vs TC avg
Strong +46% interview lift
Without
With
+46.2%
Interview Lift
resolved cases with interview
Typical timeline
3y 1m
Avg Prosecution
26 currently pending
Career history
448
Total Applications
across all art units

Statute-Specific Performance

§101
23.2%
-16.8% vs TC avg
§103
51.0%
+11.0% vs TC avg
§102
3.9%
-36.1% vs TC avg
§112
14.0%
-26.0% vs TC avg
Black line = Tech Center average estimate • Based on career data from 412 resolved cases

Office Action

§101 §103
Notice of Pre-AIA or AIA Status The present application, filed on or after March 16, 2013, is being examined under the first inventor to file provisions of the AIA . DETAILED ACTION 1. This Office Action is in response to the application filed on 03/03/2025. Claims 36-55 are pending in this application. Claims 36, 43 and 50 are independent claims while claims 1-35 are canceled. Claim Rejections - 35 USC § 101 2. 35 U.S.C. 101 reads as follows: Whoever invents or discovers any new and useful process, machine, manufacture, or composition of matter, or any new and useful improvement thereof, may obtain a patent therefor, subject to the conditions and requirements of this title. 3. Claims 36-55 are rejected under 35 U.S.C. 101 because the claimed invention is directed to an abstract idea without significantly more. The independent claims 36, 43 and 50 are corresponding to one of four statutory categories including method, system, and method respectively under step 1. The claims 36, 43 and 50 similarly recites “One or more processors, comprising: processing circuitry to identify one or more copy operations in a representation of a computer program and cause graphics processing unit (GPU) code to be speculatively performed based, at least in part, on the identified one or more copy operations”. The limitation of the claims 36, 43 and 50 of “processing circuitry to identify one or more copy operations in a representation of a computer program” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “identifying” in the context of this claim encompasses the user may identify one or more copy operations in a representation of a computer program (i.e. a readable source code) with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. This judicial exception is not integrated into a practical application. In particular, the claims 36, 43 and 50 recite additional elements such as “cause graphics processing unit (GPU) code to be speculatively performed based, at least in part, on the identified one or more copy operations”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to apply it under MPEP § 2106.05(f): Mere Instructions to Apply an Exception, which does not impose any meaningful limits on practicing the mental process (insignificant additional element). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to insignificant additional elements under Step 2A Prong 2 and Step 2B. This judicial exception is not integrated into a practical application. In particular, the claim 37 recites additional elements such as “the one or more copy operations are device-to-host copy operations”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to field of use under MPEP § 2106.05(h): Field of Use and Technological Environment, which does not impose any meaningful limits on practicing the mental process. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea under Step 2A Prong 2 and 2B. The limitation of the claim 38 of “identifying copy operations between a parallel processing unit and a host computer system, and labeling safe operations following one or more identified copy operations” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “identifying” in the context of this claim encompasses the user may identify copy operations between a parallel processing unit and a host computer system, and labeling safe operations following one or more identified copy operations with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. The limitation of the claim 39 of “processing circuitry is to generate a memory allocation structure that includes one or more indications of extended live ranges of variables to be used with instructions of the GPU code to be speculatively performed” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “generating structure” in the context of this claim encompasses the user may generate a memory allocation structure that includes one or more indications of extended live ranges of variables to be used with instructions of the GPU code to be speculatively performed with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. This judicial exception is not integrated into a practical application. In particular, the claim 40 recites additional elements such as “perform one or more kernel launch commands that cause one or more GPUs to speculatively perform the GPU code”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to field of use under MPEP § 2106.05(h): Field of Use and Technological Environment, which does not impose any meaningful limits on practicing the mental process. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea under Step 2A Prong 2 and 2B. This judicial exception is not integrated into a practical application. In particular, the claim 41 recites additional elements such as “the GPU code to be speculatively performed is part of a loop”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to field of use under MPEP § 2106.05(h): Field of Use and Technological Environment, which does not impose any meaningful limits on practicing the mental process. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea under Step 2A Prong 2 and 2B. This judicial exception is not integrated into a practical application. In particular, the claim 42 recites additional elements such as “the GPU code is in a set of instructions that follows a first value of a branch condition in central processing unit (CPU) code, and that does not follow a second value of the branch condition”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to field of use under MPEP § 2106.05(h): Field of Use and Technological Environment, which does not impose any meaningful limits on practicing the mental process. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea under Step 2A Prong 2 and 2B. This judicial exception is not integrated into a practical application. In particular, the claim 44 recites additional elements such as “the one or more copy operations are asynchronous device-to-host copy operations”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to field of use under MPEP § 2106.05(h): Field of Use and Technological Environment, which does not impose any meaningful limits on practicing the mental process. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea under Step 2A Prong 2 and 2B. The limitation of the claim 45 of “one or more processors are to identify one or more instructions of the GPU code to be speculatively performed based, at least in part, on finding one or more conditional branches in the representation of a computer program” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “finding” in the context of this claim encompasses the user may find one or more conditional branches in the representation of a computer program with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. This judicial exception is not integrated into a practical application. In particular, the claim 46 recites additional elements such as “the one or more processors are to speculatively launch one or more instructions of the GPU code to be performed by one or more GPUs, and are to stop launching instructions speculatively in response to receiving a value via a copy operation that satisfies a condition preceding the one or more instructions in the representation of a computer program”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to apply it under MPEP § 2106.05(f): Mere Instructions to Apply an Exception, which does not impose any meaningful limits on practicing the mental process (insignificant additional element). Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to insignificant additional elements under Step 2A Prong 2 and Step 2B. The limitation of the claim 47 of “identify one or more instructions of the GPU code to be speculatively performed based, at least in part, on labeling operations that are safe to be speculatively performed” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “identifying” in the context of this claim encompasses the user may identify one or more instructions of the GPU code to be speculatively performed based, at least in part, on labeling operations that are safe to be speculatively performed with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. The limitation of the claim 48 of “searching the representation of a computer program to find copy operations, and identifying operations that follow the copy operations that are safe to be speculatively performed” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “searching” in the context of this claim encompasses the user may search the representation of a computer program to find copy operations, and identifying operations that follow the copy operations that are safe to be speculatively performed with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. This judicial exception is not integrated into a practical application. In particular, the claim 49 recites additional elements such as “the GPU code is part of a loop that implements a portion of an inferencing operation using a neural network”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to field of use under MPEP § 2106.05(h): Field of Use and Technological Environment, which does not impose any meaningful limits on practicing the mental process. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea under Step 2A Prong 2 and 2B. The limitation of the claim 51 of “identifying the GPU code to be speculatively performed based, at least in part, on identifying operations in the representation of the computer program that do not change a random state, overwrite outputs, use a signal instruction, or use a wait instruction” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “identifying” in the context of this claim encompasses the user may identify the GPU code to be speculatively performed based, at least in part, on identifying operations in the representation of the computer program that do not change a random state, overwrite outputs, use a signal instruction, or use a wait instruction with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. The limitation of the claim 52 of “identifying a conditional branch in central processing unit (CPU) code and selecting a path from a plurality of paths following the conditional branch” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “identifying” in the context of this claim encompasses the user may identify a conditional branch in central processing unit (CPU) code and selecting a path from a plurality of paths following the conditional branch with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. The limitation of the claim 53 of “one or more instructions of the GPU code have been identified to be speculatively performed based, at least in part, on identifying copy operations in the GPU code” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “identifying” in the context of this claim encompasses the user may identify one or more instructions of the GPU code to be speculatively performed based, at least in part, on identifying copy operations in the GPU code with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. This judicial exception is not integrated into a practical application. In particular, the claim 54 recites additional elements such as “the GPU code includes extended live ranges of variables used in speculatively performed operations”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to field of use under MPEP § 2106.05(h): Field of Use and Technological Environment, which does not impose any meaningful limits on practicing the mental process. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea under Step 2A Prong 2 and 2B. The limitation of the claim 55 of “the GPU code has been identified to be speculatively performed based, at least in part, on identifying copy operations from a GPU to a host computer system in the GPU code” as drafted, is a mental process that, under its broadest reasonable interpretation, covers a mental process but for the recitation of generic computer components. For example, but for the “identifying” in the context of this claim encompasses the user may identify the GPU code to be speculatively performed based, at least in part, on identifying copy operations from a GPU to a host computer system in the GPU code with a pen and paper or in a human mind. If a claim limitation, under its broadest reasonable interpretation, covers performance of the limitation in the mind but for the recitation of generic computer components, then it falls within the “Mental Processes” grouping of abstract ideas. Accordingly, the claim recites an abstract idea under Step 2A Prong 1. This judicial exception is not integrated into a practical application. In particular, the claim 55 recites additional elements such as “the GPU code implements a portion of an inferencing operation using a neural network”. Examiner would like to point out that with the broad reasonable interpretation, this element amounts to field of use under MPEP § 2106.05(h): Field of Use and Technological Environment, which does not impose any meaningful limits on practicing the mental process. Accordingly, this additional element does not integrate the abstract idea into a practical application because it does not impose any meaningful limits on practicing the abstract idea. The claim is directed to an abstract idea under Step 2A Prong 2 and 2B. Dependent claims 37-42, 44-49 and 51-55 are also similar rejected under same rationale as cited above wherein these claims do not include additional elements that are sufficient to amount to significantly more than the judicial exception. These claims are merely further elaborate the mental process itself or providing additional definition of process which does not impose any meaningful limits on practicing the abstract idea. Claims 37-42, 44-49 and 51-55 are also rejected for incorporating the deficiency of their independent claims 36, 43 and 50 respectively. Claim Rejections - 35 USC § 103 4. In the event the determination of the status of the application as subject to AIA 35 U.S.C. 102 and 103 (or as subject to pre-AIA 35 U.S.C. 102 and 103) is incorrect, any correction of the statutory basis (i.e., changing from AIA to pre-AIA ) for the rejection will not be considered a new ground of rejection if the prior art relied upon, and the rationale supporting the rejection, would be the same under either status. 5. The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 6. Claims 36, 38, 43, 47, 48, 50 and 53 are rejected under 35 U.S.C. 103 as being unpatentable over Fortino (US PGPub 20180121273), in view of Sandberg (US PGPub 20210042227). As per Claim 36, Fortino teaches of one or more processors, comprising: processing circuitry to identify one or more copy operations in a representation of a computer program and cause … code to be speculatively performed based, at least in part, on the identified one or more copy operations (Fig. 4 and Par 49-50, The conversion component 420 can include a functional duplicate component 421 that creates multiple functional duplicates in the native machine language or micro code. The native machine instruction regions are speculatively executed by execution component 430 (e.g., in a processor, etc.) Par 51, Fault checking can include speculatively executing the duplicate functionality based on the multiple sets of native machine instructions. The speculative executions start from a known valid or good state and proceed down the speculative path. Par 55, It is appreciated that there can be more than two duplicate speculative executions of the functionality, which can also increase the likelihood of identifying errors. The speculative executions can occur substantially simultaneously. Each set of operations is a functional duplicate of the higher level code portion functionality.) Fortino does not specifically teach, however Sandberg teaches to cause graphics processing unit (GPU) code to be speculatively performed. (Par 81, The CPU 4 or GPU 6 or other masters may support speculative execution of instructions including instructions which can trigger read operations to read data from memory. Speculative loads executed in a trusted application can leak data to untrusted applications using cache timing side channels. One example of such a side-channel attack was described in the “Spectre” attack. The attack uses the following (or similar) code sequence:) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add causing graphics processing unit (GPU) code to be speculatively performed, as conceptually seen from the teaching of Sandberg, into that of Fortino because this modification can help overcome the memory wall from the memory structure in AL inference and matrix operations while increasing hardware utilization by fully activating CUDA cores. As per Claim 38, Fortino further teaches of the one or more processors of claim 36, wherein one or more instructions of the GPU code have been identified to be speculatively performed based, at least in part, on identifying copy operations between a parallel processing unit and a host computer system, and labeling safe operations following one or more identified copy operations. (Par 34, The reliability enhancement systems and methods can include automatic implementation of efficient fault checking utilizing speculative execution. In one embodiment, a particular portion of higher level instructions (e.g., critical code, safety related code, etc.) is associated with a functionality which is duplicated by multiple sets of operations that are speculatively executed. The multiple speculative executions can include multiple sets of operations or controls that are considered functional duplicates (even though they may not be exact literal duplicates). In one exemplary implementation, while the multiple sets of operations may perform or achieve similar functionality, they can include directions to load results in different registers (e.g., to allow parallel execution, to enable comparison, etc.). The multiple sets of operations can correspond to multiple sets of native hardware instructions. In one embodiment, the native hardware instructions or microcode are the instructions a processor actually executes.) Re Claim 43, it is the system claim, having similar limitations of claim 36. Thus, claim 43 is also rejected under the similar rationale as cited in the rejection of claim 36. As per Claim 47, Fortino further teaches of the system of claim 43, wherein the one or more processors are to identify one or more instructions of the GPU code to be speculatively performed based, at least in part, on labeling operations that are safe to be speculatively performed. (Par 34, The reliability enhancement systems and methods can include automatic implementation of efficient fault checking utilizing speculative execution. In one embodiment, a particular portion of higher level instructions (e.g., critical code, safety related code, etc.) is associated with a functionality which is duplicated by multiple sets of operations that are speculatively executed. The multiple speculative executions can include multiple sets of operations or controls that are considered functional duplicates (even though they may not be exact literal duplicates). In one exemplary implementation, while the multiple sets of operations may perform or achieve similar functionality, they can include directions to load results in different registers (e.g., to allow parallel execution, to enable comparison, etc.). The multiple sets of operations can correspond to multiple sets of native hardware instructions. In one embodiment, the native hardware instructions or microcode are the instructions a processor actually executes.) As per Claim 48, Fortino further teaches of the system of claim 43, wherein the one or more processors are to identify one or more instructions of the GPU code to be speculatively performed based, at least in part, on searching the representation of a computer program to find copy operations, and identifying operations that follow the copy operations that are safe to be speculatively performed. (Par 34, The reliability enhancement systems and methods can include automatic implementation of efficient fault checking utilizing speculative execution. In one embodiment, a particular portion of higher level instructions (e.g., critical code, safety related code, etc.) is associated with a functionality which is duplicated by multiple sets of operations that are speculatively executed. The multiple speculative executions can include multiple sets of operations or controls that are considered functional duplicates (even though they may not be exact literal duplicates). In one exemplary implementation, while the multiple sets of operations may perform or achieve similar functionality, they can include directions to load results in different registers (e.g., to allow parallel execution, to enable comparison, etc.). The multiple sets of operations can correspond to multiple sets of native hardware instructions. In one embodiment, the native hardware instructions or microcode are the instructions a processor actually executes.) Re Claim 50, it is the method claim, having similar limitations of claim 36. Thus, claim 50 is also rejected under the similar rationale as cited in the rejection of claim 36. As per Claim 53, Fortino further teaches of the method of claim 50, wherein one or more instructions of the GPU code have been identified to be speculatively performed based, at least in part, on identifying copy operations in the GPU code. (Fig. 4 and Par 49-50, The conversion component 420 can include a functional duplicate component 421 that creates multiple functional duplicates in the native machine language or micro code. The native machine instruction regions are speculatively executed by execution component 430 (e.g., in a processor, etc.) Par 51, Fault checking can include speculatively executing the duplicate functionality based on the multiple sets of native machine instructions. The speculative executions start from a known valid or good state and proceed down the speculative path. Par 55, It is appreciated that there can be more than two duplicate speculative executions of the functionality, which can also increase the likelihood of identifying errors. The speculative executions can occur substantially simultaneously. Each set of operations is a functional duplicate of the higher level code portion functionality.) 7. Claims 37, 39, 41, 44, 46 and 54 are rejected under 35 U.S.C. 103 as being unpatentable over Fortino (US PGPub 20180121273), in view of Sandberg (US PGPub 20210042227), and further in view of Eltantawy (US Patent 11768715). As per Claim 37, neither Fortino nor Sandberg specifically teaches, however Eltantawy teaches of the one or more processors of claim 36, wherein the one or more copy operations are device-to-host copy operations. (Col 8, lines 6-9, The host side (i.e., the CPU) executes the serial portion of the code, allocates the required memory on the device side (i.e., the GPU) and copies the required data from/to the host to/from the device.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add the one or more copy operations are device-to-host copy operations, as conceptually seen from the teaching of Eltantawy, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. As per Claim 39, neither Fortino nor Sandberg specifically teaches, however Eltantawy teaches of the one or more processors of claim 36, wherein processing circuitry is to generate a memory allocation structure that includes one or more indications of extended live ranges of variables to be used with instructions of the GPU code to be speculatively performed. (Col 6, lines 39-61, FIGS. 4, 5 and 6 illustrate dynamic instruction overhead, memory traffic overhead and divergence overhead respectively as a function of the number of hashtable buckets. These measurements constitute the resources consumed in spinning in an attempt to acquire a lock (synchronization overhead) versus the useful portion of the execution. Overheads are measured on Pascal GTX1080 using Nvidia profiler (nvprof) launching 120 blocks each of 256 threads (74.88% occupancy of the pascal GPU). FIGS. 4, 5 and 6 show significant synchronization overheads that are still persistent in the Pascal architecture. Specifically FIG. 4 shows that instruction count overhead ranges from 61.0% at low contention to 98.3% at high contention. Similarly, FIG. 5 shows that 41.5% to 95.6% of memory operations are due to synchronization. A significant portion of both overheads are due to failed lock acquire attempts. Another source of synchronization overhead, unique to GPUs, is control-flow divergence. FIG. 6 shows that if the code is executed by a single warp, the SIMD utilization (fraction of active lanes) ranges between 87.1%-98.6% but drops to 16.4%-47.1% when executing multiple warps. This is due to inter-warp lock conflicts, which can be impacted by warp scheduling. Col 11, lines 27-31, CAWA estimates warp criticality using a criticality metric that predicts which warp will take longer time to finish. CAWA is reported to outperform greedy-than-oldest (GTO) warp scheduling across a range of traditional GPGPU workloads.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add generating a memory allocation structure that includes one or more indications of extended live ranges of variables to be used with instructions of the GPU code to be speculatively performed, as conceptually seen from the teaching of Eltantawy, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. As per Claim 41, neither Fortino nor Sandberg specifically teaches, however Eltantawy teaches of the one or more processors of claim 36, wherein the GPU code to be speculatively performed is part of a loop. (Col 17, lines 16-27, DDOS detects busy-wait loops in two steps. First, it detect the presence of a loop. DDOS does this by tracking the sequence of program counter values of a warp. Second, DDOS speculates whether a loop identified in the first step is a busy-wait loop or a normal loop. To distinguish these cases it leverages the observation that typically in normal loops found in GPU code an induction variable changes every iteration. Moreover, this induction variable typically contributes to the computation of the loop exit condition. In NVIDIA GPUs the loop exit condition and the divergence behavior of a thread are typically determined using a set predicate instruction (available both in PTX and SASS).) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add the GPU code to be speculatively performed is part of a loop, as conceptually seen from the teaching of Eltantawy, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. As per Claim 44, neither Fortino nor Sandberg specifically teaches, however Eltantawy teaches of the system of claim 43. wherein the one or more copy operations are asynchronous device-to-host copy operations. (Col 8, lines 6-9, The host side (i.e., the CPU) executes the serial portion of the code, allocates the required memory on the device side (i.e., the GPU) and copies the required data from/to the host to/from the device.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add the one or more copy operations are asynchronous device-to-host copy operations, as conceptually seen from the teaching of Eltantawy, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. As per Claim 46, neither Fortino nor Sandberg specifically teaches, however Eltantawy teaches of the system of claim 43, wherein the one or more processors are to speculatively launch one or more instructions of the GPU code to be performed by one or more GPUs, and are to stop launching instructions speculatively in response to receiving a value via a copy operation that satisfies a condition preceding the one or more instructions in the representation of a computer program. (Col 6, lines 43-60, Overheads are measured on Pascal GTX1080 using Nvidia profiler (nvprof) launching 120 blocks each of 256 threads (74.88% occupancy of the pascal GPU). FIGS. 4, 5 and 6 show significant synchronization overheads that are still persistent in the Pascal architecture. Specifically FIG. 4 shows that instruction count overhead ranges from 61.0% at low contention to 98.3% at high contention. Similarly, FIG. 5 shows that 41.5% to 95.6% of memory operations are due to synchronization. A significant portion of both overheads are due to failed lock acquire attempts. Another source of synchronization overhead, unique to GPUs, is control-flow divergence. FIG. 6 shows that if the code is executed by a single warp, the SIMD utilization (fraction of active lanes) ranges between 87.1%-98.6% but drops to 16.4%-47.1% when executing multiple warps. This is due to inter-warp lock conflicts, which can be impacted by warp scheduling.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add the one or more processors are to speculatively launch one or more instructions of the GPU code to be performed by one or more GPUs, and are to stop launching instructions speculatively in response to receiving a value via a copy operation that satisfies a condition preceding the one or more instructions in the representation of a computer program, as conceptually seen from the teaching of Eltantawy, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. As per Claim 54, neither Fortino nor Sandberg specifically teaches, however Eltantawy teaches of the method of claim 50, wherein the GPU code includes extended live ranges of variables used in speculatively performed operations. (Col 6, lines 39-61, FIGS. 4, 5 and 6 illustrate dynamic instruction overhead, memory traffic overhead and divergence overhead respectively as a function of the number of hashtable buckets. These measurements constitute the resources consumed in spinning in an attempt to acquire a lock (synchronization overhead) versus the useful portion of the execution. Overheads are measured on Pascal GTX1080 using Nvidia profiler (nvprof) launching 120 blocks each of 256 threads (74.88% occupancy of the pascal GPU). FIGS. 4, 5 and 6 show significant synchronization overheads that are still persistent in the Pascal architecture. Specifically FIG. 4 shows that instruction count overhead ranges from 61.0% at low contention to 98.3% at high contention. Similarly, FIG. 5 shows that 41.5% to 95.6% of memory operations are due to synchronization. A significant portion of both overheads are due to failed lock acquire attempts. Another source of synchronization overhead, unique to GPUs, is control-flow divergence. FIG. 6 shows that if the code is executed by a single warp, the SIMD utilization (fraction of active lanes) ranges between 87.1%-98.6% but drops to 16.4%-47.1% when executing multiple warps. This is due to inter-warp lock conflicts, which can be impacted by warp scheduling. Col 11, lines 27-31, CAWA estimates warp criticality using a criticality metric that predicts which warp will take longer time to finish. CAWA is reported to outperform greedy-than-oldest (GTO) warp scheduling across a range of traditional GPGPU workloads.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add the GPU code includes extended live ranges of variables used in speculatively performed operations, as conceptually seen from the teaching of Eltantawy, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. 8. Claim 40 is rejected under 35 U.S.C. 103 as being unpatentable over Fortino (US PGPub 20180121273), in view of Sandberg (US PGPub 20210042227), and further in view of Mehalwal (US PGPub 20200380761). As per Claim 40, neither Fortino nor Sandberg specifically teaches, however Mehalwal teaches of the one or more processors of claim 36, wherein the processing circuitry is to perform one or more kernel launch commands that cause one or more GPUs to speculatively perform the GPU code. (Par 42, The command processor 137 performs the operations described with respect to Table 1 for each ray generation shader kernel launch. In various implementations, the command processor 137 concurrently performs multiple iterations of these operations, each for a different ray generation shader kernel, to allow multiple ray generation shader kernels to launch concurrently in the APD 116. The command processor 137 may use any technically feasible mechanism for such concurrent execution, such as using multiple hardware execution units in the command processor 137, using preemptive multitasking, using a combination thereof, or using any other technically feasible technique.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add performing one or more kernel launch commands that cause one or more GPUs to speculatively perform the GPU code, as conceptually seen from the teaching of Mehalwal, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. 9. Claims 42, 45 and 52 are rejected under 35 U.S.C. 103 as being unpatentable over Fortino (US PGPub 20180121273), in view of Sandberg (US PGPub 20210042227), and further in view of Rodriguez (US PGPub 20110143811). As per Claim 42, neither Fortino nor Sandberg specifically teaches, however Rodriguez teaches of the one or more processors of claim 36, wherein the GPU code is in a set of instructions that follows a first value of a branch condition in central processing unit (CPU) code, and that does not follow a second value of the branch condition. (Par 383, Branch prediction can be handled by not predicting: instead, the GPU processes for all of the potential outcomes of a branch in parallel, and the system uses whatever output corresponds to the actual branch condition when it becomes known. Par 226, Parallelism is widely employed in graphics processing units (GPUs) Two branch conditions are executed by GPU in parallel) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add the GPU code is in a set of instructions that follows a first value of a branch condition in central processing unit (CPU) code, and that does not follow a second value of the branch condition, as conceptually seen from the teaching of Rodriguez, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. As per Claim 45, neither Fortino nor Sandberg specifically teaches, however Rodriguez teaches of the system of claim 43, wherein the one or more processors are to identify one or more instructions of the GPU code to be speculatively performed based, at least in part, on finding one or more conditional branches in the representation of a computer program. (Par 213, While the discussion has focused on serial data processing, image or other data may be processed in two or more parallel paths. For example, the output of stage 38d may be applied to two subsequent stages, each of which starts a respective branch of a fork in the processing. Those two chains can be processed independently thereafter, or data resulting from such processing can be combined--or used in conjunction--in a subsequent stage. (Each of those processing chains, in turn, can be forked, etc.) Par 383, Branch prediction can be handled by not predicting: instead, the GPU processes for all of the potential outcomes of a branch in parallel, and the system uses whatever output corresponds to the actual branch condition when it becomes known. Par 226, Parallelism is widely employed in graphics processing units (GPUs) Two branch conditions are executed by GPU in parallel) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add identifying one or more instructions of the GPU code to be speculatively performed based, at least in part, on finding one or more conditional branches in the representation of a computer program, as conceptually seen from the teaching of Rodriguez, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. As per Claim 52, neither Fortino nor Sandberg specifically teaches, however Rodriguez teaches of the method of claim 50, wherein GPU code has been identified to be speculatively performed based, at least in part, on identifying a conditional branch in central processing unit (CPU) code and selecting a path from a plurality of paths following the conditional branch. (Par 213, While the discussion has focused on serial data processing, image or other data may be processed in two or more parallel paths. For example, the output of stage 38d may be applied to two subsequent stages, each of which starts a respective branch of a fork in the processing. Those two chains can be processed independently thereafter, or data resulting from such processing can be combined--or used in conjunction--in a subsequent stage. (Each of those processing chains, in turn, can be forked, etc.) Par 383, Branch prediction can be handled by not predicting: instead, the GPU processes for all of the potential outcomes of a branch in parallel, and the system uses whatever output corresponds to the actual branch condition when it becomes known. Par 226, Parallelism is widely employed in graphics processing units (GPUs) Two branch conditions are executed by GPU in parallel) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add GPU code has been identified to be speculatively performed based, at least in part, on identifying a conditional branch in central processing unit (CPU) code and selecting a path from a plurality of paths following the conditional branch, as conceptually seen from the teaching of Rodriguez, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. 10. Claims 49 and 55 are rejected under 35 U.S.C. 103 as being unpatentable over Fortino (US PGPub 20180121273), in view of Sandberg (US PGPub 20210042227), and further in view of Hoang (US PGPub 20210397931). As per Claim 49, neither Fortino nor Sandberg specifically teaches, however Hoang teaches of the system of claim 43, wherein the GPU code is part of a loop that implements a portion of an inferencing operation using a neural network. (Par 19, To reduce the amount of data transfer needed to perform inferencing operations for a recurrent neural network, or RNN, techniques and memory structures are presented that allow for inferencing operations to be performed through in-array multiplications within the memory arrays of a non-volatile memory device. Par 65, FIG. 10B is a flowchart describing a process for the inference phase of supervised learning using a neural network to predict the “meaning” of the input data using an estimated accuracy. Depending on the case, the neural network may be inferenced both in the cloud and by an edge device's (e.g., smart phone, automobile process, hardware accelerator) processor. Par 73, FIG. 12 is a block diagram of one embodiment of an architecture for a GRU-based process-in-memory RNN inference accelerator.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add the GPU code is part of a loop that implements a portion of an inferencing operation using a neural network, as conceptually seen from the teaching of Hoang, into that of Fortino and Sandberg because this modification can help reduce memory bottleneck and latency as well as increase AI inference and output quality. As per Claim 55, neither Fortino nor Sandberg specifically teaches, however Hoang teaches of the method of claim 50, wherein the GPU code has been identified to be speculatively performed based, at least in part, on identifying copy operations from a GPU to a host computer system in the GPU code, and wherein the GPU code implements a portion of an inferencing operation using a neural network. (Par 19, To reduce the amount of data transfer needed to perform inferencing operations for a recurrent neural network, or RNN, techniques and memory structures are presented that allow for inferencing operations to be performed through in-array multiplications within the memory arrays of a non-volatile memory device. Par 65, FIG. 10B is a flowchart describing a process for the inference phase of supervised learning using a neural network to predict the “meaning” of the input data using an estimated accuracy. Depending on the case, the neural network may be inferenced both in the cloud and by an edge device's (e.g., smart phone, automobile process, hardware accelerator) processor. Par 73, FIG. 12 is a block diagram of one embodiment of an architecture for a GRU-based process-in-memory RNN inference accelerator.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add a portion of an inferencing operation using a recurrent neural network, as conceptually seen from the teaching of Hoang, into that of Ngai because this modification can help optimize the resources such as processors and memories to predict which paths/tasks are to be speculatively performed. 11. Claim 51 is rejected under 35 U.S.C. 103 as being unpatentable over Fortino (US PGPub 20180121273), in view of Sandberg (US PGPub 20210042227), and further in view of Vincent (US PGPub 20170060579). As per Claim 51, neither Fortino nor Sandberg specifically teaches, however Vincent teaches of the method of claim 50, further comprising: identifying the GPU code to be speculatively performed based, at least in part, on identifying operations in the representation of the computer program that do not change a random state, overwrite outputs, use a signal instruction, or use a wait instruction. (Par 67, By holding instructions the processor architecture is able to remove NOPs from the instruction memory, for example. Further, implementations of the processor architecture allow for parallel execution pipelines combined with speculative execution to improve the overall performance when instructions are held waiting for inputs from other instructions.) Therefore, it would have been obvious for one of the ordinary skill in the art before the effective filing date of the claimed invention to add based, at least in part, on identifying operations that do not change a random state, overwrite outputs, use a signal instruction, or use a wait instruction, as conceptually seen from the teaching of Vincent, into that of Ngai because this modification can help optimize the resources such as processors and memories to predict which paths/tasks are to be speculatively performed. Pertinent Prior Art The prior art made of record and not relied upon is considered pertinent to applicant's disclosure. Morys (CN 1184564) This is achieved is called "recovery" of the processing process is finished. recovery may be related with additional expanded program code generated by a compiler, the code is a dependent non speculative instruction group for estimation of copy, such that, after executing all the abnormal conditions generate abnormal and all of previously written destination by rewriting the correct result. recovery code is not always accurate copy of instruction sequence, but it can be achieve the same result when executing the code. Further, in one embodiment, a new instruction is specified as a specific purpose: checking the presence of DCT and detecting the DET of condition related to activation of the recovery code. Bishop (US PGPub 20090106538) Once lookahead prefetch mode is activated, instructions that are not supported by the out-of-order execution mechanisms of the processor (if any), identified herein as "speculative instructions", are allowed to be written back to a copy of the architected registers of the microprocessors stored in an inactive thread. The copy of the architected registers and writeback are discussed in more detail in conjunction with FIG. 4B. Conclusion Any inquiry concerning this communication or earlier communications from the examiner should be directed to JAE UK JEON whose telephone number is (571)270-3649. The examiner can normally be reached 10am-6pm. Examiner interviews are available via telephone, in-person, and video conferencing using a USPTO supplied web-based collaboration tool. To schedule an interview, applicant is encouraged to use the USPTO Automated Interview Request (AIR) at http://www.uspto.gov/interviewpractice. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Chat Do can be reached at 571-272-3721. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of published or unpublished applications may be obtained from Patent Center. Unpublished application information in Patent Center is available to registered users. To file and manage patent submissions in Patent Center, visit: https://patentcenter.uspto.gov. Visit https://www.uspto.gov/patents/apply/patent-center for more information about Patent Center and https://www.uspto.gov/patents/docx for information about filing in DOCX format. For additional questions, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). If you would like assistance from a USPTO Customer Service Representative, call 800-786-9199 (IN USA OR CANADA) or 571-272-1000. /JAE U JEON/Primary Examiner, Art Unit 2193
Read full office action

Prosecution Timeline

Jul 15, 2024
Application Filed
Jul 15, 2026
Non-Final Rejection mailed — §101, §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12699562
WORKFLOW TEMPLATES FOR CONFIGURATION PACKAGES
3y 3m to grant Granted Aug 04, 2026
Patent 12697982
TECHNIQUES FOR CALCULATING SURFACE BREAKPOINTS FOR SECONDARY SAFETY VERIFICATIONS IN VEHICLE CONTROLS SYSTEMS
2y 9m to grant Granted Aug 04, 2026
Patent 12693845
APPLICATION OF DATA DESCRIPTOR MAPS IN MANAGEMENT OF FIRMWARE CONTROL DATA
2y 8m to grant Granted Jul 28, 2026
Patent 12675264
INSERTING A MEMORY FENCE IN A PROGRAM IN RESPONSE TO A DETERMINATION THAT PREDETERMINED PATTERN(S) DO NOT EXIST IN THE PROGRAM
4y 1m to grant Granted Jul 07, 2026
Patent 12676030
IN-VEHICLE COMMUNICATION SYSTEM, DATA STRUCTURE OF REPROGRAMMING POLICY METADATA, AND DATA STRUCTURE OF DOWNLOAD METADATA
2y 6m to grant Granted Jul 07, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

1-2
Expected OA Rounds
75%
Grant Probability
99%
With Interview (+46.2%)
3y 1m (~1y 0m remaining)
Median Time to Grant
Low
PTA Risk
Based on 412 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month