Prosecution Insights
Last updated: October 02, 2026
Application No. 18/689,861

MEMORY MANAGEMENT METHOD FOR PSEUDO-FUNCTIONAL DIFFERENTIABLE PROGRAMMING

Final Rejection §103
Filed
Mar 06, 2024
Priority
Sep 10, 2021 — provisional 63/242,963 +2 more
Examiner
TSAI, SHENG JEN
Art Unit
2139
Tech Center
2100 — Computer Architecture & Software
Assignee
Purdue Research Foundation
OA Round
4 (Final)
70%
Grant Probability
Favorable
5-6
OA Rounds
9m
Est. Remaining
84%
With Interview

Examiner Intelligence

Grants 70% — above average
70%
Career Allowance Rate
567 granted / 805 resolved
+15.4% vs TC avg
Moderate +14% lift
Without
With
+13.8%
Interview Lift
resolved cases with interview
Typical timeline
3y 4m
Avg Prosecution
19 currently pending
Career history
829
Total Applications
across all art units

Statute-Specific Performance

§101
2.7%
-37.3% vs TC avg
§103
54.2%
+14.2% vs TC avg
§102
26.6%
-13.4% vs TC avg
§112
13.4%
-26.6% vs TC avg
Black line = Tech Center average estimate • Based on career data from 805 resolved cases

Office Action

§103
DETAILED ACTION 1. This Office Action is taken in response to Applicants’ Amendments and Remarks filed on 6/21/2026 regarding application 18/689,861 filed on 3/6/2024. Claims 1, 6, 8-17, 19-21, and 23-27 are pending for consideration. 2. Response to Amendments and Remarks Applicants’ amendments and remarks have been fully and carefully considered, with the Examiner’s response set forth below. (1) In view of the amendments and remarks, rejections under 112(a) and 112(b) have been withdrawn. (2) Applicant contends that, regarding claim 1, Patrick in view of Becchi fails to teach the limitation “generating a streaming plan for usage of the streaming tensor.” The Examiner disagrees. First, the limitation merely recites “a streaming plan for usage of the streaming tensor,” and is otherwise completely silent regarding on the scope and definition of what constitute, and what does not constitute, “a streaming plan for usage of the streaming tensor.” As such, the term “a streaming plan for usage of the streaming tensor” must be given the broadest, reasonable interpretation according to the guidelines set forth by the MPEP. Within the context of the limitations recited in claim 1, and as far as claim 1 is concerned, any streaming plan that uses the streaming tensors in any manner qualifies as “a streaming plan for usage of the streaming tensor.” Second, Patrick specifically teaches a streaming plan that utilizes streaming tensors [a streaming plan as shown in figure 5; The operations of method 500 can be initiated by the compiler operating on at least one computer processor and/or on a host server integrated into the deterministic streaming system or separate from the deterministic streaming system … The deterministic streaming system evaluates 505 (e.g., by the scheduler) a latency for each task of a plurality of tasks to be run at the deterministic streaming system. The deterministic streaming system adjusts 510 (e.g., by the scheduler) at least one of an accuracy metric and a quality metric for an output of each of the plurality of tasks based on the evaluated latency until the plurality of tasks can be completed before expiration of one or more contractual deadlines. The deterministic streaming system runs 515, by at least a subset of the plurality of deterministic streaming processors of the deterministic streaming system, the plurality of tasks each having the output with at least one of the adjusted accuracy metric and the adjusted quality metric. The deterministic streaming system selects (e.g., by the scheduler) a pre-compiled model variation for compilation (e.g., by the compiler). The deterministic streaming system selects (e.g., by the scheduler) quality and accuracy information during a static capacity planning process for when the scheduler decides which model variations should be compiled … In some embodiments, the deterministic streaming system compiles (e.g., by the compiler) source code of each model of a plurality of models associated with the plurality of tasks into an intermediate representation … The deterministic streaming system calculates (e.g., by the compiler) an amount of computation that can be performed within a period of time for each of the plurality of tasks, and provides information about the calculated amount of computation to the scheduler for the evaluation of latency for each task … The deterministic streaming system meets a defined QoE metric based on at least the subset of the plurality of deterministic streaming processors running the plurality of tasks each having the output with at least one of the adjusted accuracy metric and the adjusted quality metric … (¶ 0119-0124); “capacity planning” as shown in figure 4B, 465; The process for ensuring the drainage condition is simplified due to the deterministic nature of the TSP farm 420. A non-real-time subcomponent of the scheduler 425 that can be referred to as a “capacity planner” … If the simulation can drain the leaky buckets within all registered contractual agreements, then the capacity planner would proceed with the registration. Otherwise, the capacity planner determines the new registration to be infeasible and would require a user 435 to change their registration parameters to be less intensive on the TSP farm 420 … (¶ 0111-0112); The benefits of involving the scheduler 425 in the compilation process arise from the fact that the scheduler 425 supports a plurality of models 415 for a plurality of users 435. If a new model 415 belonging to an arbitrary user 435 is registered to the TSP farm 420 with pre-existing registered models 415, the scheduler 425 can elect to change which binary variations would be utilized for any subset of existing pre-registered models 415 as part of its optimization routine (e.g., when ensuring the drainage condition for capacity planning, as discussed in more detail in the section below). Partial compilation is useful to expedite this process because, otherwise, recompilation of models 415 would be required. Additionally, the scheduler 415 can perform its part of compilation process outside the critical path of incoming requests as, e.g., a background job. Otherwise, the non-determinism and an additional latency would be introduced to the incoming requests as model compilation itself is not deterministic … The benefit of splitting the compilation of model 415 between the compiler 410 and the scheduler 425 is that the scheduler 425 can dynamically modify a manner of running the compiled model 415 during runtime in the background … (¶ 0103-0105)]. Therefore, Patrick clearly teaches the limitation “generating a streaming plan for usage of the streaming tensor.” (3) In response to the amendments and remarks, an updated claim analysis with a newly identified reference has been made. Refer to the corresponding sections of the following Office Action for details. 3. Examiner’s Note (1) In the case of amending the Claimed invention, Applicant is respectfully requested to indicate the portion(s) of the specification which dictate(s) the structure relied on for proper interpretation and also to verify and ascertain the metes and bounds of the claimed invention. This will assist in expediting compact prosecution. MPEP 714.02 recites: “Applicant should also specifically point out the support for any amendments made to the disclosure. See MPEP § 2163.06. An amendment which does not comply with the provisions of 37 CFR 1.121(b), (c), (d), and (h) may be held not fully responsive. See MPEP § 714.” Amendments not pointing to specific support in the disclosure may be deemed as not complying with provisions of 37 C.F.R. 1.131(b), (c), (d), and (h) and therefore held not fully responsive. Generic statements such as “Applicants believe no new matter has been introduced” may be deemed insufficient. (2) Examiner has cited particular columns/paragraph and line numbers in the references applied to the claims above for the convenience of the applicant. Although the specified citations are representative of the teachings of the art and are applied to specific limitations within the individual claim, other passages and figures may apply as well. It is respectfully requested from the applicant in preparing responses, to fully consider the references in entirety as potentially teaching all or part of the claimed invention, as well as the context of the passage as taught by the prior art or disclosed by the Examiner. Claim Rejections - 35 USC § 103 The following is a quotation of 35 U.S.C. 103 which forms the basis for all obviousness rejections set forth in this Office action: A patent for a claimed invention may not be obtained, notwithstanding that the claimed invention is not identically disclosed as set forth in section 102, if the differences between the claimed invention and the prior art are such that the claimed invention as a whole would have been obvious before the effective filing date of the claimed invention to a person having ordinary skill in the art to which the claimed invention pertains. Patentability shall not be negated by the manner in which the invention was made. 4. Claim 1, 6-10, 12-17, 19-21, and 23 are rejected under 35 U.S.C. 103 as being anticipated by Patrick et al. (US Patent Application Publication 2024/0370302, hereinafter Patrick), in view of Becchi et al. (US Patent Application Publication 2011/0173155, hereinafter Becchi), and further in view of Fontijn (US Patent Application Publication 2009/0157957). As to claim 1, Becchi teaches A computer-implemented method of operating on a program, comprising: executing at least one instruction towards method of operating a functional program [as shown in figure 7, where program instructions (724) are stored in a main memory (704), then loaded into and executed by a processor (702); Some portions of this description describe the embodiments of the disclosure in terms of algorithms and symbolic representations of operations on information. These algorithmic descriptions and representations are commonly used by those skilled in the data processing arts to convey the substance of their work effectively to others skilled in the art. These operations, while described functionally, computationally, or logically, are understood to be implemented by computer programs or equivalent electrical circuits, microcode, or the like. Furthermore, it has also proven convenient at times, to refer to these arrangements of operations as modules, without loss of generality. The described operations and their associated modules can be embodied in software, firmware, hardware, or any combinations thereof … (¶ 0152-0154); Becchi also teaches this limitation -- Embodiments may include a computer program product accessible from a computer-usable or computer-readable medium providing program code for use by or in connection with a computer or any instruction execution system. A computer-usable or computer readable medium may include any apparatus that stores, communicates, propagates, or transports the program for use by or in connection with the instruction execution system, apparatus, or device … A data processing system suitable for storing and/or executing program code may include at least one processor coupled directly or indirectly to memory elements through a system bus. The memory elements can include local memory employed during actual execution of the program code, bulk storage, and cache memories which provide temporary storage of at least some program code to reduce the number of times code is retrieved from bulk storage during execution … (¶ 0023-0024)], wherein the at least one instruction operates on a streaming tensor [Disclosed are configurations that include a deterministic streaming system with one or more deterministic streaming processors (e.g., tensor streaming processors (TSPs) or artificial intelligence processors) … The disclosed embodiments are directed to one or more deterministic streaming processors each having a functional slicing architecture. In some embodiments, each deterministic streaming processor comprises a tensor streaming processor (TSP) having a functional slicing architecture, which can be used for hardware-accelerated machine learning (ML) applications … The computational elements of the deterministic streaming processor can be divided between different functionalities (e.g., memory, arithmetic operation, etc.), and can be organized as functional slices which operate on multi-dimensional data (e.g., tensors) … (¶ 0037-0039)], streaming tensor defined as a block of immutable data wherein residence of said block of immutable data is decoupled from manipulation of said block of immutable data [read and transfer operations performed on the data do not mutate the data, hence immutable -- FIG. 1A illustrates an arrangement of functional slices in a tensor streaming processor (TSP), in accordance with some embodiments (¶ 0023); FIG. 2 depicts stream registers of a TSP that are numbered to show their locations between functional slices within a superlane, in accordance with some embodiments (¶ 0026)], the streaming tensors operated on a processor of a second class [Disclosed are configurations that include a deterministic streaming system with one or more deterministic streaming processors (e.g., tensor streaming processors (TSPs) or artificial intelligence processors) … The disclosed embodiments are directed to one or more deterministic streaming processors each having a functional slicing architecture. In some embodiments, each deterministic streaming processor comprises a tensor streaming processor (TSP) having a functional slicing architecture, which can be used for hardware-accelerated machine learning (ML) applications … The computational elements of the deterministic streaming processor can be divided between different functionalities (e.g., memory, arithmetic operation, etc.), and can be organized as functional slices which operate on multi-dimensional data (e.g., tensors) … (¶ 0037-0039)] and wherein a copy of the streaming tensor is reliably available in memory of a processor of a first class and may be optionally available in a memory of the processor of a second class [The deterministic cloud system 400 can run a workload (i.e., a stream of incoming tasks 430) that is otherwise very expensive to process using the traditional CPU or GPU computational resources. The workload can vary and the request patterns of users 435 can be unknown … (¶ 0089); Becchi more expressively teaches this limitation -- If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 … (¶ 0043-0045)]; the execution of the at least one instruction includes: initially analyzing the functional program to determine when and for how long the streaming tensors are needed by the processor of the second class [The deterministic cloud system 400 can run a workload (i.e., a stream of incoming tasks 430) that is otherwise very expensive to process using the traditional CPU or GPU computational resources. The workload can vary and the request patterns of users 435 can be unknown. By employing the TSP farm 420, it is possible to dynamically change the quality of output results. For example, the TSP farm 420 is configured to process 200 tasks at a first quality level or 400 tasks at a second quality level that is lower than the first quality level. Details about dynamically changing the quality of output results are described below in relation to FIG. 4B (¶ 0089); Fontijn more expressively teaches this limitation – as shown in figures 3a, 3b, and 3c, where a plurality of time intervals each marked by a starting time point X and an ending time point X denote when and how long a data block is needed; More advanced cache management policies use a profile to predict the data blocks that will be used during execution of a program. Profiling, i.e. the establishment of a profile, involves recording which data blocks are used and in which sequence the data blocks are used during execution of the program, and computing statistical data from repeated executions of the program. In such a profile based cache management policy the discardable data blocks are selected so as to minimize the probability of cache misses, i.e. instances during execution of the program when a data block is not in cache when needed by the program … To realize this, the cache management unit makes use of a prediction when data blocks will be used. When selecting a particular data block that will not be retained the cache management unit takes account of whether other data blocks from the same fetch unit as the particular data block are in the cache memory and when these other data blocks are expected to be used (¶ 0004-0008); In order to determine which data blocks have to be copied to or retained in cache memory 12, cache management unit 16 uses a prediction of the data blocks that will be needed by processor 14 during execution of the application program … This profile may be provided with the program, or recorded during a first execution of the program. During subsequent executions the recorded order is used to predict when the data blocks will be needed (¶ 0034); FIGS. 3a,b are used to illustrate the operation of selection algorithms … Some data blocks are shown to be used more than once at different times points, making it attractive to retain these data blocks in came memory 12 in between. This is expressed in FIG. 3a by drawing the time intervals between the different time points at which data from the same data block is used as solid line segments 32a-f … Assuming for example that the cache has capacity for four blocks, it can be seen that a conflict exists from the start of interval 32e … The algorithm searches for blocks whose cache memory location may be reused during certain time intervals. The time intervals that the search algorithm select run from one time point where a block is used to a next time point where the block is used again. That is, the time intervals are the time intervals 32a-f … (¶ 0036-0045); … a cache management unit (16) arranged to select which data blocks from within the fetch units not to retain in the cache memory (12), dependent on a prediction of when which data block will be needed during execution of the program … (claim 1)]; generating a streaming plan for usage of the streaming tensor, where the streaming plan runs in a background execution environment [a streaming plan as shown in figure 5; The operations of method 500 can be initiated by the compiler operating on at least one computer processor and/or on a host server integrated into the deterministic streaming system or separate from the deterministic streaming system … The deterministic streaming system evaluates 505 (e.g., by the scheduler) a latency for each task of a plurality of tasks to be run at the deterministic streaming system. The deterministic streaming system adjusts 510 (e.g., by the scheduler) at least one of an accuracy metric and a quality metric for an output of each of the plurality of tasks based on the evaluated latency until the plurality of tasks can be completed before expiration of one or more contractual deadlines. The deterministic streaming system runs 515, by at least a subset of the plurality of deterministic streaming processors of the deterministic streaming system, the plurality of tasks each having the output with at least one of the adjusted accuracy metric and the adjusted quality metric. The deterministic streaming system selects (e.g., by the scheduler) a pre-compiled model variation for compilation (e.g., by the compiler). The deterministic streaming system selects (e.g., by the scheduler) quality and accuracy information during a static capacity planning process for when the scheduler decides which model variations should be compiled … In some embodiments, the deterministic streaming system compiles (e.g., by the compiler) source code of each model of a plurality of models associated with the plurality of tasks into an intermediate representation … The deterministic streaming system calculates (e.g., by the compiler) an amount of computation that can be performed within a period of time for each of the plurality of tasks, and provides information about the calculated amount of computation to the scheduler for the evaluation of latency for each task … The deterministic streaming system meets a defined QoE metric based on at least the subset of the plurality of deterministic streaming processors running the plurality of tasks each having the output with at least one of the adjusted accuracy metric and the adjusted quality metric … (¶ 0119-0124); “capacity planning” as shown in figure 4B, 465; The process for ensuring the drainage condition is simplified due to the deterministic nature of the TSP farm 420. A non-real-time subcomponent of the scheduler 425 that can be referred to as a “capacity planner” … If the simulation can drain the leaky buckets within all registered contractual agreements, then the capacity planner would proceed with the registration. Otherwise, the capacity planner determines the new registration to be infeasible and would require a user 435 to change their registration parameters to be less intensive on the TSP farm 420 … (¶ 0111-0112); The benefits of involving the scheduler 425 in the compilation process arise from the fact that the scheduler 425 supports a plurality of models 415 for a plurality of users 435. If a new model 415 belonging to an arbitrary user 435 is registered to the TSP farm 420 with pre-existing registered models 415, the scheduler 425 can elect to change which binary variations would be utilized for any subset of existing pre-registered models 415 as part of its optimization routine (e.g., when ensuring the drainage condition for capacity planning, as discussed in more detail in the section below). Partial compilation is useful to expedite this process because, otherwise, recompilation of models 415 would be required. Additionally, the scheduler 415 can perform its part of compilation process outside the critical path of incoming requests as, e.g., a background job. Otherwise, the non-determinism and an additional latency would be introduced to the incoming requests as model compilation itself is not deterministic … The benefit of splitting the compilation of model 415 between the compiler 410 and the scheduler 425 is that the scheduler 425 can dynamically modify a manner of running the compiled model 415 during runtime in the background … (¶ 0103-0105)]; implementing memory management based on the streaming plan [ … Each deterministic streaming processor is divided into a plurality of functional units organized into a plurality of functional slices. Each functional slice is configured to perform specific functions within the deterministic streaming processor, which can include memory functional slices (MEMs) for storing operand data, arithmetic functional slices for performing operations on received operand data (e.g., vector processing, matrix manipulation), and/or the like … (¶ 0012-0017)], wherein the memory management includes: determining when the streaming tensor is needed by the processor of a second class [as shown in figure 4A, where users (user 1 to user n, 435) request for services, which are translated into the corresponding tasks (task 1 to task n, 430), and in response, Tensor Streaming Processors (TSP 1 to TSP n, 420) are activated to serve users’ requests; FIG. 4A illustrates an example deterministic cloud system 400, in accordance with some embodiments. The deterministic cloud system 400 is implemented as a serverless cloud configuration with multiple TSPs configured to manage, e.g., Deep Neural Network (DNN) inference workloads … (¶ 0083-0089)] receiving request for the streaming tensor [as shown in figure 4A, where users (user 1 to user n, 435) request for services, which are translated into the corresponding tasks (task 1 to task n, 430), and in response, Tensor Streaming Processors (TSP 1 to TSP n, 420) are activated to serve users’ requests; FIG. 4A illustrates an example deterministic cloud system 400, in accordance with some embodiments. The deterministic cloud system 400 is implemented as a serverless cloud configuration with multiple TSPs configured to manage, e.g., Deep Neural Network (DNN) inference workloads … (¶ 0083-0089); In one or more embodiments, the deterministic cloud system 400 would offer reserved execution of models 415 for customers with strict SLA requirements. This requires registering their model 415 before issuing tasks 430 by providing a variety of constraints of the TSP farm 420 and constraints of users 435. The constraints of the TSP farm 420 can be, e.g., required latency SLAs of registered models 415, quality SLAs, and accuracy SLAs. The users 435 can be constrained to issuing a maximum inferences per second (IPS) (i.e., constraining an average request load), and a maximum request queue size (i.e., constraining a peak request load) … (¶ 0109-0112); Becchi also teaches this limitation -- figure 6, step 602, “determine size and location of requested data”]; at run time determining if the streaming tensor is resident on the memory of the processor of a second class [this limitation is taught by Becchi -- Referring now to FIG. 3b, if the block B 308 resides in GPU memory, no memory allocation is performed and the content of the data block list is used to return the proper GPU address … (¶ 0044)]; if the streaming tensor is resident on the memory of the processor of a second class, i) retrieving the streaming tensor from the memory of the processor of a second class, ii) using the retrieved streaming tensor in the execution of the at last one instruction, and iii) generating an output [this limitation is taught by Becchi -- as shown in figure 6; Referring now to FIG. 3b, if the block B 308 resides in GPU memory, no memory allocation is performed and the content of the data block list is used to return the proper GPU address … (¶ 0044-0045)]; if not resident on the memory of the processor of a second class, [this limitation is taught by Becchi -- If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 … (¶ 0043-0045)], i) retrieving the streaming tensor from the memory of the processor of a first class, ii) copying the streaming tensor onto the memory of the processor of a second class, iii) using the retrieved input data in the execution of the at last one instruction, and iv) generating an output [this limitation is taught by Becchi -- as shown in figure 6; If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 (¶ 0043); When a kernel is invoked on CPU, the runtime must ensure that the CPU memory has an up-to-date copy of all input parameters … After execution of a CPU kernel call, output parameters are marked as residing on the CPU memory … (¶ 0043-0045)]. Regarding claim 1, Patrick does not teach coping the needed tensor data into the tensor streaming processor (TSP) if it does not reside in the memory of the TSP. However, Becchi specifically teaches coping the needed data into a GPU from a CPU if it does not reside in the memory of the GPU [as shown in figure 6; If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 (¶ 0043); When a kernel is invoked on CPU, the runtime must ensure that the CPU memory has an up-to-date copy of all input parameters … After execution of a CPU kernel call, output parameters are marked as residing on the CPU memory … (¶ 0043-0045)]. Therefore, it would have been obvious for one of ordinary skills in the art before the effective filing date of the claimed invention to copy the needed data into a GPU/TSP from a CPU if it does not reside in the memory of the GPU/TSP, as specifically demonstrated by Becchi, and to incorporate it into the existing scheme disclosed by Patrick because Becchi teaches doing so allows better utilizing the performance of both CPU and GPU [If an application has three candidate kernels with both CPU and GPU implementations and, during a certain execution path, the first kernel is estimated to be much faster, but the second and third much slower on the GPU (based on the sizes of their parameters), a data-agnostic scheduler is likely to run the first kernel on the GPU, and the rest on the CPU … (¶ 0028)]. Further regarding claim 1, Patrick in view of Becchi dose not expressively teach determining when and for how long the streaming tensors are needed by the processor of the second class. However, Fontijn specifically teaches determining when and for how long data blocks are needed by the processor executing a functional program [as shown in figures 3a, 3b, and 3c, where a plurality of time intervals each marked by a starting time point X and an ending time point X denote when and how long a data block is needed; More advanced cache management policies use a profile to predict the data blocks that will be used during execution of a program. Profiling, i.e. the establishment of a profile, involves recording which data blocks are used and in which sequence the data blocks are used during execution of the program, and computing statistical data from repeated executions of the program. In such a profile based cache management policy the discardable data blocks are selected so as to minimize the probability of cache misses, i.e. instances during execution of the program when a data block is not in cache when needed by the program … To realize this, the cache management unit makes use of a prediction when data blocks will be used. When selecting a particular data block that will not be retained the cache management unit takes account of whether other data blocks from the same fetch unit as the particular data block are in the cache memory and when these other data blocks are expected to be used (¶ 0004-0008); In order to determine which data blocks have to be copied to or retained in cache memory 12, cache management unit 16 uses a prediction of the data blocks that will be needed by processor 14 during execution of the application program … This profile may be provided with the program, or recorded during a first execution of the program. During subsequent executions the recorded order is used to predict when the data blocks will be needed (¶ 0034); FIGS. 3a,b are used to illustrate the operation of selection algorithms … Some data blocks are shown to be used more than once at different times points, making it attractive to retain these data blocks in came memory 12 in between. This is expressed in FIG. 3a by drawing the time intervals between the different time points at which data from the same data block is used as solid line segments 32a-f … Assuming for example that the cache has capacity for four blocks, it can be seen that a conflict exists from the start of interval 32e … The algorithm searches for blocks whose cache memory location may be reused during certain time intervals. The time intervals that the search algorithm select run from one time point where a block is used to a next time point where the block is used again. That is, the time intervals are the time intervals 32a-f … (¶ 0036-0045); … a cache management unit (16) arranged to select which data blocks from within the fetch units not to retain in the cache memory (12), dependent on a prediction of when which data block will be needed during execution of the program … (claim 1)]. Therefore, it would have been obvious for one of ordinary skills in the art before the effective filing date of the claimed invention to determine when and for how long data blocks are needed by the processor executing a functional program, as specifically demonstrated by Fontijn, and to incorporate it into the existing scheme disclosed by Patrick in view of Becchi, because Fontijn teaches doing so minimize the number of times that data be fetched from a disk into the cache memory, and better utilize the cache memory space [More advanced cache management policies use a profile to predict the data blocks that will be used during execution of a program. Profiling, i.e. the establishment of a profile, involves recording which data blocks are used and in which sequence the data blocks are used during execution of the program, and computing statistical data from repeated executions of the program. In such a profile based cache management policy the discardable data blocks are selected so as to minimize the probability of cache misses, i.e. instances during execution of the program when a data block is not in cache when needed by the program. Typically, a data block is discarded from memory if it is predicted from the profile that the data block will not be reused, or, if all data blocks in the memory are predicted to be reused, a data block that is predicted to be used last is discarded, because it leaves most room over time. Thus the number times a data block has to be fetched from disc is minimized … (¶ 0004-0005)]. As to claim 6, Patrick in view of Becchi & Fontijn teaches The method of claim 1, further comprising: initially analyzing the program; and generating a streaming plan for usage of the streaming tensor, where the streaming plan runs in a background execution environment and includes determining when the streaming tensor is needed by the processor of a second class, if not already in the memory of the processor of a second class, prefetching the streaming tensors into the memory of the processor of a second class ahead of when the streaming tensor is needed by the processor of a second class to avoid the processor of a second class waiting for the streaming tensor [Becchi -- The eval_loc routine (line 3) is also defined within the function call handler 204, and determines the best target for the intercepted function call at 208. This decision is made by estimating the data transfer time of the input parameters and the kernel execution time on both CPU and GPU. The runtime transfers data only when such data do not reside on the memory module where they are needed for execution … (¶ 0036); If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 (¶ 0043); Each data block consists of a linear start address, a byte size, a device address, a location (identifier of the device where the data have been allocated), the timestamp of the last access to the block on device, and a synchronization status, indicating whether the content of CPU and device memory is synchronized or whether the up-to-date copy of the data in the block resides in CPU or device memory. Additionally, in the case of integrated devices, an additional field indicates the address in page-locked memory where the data block has been mapped (¶ 0067)]. As to claim 8, Patrick in view of Becchi & Fontijn teaches The method of claim 6, wherein the step of analyzing the program includes performing a profile run at run time to determine structure the program [Becchi -- If an application has three candidate kernels with both CPU and GPU implementations and, during a certain execution path, the first kernel is estimated to be much faster, but the second and third much slower on the GPU (based on the sizes of their parameters), a data-agnostic scheduler is likely to run the first kernel on the GPU, and the rest on the CPU. However if the runtime discovers that the first kernel produces a large amount of data that is consumed by the second kernel, a better schedule may be to run the second kernel also on the GPU … A runtime according to the present principles analyzes such situations using history-based models to predict processing as well as data transfer time and uses these to guide the scheduling policy … The runtime has mechanisms to ensure coherent access to multiple copies of the same data residing in different memories (e.g., CPU and GPU memory) … (¶ 0028-0031)]. As to claim 9, Patrick in view of Becchi & Fontijn teaches The method of claim 8, wherein the program is a functional differentiable program [Becchi -- Systems and method for data-aware scheduling of applications on a heterogeneous platform having at least one central processing unit (CPU) and at least one accelerator. Such systems and methods include a function call handling module configured to intercept, analyze, and schedule library calls on a processing element. The function call handling module further includes a function call interception module configured to intercept function calls to predefined libraries, a function call analysis module configured to analyze argument size and location, and a function call redirection module configured to schedule library calls and data transfers. The systems and methods also use a memory unification module, configured to keep data coherent between memories associated with the at least one CPU and the at least one accelerator based on the output of the function call redirection module (abstract)]. As to claim 10, Patrick in view of Becchi & Fontijn teaches The method of claim 9, wherein the functional differentiable program is a structured network [Patrick -- FIG. 4A illustrates an example deterministic cloud system 400, in accordance with some embodiments. The deterministic cloud system 400 is implemented as a serverless cloud configuration with multiple TSPs configured to manage, e.g., Deep Neural Network (DNN) inference workloads … (¶ 0083); Becchi -- Systems and method for data-aware scheduling of applications on a heterogeneous platform having at least one central processing unit (CPU) and at least one accelerator. Such systems and methods include a function call handling module configured to intercept, analyze, and schedule library calls on a processing element. The function call handling module further includes a function call interception module configured to intercept function calls to predefined libraries, a function call analysis module configured to analyze argument size and location, and a function call redirection module configured to schedule library calls and data transfers. The systems and methods also use a memory unification module, configured to keep data coherent between memories associated with the at least one CPU and the at least one accelerator based on the output of the function call redirection module (abstract); Network adapters may also be coupled to the system to enable the data processing system to become coupled to other data processing systems or remote printers or storage devices through intervening private or public networks. Modems, cable modem and Ethernet cards are just a few of the currently available types of network adapters (¶ 0025)]. As to claim 11, Patrick in view of Becchi & Fontijn teaches The method of claim 10, wherein the structured network is a neural network [Patrick -- FIG. 4A illustrates an example deterministic cloud system 400, in accordance with some embodiments. The deterministic cloud system 400 is implemented as a serverless cloud configuration with multiple TSPs configured to manage, e.g., Deep Neural Network (DNN) inference workloads … (¶ 0083)]. As to claim 12, Patrick in view of Becchi & Fontijn teaches The method of claim 6, wherein the streaming plan includes pre-fetching data for the memories of the processor of a second class based on a window of future cycles of the one or more processors of a second class [Becchi -- A runtime operates at the granularity of a function call. An application runs by default on the CPU and may perform calls to well known kernels for which multiple implementations--either targeting CPU or GPU--are provided. When one of these computational kernels is invoked, the runtime determines the implementation to instantiate. This decision depends on two factors: the kernel execution time and the data transfer time. In turn, these factors depend on the size of the function call parameters and on the location of the corresponding data. GPU kernel implementations assume that their parameters reside on the GPU memory. It is therefore the responsibility of the runtime to hide this fact from the calling application and to maintain a mapping between data structures residing on CPU and on GPU memories. Data is not transferred to the CPU memory at the end of each GPU kernel invocation, but only when that data is used (¶ 0031); If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 (¶ 0043)]. As to claim 13, Patrick in view of Becchi & Fontijn teaches The method of claim 12, wherein size of the window is predefined [Becchi -- A runtime operates at the granularity of a function call. An application runs by default on the CPU and may perform calls to well known kernels for which multiple implementations--either targeting CPU or GPU--are provided. When one of these computational kernels is invoked, the runtime determines the implementation to instantiate. This decision depends on two factors: the kernel execution time and the data transfer time. In turn, these factors depend on the size of the function call parameters and on the location of the corresponding data. GPU kernel implementations assume that their parameters reside on the GPU memory. It is therefore the responsibility of the runtime to hide this fact from the calling application and to maintain a mapping between data structures residing on CPU and on GPU memories. Data is not transferred to the CPU memory at the end of each GPU kernel invocation, but only when that data is used (¶ 0031)]. As to claim 14, Patrick in view of Becchi & Fontijn teaches The method of claim 12, wherein size of the window is provided by a user [Becchi -- Heterogeneous platforms are those with both a multi-core central processing unit (CPU) and a many-core accelerated processor such as a graphics processing unit (GPU). To realize the higher performance that such platforms can deliver, however, programmers need intimate knowledge of the GPU architecture. In order to help the common programmer develop code for such platforms, GPU implementations of several "kernels" are made available as libraries. Thus each library kernel has both a CPU and GPU implementation (¶ 0005); The present principles enable legacy applications to automatically run on heterogeneous platforms with minimal data transfers and with full data coherence. The operating system and runtime may be used to provide the programmer with a unified memory view of possibly discrete underlying memory sub-systems … (¶ 0052)]. As to claim 15, Patrick in view of Becchi & Fontijn teaches The method of claim 12, wherein size of the window is adaptive based on the out-of-memory errors [Becchi -- … The runtime transfers data only when such data do not reside on the memory module where they are needed for execution. eval_loc queries the memory access module 206 for the location of each input parameter, and estimates the data transfer time based on the parameter size. In case of GPU execution, eval_loc considers the size and the location of the output parameters to determine whether the GPU has enough free memory to allocate them … (¶ 0036)]. As to claim 16, Patrick in view of Becchi & Fontijn teaches The method of claim 6, wherein the step of analyzing the program includes performing a static analysis at compiler stage to determine variable control flow of the program [Becchi -- Ideally, a heterogeneous system should enable any legacy code written for homogeneous systems to run faster, transparently to the programmer. Library-based programming, where pre-compiled assembly-level libraries for common kernels on the accelerators are made available, eases the burden of parallelizing applications on heterogeneous systems such as that shown in FIG. 1 … (¶ 0027)]. As to claim 17, Patrick in view of Becchi & Fontijn teaches The method of claim 16, wherein the static analysis is adaptive based on a speculative execution scheme [Becchi -- Systems and method for data-aware scheduling of applications on a heterogeneous platform having at least one central processing unit (CPU) and at least one accelerator. Such systems and methods include a function call handling module configured to intercept, analyze, and schedule library calls on a processing element. The function call handling module further includes a function call interception module configured to intercept function calls to predefined libraries, a function call analysis module configured to analyze argument size and location, and a function call redirection module configured to schedule library calls and data transfers. The systems and methods also use a memory unification module, configured to keep data coherent between memories associated with the at least one CPU and the at least one accelerator based on the output of the function call redirection module (abstract); A runtime operates at the granularity of a function call. An application runs by default on the CPU and may perform calls to well known kernels for which multiple implementations--either targeting CPU or GPU--are provided. When one of these computational kernels is invoked, the runtime determines the implementation to instantiate. This decision depends on two factors: the kernel execution time and the data transfer time. In turn, these factors depend on the size of the function call parameters and on the location of the corresponding data. GPU kernel implementations assume that their parameters reside on the GPU memory. It is therefore the responsibility of the runtime to hide this fact from the calling application and to maintain a mapping between data structures residing on CPU and on GPU memories. Data is not transferred to the CPU memory at the end of each GPU kernel invocation, but only when that data is used (¶ 0031)]. As to claim 19, Patrick in view of Becchi & Fontijn teaches The method of claim 1, wherein the one or more processors of a second class includes a coprocessor [Becchi -- Heterogeneous platforms are those with both a multi-core central processing unit (CPU) and a many-core accelerated processor such as a graphics processing unit (GPU) … (¶ 0005-0006)]. As to claim 20, Patrick in view of Becchi & Fontijn teaches The method of claim 19, wherein the coprocessor include graphics processing units [Becchi -- Heterogeneous platforms are those with both a multi-core central processing unit (CPU) and a many-core accelerated processor such as a graphics processing unit (GPU) … (¶ 0005-0006)]. As to claim 21, Patrick in view of Becchi & Fontijn teaches The method of claim 19, wherein the coprocessor include tensor processing units [Patrick -- Disclosed are configurations that include a deterministic streaming system with one or more deterministic streaming processors (e.g., tensor streaming processors (TSPs) or artificial intelligence processors) … The disclosed embodiments are directed to one or more deterministic streaming processors each having a functional slicing architecture. In some embodiments, each deterministic streaming processor comprises a tensor streaming processor (TSP) having a functional slicing architecture, which can be used for hardware-accelerated machine learning (ML) applications … The computational elements of the deterministic streaming processor can be divided between different functionalities (e.g., memory, arithmetic operation, etc.), and can be organized as functional slices which operate on multi-dimensional data (e.g., tensors) … (¶ 0037-0039)]. As to claim 23, Patrick in view of Becchi & Fontijn teaches The method of claim 1, the copying steps, includes: determining if there is sufficient contiguous memory available in the memory of the processor of a second class; if there is sufficient contiguous memory, proceeding with the copying step [Becchi -- … The runtime transfers data only when such data do not reside on the memory module where they are needed for execution. eval_loc queries the memory access module 206 for the location of each input parameter, and estimates the data transfer time based on the parameter size. In case of GPU execution, eval_loc considers the size and the location of the output parameters to determine whether the GPU has enough free memory to allocate them … (¶ 0036); If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 (¶ 0043); When a kernel is invoked on CPU, the runtime must ensure that the CPU memory has an up-to-date copy of all input parameters … After execution of a CPU kernel call, output parameters are marked as residing on the CPU memory … (¶ 0047-0048)]. 5. Claims 24-27 are rejected under 35 U.S.C. 103 as being unpatentable over Patrick in view of Becchi & Fontijn, and further in view of Payer et al. (US Patent 9,448,929, hereinafter Payer). Regarding claim 24, Patrick in view of Becchi & Fontijn does not teach if there is insufficient contiguous memory in the memory of the processor of the second class, calling a garbage collector adapted to remove unneeded data in the memories of the one or more processors of a second class. However, Payer specifically teaches if there is sufficient contiguous memory, performing the copying or writing operation to the memory of the one or more processors of a second class; if there is not sufficient contiguous memory, calling a garbage collector adapted to remove unneeded data in the memories [… the inserted instruction being configured to determine if a contiguous memory block of sufficient size to allocate the first amount of memory and the second amount of memory is available and, if the contiguous memory block of sufficient size is not available, trigger garbage collection … (c5 L35-65)]. Therefore, it would have been obvious for one of ordinary skills in the art before the effective filing date of the claimed invention to determine if there is sufficient contiguous memory, performing the copying or writing operation to the memory of the one or more processors of a second class; if there is not sufficient contiguous memory, calling a garbage collector adapted to remove unneeded data in the memories, as expressively demonstrated by Payer, and to incorporate it into the existing scheme disclosed by Patrick in view of Becchi & Fontijn, in order to obtain enough memory space to accommodate the desired data. As to claim 25, Patrick in view of Becchi & Fontijn & Payer teaches The method of claim 24, further comprising if there is still insufficient contiguous memory in the memories of the processor of a second class and collectively there is sufficient contiguous and non-contiguous memory, then compacting data in the memories memory of the processor of a second class, and if there is sufficient contiguous memory in the memory of the processor of a second class after the compacting data step, then proceeding with the copying step [Payer -- … Garbage collection can also be used to defragment a block of memory that is associated with a given program. Such a defragmentation process can group (move) live (active) memory objects together in memory and, as a result, free up larger blocks (sections, chunks, etc.) of available (free, unassigned, and so forth) memory space by eliminating portions of free memory that located between live objects (fragmented memory). These portions of unused (fragmented) memory may, for instance, be associated with objects that are no longer being used by the given application (c1 L30-46); Becchi -- … The runtime transfers data only when such data do not reside on the memory module where they are needed for execution. eval_loc queries the memory access module 206 for the location of each input parameter, and estimates the data transfer time based on the parameter size. In case of GPU execution, eval_loc considers the size and the location of the output parameters to determine whether the GPU has enough free memory to allocate them … (¶ 0036); If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 (¶ 0043); When a kernel is invoked on CPU, the runtime must ensure that the CPU memory has an up-to-date copy of all input parameters … After execution of a CPU kernel call, output parameters are marked as residing on the CPU memory … (¶ 0047-0048)]. As to claim 26, Patrick in view of Becchi & Fontijn & Payer teaches The method of claim 25, further comprising if there is still insufficient contiguous memory in the memories of the processor of a second class, then removing one or more recently used data in the memories of the processor of a second class [Payer -- The garbage collector 147 may be implemented as a generational garbage collector, though other types of garbage collectors may be used. The garbage collector 147 can use a semi-space strategy that classifies objects as “young generation” objects, which have not yet been observed and/or moved by the garbage collector 147, and “old generation” objects, which have been previously observed and/or moved by the garbage collector. In such approaches, the garbage collector 147 can be configured to perform frequent minor collections of the young generation objects, and may also implement a mark-and-sweep collector with incremental marking for major collections of the old generation objects (c7 L27-39)], and if there is sufficient contiguous memory in the memory of the processor of a second class after the removing the one or more least recently used streaming tensors step, then proceeding with the copying step, and if there is still insufficient contiguous memory in the memory of the processor of a second class and collectively there is sufficient contiguous and non-contiguous memory, then re- compacting data in the memory of the processor of a second class, and if there is sufficient contiguous memory in the memory of the processor of a second class after the re-compacting data step, then proceeding with the copying step [Payer -- … Garbage collection can also be used to defragment a block of memory that is associated with a given program. Such a defragmentation process can group (move) live (active) memory objects together in memory and, as a result, free up larger blocks (sections, chunks, etc.) of available (free, unassigned, and so forth) memory space by eliminating portions of free memory that located between live objects (fragmented memory). These portions of unused (fragmented) memory may, for instance, be associated with objects that are no longer being used by the given application (c1 L30-46); Becchi -- … The runtime transfers data only when such data do not reside on the memory module where they are needed for execution. eval_loc queries the memory access module 206 for the location of each input parameter, and estimates the data transfer time based on the parameter size. In case of GPU execution, eval_loc considers the size and the location of the output parameters to determine whether the GPU has enough free memory to allocate them … (¶ 0036); If the requested block does not reside in GPU memory, as shown in FIG. 3a, a GPU memory allocation is performed and a new entry is added to the data block list. FIG. 3a shows a block list 302 that includes a block A 306, but not the requested block B 308. The memory allocation is followed by a data transfer (from CPU to GPU) 216 only if the update parameter of the get call is set to true. The resulting data block list 304 then includes the synced block B 308 (¶ 0043); When a kernel is invoked on CPU, the runtime must ensure that the CPU memory has an up-to-date copy of all input parameters … After execution of a CPU kernel call, output parameters are marked as residing on the CPU memory … (¶ 0047-0048)]. As to claim 27, Patrick in view of Becchi & Fontijn & Payer teaches The method of claim 26, further comprising determining if there is still insufficient contiguous memory in the memories of the processor of a second class, then halt execution of the at least one instruction and issuing an out-of- memory error [Becchi -- … The runtime transfers data only when such data do not reside on the memory module where they are needed for execution. eval_loc queries the memory access module 206 for the location of each input parameter, and estimates the data transfer time based on the parameter size. In case of GPU execution, eval_loc considers the size and the location of the output parameters to determine whether the GPU has enough free memory to allocate them … (¶ 0036)]. Conclusion 6. Claims 1, 6, 8-17, 19-21, and 23-27 are rejected as explained above. 7. THIS ACTION IS MADE FINAL. Applicant is reminded of the extension of time policy as set forth in 37 CFR 1.136(a). A shortened statutory period for reply to this final action is set to expire THREE MONTHS from the mailing date of this action. In the event a first reply is filed within TWO MONTHS of the mailing date of this final action and the advisory action is not mailed until after the end of the THREE-MONTH shortened statutory period, then the shortened statutory period will expire on the date the advisory action is mailed, and any extension fee pursuant to 37 CFR 1.136(a) will be calculated from the mailing date of the advisory action. In no event, however, will the statutory period for reply expire later than SIX MONTHS from the mailing date of this final action. 8. Any inquiry concerning this communication or earlier communications from the examiner should be directed to SHENG JEN TSAI whose telephone number is 571-272-4244. The examiner can normally be reached on Monday-Friday, 9-6. If attempts to reach the examiner by telephone are unsuccessful, the examiner’s supervisor, Reginald Bragdon can be reached on 571-272-4204. The fax phone number for the organization where this application or proceeding is assigned is 571-273-8300. Information regarding the status of an application may be obtained from the Patent Application Information Retrieval (PAIR) system. Status information for published applications may be obtained from either Private PAIR or Public PAIR. Status information for unpublished applications is available through Private PAIR only. For more information about the PAIR system, see http://pair-direct.uspto.gov. Should you have questions on access to the Private PAIR system, contact the Electronic Business Center (EBC) at 866-217-9197 (toll-free). /SHENG JEN TSAI/Primary Examiner, Art Unit 2139
Read full office action

Prosecution Timeline

Show 3 earlier events
Sep 26, 2025
Final Rejection mailed — §103
Feb 17, 2026
Request for Continued Examination
Feb 20, 2026
Response after Non-Final Action
Mar 13, 2026
Non-Final Rejection mailed — §103
Jun 11, 2026
Applicant Interview (Telephonic)
Jun 11, 2026
Examiner Interview Summary
Jun 21, 2026
Response Filed
Aug 26, 2026
Final Rejection mailed — §103 (current)

Precedent Cases

Applications granted by this same examiner with similar technology

Patent 12730723
SELECTIVE PROCESSING OF FILE SYSTEM OBJECTS FOR IMAGE LEVEL BACKUPS
3y 0m to grant Granted Sep 08, 2026
Patent 12730578
DATA MIGRATION AND ASYNCHRONOUS REPLICATION WITH TRUSTWORTHY ENERGY AWARENESS
2y 7m to grant Granted Sep 08, 2026
Patent 12705184
USING RETIRED PAGES HISTORY FOR INSTRUCTION TRANSLATION LOOKASIDE BUFFER (TLB) PREFETCHING IN PROCESSOR-BASED DEVICES
2y 4m to grant Granted Aug 11, 2026
Patent 12670072
LOW IMPACT MIGRATION OF LARGE DATA TO CLOUD AND VIRTUALIZED ENVIRONMENTS
3y 0m to grant Granted Jun 30, 2026
Patent 12656954
COMPUTE EXPRESS LINK DRAM + NAND SYSTEM SOLUTION
2y 3m to grant Granted Jun 16, 2026
Study what changed to get past this examiner. Based on 5 most recent grants.

Strategy Recommendation AI-generated — please review before filing

Get a prosecution strategy drawn from examiner precedents, rejection analysis, and claim mapping.
Typically takes 5-10 seconds — AI-generated, attorney review required before filing

Prosecution Projections

5-6
Expected OA Rounds
70%
Grant Probability
84%
With Interview (+13.8%)
3y 4m (~9m remaining)
Median Time to Grant
High
PTA Risk
Based on 805 resolved cases by this examiner. Grant probability derived from career allowance rate.

Sign in with your work email

Enter your email to receive a magic link. No password needed.

Personal email addresses (Gmail, Yahoo, etc.) are not accepted.

Free tier: 3 strategy analyses per month